TOPIK to be graded by AI, trained on test-taker answers: Pilot exams in 2028
TOPIK moves to AI grading in 2029, trained on past test-takers' answers. Pilot exams in 2028, home tests with face recognition in 2029. ₩23.5 billion, no contractor or AI model named.

The TOPIK Korean language exam is set to hand its grading to artificial intelligence. Under a plan published by the Ministry of Education on September 29, AI will write TOPIK questions and score test-takers' writing and speaking from 2029, with pilot exams of the redesigned test in 2028.
Reports last year described the government negotiating with Korean tech conglomerate Naver to run the test. This time it is an official plan, with a committee, a ₩23.5 billion budget and a year-by-year schedule. More than 560,000 people applied for TOPIK in 2025, many of them for visas, university admission and jobs in Korea.
Let's look at what the plan contains, from AI grading and the data it will learn from to who will build it.
TL;DR
- From 2029, AI will write TOPIK questions and give writing and speaking answers their first score, with human graders checking its work
- The AI will be trained on past test-takers' handwritten and spoken answers
- Pilot exams of the redesigned test will run in 2028, followed by home-test pilots with face recognition and gaze tracking in 2029
- The plan cites no research showing AI can grade TOPIK reliably, and in research by the institute that sets the 수능, teachers put current accuracy at 60-70%
- A private contractor will build the system, and the plan names no company or AI model
- The plan cites Korean officials and administrators, but no foreign Korean-language experts
The 2029 Overhaul
The plan rebuilds TOPIK around an AI system with two jobs. A large language model trained on the test's question bank will write the questions for each level. The system also checks new questions against past ones and predicts how difficult they are.
Currently, the questions for each sitting are written by people kept in sealed-off sessions. From 2029 they will be drawn from a stored bank, which the plan expands by 4,000 questions a year. TOPIK II's writing section has four question types, from filling in sentences to a 600-700 character essay. The speaking test is a separate exam taken on a computer. For both, handwritten and recorded answers will be converted to text and scored on grammar, vocabulary and logical flow. The AI gives the first score and human graders then review it, a step the plan calls "Human-in-the-Loop."
Home-test pilots on personal computers follow in 2029. They are limited to lower levels whose results aren't used for study or visa purposes, while paper tests remain for now. Six of the 15 sittings for 2027 are on paper, with the other nine being computer-based: six IBT exams and three speaking tests.
A reform committee drawn from Korean language education, testing and AI will form this October and then publish a new test design in June 2027. Pilot exams are scheduled for 2028, and together with online practice questions they are meant to let test-takers try the new format before it takes effect. The government has budgeted ₩23.5 billion for the AI and digital work and set a target of one million applicants by 2030.
The committee is to be balanced by region, gender and affiliation. For the pool of question writers, region means whether members are based in the Seoul area or elsewhere in Korea, rather than something more global. Foreign Korean-language experts, people who learned Korean as a foreign language and could speak from the test-taker's side, appear nowhere in the plan. All four field opinions it cites come from Korean officials and administrators: two directors of the ministry's Korean Education Centers abroad, a diplomatic mission and the head of a university's international affairs office.
Unproven AI Grading
The reform committee's agenda covers level structure, matching paper and computer scores, writing and speaking tasks and a version of TOPIK for teenagers. Whether AI should grade the test isn't on this agenda, because that had already been decided before the plan was announced. It's also unclear whether the 2028 pilot exams will use AI grading, since the plan timeline suggests the system may still be under development that year. If so, the implicit scenario would be that the AI grading isn't tested through pilots at all.
Another thing missing from the plan is any cited research on how well AI grades TOPIK writing or speaking. In studies by the institute that sets the 수능 college entrance exam, teachers put current AI grading accuracy at 60-70%, and most said at least 80% would be needed before it could be used in practice. The researchers found that performance varied widely by subject and question. They also called for AI to check human graders' work, which is the reverse of TOPIK's design where human graders review the AI's output. Separately, a year-long study at a Korean science high school found that the same answer, graded again, could receive a different score.
Automated scoring isn't entirely new to language exams. ETS, which runs the TOEFL, has long used a scoring engine alongside human raters. Its own research found that for most groups, the engine and human raters gave almost identical average scores. But there was a notable deviation: essays from mainland China were the exception, receiving much higher scores from the machine than from people. TOPIK is held in 94 countries, so there is a real risk of similar variability here.
The plan's safeguard also has a built-in tension. If human graders check every answer closely, the AI saves little. If they don't, the AI is effectively the grader.
Last year, Lee Chang-yong, head of the Korean Language Teachers' Branch at the Workplace Gapjil 119 Online Union, set out what should come first: "Before AI can be used in a high-stakes test that seriously affects people's lives, research and verification to secure the validity, reliability and fairness of the assessment must come first. That verification has not yet been done sufficiently."
Caution Elsewhere
Korea is weighing the same idea for its own students, far more cautiously. AI grading of essay questions on the 수능 is still under public deliberation, and the head of the National Education Commission has said that even if it is introduced, human grading must remain central, with AI only as an aid.
Other countries have pulled back. Japan put planned essay questions for its common university entrance test on hold in 2019 because reliable grading was close to impossible to guarantee. Australia considered AI grading before cancelling it.
That leaves TOPIK as the only Korean national exam with a budgeted and scheduled plan for AI to give the first score on written and spoken answers.
Handwriting, Voices and Faces
In 2027 the government will start gathering training data for the AI, a year before the system itself is built. This training data includes the government's archive of real test-takers' handwritten and spoken answers from TOPIK's writing and speaking sections.
For home tests the plan adds AI proctoring that uses face recognition and gaze tracking to monitor test-takers in their own homes. At registration, test-takers will also have to consent to sharing personal data with third parties, which the plan ties to cooperation with "related institutions." Universities designated as TOPIK research institutes will be given test data for their research, under a legal basis the government aims to set up this year.
The plan doesn't say whether the company that builds the system will see test-takers' answers, recordings or face data, or which third parties the mandatory consent covers.
Korea has a precedent here. In a project that ended in 2021, the justice and science ministries let private companies train AI on immigration face photos. The privacy regulator later found that the companies had used nearly 180 million records, 120 million of them foreigners', ruled the arrangement a legitimate outsourcing of data processing, and fined the justice ministry ₩1 million, only for failing to disclose the companies' names. In February 2026 the Constitutional Court dismissed a challenge because the project had ended, adding that the Constitution creates no duty to ban using such data to train AI.
Who Builds the Grader
The bulk of the budget goes to building the system and developing its AI from 2028. That accounts for ₩23 billion of the ₩23.5 billion, with a strategy plan coming first in 2027. In Korea, systems like these are built by private contractors through public tenders.
Until December 2025, the plan was for a company to fund and run the test. In 2024 the government chose a consortium led by Naver for a privately financed project in which the company would have covered about ₩344 billion in costs, recovering them from the test's revenue. Korean language teachers and unions campaigned against it, and in December 2025 the government cancelled the arrangement and moved the project onto its own budget. While the full privatization is now gone, the AI grading, to be developed by a private company, has survived the switch.
Missing from the plan is any mention of who will build the system or which AI model it will use. The leading models are American, but the plan puts the platform on the government's own cloud, which makes sending unreleased questions and test-takers' answers to a foreign AI service unlikely. Most of the strongest models that can be run in-house are Chinese, and politically toxic: government ministries blocked DeepSeek in 2025, and Naver was dropped from the government's national AI model project in January 2026 amid controversy over its use of Chinese model components. That leaves Korean models. The institute that sets the 수능 has already hired Korean AI company Upstage to build question writing for 수능 English, but grading essays and speech is a much more demanding task than writing new questions.
Model choice matters more than it might seem. Through building TOPIK Easy6, we've found that no single AI model is best at every task in assessing Korean, and that only the most expensive proprietary models are accurate enough for fair grading. On top of that, a model's fluency in Korean tells surprisingly little about how well it grades a learner's essay.
The Ministry of Education's last major AI project offers a warning. After just one semester in 2025, AI digital textbooks for schools were downgraded to "educational material." The Board of Audit later found the project had been rushed without consultation or pilot runs, with average usage of 8.1%.
Timeline and Implementation
The reform committee is due to form this month, and a call for new test questions is planned for later this year. In 2027 the committee will publish the new test design in June. The government will also draw up its AI strategy that year and start building the question bank and gathering AI training data.
In 2028 the AI platform will be built on the government's cloud, and pilot exams of the redesigned test will run in Korea and abroad. The new TOPIK is set to take effect in 2029, with AI writing questions and grading answers, and home-test pilots beginning at lower levels.
Sources
- 교육부 - 국가 공인 한국어시험 토픽, 2029년 평가 개편 및 인공지능·디지털 기반 체제로 전환
- 교육부·국립국제교육원 - TOPIK 개편 및 시행 확대 방안
- 교육플러스 - [심층] AI가 수능 논술 채점한다면? 시간 '병목'은 풀렸지만 '공정성'은...
- ETS - Comparison of Human and Machine Scoring of Essays: Differences by Gender, Ethnicity, and Country
- 연합뉴스TV - 수능 서논술형·AI 채점 논란…해외선 '백지화'
- 뉴시스 - 수능 팔아넘기는 것과 같아…TOPIK 민간투자형 사업 반대 연서명 1만명↑
- 법조신문 - [결정] 헌재, 출입국 안면데이터 AI 학습 이전 사건 각하
