Evaluation platform and evaluation method for Chinese reading and spoken language annotation
By constructing a Chinese reading and speaking annotation evaluation platform with B/S architecture and microservice architecture, and combining AI automatic speech recognition and human review, multi-level time domain annotations are generated. This solves the problems of poor robustness to different languages and lack of fine-grained feedback in existing technologies, and achieves efficient and professional multi-dimensional speaking evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-02-28
- Publication Date
- 2026-05-15
AI Technical Summary
Existing Chinese oral assessment systems are not robust to different languages, lack a closed-loop mechanism of automatic assessment and manual review, cannot provide fine-grained feedback, and cannot effectively evaluate open-ended oral expressions, thus failing to meet the needs of international students to advance their abilities from pronunciation correction to free expression.
The platform, which adopts a B/S architecture and a microservice architecture, combines AI-powered automatic speech recognition to generate Praat annotation files and forms a closed-loop feedback mechanism with teacher manual review. Through the ASR engine and forced alignment algorithm, multi-level time domain annotations are generated. Teachers can perform visual verification and correction in the Praat software to build a data closed loop to optimize the model.
It significantly improves the annotation efficiency and professional accuracy of speech evaluation, reduces labor costs, achieves effective integration of automatic scoring and human review, provides multi-dimensional fine-grained feedback, supports the evaluation of open-ended spoken expressions, and enhances the ASR system's adaptability to non-native speaker accents.
Smart Images

Figure CN122050362A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent education assessment technology, and relates to an audio grading platform for international students' oral practice that integrates automatic speech recognition and manual review mechanisms. It is used to achieve multi-dimensional scoring and feedback on the audio of non-native language learners' reading and free expression. Specifically, it is a Chinese reading and oral annotation assessment platform and assessment method. Background Technology
[0002] With the rapid development of Automatic Speech Recognition (ASR) and deep learning technologies, Computer-Assisted Pronunciation Training (CAPT) systems have been widely used in the domestic education sector. This system integrates advanced technologies such as speech recognition, test administration, and psychology, automatically assessing learners' pronunciation accuracy, prosodic accuracy, and fluency through techniques including pronunciation and prosody evaluation.
[0003] In the field of Chinese language teaching for non-native speakers, a series of automatic assessment schemes based on speech recognition technology have been proposed. Patent "A Speech Assessment Method and Device Based on Pronunciation Rhythm," patent number: ZL 2012 10473420.2, authorized on February 11, 2015, discloses a speech assessment method based on pronunciation rhythm. This method uses a Gaussian Mixture Model (GMM) to extract rhythmic feature parameters from the speech being assessed, obtains likelihood values through model matching, and then automatically assesses the pronunciation rhythm of foreign students in Chinese. Another example is patent "An Automatic Method for Detecting Spoken Chinese Stress," patent number: CN101751919B, which discloses a method for automatically testing and evaluating Mandarin using speech recognition technology. It employs data-driven technology to build a model for detecting and diagnosing spoken stress.
[0004] In recent years, advancements have been made in technologies for multi-dimensional pronunciation evaluation. The patent application CN114283828A, titled "Method for Evaluating Spoken Pronunciation of Minor Languages," proposes calculating and analyzing speech from multiple dimensions, including accuracy, completeness, fluency, phrasing, tone, and intonation. It obtains multi-dimensional scores for the read-aloud audio by comparing phoneme decoding and alignment results. Furthermore, the patent CN116403567B, titled "Scoring Method for Speech Recognition Text," establishes a high-precision recognition model using a dynamic time-warped continuous speech recognition algorithm and leverages acoustic feature extraction technology to obtain multi-dimensional features such as pronunciation clarity, stress, and intonation.
[0005] However, the aforementioned existing technologies still have the following problems and drawbacks:
[0006] 1. Lack of robust recognition optimization for different languages. Existing Chinese spoken language assessment systems, such as the patents CN101751919B "An Automatic Method for Detecting Stress in Spoken Chinese" and ZL201210473420.2 "A Speech Assessment Method Based on Pronunciation Rhythm," are mainly designed for standard Mandarin pronunciation, and their acoustic models are based on native speaker speech training. However, the pronunciation of foreign students and other speakers of different languages generally exhibits "non-standard accent" characteristics such as phoneme substitution, tone errors, and prosodic abnormalities caused by negative transfer from the mother tongue. Existing ASR systems show a significantly higher word error rate (WER) for such speech, leading to a decrease in the reliability of subsequent scoring.
[0007] 2. Lack of a closed-loop mechanism integrating automatic evaluation and manual review. Existing technologies mostly adopt a single automatic scoring mode, such as the patent CN116403567B "Scoring Method for Speech Recognition Text" and the patent CN202010583611.9 "A Method and Device for Speech Scoring", or rely solely on manual review. There is a lack of a collaborative calibration mechanism between the two. Automatic scoring cannot handle audio segments with ASR recognition failure or low confidence, while manual review faces problems such as strong subjectivity, high cost, and feedback delay.
[0008] 3. Lack of fine-grained feedback tailored to practice scenarios. Existing assessment systems, such as the scoring methods, devices, electronic equipment, and storage media for speech recognition text (CN116403567B), primarily focus on providing a final score or grade, lacking visual annotations at the phoneme, tone, and stress levels. This fails to provide learners with targeted feedback on "where they misread and how to correct it." For international students, simply obtaining a score is insufficient to effectively guide targeted practice; an interactive feedback mechanism that provides fine-grained error localization and correction suggestions is urgently needed.
[0009] 4. Inability to handle the assessment of open-ended spoken expressions. Existing technologies, such as the automatic detection method for spoken Chinese stress in CN101751919B, mainly rely on forced alignment algorithms, requiring learners to read fixed texts and scoring based on phoneme-level alignment. For open-ended spoken expression tasks such as picture description and impromptu question-and-answer, existing technologies, such as the testing method and system for semi-open-ended spoken test questions in CN201110254211.4, mainly achieve reading assessment through forced text alignment or keyword matching. They lack the ability to evaluate dimensions such as content logic and pragmatic appropriateness in free expression, such as picture description and impromptu question-and-answer, and cannot meet the needs of cultivating advanced spoken language skills.
[0010] Therefore, there is an urgent need for an audio correction platform for international students' oral practice that can integrate the advantages of automatic speech recognition and human review. This platform should utilize ASR to achieve real-time automatic scoring and allow human intervention to optimize recognition blind spots, providing international students with multi-dimensional and fine-grained feedback, from pronunciation correction to free expression. Summary of the Invention
[0011] The technical problem to be solved by this invention is as follows:
[0012] 1. To address the poor robustness of existing ASR systems in recognizing foreign accents from different languages. Therefore, it is urgent to improve the adaptability of ASR systems to non-native speakers' accents;
[0013] 2. Address the lack of an effective integration mechanism between automatic scoring and human review. Existing technologies cannot manually annotate and re-enter incorrectly identified audio for model optimization, making it difficult to form a closed loop for continuous improvement.
[0014] 3. To address the problem of existing evaluation dimensions being too singular and lacking a comprehensive evaluation of suprasegmental features.
[0015] 4. Addressing the problem that existing technologies cannot effectively support the assessment of open-ended oral expression. Existing systems cannot evaluate the logical coherence, semantic completeness, and pragmatic appropriateness of responses, making it difficult to meet the needs of international students' progressive abilities from "pronunciation correction" to "free expression."
[0016] 5. Address the issue of existing systems lacking fine-grained visual feedback and personalized error correction guidance.
[0017] This invention provides a Chinese reading aloud and spoken language annotation and evaluation platform. It adopts a system design combining a B / S architecture and a microservice architecture, and uses AI to automatically generate Praat annotation files through speech recognition, forming a closed-loop feedback mechanism with manual review by teachers. Specifically, it includes:
[0018] The client layer includes teacher terminals and student terminals. The teacher terminals are equipped with Praat speech analysis software, which is used to open, edit, and save TextGrid format annotation files. The student terminals are used to collect and upload audio recordings.
[0019] The access layer includes an Nginx reverse proxy server, configured with a load balancing module and an SSL encryption module, used to receive client requests and route them to backend services;
[0020] The service layer includes a RESTful API service cluster built on the Flask framework, the service cluster comprising:
[0021] The voice receiving module is used to receive audio streams uploaded by student terminals and temporarily store them in a cache;
[0022] The ASR inference module is used for speech recognition and forced alignment of audio, generating phoneme-level time boundary information;
[0023] The annotation generation module is used to convert the time boundary information into a Praat standard TextGrid format file, which contains three layers of annotation: phoneme layer, tone layer, and error layer.
[0024] The difference comparison module is used to compare the AI-generated TextGrid file with the teacher-corrected TextGrid file, and calculate the time offset error and label error rate.
[0025] The data processing layer includes an ASR engine and a forced alignment module. The ASR engine is based on an end-to-end deep learning model, and the forced alignment module is used to align the recognized text with audio features on the time axis.
[0026] The storage layer includes a Redis cache database and a MySQL relational database. The Redis database is used to cache audio streams, MFCC feature vectors, and ASR intermediate results, while the MySQL database is used to store user identity information, audio file metadata, annotation version records, and rating history.
[0027] Furthermore, the three-level annotation of the TextGrid file specifically includes:
[0028] The phoneme layer records the actual phoneme sequence of pronunciation, including phoneme labels, start time, end time, and log likelihood score. Phonemes with a likelihood score below a preset threshold are marked as suspected pronunciation errors.
[0029] The tone layer records the results of the fundamental frequency F0 curve extraction, including tone type determination and tone deviation marking. Syllables whose predicted tone does not match the standard tone or whose F0 curve variation coefficient exceeds the threshold are marked with tone deviation indicators.
[0030] The error layer combines the confidence level of the phoneme layer and the deviation level of the tone layer. Intervals that simultaneously satisfy both a phoneme likelihood below a first threshold and a tone DTW distance greater than a second threshold are labeled as severely erroneous; intervals satisfying only one of these conditions are labeled as generally erroneous. The first threshold refers to the lower limit of the log-likelihood score used to determine the reliability of individual phoneme recognition; the value range of the first threshold is -5.0 to -2.0, with -3.0 yielding the best results. The second threshold refers to the upper limit of the DTW distance used to determine the accuracy of tone recognition; the value range of the second threshold is 0.15 to 0.35, with 0.25 yielding the best results.
[0031] Furthermore, the specific steps for the annotation generation module to generate a TextGrid format file include:
[0032] Extract the Mel frequency cepstral coefficients (MFCC) feature vector of the audio. The default frame length is 25ms, the frame shift is 10ms, and the number of Mel filter banks is 40. The configuration can be adjusted according to the audio frequency and length.
[0033] The audio is decoded using an ASR engine trained on WeNet or Kaldi to obtain the text recognition results;
[0034] A forced alignment algorithm based on Hidden Markov Model (HMM) or Deep Neural Network (DNN) is used to align the recognized text with MFCC features, generating phoneme-level, syllable-level, and word-level temporal boundary information.
[0035] The time boundary information is serialized according to the Praat TextGrid file format specification. Three IntervalTiers are set to correspond to the phoneme layer, tone layer and error layer respectively. Each IntervalTier contains a list of time intervals, and each time interval contains the start time, end time and marker text.
[0036] Output a UTF-8 encoded TextGrid file. The filename should include the student identifier, text identifier, version type, and timestamp.
[0037] Furthermore, the workflow of the difference comparison module includes:
[0038] Receive the TextGrid file returned by the teacher after correction using Praat software, and parse it into structured data;
[0039] Align and match AI-generated annotations and manually corrected annotations according to time intervals, and calculate the time offset error for each corresponding interval. The time offset error is the absolute value of the difference between the start time of AI annotation and the start time of manual annotation plus the absolute value of the difference between the end time.
[0040] The number of intervals with inconsistent phoneme labels is counted, and the label error rate is calculated, where the label error rate is the proportion of inconsistent intervals to the total number of intervals.
[0041] Count the number of syllables with consistent tone labels and calculate the tone consistency rate.
[0042] Generate a discrepancy report, including average time error, label error rate, tone consistency rate, and a list of specific error intervals.
[0043] Furthermore, the platform of the present invention also includes a model management module and a system control module:
[0044] The model management module is used to trigger the incremental learning process of the ASR model when the labeling error rate of a specific phoneme or a student group with a specific native language background exceeds a preset threshold or the tone consistency rate is lower than a preset threshold.
[0045] The incremental learning process includes: querying historical correction data of the phoneme or group from the MySQL database to construct a fine-tuned training set; performing 10-20 rounds of incremental training based on the original ASR model parameters; setting the learning rate to 0.05-0.2 times the original learning rate; and updating the online model when the test word error rate on the validation set decreases by more than 2%.
[0046] The system control module includes monitoring and managing system resources, as well as comprehensively managing entities stored in the system, including teacher accounts, student accounts, audio files, and annotation files.
[0047] Furthermore, the Nginx reverse proxy server is configured with:
[0048] The static resource caching module is used to cache the JS / CSS files built by the Angular frontend;
[0049] The rate limiting module is used to limit the frequency of audio upload requests from student terminals to prevent malicious attacks.
[0050] The WebSocket support module is used to push AI annotation generation progress and teacher review status in real time.
[0051] This invention also relates to a method for evaluating Chinese reading aloud and spoken language annotation, comprising the following steps:
[0052] Step A: The student terminal records audio of reading aloud through the Angular frontend, uploads it to the Flask service via Nginx, extracts the MFCC feature vector, and temporarily stores it in the Redis cache;
[0053] Step B: The ASR engine performs speech recognition decoding on the audio to obtain the text recognition result; a forced alignment algorithm is used to align the recognized text with the audio features to generate phoneme-level time boundary information;
[0054] The forced alignment algorithm includes:
[0055] Based on the Montreal Forced Aligner tool, using a standard Mandarin pronunciation dictionary;
[0056] The alignment results will be output in CSV format, including the fields: phoneme symbol, start time, end time, duration, and log-likelihood.
[0057] Phonemes with a log-likelihood below -3.0 are marked as low confidence for subsequent error layer labeling;
[0058] Step C: Generate three layers of annotation content based on the alignment results: The first layer is the phoneme layer, which records the actual pronunciation phoneme sequence and suspected error markers; the second layer is the tone layer, which extracts the fundamental frequency F0 curve and marks tone deviations; the third layer is the error layer, which automatically marks the error level by combining phoneme confidence and tone deviation.
[0059] Step D: Serialize the three-layer annotation data into the Praat standard TextGrid file format and push it to the teacher's terminal via the Nginx download interface.
[0060] Step E: The teacher opens the AI-generated TextGrid file in Praat software, manually verifies and corrects the phoneme boundaries, tone markings, and error markers, saves it as a new version of the TextGrid file, and uploads it back to the Flask service through the Angular frontend;
[0061] Step F: The difference comparison module parses the AI-annotated and manually annotated TextGrid files, calculates the time offset error and label error rate, and generates a difference matrix;
[0062] Step G: When the error rate of a specific phoneme or a student group with a specific native language background exceeds a preset threshold, the model fine-tuning process is triggered, the corrected alignment data is added to the training set, and the ASR model parameters are updated through incremental learning.
[0063] Step H: The MySQL database records the version number, annotator identity, timestamp, and confidence score for each annotation, supporting annotation history backtracking and multi-version comparison.
[0064] Furthermore, the specific operations for teacher manual correction in step E include:
[0065] Drag the Interval boundary on the timeline to adjust the phoneme segmentation position and correct the boundary where the time deviation exceeds 30ms;
[0066] Modify the mark text of the phoneme layer Interval to correct phoneme recognition errors;
[0067] Modify the tone type marker in the tone layer to correct tone determination errors;
[0068] Adjust the error level label at the error level, and upgrade, downgrade or cancel the error label based on professional judgment.
[0069] The above technical solution realizes a closed-loop technology for the entire chain of foreign students' spoken audio, from AI automatic annotation, teacher manual verification to continuous model optimization, and solves the technical problems of low ASR recognition accuracy, high manual review cost and lack of fine-grained feedback in the existing technology.
[0070] The present invention has the following beneficial effects
[0071] This invention significantly improves the efficiency and professional accuracy of speech evaluation annotation. It automatically generates TextGrid annotation files conforming to the Praat specification using AI, converting ASR decoding results into multi-level time-domain annotations at the phoneme, tone, and error levels. Teachers can directly perform visual verification and fine-tuning in the professional speech analysis software Praat. Compared to traditional purely manual annotation methods that require sentence-by-sentence dictation and manual boundary segmentation, this invention reduces annotation preparation time by approximately 60%-80% while maintaining the accuracy requirements of expert-level phonetics annotation.
[0072] This invention achieves an effective closed-loop integration of automatic scoring and human review, solving the problem of the separation between the two in existing technologies. It constructs a data closed loop of AI pre-annotation, teacher correction, difference comparison, and model micro-evaluation. By comparing the time offset error and label error rate between AI-generated annotations and teacher-corrected annotations, the system can automatically identify blind spots in ASR recognition for learners with specific phonemes, tones, or native language backgrounds, and trigger an incremental learning mechanism to continuously optimize the acoustic model. After 3-5 iterations, the system can reduce the word error rate (WER) for non-standard accents by 15%-25%, significantly outperforming the evaluation reliability of existing static ASR systems.
[0073] This invention provides fine-grained evaluation capabilities for suprasegmental features, overcoming the shortcomings of existing technologies in prosodic evaluation. Utilizing timestamp information generated by a forced alignment algorithm, the invention simultaneously annotates phoneme accuracy, tone contour deviation (calculated via the fundamental frequency F0 curve), and stress position in a TextGrid, achieving multi-dimensional parallel evaluation of pronunciation accuracy, tone accuracy, and prosodic fluency. Teachers can visually view students' fundamental frequency curve deviations in specific segments through the Praat visualization interface, providing data support for targeted pronunciation correction.
[0074] This invention adopts a standardized open-source format and a B / S architecture, improving the system's compatibility and scalability. Unlike the closed architecture of existing commercial evaluation systems (such as patent CN202010583611.9), this invention generates a standard PraatTextGrid file format (UTF-8 encoding), which is seamlessly compatible with speech analysis toolchains widely used in linguistic research, supporting subsequent acoustic feature extraction, statistical analysis, and cross-platform sharing. Simultaneously, based on the Nginx+Angular+Flask+Redis+MySQL technology stack, it achieves front-end and back-end separation and microservice deployment, supporting high-concurrency access. A single node can support 500+ students submitting audio simultaneously online, and the Redis caching mechanism can control the ASR inference response time for similar audio samples to within 500ms, significantly outperforming the deployment flexibility and response speed of traditional C / S architecture evaluation software.
[0075] This invention reduces the labor costs of large-scale oral assessments while ensuring the consistency of scoring standards. AI handles over 80% of the routine annotation work (automatic segmentation and alignment of standard pronunciation), allowing teachers to focus only on reviewing potentially erroneous areas marked by the AI, which account for approximately 20% of all audio segments. This reduces the workload of reviewing each audio unit by about 70%. Furthermore, by combining the objective consistency of AI pre-annotation with the subjective professionalism of teachers' final review, a balance is achieved between meeting objective phonetic standards and taking into account the subjective needs of the teaching scenario. This solves the dilemma in existing technologies where purely automatic scoring is too mechanical and purely manual scoring is too costly. Attached Figure Description
[0076] Figure 1 This is a diagram of the overall platform architecture of the present invention;
[0077] Figure 2 Flowchart for generating Praat annotation files for AI;
[0078] Figure 3 This is a diagram illustrating the structure and annotation hierarchy of a TextGrid file.
[0079] Figure 4 A flowchart for teacher manual review and closed-loop optimization;
[0080] Figure 5 This is a schematic diagram of the platform's interactive interface. Detailed Implementation
[0081] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, through the technology stack and data flow.
[0082] A Chinese reading aloud and spoken language annotation and evaluation platform, the overall architecture of the platform and its hardware connection relationship are as follows:
[0083] The platform consists of a five-layer architecture: client layer, access layer, service layer, data processing layer, and storage layer.
[0084] Client layer: This includes teacher terminals, which refer to personal computer terminals running the Praat annotation software, and student terminals, which refer to browsers or mobile devices running the Angular front-end application; teacher terminals communicate with the access layer via HTTP / HTTPS protocols and support uploading and downloading Praat TextGrid files;
[0085] The TextGrid file has three levels of annotations, specifically including:
[0086] The phoneme layer records the actual phoneme sequence of pronunciation, including phoneme labels, start time, end time, and log likelihood score. Phonemes with a likelihood score below a preset threshold are marked as suspected pronunciation errors.
[0087] The tone layer records the results of the fundamental frequency F0 curve extraction, including tone type determination and tone deviation marking. Syllables whose predicted tone does not match the standard tone or whose F0 curve variation coefficient exceeds the threshold are marked with tone deviation indicators.
[0088] The error layer combines the confidence level of the phoneme layer and the deviation level of the tone layer. Intervals that simultaneously meet the conditions of a phoneme likelihood below a first threshold and a dynamic time warping (DTW) distance greater than a second threshold are classified as severely erroneous. Intervals that meet only one condition are classified as moderately erroneous. The first threshold ranges from -5.0 to -2.0; the second threshold ranges from 0.15 to 0.35.
[0089] Access layer: Deploy Nginx reverse proxy server, configure load balancing module and SSL terminal to route front-end requests to back-end Flask service cluster, and provide static resource caching and rate limiting control;
[0090] The service cluster includes:
[0091] The voice receiving module is used to receive audio streams uploaded by student terminals and temporarily store them in a cache;
[0092] The ASR inference module is used for speech recognition and forced alignment of audio, generating phoneme-level time boundary information;
[0093] The annotation generation module is used to convert the time boundary information into a Praat standard TextGrid format file;
[0094] The difference comparison module is used to compare the AI-generated TextGrid file with the teacher-corrected TextGrid file, and calculate the time offset error and label error rate.
[0095] The specific steps for the annotation generation module to generate a TextGrid file include:
[0096] Extract the Mel frequency cepstral coefficients (MFCC) feature vector of the audio, and configure the frame length, frame shift, and number of Mel filter banks according to the audio frequency and length;
[0097] The audio is decoded using an ASR engine trained on WeNet or Kaldi to obtain the text recognition results;
[0098] A forced alignment algorithm based on Hidden Markov Model (HMM) or Deep Neural Network (DNN) is used to align the recognized text with MFCC features, generating phoneme-level, syllable-level, and word-level temporal boundary information.
[0099] The time boundary information is serialized according to the Praat TextGrid file format specification. Three IntervalTiers are set to correspond to the phoneme layer, tone layer and error layer respectively. Each IntervalTier contains a list of time intervals, and each time interval contains the start time, end time and marker text.
[0100] Output a UTF-8 encoded TextGrid file. The filename should include the student identifier, text identifier, version type, and timestamp.
[0101] The workflow of the difference comparison module includes:
[0102] Receive the TextGrid file returned by the teacher after correction using Praat software, and parse it into structured data;
[0103] Align and match AI-generated annotations and manually corrected annotations according to time intervals, and calculate the time offset error for each corresponding interval. The time offset error is the absolute value of the difference between the start time of AI annotation and the start time of manual annotation plus the absolute value of the difference between the end time.
[0104] The number of intervals with inconsistent phoneme labels is counted, and the label error rate is calculated, where the label error rate is the proportion of inconsistent intervals to the total number of intervals.
[0105] Count the number of syllables with consistent tone labels and calculate the tone consistency rate.
[0106] Generate a discrepancy report, including average time error, label error rate, tone consistency rate, and a list of specific error intervals.
[0107] Service Layer: A RESTful API service cluster built on the Flask framework, including a voice reception module, an ASR inference module, an annotation generation module, a difference comparison module, and a model management module; each module uses a Redis message queue for asynchronous task scheduling.
[0108] Data processing layer: Integrates a deep learning-based ASR engine and a forced alignment algorithm to convert audio streams into phoneme-level timestamp data;
[0109] Storage layer: MySQL relational database is used to store user identity information, audio file metadata, annotation version records and rating history; Redis non-relational database is used to cache hotspot annotation data and intermediate results of ASR model inference.
[0110] It also includes a model management module and a system control module:
[0111] The model management module is used to trigger the incremental learning process of the ASR model when the error rate of the labels of international students of different nationalities and native language backgrounds exceeds a preset threshold or the tone consistency rate is lower than a preset threshold. The incremental learning process includes: querying the historical correction data of the phoneme or group from the MySQL database to build a fine-tuned training set, performing 10-20 rounds of incremental training based on the original ASR model parameters, setting the learning rate to 0.05-0.2 times the original learning rate, and updating the online model when the test word error rate on the validation set decreases by more than 2%.
[0112] The system control module includes monitoring and managing system resources, as well as comprehensively managing entities stored in the system, including teacher accounts, student accounts, audio files, and annotation files.
[0113] The present invention provides a method for evaluating Chinese reading aloud and spoken language annotation, comprising the following steps:
[0114] Steps A through D are the technical process for AI to generate Praat annotation files, and steps E through H are the teacher manual review and closed-loop optimization methods.
[0115] Step A: Audio Preprocessing and Feature Extraction
[0116] The student terminal records audio readings via the Angular frontend and uploads them to the Flask service via Nginx; the Flask service stores the audio stream in a temporary Redis cache and extracts the Mel-frequency cepstral coefficients (MFCC) feature vector.
[0117] Step B: Speech Recognition and Forced Alignment
[0118] The ASR engine decodes the audio to obtain the text recognition result; it uses a forced alignment algorithm based on a deep neural network to force the recognition text and audio features to align on the time axis, generating time boundary information at the phoneme, syllable, and word levels, indicating the start time, end time, and duration.
[0119] Step C: Generation of multi-level annotation content
[0120] The system automatically generates three layers of annotation content based on the alignment results:
[0121] The first layer, the phoneme layer, records the actual sequence of phonemes produced and marks errors such as replacements, deletions, and insertions.
[0122] The second layer, the tone layer, extracts the fundamental frequency (F0) curve, calculates the similarity between the tone profile and the standard model, and marks tone errors.
[0123] The third layer, the error layer: combines phoneme confidence scores and tone deviation to automatically mark suspected pronunciation error areas;
[0124] Step D: TextGrid Format Conversion and Output
[0125] The three layers of labeled data are serialized according to the Praat TextGrid file format specification to generate a labeled file containing the time domain (IntervalTier); the TextGrid file is pushed to the teacher's terminal through the Nginx download interface so that the teacher can open it in the local Praat software.
[0126] Step E: Manual Correction Process
[0127] Teachers open the AI-generated TextGrid file in Praat software and manually check and correct phoneme boundaries, tone markings, and error markers. The specific manual correction operations for teachers include:
[0128] Drag the Interval boundaries on the timeline to adjust the phoneme segmentation positions, correcting boundaries where the time deviation exceeds 30ms, and aligning the actual pronunciation time of the phoneme with the Interval boundaries for subsequent model fine-tuning. After aligning the three layers of boundaries, the following adjustments need to be made to each layer:
[0129] Modify the mark text of the phoneme layer Interval to correct phoneme recognition errors;
[0130] Modify the tone type marker in the tone layer to correct tone determination errors;
[0131] Adjust the error level label at the error level, and upgrade, downgrade or cancel the error label based on professional judgment.
[0132] After making the modifications and annotations, save it as a new TextGrid file; then use the Angular frontend to upload the corrected TextGrid file back to the Flask service.
[0133] Step F: Difference Comparison and Model Fine-tuning
[0134] Flask's discrepancy comparison module parses the AI-annotated and manually annotated TextGrid files, calculates the time offset error and label error rate, and generates a discrepancy matrix.
[0135] Step G: When the error rate of a specific phoneme exceeds a preset threshold, the model fine-tuning process is triggered, the corrected alignment data is added to the training set, and the ASR model parameters are updated through incremental learning.
[0136] Step H: Data Persistence
[0137] The MySQL database records the version number of each annotation, the annotator's identity (including AI and teacher), timestamp, and confidence score, supporting annotation history backtracking and multi-version comparison.
[0138] Platform signal transmission relationships and interaction processes
[0139] Audio stream transmission path: microphone acquisition on student terminal → WebRTC encoding of audio data in Angular frontend → Nginx SSL encryption → Flask service → Redis → ASR engine for feature extraction;
[0140] The data transmission path for annotation is as follows: ASR engine exports timestamp data → annotation generation module generates TextGrid standard file → MySQL stores metadata + Redis caches hot data → Nginx provides download service → teacher terminal runs Praat software;
[0141] Control signal transmission path: Teacher terminal sends correction confirmation signal → Flask service triggers difference comparison → Redis task queue → Model management module triggers fine-tuning → ASR engine updates parameters.
[0142] Example 1: System Hardware Deployment and Network Architecture Implementation
[0143] like Figure 1 As shown, the hardware deployment in this embodiment includes the following components:
[0144] Client hardware configuration
[0145] Teacher terminal: A PC workstation with an Intel i7 processor and 16GB of memory, running Ubuntu 20.04 operating system, and pre-installed with Praat 6.3.0 or later version speech analysis software; connected to the campus network via gigabit Ethernet.
[0146] Student terminal: PC or Android / iOS mobile device that supports Chrome / Edge / Firefox browsers, equipped with microphone array (signal-to-noise ratio ≥60dB), and network bandwidth requirement ≥2Mbps.
[0147] Server cluster configuration
[0148] Nginx Servers: Deploy two Nginx 1.20.0 servers in a master-slave configuration, with SSL certificates (HTTPS protocol) and load balancing strategies (weighted round-robin algorithm) configured; each server is configured with a 4-core CPU, 8GB of memory, and 100GB of SSD storage.
[0149] Flask application server: Deploy 4 Flask 2.0 servers, each configured with an 8-core CPU, 32GB of memory, and a Tesla T4 GPU accelerator card (for ASR inference); start via Gunicorn WSGI server, with the number of worker processes set to 4 × the number of CPU cores.
[0150] Redis caching server: Deploy a Redis 6.2 cluster (1 master 2 slave mode), configure the persistence strategy (RDB+AOF), and use 64GB of memory to store audio stream cache and ASR intermediate results.
[0151] MySQL database server: Deploy MySQL 8.0 master-slave architecture, with the master database responsible for write operations and the slave database responsible for read operations; Configure InnoDB storage engine, and partition the table when the data volume of a single table exceeds 10 million rows (sharding by student ID hash).
[0152] Network connectivity
[0153] All servers are deployed in the campus network computer room and interconnected through gigabit switches; the Nginx server exposes port 443 to the public network through firewall NAT mapping; the Flask server communicates with Redis and MySQL through internal network IPs and does not directly expose to the public network.
[0154] Example 2: Specific Implementation of AI-Generated Praat Annotation Files
[0155] like Figure 2 and Figure 3 As shown in the figure, this embodiment details the specific steps for generating a TextGrid annotation file using AI:
[0156] Step S1: Audio Acquisition and Upload
[0157] The student clicks the "Record" button on the Angular frontend. The frontend uses the WebRTC API to access the microphone and capture audio data with a sampling rate of 16kHz, a bit depth of 16bit, and mono. After recording, the frontend converts the WAV audio to MP3 with a bit rate of 128kbps to reduce transmission bandwidth and uploads it to the Nginx server via an HTTPS POST request. Nginx load balances the request and forwards it to the Flask audio receiving module. The module temporarily stores the audio stream in a Redis temporary cache with an expiration time of 30 minutes and a key name format of audio:{student_id}:{timestamp}.
[0158] Step S2: MFCC Feature Extraction
[0159] The ASR inference module reads the audio stream from Redis and uses the librosa library to extract MFCC features: the frame length is set to 25ms, the frame shift to 10ms, the number of Mel filter banks is 40, and log energy is preprocessed; the extracted 80-dimensional feature vector (40 MFCC + 40 ΔMFCC) is cached in Redis with the key name format mfcc:{audio_id}.
[0160] Step S3: Speech Recognition and Forced Alignment
[0161] An end-to-end ASR model trained using the WeNet toolkit is employed. The model structure is a hybrid structure of Conformer + CTC / Attention. The training data includes 10,000 hours of Mandarin native speaker speech and 2,000 hours of foreign student accent speech. During decoding, a CTC greedy search is used to output the text recognition results.
[0162] Forced alignment is achieved using the Montreal Forced Aligner (MFA) tool, which aligns the recognized text with the audio MFCC features to generate phoneme-level time boundaries. The alignment dictionary uses a general Mandarin pronunciation dictionary, and the alignment model is a triphone HMM model based on Kaldi. The output format is CSV, which includes the following fields: phoneme, start time (s), end time (s), duration (ms), and log-likelihood.
[0163] Step S4: Generation of three-layer annotations
[0164] Phoneme layer generation: The phoneme sequence in the alignment result is directly mapped to the first level of TextGrid. The Interval boundary strictly adopts the start and end time of the forced alignment output. If the log likelihood of a phoneme is lower than the -3.0 threshold, a mark [*] is added after the phoneme text to indicate a suspected pronunciation error.
[0165] Tone layer generation: The audio fundamental frequency F0 curve is extracted using a WORLD vocoder. The duration of each syllable is divided into 5 equal parts, and the F0 value at each point is calculated and normalized. The DTW distance is calculated with the standard four-tone template, namely, first tone, second tone, third tone, and fourth tone. The one with the smallest distance is determined as the predicted tone. If the predicted tone does not match the standard tone, or the coefficient of variation of the F0 curve is >0.15, the actual tone type is marked in the tone layer and a [bias] mark is added.
[0166] Error layer generation: Combining the confidence of the phoneme layer and the deviation of the tone layer, the intervals that simultaneously satisfy "phoneme likelihood is lower than the first threshold, <-3.0" and "tone DTW distance is greater than the second threshold, >0.25" are marked as [serious error]; those that satisfy only one of them are marked as [general error]; the rest are marked as [correct].
[0167] The first threshold value can also be between -5.0 and -2.0, with -3.0 yielding the best results; the second threshold value can be between 0.15 and 0.35, with 0.25 yielding the best results.
[0168] Step S5: TextGrid Format Serialization and Output
[0169] The TextGrid library (Python) is used to create an object, setting three IntervalTiers (corresponding to three levels of annotation). Each IntervalTier contains name, xmin (global minimum time), xmax (global maximum time), and a list of intervals (each interval contains xmin, xmax, and mark fields). The text is serialized into a UTF-8 encoded TextGrid text file with the filename format {student_id}_{text_id}_AI_{timestamp}.TextGrid. The file is provided via an Nginx download interface with the URL format https: / / platform.edu / download / {file_id}, and the validity period is set to 24 hours.
[0170] Example 3: Teacher Manual Review and Closed-Loop Optimization Implementation
[0171] like Figure 4 As shown in this embodiment, the process of teachers using Praat for manual review and system closed-loop optimization is described in detail:
[0172] Teachers download and open AI-annotated files
[0173] After logging in on the Angular frontend, the teacher enters the "Pending Review" list and selects the audio recordings submitted by students. The frontend calls the Nginx download interface to obtain the TextGrid file and automatically triggers the Praat software to open. Praat loads the student's audio WAV file and the AI-generated TextGrid file. The interface displays three layers of labeled tracks: phoneme layer, tone layer, and error layer. The teacher can zoom in and out to view the waveform, spectrogram, and labeled content for a specific time period.
[0174] Manual correction operation
[0175] Teachers should correct the following in the Praat interface:
[0176] Time boundary correction: If the phoneme boundary segmented by AI deviates from the actual start and end points of pronunciation by more than 30ms, the teacher should manually drag the Interval boundary to adjust it;
[0177] Phoneme label correction: If the phoneme recognized by AI does not match the actual pronunciation, for example, misidentifying / f / as / h / , the teacher double-clicks the Interval to modify the mark text;
[0178] Tone labeling correction: If the AI determines the tone incorrectly, the teacher can modify the tone type label in the tone layer;
[0179] Error level review: Based on professional judgment, teachers may downgrade an AI-marked [Critical Error] to [General Error] or upgrade it to [Critical Error], or remove the error mark.
[0180] After making the corrections, the teacher selects "File→Save TextGrid as...", with the file name formatted as {student_id}_{text_id}_Manual_{timestamp}.TextGrid, and uploads it back to the Flask service via the "Upload Correction" button in the Angular frontend.
[0181] Comparison and Quantitative Analysis
[0182] The difference comparison module parses the AI version and the human version of the TextGrid file and performs the following calculations:
[0183] Time offset error (Δt): For each corresponding interval, calculate |t_AI_start - t_Manual_start| + |t_AI_end - t_Manual_end|, in ms;
[0184] Error Rate (LER): The percentage of intervals with inconsistent phoneme labels out of the total number of intervals;
[0185] Tone consistency rate: The percentage of syllables with consistent tone labels out of the total number of syllables.
[0186] Generate a difference report in JSON format: {"audio_id": "xxx", "avg_time_error": 12.5, "ler": 0.08, "tone_accuracy": 0.92, "error_segments": [{"tier": "phones", "start": 1.25, "error_type": "boundary_shift"}]}.
[0187] Model fine-tuning triggering and execution
[0188] The threshold judgment unit sets trigger conditions: when the LER of a certain type of phoneme, such as / ʂ / vs / s / , is greater than 0.15, or when the tone consistency rate of a certain native language student group is less than 0.80, the incremental learning process is triggered.
[0189] Data preparation: Query all historical correction data for this phoneme / group from MySQL to build a fine-tuning training set;
[0190] Model fine-tuning: Based on the original ASR model parameters, perform 10 rounds of incremental training using the Kaldi chain model or WeNet model, with the learning rate set to 0.1 times the original learning rate and the batch size 32;
[0191] Model evaluation: Test on the validation set. If the WER decreases by more than 2%, update the online model; otherwise, roll back to the original model.
[0192] Example 4: Multi-Scenario Application Example
[0193] Scenario A: Assessment of Reading Aloud
[0194] Students read aloud the text from the HSK Standard Course, with the text fixed; the system uses a forced alignment mode, and the dictionary only contains vocabulary from the text; AI annotation focuses on detecting phoneme substitution errors, such as native English speakers pronouncing / ʂ / as / s / and tone errors, such as the first tone not being high enough; when reviewing, teachers focus on whether the AI correctly marks the "foreign accent" features.
[0195] Scene B: Describe the picture freely
[0196] Students give a 2-minute free description based on the picture, with the text open; the system uses large-vocabulary continuous speech recognition (LVCSR) mode, generates text and scores it with language model; in addition to pronunciation features, AI annotation adds a "fluency layer" (detecting pauses >2 seconds); teachers supplement the content with human scoring for logical and grammatical correctness during review.
[0197] Scenario C: Computer-aided scoring for oral exams
[0198] In the HSK speaking test scenario, the system serves as the "computer-assisted initial assessment" stage. After the AI generates annotations, they are independently reviewed by two teachers. If the difference between the teachers' scores is greater than 1 point (out of a possible 5 points), a third teacher is introduced for arbitration. All corrected data is used for model optimization after the test to improve the accuracy of automatic scoring in the next round of tests.
[0199] Example 5: Collecting and Annotating Student Audio Data
[0200] Audio Acquisition and Metadata Input
[0201] Students record audio readings using the Angular frontend. After recording, a pop-up window prompts students to fill in the following metadata: native language background (e.g., Japanese, English, Korean), Chinese proficiency level (e.g., HSK 1-6), the text number being read, and the practice mode. This metadata and the audio file are packaged together in JSON format and uploaded to the Flask speech receiving module via Nginx.
[0202] Audio and metadata structured storage
[0203] The Flask service parses the received data and stores it in the storage layer:
[0204] MySQL database: Insert records into the audio_metadata table, with fields including audio_id (primary key, generated by UUID), student_id, native_language, hsk_level, text_id, practice_mode, upload_time, file_path (storage server path), duration (audio duration, seconds), and status (status: pending / AI annotation / pending review / reviewed).
[0205] File storage: Audio files are named {audio_id}.mp3 and stored in the / data / audio / {year} / {month} / directory of NFS shared storage;
[0206] Redis caching: The audio_id and metadata are mapped and stored in a hash structure with the key audio:meta:{audio_id} and an expiration time of 2 hours, so that the ASR module can quickly read it later.
[0207] Structured storage of labeled data
[0208] Both AI-generated annotations and teacher-corrected annotations are persisted in a structured format:
[0209] AI annotation storage: The annotation_ai table stores the TextGrid file path (textgrid_ai_path), annotation generation time (ai_time), and ASR model version (model_version); the detailed data of the three-layer annotations is serialized into JSON and stored in the annotation_detail field, which includes the start and end time of each layer's interval, the labeled text, and the confidence score.
[0210] Manual annotation storage: After the TextGrid file is parsed after being corrected by the teacher, the textgrid_manual_path, correction time manual_time, and reviewing teacher ID teacher_id are stored in the annotation_manual table; the difference comparison results are stored in the annotation_diff field, which includes time offset error, label error rate, and list of modification intervals.
[0211] Version association: The audio metadata table is associated through the audio_id foreign key, and the version field is used to distinguish multiple annotation iterations of the same audio. For example, v1.0 is AI annotation, v2.0 is the first manual correction, and v2.1 is the second review.
[0212] Corpus Construction and Retrieval
[0213] The platform supports searching annotated corpora based on multiple criteria:
[0214] The Angular front-end provides an advanced search interface that supports filtering by native language background, HSK level, text type, error type, and review status.
[0215] The Flask service converts the search criteria into an SQL query, returning a list of matching audio files and a preview of the tags.
[0216] Search results support batch export: After selecting multiple records, you can download the audio file, the corresponding TextGrid annotation file, and the metadata CSV table for offline research or analysis by third-party tools.
[0217] Data security and privacy protection
[0218] Audio files and annotation data are stored on servers within the education network and are not directly exposed for public network access.
[0219] The student_id field in the MySQL database is pseudonymous and stored separately from the real identity information;
[0220] Teachers' access permissions are managed through the RBAC (Role-Based Access Control) model, and only the data owner and authorized teachers can view the original audio.
Claims
1. A Chinese reading aloud and spoken language annotation and evaluation platform, characterized in that, include: The client layer includes teacher terminals and student terminals. The teacher terminals are equipped with Praat speech analysis software, which is used to open, edit, and save TextGrid format annotation files. The student terminals are used to collect and upload audio recordings. The access layer includes an Nginx reverse proxy server, configured with a load balancing module and an SSL encryption module, used to receive client requests and route them to backend services; The service layer includes a cluster of RESTful API services built on the Flask framework; The data processing layer includes an ASR engine and a forced alignment module. The ASR engine is based on an end-to-end deep learning model, and the forced alignment module is used to align the recognized text with audio features on the time axis. The storage layer includes a Redis cache database and a MySQL relational database. The Redis database is used to cache audio streams, MFCC feature vectors, and ASR intermediate results, while the MySQL database is used to store user identity information, audio file metadata, annotation version records, and rating history.
2. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 1, characterized in that, The TextGrid file has three levels of annotations, specifically including: The phoneme layer records the actual phoneme sequence of pronunciation, including phoneme labels, start time, end time, and log likelihood score. Phonemes with a likelihood score below a preset threshold are marked as suspected pronunciation errors. The tone layer records the results of the fundamental frequency F0 curve extraction, including tone type determination and tone deviation marking. Syllables whose predicted tone does not match the standard tone or whose F0 curve variation coefficient exceeds the threshold are marked with tone deviation indicators. The error layer combines the confidence level of the phoneme layer and the deviation level of the tone layer. Intervals that simultaneously satisfy the condition of having a phoneme likelihood lower than the first threshold and a tone dynamic time warping (DTW) distance greater than the second threshold are classified as severely erroneous. Intervals that satisfy only one condition are classified as generally erroneous. The first threshold ranges from -5.0 to -2.0, and the second threshold ranges from 0.15 to 0.
35.
3. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 1, characterized in that, The service cluster includes: The voice receiving module is used to receive audio streams uploaded by student terminals and temporarily store them in a cache; The ASR inference module is used for speech recognition and forced alignment of audio, generating phoneme-level time boundary information; The annotation generation module is used to convert the time boundary information into a Praat standard TextGrid format file; The difference comparison module is used to compare the AI-generated TextGrid file with the teacher-corrected TextGrid file, and calculate the time offset error and label error rate.
4. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 3, characterized in that, The specific steps for the annotation generation module to generate a TextGrid file include: Extract the Mel frequency cepstral coefficients (MFCC) feature vector of the audio, and configure the frame length, frame shift, and number of Mel filter banks according to the audio frequency and length; The audio is decoded using an ASR engine trained on WeNet or Kaldi to obtain the text recognition results; A forced alignment algorithm based on Hidden Markov Model (HMM) or Deep Neural Network (DNN) is used to align the recognized text with MFCC features, generating phoneme-level, syllable-level, and word-level temporal boundary information. The time boundary information is serialized according to the Praat TextGrid file format specification. Three IntervalTiers are set to correspond to the phoneme layer, tone layer and error layer respectively. Each IntervalTier contains a list of time intervals, and each time interval contains the start time, end time and marker text. Output a UTF-8 encoded TextGrid file. The filename should include the student identifier, text identifier, version type, and timestamp.
5. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 3, characterized in that, The workflow of the difference comparison module includes: Receive the TextGrid file returned by the teacher after correction using Praat software, and parse it into structured data; Align and match AI-generated annotations and manually corrected annotations according to time intervals, and calculate the time offset error for each corresponding interval. The time offset error is the absolute value of the difference between the start time of AI annotation and the start time of manual annotation plus the absolute value of the difference between the end time. The number of intervals with inconsistent phoneme labels is counted, and the label error rate is calculated, where the label error rate is the proportion of inconsistent intervals to the total number of intervals. Count the number of syllables with consistent tone labels and calculate the tone consistency rate. Generate a discrepancy report, including average time error, label error rate, tone consistency rate, and a list of specific error intervals.
6. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 1, characterized in that, It also includes a model management module and a system control module: The model management module is used to trigger the incremental learning process of the ASR model when the error rate of the labels of international students of different nationalities and native language backgrounds exceeds a preset threshold or the tone consistency rate is lower than a preset threshold. The incremental learning process includes: querying historical correction data of the phoneme or group from the MySQL database to construct a fine-tuned training set; performing 10-20 rounds of incremental training based on the original ASR model parameters; setting the learning rate to 0.05-0.2 times the original learning rate; and updating the online model when the test word error rate on the validation set decreases by more than 2%. The system control module includes monitoring and managing system resources, as well as comprehensively managing entities stored in the system, including teacher accounts, student accounts, audio files, and annotation files.
7. The Chinese reading aloud and spoken language annotation and evaluation platform according to claim 1, characterized in that, The Nginx reverse proxy server is configured with: The static resource caching module is used to cache the JS / CSS files built by the Angular frontend; The rate limiting module is used to limit the frequency of audio upload requests from student terminals to prevent malicious attacks. The WebSocket support module is used to push AI annotation generation progress and teacher review status in real time.
8. A method for evaluating Chinese reading aloud and spoken language using the platform described in any one of claims 1-6, characterized in that, Includes the following steps: Step A: The student terminal records audio of reading aloud through the Angular frontend, uploads it to the Flask service via Nginx, extracts the MFCC feature vector, and temporarily stores it in the Redis cache; Step B: The ASR engine performs speech recognition decoding on the audio to obtain the text recognition result; A forced alignment algorithm is used to align the recognized text with audio features, generating phoneme-level temporal boundary information; Step C: Generate three layers of annotation content based on the alignment results: the first layer is the phoneme layer, which records the actual pronunciation phoneme sequence and suspected error marks; The second layer is the tone layer, which extracts the fundamental frequency F0 curve and marks tone deviations; the third layer is the error layer, which automatically marks the error level by combining phoneme confidence and tone deviation. Step D: Serialize the three-layer annotation data into the Praat standard TextGrid file format and push it to the teacher's terminal via the Nginx download interface; Step E: The teacher opens the AI-generated TextGrid file in Praat software, manually verifies and corrects the phoneme boundaries, tone markings, and error markers, saves it as a new version of the TextGrid file, and uploads it back to the Flask service through the Angular frontend; Step F: The difference comparison module parses the AI-annotated and manually annotated TextGrid files, calculates the time offset error and label error rate, and generates a difference matrix; Step G: When the error rate of a specific phoneme or a student group with a specific native language background exceeds a preset threshold, the model fine-tuning process is triggered, the corrected alignment data is added to the training set, and the ASR model parameters are updated through incremental learning. Step H: The MySQL database records the version number, annotator identity, timestamp, and confidence score for each annotation, supporting annotation history backtracking and multi-version comparison.
9. The method according to claim 8, characterized in that, The forced alignment algorithm in step B includes: Based on the Montreal Forced Aligner tool, using a standard Mandarin pronunciation dictionary; The alignment results will be output in CSV format, including the fields: phoneme symbol, start time, end time, duration, and log-likelihood. Phonemes with a log-likelihood below -3.0 are marked as low confidence for subsequent error layer labeling.
10. The method according to claim 8, characterized in that, The specific steps for teacher manual correction in step E include: Drag the Interval boundary on the timeline to adjust the phoneme segmentation position and correct the boundary where the time deviation exceeds 30ms; Modify the mark text of the phoneme layer Interval to correct phoneme recognition errors; Modify the tone type marker in the tone layer to correct tone determination errors; Adjust the error level label at the error level, and upgrade, downgrade or cancel the error label based on professional judgment.