A language corpus acquisition, quality inspection, labeling and verification method and device based on a cloud platform

CN122819784APending Publication Date: 2026-09-25GUANGXI UNIV FOR NATITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610990429.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于云平台的语言语料采集质检标注校验方法及装置,以解决现有语言语料处理方式中采集不便、语料流转割裂、质检补录重录协同效率低、线上标注和校验规范性不足以及云端数据池管理能力弱的问题

Benefits of technology

[0016]与现有技术相比,本发明至少具有如下有益效果:本发明通过云平台统一组织语料采集任务、质检任务、标注任务和校验任务,能够将采集终端、质检端、标注端和校验端形成协同处理流程,减少本地化软件之间的数据导入导出和人工汇总环节。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819784A_ABST
    Figure CN122819784A_ABST
Patent Text Reader

Abstract

The application relates to the field of cloud computing and natural language processing, and discloses a language corpus collection quality inspection, marking and checking method and device based on a cloud platform. The method generates a corpus collection task through the cloud platform and distributes the corpus collection task to a collection terminal, acquires collection object information and target corpus data, associates the collection object information and the target corpus data with corpus attribute information to form to-be-inspected corpus data, creates an inspection task and generates an inspection state through an inspection terminal, shunts the data to a marking process or a collection terminal according to the inspection state, supplements or re-collects the data, creates a marking task and a checking task, acquires marking data through a marking terminal, completes checking through a checking terminal, and stores the checked corpus data into a checked data pool.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cloud computing and natural language processing, and in particular to a method and apparatus for quality inspection, annotation and verification of language corpus based on a cloud platform. Background Technology

[0002] The collection, organization, and annotation of language resources are crucial foundations for language preservation, speech research, language and cultural heritage transmission, and the training of natural language processing models. For low-resource languages ​​such as Zhuang and Yao, corpora are typically scattered across different regions, speakers, and collection scenarios. The corpus collection process also requires recording multi-dimensional information such as the speaker's accent region, language, phonology, vocabulary, and entries to facilitate subsequent quality inspection, annotation, verification, and reuse.

[0003] Existing language processing solutions mostly employ localized software or local server deployment, using localized speech recognition, manual annotation tools, or offline tables to complete language acquisition, transcription, and annotation. These solutions typically deploy computing resources and language processing models on local devices or servers, and then export the results into standard format files for storage after processing.

[0004] However, localized processing methods have significant shortcomings when dealing with large-scale language corpora: First, local computing resources are limited by single-machine performance, resulting in low processing efficiency when handling massive amounts of ethnic language corpora; second, language models or business rules updates require local redeployment, making it difficult to utilize the elastic computing power and unified update capabilities of the cloud; third, research team members cannot share processing progress and intermediate results in real time, and data conflicts are prone to occur when multiple people collaborate on collection, quality inspection, and annotation; fourth, the maintenance costs of local servers and software environments are high, which is not conducive to the long-term operation of language protection projects with limited funding; fifth, existing solutions lack dedicated collection and annotation processes for low-resource languages, making it difficult to effectively handle the needs of special characters, dialect variations, regional accent differences, and International Phonetic Alphabet annotation; sixth, the granularity of corpus resource management is too coarse, failing to organize corpora in multiple dimensions around speakers, phonological systems, word lists, entries, collection status, quality inspection status, annotation status, and verification status, leading to difficulties in corpus retrieval, reuse, and subsequent applications. Summary of the Invention

[0005] The purpose of this invention is to provide a method and apparatus for quality inspection, annotation and verification of language corpus based on a cloud platform, so as to solve the problems of inconvenient collection, fragmented corpus flow, low efficiency of collaborative quality inspection, supplementation and re-recording, insufficient standardization of online annotation and verification, and weak cloud data pool management capabilities in existing language corpus processing methods.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for quality inspection, annotation, and verification of language corpus collection based on a cloud platform, comprising the following steps: acquiring corpus collection information through a cloud platform and generating a corpus collection task based on the corpus collection information; sending the corpus collection task to a collection terminal and acquiring collection object information and target corpus data corresponding to the corpus collection task through the collection terminal; associating the target corpus data with the collection object information and corpus attribute information to obtain corpus data to be inspected, and uploading the corpus data to be inspected to the cloud platform; creating a quality inspection task through the cloud platform and assigning the quality inspection task to a quality inspection terminal; performing quality inspection on the corpus data to be inspected through the quality inspection terminal according to the quality inspection task to obtain the quality inspection status corresponding to the corpus data to be inspected; and performing quality inspection on the corpus data to be inspected according to the quality inspection status. Data is processed by stream splitting. Data in the quality inspection state that is "passed" is identified as data to be labeled. Collection entry information corresponding to data in the quality inspection state that is "to be supplemented" or "to be re-collected" is sent to the collection terminal, enabling the collection terminal to supplement or re-collect the target data corresponding to the collection entry information. Labeling tasks are created through the cloud platform and assigned to labeling terminals. Labeling terminals receive labeled data for the data to be labeled according to the labeling tasks. Verification tasks are created through the cloud platform and assigned to verification terminals. Verification terminals verify the labeled data according to the verification tasks to obtain verified data. The verified data is stored in a verified data pool. The cloud platform also includes a quality-inspected data pool and a labeled data pool.

[0007] Furthermore, the corpus collection information includes at least one of the following: collection vocabulary information, collection language category information, collection phonological information, and collection item information; the collection object information includes at least one of the following: collection object name, accent type, language mastery information, and accent region information; the corpus attribute information includes at least one of the following: corpus vocabulary, corpus phonological system, corpus item, corpus collection time, corpus collection status, and corpus identifier; the target corpus data includes at least speech corpus data, and the speech corpus data includes recording data corresponding to the collection items.

[0008] Further, the corpus collection task is sent to the collection terminal, and the collection terminal is used to obtain the collection object information and the target corpus data corresponding to the corpus collection task. This includes: receiving the collection object information entered by the collection object through the collection terminal and uploading the collection object information to the cloud platform; displaying the current item to be collected on the collection terminal according to the corpus collection task; receiving the target corpus data uploaded by the collection terminal after completing the collection operation of the current item to be collected, and associating and saving the target corpus data with the item identifier of the current item to be collected; when the target corpus data corresponding to the current item to be collected is saved and there is a next item to be collected, determining the next item to be collected as the new current item to be collected, and displaying the new current item to be collected on the collection terminal; and sequentially receiving the target corpus data corresponding to each item to be collected until the target corpus data collection corresponding to all items to be collected in the corpus collection task is completed.

[0009] Further, the data to be inspected is processed by splitting according to the quality inspection status, including: when the quality inspection status is "passed", the corresponding data to be inspected is written into the already inspected data pool and the corresponding data to be inspected is identified as data to be labeled; when the quality inspection status is "to be supplemented", the collection status of the corresponding data to be inspected is updated to "to be supplemented", and a supplemented collection instruction is sent to the collection terminal; when the quality inspection status is "to be re-collected", the collection status of the corresponding data to be inspected is updated to "to be re-collected", and a re-collection instruction is sent to the collection terminal; when the collection terminal uploads supplemented data according to the supplemented collection instruction, the supplemented data is added to the data record of the corresponding data to be collected; when the collection terminal uploads re-collected data according to the re-collection instruction, the original data record of the corresponding data to be collected is replaced based on the re-collected data.

[0010] Furthermore, creating a quality inspection task through the cloud platform includes: obtaining a vocabulary list to be inspected; receiving at least one search condition for the target audience's accent type, language proficiency, accent region information, target audience name, and current quality inspection status; filtering target audiences from the cloud platform based on the search conditions; determining the range of corpus data to be inspected based on the vocabulary list and the target audience; and generating the quality inspection task based on the range of corpus data to be inspected.

[0011] Furthermore, based on the quality inspection task, the quality inspection data to be inspected is subjected to quality detection to obtain the quality inspection status corresponding to the data, including: retrieving the data to be inspected corresponding to the quality inspection task through the quality inspection terminal; displaying the data to be inspected according to the collection items; when the data to be inspected is speech data, playing the speech data corresponding to the collection item; receiving the quality inspection results input by the quality inspector for the data to be inspected; generating the corresponding quality inspection status based on the quality inspection results; and synchronizing the quality inspection status to the cloud platform after completing the quality inspection of the data to be inspected corresponding to the same collection object.

[0012] Furthermore, creating a labeling task through the cloud platform includes: obtaining the labeling task name, labeler information, accent region information, labeling vocabulary information, and collection object information; filtering the corpus data to be labeled from the quality-checked data pool based on the accent region information, labeling vocabulary information, and collection object information; associating the filtered corpus data to be labeled with the labeler information; generating the labeling task based on the associated corpus data to be labeled and labeler information, and publishing the labeling task to the labeling terminal corresponding to the labeler information.

[0013] Furthermore, receiving annotation data for the corpus data to be annotated based on the annotation task includes: retrieving the corpus data to be annotated corresponding to the annotation task through the annotation terminal; displaying the corpus data to be annotated according to the collection items; when the corpus data to be annotated is speech corpus data, playing the speech corpus data corresponding to the collection item; receiving annotation content input by the annotator for the corpus data to be annotated, wherein when the corpus data to be annotated is speech corpus data, the annotation content includes the International Phonetic Alphabet annotation content corresponding to the speech corpus data; associating the annotation content with the corresponding corpus data to be annotated, collection object information, and corpus attribute information to obtain the annotation data; and uploading the annotation data to the cloud platform and storing it in the annotated data pool.

[0014] Further, the annotation data is verified based on the verification task to obtain verified corpus data, including: retrieving the annotation data corresponding to the verification task from the labeled data pool through the verification terminal; performing at least one of morpheme verification, syllable verification, segment verification, and word list verification on the annotation data; configuring the verification status of the annotation data according to the verification result; after the annotation data corresponding to the same collection object has been verified, determining the annotation data with the verification status as passed as the verified corpus data, and pushing the verified corpus data to the verified data pool.

[0015] This invention also provides a cloud-based language corpus acquisition quality inspection, annotation, and verification device, comprising: The system includes a task generation module for acquiring corpus acquisition information through a cloud platform and generating corpus acquisition tasks based on that information; a corpus acquisition module for sending the corpus acquisition tasks to an acquisition terminal and acquiring acquisition object information and target corpus data corresponding to the acquisition tasks through the acquisition terminal; a corpus association module for associating the target corpus data with the acquisition object information and corpus attribute information to obtain corpus data to be inspected and uploading it to the cloud platform; a quality inspection task processing module for creating quality inspection tasks through the cloud platform, assigning the quality inspection tasks to quality inspection terminals, and performing quality inspection on the corpus data to be inspected based on the quality inspection tasks to obtain the quality inspection status of the corpus data to be inspected; and a quality inspection routing module for processing the corpus data to be inspected based on the quality inspection status. The process involves a split-processing module, which identifies corpus data with a "passed" quality inspection status as corpus data to be labeled, and sends the collection entry information corresponding to corpus data with a "to be supplemented" or "to be re-collected" quality inspection status to the collection terminal, enabling the collection terminal to supplement or re-collect the target corpus data corresponding to the collection entry information. A labeling task processing module creates labeling tasks through the cloud platform, assigns these tasks to labeling terminals, and receives labeled data for the corpus data to be labeled through the labeling tasks. A verification task processing module creates verification tasks through the cloud platform, assigns these tasks to verification terminals, and verifies the labeled data through the verification tasks to obtain verified corpus data. A data storage module stores the verified corpus data into a verified data pool.

[0016] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention organizes the corpus collection task, quality inspection task, annotation task and verification task in a unified manner through the cloud platform, which can form a collaborative processing flow between the collection terminal, quality inspection terminal, annotation terminal and verification terminal, reducing the data import and export between local software and the manual summary steps. Attached Figure Description

[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. Wherein: Figure 1 This is a schematic diagram of the language acquisition process in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the overall process of an embodiment of the present invention. Detailed Implementation

[0019] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Various examples are provided by way of explanation and not by way of limitation. Indeed, those skilled in the art will recognize that modifications and variations can be made to the invention without departing from its scope or spirit. For example, a feature shown or described as part of one embodiment may be used in another embodiment to produce yet another embodiment. Therefore, it is desirable that the invention encompass such modifications and variations falling within the scope of the appended claims and their equivalents.

[0020] The accompanying drawings illustrate one or more examples of the invention. The detailed description uses numerals and letters to refer to features in the drawings. Similar or analogous reference numerals in the drawings and description have been used to refer to similar or analogous parts of the invention. As used herein, the terms “first,” “second,” “third,” and “fourth,” etc., are used interchangeably to distinguish one component from another and are not intended to indicate the location or importance of individual components.

[0021] The present invention will be further described below with reference to specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Where there is no conflict, the technical features in the following embodiments can be combined with each other.

[0022] The cloud platform in this invention refers to a server or server cluster used for unified management of corpus collection tasks, quality inspection tasks, annotation tasks, verification tasks, corpus data, task status, and data pools. The cloud platform may include functional systems such as a management system, a collection system, an annotation system, and a map system, or it may be deployed as logical modules within the same cloud service environment. The management system can be used to publish tasks, design vocabulary lists, upload and manage recordings, and export audio recordings; the collection system can collect corpus data through collection terminals such as WeChat mini-programs; the annotation system can be used for audio recording quality inspection, text annotation, phonetic transcription verification, and corpus output; and the map system can be used to organize and display the geographical information of the collected objects or corpus data.

[0023] In this invention, the data acquisition terminal refers to a terminal device or client program that allows the data acquisition target to perform corpus acquisition operations. The data acquisition terminal can be a mini-program, mobile terminal application, web client, or other terminal capable of communicating with a cloud platform. In a specific scenario, after entering the Zhuang-Yao language database mini-program, the data acquisition target can fill in the speaker information in the language acquisition module at the bottom of the main interface, select the phonological system to be recorded, and record, save, and upload the audio according to the entries.

[0024] In this invention, the quality inspection terminal refers to a terminal or functional interface used by quality inspectors to perform quality checks on the speech data to be inspected. The quality inspection terminal can retrieve collection items from the quality inspection task, play the corresponding speech data, and receive the quality inspection results input by the quality inspectors. The quality inspection results can be converted into a "passed," "to be supplemented," or "to be re-collected" status.

[0025] In this invention, the annotation terminal refers to a terminal or functional interface for annotators to annotate the corpus data to be annotated. The annotation terminal can display the corpus data to be annotated according to the collected items, play the corresponding recording when the corpus data to be annotated is speech data, and receive the International Phonetic Alphabet annotation content or other annotation content input by the annotator.

[0026] In this invention, the verification terminal refers to a terminal or functional interface used by verification personnel or verification function modules to verify the labeled data. The verification terminal can retrieve the labeled data corresponding to the verification task from the labeled data pool and perform at least one of the following on the labeled data: morpheme verification, syllable verification, phonetic segment verification, and word list verification.

[0027] In this invention, the data pool refers to a collection of data organized and stored in a cloud platform according to processing stages. The data pool can include a quality-checked data pool, a labeled data pool, and a validated data pool. Through the staged organization of the data pool, a clear data flow relationship can be established between data collection, quality check, labeling, and validation.

[0028] This embodiment provides a method for quality inspection, annotation, and verification of language corpus collection based on a cloud platform. To fully correspond to the method flow of this invention, the method of this embodiment includes the following steps: S101 obtains corpus collection information through the cloud platform and generates corpus collection tasks based on the corpus collection information.

[0029] S102, the corpus collection task is sent to the collection terminal, and the collection terminal is used to obtain the collection object information and the target corpus data corresponding to the corpus collection task.

[0030] S103. Associate the target corpus data with the information of the collection object and the corpus attribute information to obtain the corpus data to be inspected, and upload the corpus data to be inspected to the cloud platform.

[0031] S104 creates quality inspection tasks through the cloud platform and assigns them to the quality inspection terminal; the quality inspection terminal performs quality inspection on the corpus data to be inspected according to the quality inspection tasks, and obtains the corresponding quality inspection status of the corpus data to be inspected.

[0032] S105. Based on the quality inspection status, the data to be inspected is split into streams. Data to be inspected with a quality inspection status of "pass" is identified as data to be labeled. Data to be inspected with a quality inspection status of "to be supplemented for collection" or "to be re-collected" is sent to the collection terminal so that the collection terminal can supplement or re-collect the target data corresponding to the collection item information.

[0033] S106 creates annotation tasks through the cloud platform and assigns them to the annotation terminal; the annotation terminal receives annotation data for the corpus data to be annotated according to the annotation tasks.

[0034] S107 creates a verification task through the cloud platform and assigns the verification task to the verification terminal; the verification terminal verifies the labeled data according to the verification task to obtain the verified corpus data.

[0035] S108, Store the verified corpus data into the verified data pool; The cloud platform also includes a quality-inspected data pool and a labeled data pool.

[0036] In step S101, the cloud platform acquires the corpus collection information and generates a corpus collection task based on the information. The corpus collection information can be entered or configured by the administrator through the management system, or it can be read by the cloud platform from a pre-established vocabulary resource library, phonological resource library, or collection project configuration. The corpus collection information includes at least one of the following: corpus vocabulary information, corpus language category information, corpus phonological information, and corpus item information.

[0037] The vocabulary list information is used to determine the set of terms included in this collection. For example, it could be a basic vocabulary list, a thematic vocabulary list, or a dialect vocabulary list for languages ​​such as Zhuang and Yao. The language category information is used to determine the language category to which the corpus to be collected belongs. The phonological information is used to determine the phonological system that needs to be entered into the data collection. The item information is used to determine the items to be collected that will be displayed one by one on the data collection terminal.

[0038] The cloud platform can generate corpus collection tasks based on corpus collection information. These tasks can include information such as task identifiers, word list identifiers, language categories, phonological identifiers, collection item identifiers, collection target scope, task release time, and task status. By generating corpus collection tasks, the cloud platform can uniformly distribute the content to be collected from the management side to the collection side, thus avoiding the problems of task dispersion and inconsistent recording that occur when collection personnel use voice recorders, offline spreadsheets, or local files.

[0039] In one implementation, the management system can provide functions such as task publishing, vocabulary design, upload management, and audio recording export. Administrators can configure vocabulary, phonology, and the scope of data collection within the management system. The cloud platform then generates corpus collection tasks that can be read by the collection terminals based on these configurations. For scenarios involving a large number of speakers received daily or with complex geographical distributions, the cloud platform can also classify or organize collection tasks in batches based on accent region, language category, or vocabulary information.

[0040] In step S102, the cloud platform sends the corpus collection task to the collection terminal, and obtains the collection object information and the target corpus data corresponding to the corpus collection task through the collection terminal. The collection terminal can be a mini-program, a mobile application, or a web page. The collection object can be a speaker or other objects that provide corpus data according to the corpus collection task.

[0041] The information collected includes at least one of the following: subject name, accent type, language proficiency, and accent region. In a speech corpus scenario, the subject information can specifically include the speaker's name, accent type, language proficiency, and the region of the recorded accent. This information can serve as filtering criteria in subsequent quality control, annotation, and validation tasks, and can also be used as metadata for corpus resource retrieval and reuse.

[0042] The target corpus data includes at least speech corpus data, which includes recording data corresponding to the collected items. In an scalable implementation, the target corpus data can also be expanded to text corpus data, image corpus data, or video corpus data according to language resource management needs, but this embodiment focuses on the segmented collection of speech corpus data as an example.

[0043] The data acquisition terminal displays the current entry to be acquired based on the corpus acquisition task. After seeing the entry to be entered, the user records an audio message and reads the entry aloud in their local or designated language; after reading, recording stops, and the message is saved and uploaded. After completing the acquisition of the current entry, the terminal uploads the target corpus data to the cloud platform and associates the target corpus data with the entry identifier of the current entry.

[0044] When the target corpus data corresponding to the current item to be collected is saved and a next item to be collected exists, the collection terminal identifies the next item to be collected as the new current item to be collected and displays the new current item to be collected on the collection terminal. The collection object continues to record, stop, save, and upload in the above manner, and the cloud platform sequentially receives the target corpus data corresponding to each item to be collected until the target corpus data corresponding to all items to be collected in the corpus collection task is collected.

[0045] The segmented recording method described above allows the user to input voice data item by item, saving and uploading each recorded item after completion. Compared to continuous audio files, segmented recording data establishes a one-to-one correspondence with entries, phonological systems, word lists, and speaker information, facilitating subsequent playback verification, supplementary recording, re-recording, IPA annotation, and validation. This method also supports fragmented recording, allowing the user to complete the recording task in segments according to their schedule.

[0046] In step S103, the cloud platform associates the target corpus data with the collection object information and corpus attribute information to obtain the corpus data to be inspected, and uploads the corpus data to be inspected to the cloud platform or stores it in the data area to be processed on the cloud platform.

[0047] The corpus attribute information includes at least one of the following: the vocabulary to which the corpus belongs, the phonological system to which the corpus belongs, the entry to which the corpus belongs, the corpus collection time, the corpus collection status, and the corpus identifier. For speech corpus data, the cloud platform can associate the recording data with the entry identifier, vocabulary identifier, phonological system identifier, speaker name, accent region information, and collection time of the current entry to be collected.

[0048] Through the above association, a single corpus no longer exists merely as an independent audio file, but is organized into a structured data record containing the collection object, corpus attributes, and processing status. This structured data record, as the corpus data to be quality inspected, can be used for subsequent quality inspection tasks such as creation, retrieval, allocation, playback verification, supplementary recording, re-recording, and status transition.

[0049] In one implementation, the cloud platform can configure corpus identifiers and corpus acquisition status for the target corpus data. The corpus identifier uniquely identifies the corpus record corresponding to the acquired item, and the corpus acquisition status can indicate whether the corpus is in a state such as acquired, awaiting quality inspection, awaiting supplementary acquisition, or awaiting re-acquisition. Through corpus identifiers and status information, the cloud platform can maintain data record consistency between the acquisition end, quality inspection end, annotation end, and verification end.

[0050] In step S104, the cloud platform creates a quality inspection task and assigns the task to the quality inspection terminal. The quality inspection terminal performs quality inspection on the corpus data to be inspected according to the quality inspection task and obtains the quality inspection status of the corpus data to be inspected.

[0051] When creating a quality inspection task through the cloud platform, you can first obtain the vocabulary list to be inspected, and then receive at least one search condition from the following: accent type of the target audience, language proficiency information, accent region information, target audience name, and current quality inspection status. The cloud platform filters the target audience based on the search conditions, determines the range of the corpus data to be inspected based on the vocabulary list and the target audience, and then generates the quality inspection task based on the range of the corpus data.

[0052] In one specific implementation, after quality inspectors enter the quality inspection task management unit, they can search for and locate the newly added quality inspection task, click to enter the quality inspection data viewing interface, and be redirected to the quality inspection page for the collected items. The quality inspection terminal retrieves the corpus data to be inspected corresponding to the quality inspection task and displays the corpus data to be inspected according to the collected items.

[0053] When the data to be inspected is audio data, the inspection terminal plays the audio data corresponding to the collected items. Inspection personnel sequentially retrieve and play the audio files corresponding to each item for verification, and input the inspection results based on the actual recording conditions. The inspection results can include "passed," "supplemented recording," and "re-recorded." The cloud platform or inspection terminal generates a corresponding inspection status based on the inspection results, which can include "passed," "pending supplementary collection," and "pending re-collection."

[0054] After completing the quality inspection of the corpus data corresponding to the same collection object, the quality inspectors can submit the inspection results through the submission function on the quality inspection page. The quality inspection terminal will then synchronize the quality inspection status to the cloud platform. Based on this, the cloud platform will standardize and archive the inspected data, and provide status information for subsequent processing, annotation task creation, and data retrieval.

[0055] In step S105, the cloud platform performs diversion processing on the corpus data to be inspected based on its quality inspection status. Diversion processing is used to send data with different quality statuses to different subsequent processes, allowing data that passes the quality inspection to enter the annotation stage, and allowing data that needs to be supplemented or re-collected to return to the collection stage.

[0056] When the quality inspection status is "passed," the cloud platform writes the corresponding corpus data to be inspected into the inspected data pool and designates it as corpus data to be labeled. Therefore, the corpus data to be labeled in the inspected data pool can serve as candidate data for subsequent labeling tasks.

[0057] When the quality inspection status is "Pending Supplementation," the cloud platform updates the collection status of the corresponding corpus data to "Pending Supplementation" and sends a supplementary collection instruction to the collection terminal. This instruction can carry information about the corresponding collection entries, allowing the user to view the entries that need to be added and perform the supplementary collection on the collection terminal. The user can view the entered words in the language collection module of the collection terminal. After the quality inspection is completed, the collection terminal displays the words that need to be added. The user clicks on the corresponding word and repeats the recording, stopping, saving, and uploading operations to complete the supplementary collection.

[0058] When the quality inspection status is "pending re-collection," the cloud platform updates the collection status of the corresponding corpus data to "pending re-collection" and sends a re-collection command to the collection terminal. The re-collection command can carry information about the corresponding collection item, allowing the user to re-enter the target corpus data for that item. After clicking on the item to be re-entered, the user repeats the recording, stopping, saving, and uploading operations to complete the re-collection.

[0059] When the cloud platform receives supplementary corpus data uploaded by the acquisition terminal according to the supplementary acquisition instruction, it adds the supplementary corpus data to the corpus record of the corresponding entry to be acquired. When the cloud platform receives re-acquired corpus data uploaded by the acquisition terminal according to the re-acquired acquisition instruction, it replaces the original corpus record of the corresponding entry to be acquired with the re-acquired corpus data.

[0060] Using the above method, when the subject misreads a word or the recording quality is unsatisfactory, the correct recording obtained from the re-collection can overwrite the previous erroneous record. When the subject skips words or does not complete a part of the entry, supplementary data can be added to the corpus record of the corresponding entry to be collected. This process can adapt to situations such as user skipping, subsequent re-recording, and fragmented recording, while also ensuring audio quality and reducing the workload of later quality inspection and annotation.

[0061] In step S106, the cloud platform creates annotation tasks and assigns them to the annotation end; the annotation end receives annotation data for the corpus data to be annotated according to the annotation tasks.

[0062] When creating an annotation task through the cloud platform, you can obtain the annotation task name, annotation personnel information, accent and region information, annotation vocabulary information, and data collection object information. Based on the accent and region information, annotation vocabulary information, and data collection object information, the cloud platform filters the corpus data to be annotated from the quality-checked data pool; it then associates the filtered corpus data with the annotation personnel information; based on the associated corpus data and annotation personnel information, it generates the annotation task and publishes it to the annotation client corresponding to the annotation personnel information.

[0063] In one specific implementation, the operator navigates to the administrator's annotation task management module and clicks the "Add Annotation Task" function entry to initiate the annotation task creation process. The annotation task is then named to clearly identify it. The operator then proceeds to the annotation personnel selection stage, using the system's built-in annotation personnel selection component to filter and determine the appropriate annotation personnel to undertake this annotation task.

[0064] In the accent region configuration stage, operators select the corresponding accent region from the system's accent region options based on the regional characteristics of the speakers collected for this corpus. In the vocabulary selection stage, operators filter and determine suitable annotation vocabulary from the system's vocabulary resource library according to the specific needs of this annotation task. In the speaker selection stage, operators use the system's search function to input search criteria such as speaker name and the corresponding accent region at the time of recording, filtering the corpus data to be annotated from the quality-checked data pool.

[0065] For scenarios involving a large number of speakers received daily and a complex geographical distribution, the cloud platform can quickly classify and batch annotation tasks according to accent region, vocabulary, language category, or target audience, and then carry out annotation work for each batch of tasks. Geographical classification and annotation makes the target audience for annotation tasks clearer, reducing offline paper recording and manual summarization.

[0066] After a new annotation task is added, the annotator can access the annotation task management module. They can search by entering the target task name, find the newly added annotation task, and then click the annotation operation entry to enter the data annotation page for the corresponding collection item. The annotation client retrieves the corpus data to be annotated for the annotation task and displays the corpus data according to the collection items.

[0067] When the corpus data to be annotated is speech data, the annotation terminal plays the speech data corresponding to the collected item. Annotators input the annotation content based on the pronunciation features in the recording. In the speech data scenario, the annotation content includes the International Phonetic Alphabet (IPA) annotation content corresponding to the speech data. After completing the annotation of a single collected item, the annotation terminal can switch to the next collected item, repeating the recording playback and IPA input operations until the IPA annotation work for all collected items under the annotation task is completed.

[0068] The cloud platform associates the annotated content with the corresponding corpus data to be annotated, the information of the data collection object, and the corpus attribute information to obtain annotated data. This annotated data is then uploaded to the cloud platform and stored in the annotated data pool. Thus, the IPA annotation results form a linked record with their corresponding recording data, speaker information, accent region, vocabulary, and entry identifiers, providing foundational data for subsequent morpheme verification, syllable verification, segment verification, and vocabulary verification.

[0069] In step S107, the cloud platform creates a verification task and assigns the verification task to the verification terminal; the verification terminal verifies the labeled data according to the verification task to obtain the verified corpus data.

[0070] In one specific implementation, the operator enters the administrator's verification task management unit and performs a new verification task creation operation. After the verification personnel enter the verification task management unit, the verification terminal retrieves the annotation data corresponding to the verification task from the labeled data pool. The annotation data can be the International Phonetic Alphabet annotation data submitted in the previous annotation stage, along with its corresponding collection object information, collection entry information, and corpus attribute information.

[0071] The verification end performs at least one of the following checks on the labeled data: morpheme verification, syllable verification, segment verification, and word list verification. In one embodiment, the verification end performs morpheme verification, syllable verification, segment verification, and word list verification on the labeled data. Morpheme verification is used to confirm the correspondence between the labeled data and the target corpus at the morpheme level; syllable verification is used to confirm the syllable structure of the labeled content; segment verification is used to confirm whether the segment records are consistent with speech features; and word list verification is used to confirm whether the labeled data matches the specified word list and collected entries.

[0072] The verification end configures the verification status of the labeled data based on the verification results. After the labeled data corresponding to the same collection object has been verified, if the verification status is "passed," the labeled data is determined as the verified corpus data. After completing the verification of all entries for the current speaker and confirming the verification results, the verifier can submit the verification results through the system's submission function.

[0073] Through the above verification process, the cloud platform can continue to standardize and verify the corpus annotation data obtained in the previous annotation stage, so that the IPA annotation content, collection entries, vocabulary and speaker information are consistent, and avoid confusion in the correspondence of the annotation data in the subsequent processing, analysis and application.

[0074] In step S108, the cloud platform stores the verified corpus data into the verified data pool. The verified data pool is used to store data that has completed the collection, quality inspection, annotation, and verification process. For all entries of the currently collected object, after the verification work is completed and the verification results are confirmed, the verification end can use the system submission function to push the complete annotation data of the collected object after verification to the verified data pool of the cloud platform.

[0075] The data in the verified data pool can be used for subsequent language resource retrieval, corpus export, speech research, dialect analysis, International Phonetic Alphabet data organization, and language and cultural resource applications. Through the phased storage of the quality-checked data pool, the labeled data pool, and the verified data pool, this invention enables standardized archiving and subsequent transfer of corpus data across different processing stages.

[0076] In one implementation, the cloud platform can also configure at least one status identifier among acquisition status, quality inspection status, annotation status, and verification status for the target corpus data according to the processing stages of the corpus data in acquisition, quality inspection, annotation, and verification, and determine the data pool corresponding to the target corpus data based on the status identifier. Thus, the cloud platform can uniformly manage the processing progress and flow path of the corpus data.

[0077] This embodiment also provides a cloud-based language corpus acquisition, quality inspection, annotation, and verification device. This device can be deployed on a cloud platform, or it can be formed collaboratively by the cloud platform, acquisition terminal, quality inspection terminal, annotation terminal, and verification terminal. The device includes an acquisition task generation module, a corpus acquisition module, a corpus association module, a quality inspection task processing module, a quality inspection routing module, an annotation task processing module, a verification task processing module, and a data storage module.

[0078] The data collection task generation module is used to acquire corpus collection information through the cloud platform and generate corpus collection tasks based on this information. This module can connect to the management system, vocabulary resource library, and phonology resource library to generate corpus collection tasks based on vocabulary information, language category information, phonology information, and item information.

[0079] The corpus acquisition module is used to send corpus acquisition tasks to the acquisition terminal and obtain the acquisition object information and the target corpus data corresponding to the corpus acquisition task through the acquisition terminal. The corpus acquisition module can receive personal information, phonetic selection information and segmented audio data entered by the acquisition object through a mini-program or mobile application, and receive, save and upload the target corpus data after the acquisition operation of each item to be acquired is completed.

[0080] The corpus association module is used to associate target corpus data with information about the collection object and corpus attributes to obtain corpus data to be inspected, and then upload the corpus data to be inspected to the cloud platform. The corpus association module can bind recording data with information such as speaker name, accent type, language proficiency information, accent region information, vocabulary identifier, phonological identifier, entry identifier, collection time, and corpus identifier.

[0081] The quality inspection task processing module is used to create quality inspection tasks through the cloud platform, allocate these tasks to the quality inspection end, and then use the quality inspection end to perform quality checks on the corpus data to be inspected based on the tasks, obtaining the corresponding quality inspection status. The module can filter target collection objects based on search criteria such as the vocabulary to be inspected, accent type, language mastery information, accent region information, collection object name, and current quality inspection status, and generate corresponding quality inspection tasks.

[0082] The quality inspection and routing module is used to route the corpus data to be inspected based on its quality inspection status. Corpus data with a "passed" status is identified as corpus data to be labeled. Collection entry information corresponding to corpus data with a "to be supplemented" or "to be re-collected" status is sent to the collection terminal, enabling the terminal to supplement or re-collect the target corpus data corresponding to the collection entry information. The quality inspection and routing module can also add supplementary corpus data to the corresponding corpus record of the item to be collected upon receiving supplementary corpus data, and replace the original corpus record of the corresponding item when re-collected corpus data is received.

[0083] The annotation task processing module is used to create annotation tasks through the cloud platform, allocate these tasks to annotation endpoints, and receive annotation data for the corpus data to be annotated through the annotation endpoints. The module can filter the corpus data to be annotated from the quality-checked data pool based on the annotation task name, annotator information, accent / region information, annotation vocabulary information, and data collection object information, and then publish the annotation task to the corresponding annotation endpoint.

[0084] The verification task processing module is used to create verification tasks through the cloud platform, assign these tasks to the verification endpoints, and then use these endpoints to verify the labeled data according to the verification tasks, obtaining the verified corpus data. The verification task processing module can retrieve the labeled data corresponding to the verification task from the labeled data pool and perform at least one of the following: morpheme verification, syllable verification, segment verification, and vocabulary verification.

[0085] The data storage module is used to store the verified corpus data into the verified data pool. The data storage module can also connect to the quality-inspected data pool, the labeled data pool, and the verified data pool to perform phased storage of the corpus data based on its acquisition, quality inspection, labeling, and verification status.

[0086] The modules described above can be software functional modules, hardware functional modules, or functional modules formed by a combination of software and hardware. For the specific processing procedures of each module in the above device embodiments, please refer to the descriptions of the corresponding steps in the foregoing method embodiments; these will not be repeated here.

[0087] Compared with the prior art, the present invention has the following advantages: This invention unifies the organization of corpus collection, quality inspection, annotation, and verification tasks through a cloud platform, enabling collaborative processing among collection terminals, quality inspection terminals, annotation terminals, and verification terminals, reducing data import / export and manual aggregation between local software. This invention supports collection of speech corpus data by collection item. The collection object can record, save, and upload the current item to be collected; when a new item to be collected exists, the collection terminal switches and displays the new current item, thus achieving segmented and continuous collection. This invention processes the corpus data to be inspected according to its quality inspection status, sending data that has passed the status into the annotation process, and feeding back the collection item information corresponding to data in the supplementary or re-collection status to the collection terminal, allowing the collection object to supplement or re-collect specific items, thus adapting to collection scenarios such as skipping, misreading, subsequent supplementary recording, and re-recording. Through the annotation task publishing and annotation terminal processing mechanism, this invention enables annotators to retrieve the corpus data to be annotated according to the collection items, play the corresponding recordings in the speech corpus data scenario, and input IPA annotation content, improving the standardization and processing efficiency of online IPA annotation. This invention connects the verification task with the labeled data pool to perform morpheme verification, syllable verification, segment verification and word list verification on the labeled data, and pushes the verified data to the verified data pool, so that the language corpus data forms a complete data closed loop from collection, quality inspection, labeling to verification.

[0088] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for quality inspection, annotation, and verification of language corpus collection based on a cloud platform, characterized in that, Includes the following steps: The system acquires corpus collection information through a cloud platform and generates corpus collection tasks based on the corpus collection information. The corpus collection task is sent to the collection terminal, and the collection terminal is used to obtain the collection object information and the target corpus data corresponding to the corpus collection task. The target corpus data is associated with the information of the collected objects and the corpus attribute information to obtain the corpus data to be inspected, and the corpus data to be inspected is uploaded to the cloud platform; Quality inspection tasks are created through the cloud platform and assigned to the quality inspection terminals; The quality inspection terminal performs quality inspection on the corpus data to be inspected according to the quality inspection task, and obtains the quality inspection status corresponding to the corpus data to be inspected. The data to be inspected is split according to the quality inspection status. The data to be inspected with the quality inspection status of "pass" is determined as the data to be labeled. The collection item information corresponding to the data to be inspected with the quality inspection status of "to be supplemented" or "to be re-collected" is sent to the collection terminal so that the collection terminal can supplement or re-collect the target data corresponding to the collection item information. Annotation tasks are created through the cloud platform and assigned to annotation endpoints; The annotation terminal receives annotation data for the corpus data to be annotated according to the annotation task. A verification task is created through the cloud platform and assigned to the verification terminal; The verification terminal verifies the labeled data according to the verification task to obtain the verified corpus data; The verified corpus data is stored in the verified data pool; The cloud platform also includes a quality-inspected data pool and a labeled data pool.

2. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, The corpus information collected includes at least one of the following: vocabulary information, language category information, phonological information, and entry information. The information collected includes at least one of the following: the name of the subject, accent type, language proficiency, and accent region. The corpus attribute information includes at least one of the following: the vocabulary to which the corpus belongs, the phonological system to which the corpus belongs, the entry to which the corpus belongs, the time of corpus collection, the status of corpus collection, and corpus identifier; The target corpus data includes at least speech corpus data, which includes recording data corresponding to the collected entries.

3. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, The corpus acquisition task is sent to the acquisition terminal, and the acquisition terminal is used to obtain the acquisition object information and the target corpus data corresponding to the corpus acquisition task, including: The data acquisition terminal receives the data acquisition object information entered by the data acquisition object and uploads the data acquisition object information to the cloud platform. According to the corpus collection task, the current entries to be collected are displayed on the collection terminal; The system receives target corpus data uploaded by the acquisition terminal after completing the acquisition operation of the current item to be acquired, and associates and saves the target corpus data with the item identifier of the current item to be acquired. When the target corpus data corresponding to the current item to be collected is saved and there is a next item to be collected, the next item to be collected is determined as the new current item to be collected, and the new current item to be collected is displayed on the collection terminal. The target corpus data corresponding to each item to be collected is received sequentially until the target corpus data corresponding to all items to be collected in the corpus collection task is collected.

4. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 3, characterized in that, The data to be inspected is split and processed according to the quality inspection status, including: When the quality inspection status is "passed", the corresponding data to be inspected is written into the inspected data pool, and the corresponding data to be inspected is determined as data to be labeled. When the quality inspection status is "to be supplemented for collection", the collection status of the corresponding corpus data to be inspected is updated to "to be supplemented for collection", and a supplementary collection instruction is sent to the collection terminal. When the quality inspection status is pending re-collection, the collection status of the corresponding corpus data to be inspected is updated to pending re-collection, and a re-collection command is sent to the collection terminal. When the acquisition terminal receives supplementary corpus data uploaded according to the supplementary acquisition instruction, the supplementary corpus data is added to the corpus record of the corresponding entry to be acquired; when the acquisition terminal receives re-acquired corpus data uploaded according to the re-acquired instruction, the original corpus record of the corresponding entry to be acquired is replaced with the re-acquired corpus data.

5. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, Creating quality inspection tasks through the cloud platform includes: Obtain the list of terms to be inspected; The system can receive at least one of the following search criteria: accent type of the subject being collected, language information, accent region information, subject name, and current quality inspection status. Target objects are selected from the cloud platform according to the search criteria; The range of the corpus data to be inspected is determined based on the vocabulary to be inspected and the target collection object. The quality inspection task is generated based on the range of the corpus data to be inspected.

6. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, Based on the quality inspection task, the quality inspection data to be inspected is performed to obtain the quality inspection status of the data to be inspected, including: The corpus data to be inspected corresponding to the quality inspection task is retrieved through the quality inspection terminal. The data of the corpus to be inspected is displayed according to the collected items; When the data to be inspected is speech data, the speech data corresponding to the collected item is played. Receive the quality inspection results input by the quality inspection personnel for the corpus data to be inspected; Generate the corresponding quality inspection status based on the quality inspection results; After completing the quality inspection of the corpus data to be inspected corresponding to the same collection object, the quality inspection status is synchronized to the cloud platform.

7. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, Creating annotation tasks through the cloud platform includes: Obtain the annotation task name, annotation personnel information, accent and region information, annotation vocabulary information, and data collection object information; Based on the accent region information, labeled vocabulary information, and data collection object information, the corpus data to be labeled is selected from the quality-checked data pool; The filtered corpus data to be annotated is associated with the annotator information; The annotation task is generated based on the associated corpus data to be annotated and the annotation personnel information, and the annotation task is published to the annotation terminal corresponding to the annotation personnel information.

8. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, Based on the annotation task, the system receives annotation data for the corpus data to be annotated, including: The annotation task's corresponding corpus data is retrieved through the annotation terminal; The corpus data to be annotated is displayed according to the collected items; When the data to be labeled is speech data, the speech data corresponding to the collected item is played. The annotation content input by the annotator for the corpus data to be annotated, when the corpus data to be annotated is speech data, includes the International Phonetic Alphabet annotation content corresponding to the speech data; The labeled content is associated with the corresponding corpus data to be labeled, the collection object information, and the corpus attribute information to obtain the labeled data; The labeled data is uploaded to the cloud platform and stored in the labeled data pool.

9. The method for quality inspection, annotation, and verification of language corpus based on a cloud platform according to claim 1, characterized in that, The labeled data is validated based on the validation task to obtain validated corpus data, including: The verification terminal retrieves the labeled data corresponding to the verification task from the labeled data pool. Perform at least one of the following on the labeled data: morpheme verification, syllable verification, phonetic segment verification, and word list verification; Configure the verification status for the labeled data based on the verification results; After the annotation data corresponding to the same collection object has been verified, the annotation data with the verification status of "passed" is determined as the verified corpus data, and the verified corpus data is pushed to the verified data pool.

10. A cloud-based language corpus acquisition quality inspection, annotation, and verification device, characterized in that, include: The data collection task generation module is used to obtain data collection information through the cloud platform and generate data collection tasks based on the data collection information. The corpus acquisition module is used to send the corpus acquisition task to the acquisition terminal, and to obtain the acquisition object information and the target corpus data corresponding to the corpus acquisition task through the acquisition terminal; The corpus association module is used to associate the target corpus data with the collection object information and corpus attribute information to obtain the corpus data to be inspected, and upload the corpus data to be inspected to the cloud platform; The quality inspection task processing module is used to create quality inspection tasks through the cloud platform, allocate the quality inspection tasks to the quality inspection terminal, and perform quality inspection on the corpus data to be inspected according to the quality inspection tasks through the quality inspection terminal to obtain the quality inspection status of the corpus data to be inspected. The quality inspection and diversion module is used to divert the corpus data to be inspected according to the quality inspection status. The corpus data to be inspected with a quality inspection status of "pass" is identified as corpus data to be labeled. The collection item information corresponding to the corpus data to be inspected with a quality inspection status of "to be supplemented for collection" or "to be re-collected" is sent to the collection terminal so that the collection terminal can supplement or re-collect the target corpus data corresponding to the collection item information. The annotation task processing module is used to create annotation tasks through the cloud platform, allocate the annotation tasks to the annotation terminal, and receive annotation data for the corpus data to be annotated through the annotation terminal according to the annotation tasks. The verification task processing module is used to create verification tasks through the cloud platform, allocate the verification tasks to the verification terminal, and verify the labeled data through the verification terminal according to the verification tasks to obtain the verified corpus data. The data storage module is used to store the verified corpus data into the verified data pool.