The application belongs to the technical field of
natural language processing, and relates to a Tibetan-Chinese
bilingual corpus collaborative labeling and versioned publishing
system and method. Through the
system architecture formed by the front-end interaction module, the business
processing module, the data storage module and the basic management and control module, relying on the Tibetan-Chinese bilingual labeling, intelligent
collaborative management and control, corpus
version management and standardized publishing core units integrated by the business
processing module, cooperating with the distributed
correlation index storage mechanism of the data storage module, the fine-grained permission and operation
log management and
control function of the basic management and control module, the core technical problems of the
Tibetan language characteristics
adaptation deficiency, the low efficiency and frequent conflicts of multi-person collaborative labeling, the lack of version tracing and quality grading control of the corpus, the non-standard publishing process and the disconnection of the corpus and the downstream model in the prior art are solved. The accuracy of Tibetan-Chinese
bilingual corpus labeling and storage is ensured through exclusive Tibetan
adaptation processing, and the orderly promotion of multi-person labeling is realized through modular
collaborative management and control.