Tibetan word segmentation method and system based on multi-language pre-training model CINO
By using the multilingual pre-training model CINO, collecting and optimizing data sets and dynamically adjusting training parameters, the problem of insufficient adaptability of the Tibetan word segmentation model was solved, achieving more efficient word segmentation accuracy and reliability, and improving the effect of Tibetan information processing.
Patent Information
- Application Number
- CN202510708316.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-26
AI Technical Summary
The existing Tibetan word segmentation model lacks verification and continuous learning mechanisms after training, which makes it difficult to adapt to new language phenomena and vocabulary updates, limiting its effectiveness and generalization ability in practical applications.
Using the multilingual pre-training model CINO, we collect datasets to be annotated for word segmentation conversion, divide the datasets into training and validation datasets, optimize the datasets and initialize the model, and dynamically adjust the learning rate and number of training rounds to ensure that the model is built based on high-quality data and scientific processes, thereby improving word segmentation accuracy and reliability.
The accuracy and model performance of Tibetan word segmentation have been improved, the application effect in the Tibetan field has been enhanced, the word segmentation results are presented intuitively through visualization, and the generalization ability and training efficiency of the model have been improved.
Smart Images

Figure CN120706424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a Tibetan word segmentation method and system based on a multilingual pre-training model CINO. Background Art
[0002] As an important component of minority languages, Tibetan holds significant research value in areas such as cultural heritage, information processing, and natural language processing. With the rapid development of the internet and big data technologies, the demand for Tibetan information processing is growing. Tibetan word segmentation, a fundamental task in Tibetan natural language processing, plays a key role in subsequent tasks such as part-of-speech tagging, syntactic analysis, machine translation, and information retrieval.
[0003] For example, the invention patent with announcement number CN114330328B announces a Tibetan word segmentation method based on Transformer-CRF. The method includes: inputting a data set, data preprocessing, syllable expansion, building a Tibetan word segmentation model based on Transformer-CRF, training and saving the model and its parameters, and inputting the data to be segmented, outputting the word segmentation results, and expanding two units to the left and right with the current syllable as the center. Using a combination of unigram and bigram methods, more feature vectors can be extracted.
[0004] For example, the invention patent with announcement number CN117556814B announces an integrated method and system for Tibetan word segmentation and part-of-speech tagging, which involves the field of electronic information. By obtaining Tibetan text information input by the user, calling the integrated model and segmenting Tibetan syllables and non-Tibetan character blocks, CRF prediction is performed to obtain the optimal label prediction. According to the results of the label prediction, the written form of each Tibetan syllable is sorted to obtain the corresponding tagging results.
[0005] However, in the process of implementing the embodiments of the present application, the present application found that the above technology has at least the following technical problems: the current training method of the Tibetan word segmentation model is a single mode, that is, the model is put into use directly after the training is completed, and there is a lack of subsequent verification, continuous learning and optimization mechanism. This training method makes it difficult for the model to adaptively adjust and improve when facing new language phenomena, vocabulary updates or changes in application scenarios, thereby limiting its effectiveness and generalization ability in practical applications. Summary of the Invention
[0006] In response to the deficiencies in the prior art, the present invention provides a Tibetan word segmentation method and system based on the multilingual pre-training model CINO, which can effectively solve the problems involved in the above-mentioned background technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: In a first aspect, the present invention provides a Tibetan word segmentation method based on a multilingual pre-training model CINO, including: step 1, collecting a data set to be labeled, and performing word segmentation conversion on the data set to be labeled, thereby obtaining a data set to be trained, obtaining and analyzing the attribute parameters of the data set to be trained, thereby determining whether to perform data division on the data set to be trained; step 2, obtaining a training data set and a verification data set through data division, using the training data set and the verification data set to train the multilingual pre-training model CINO, collecting and analyzing the training process parameters, thereby initializing the multilingual pre-training model CINO; step 3, obtaining a test data set, and using the initialized multilingual pre-training model CINO to perform Tibetan word segmentation on the test data set, thereby performing a visual display of Tibetan word segmentation.
[0008] As a further method, data optimization is performed on the training data set. The specific optimization process is as follows: Tibetan sentences with a length greater than a defined length are screened out from the training data set, and the sentences are aggregated and marked as the data set to be segmented; the data set to be segmented is subjected to secondary word segmentation conversion, and the data labeling completeness index of the data set to be segmented after the secondary word segmentation conversion is completed is obtained; the lengths of the Tibetan sentences belonging to the data set to be segmented after the secondary word segmentation conversion are performed on the mean to obtain the average length of the Tibetan sentences belonging to the data set to be segmented, and a threshold correction coefficient is matched from the training database to correct and update the data labeling completeness threshold; the data labeling completeness index of the data set to be segmented after the secondary word segmentation conversion is completed is compared with the corrected and updated data labeling completeness threshold; if the data set to be segmented after the secondary word segmentation conversion meets the first condition, the data set to be segmented is refilled into the training data set, and it is determined that the training data set is to be divided; if the data set to be segmented after the secondary word segmentation conversion meets the second condition, the data set to be segmented is subjected to secondary optimization.
[0009] As a further method, a secondary optimization is performed on the dataset to be segmented. The specific optimization process is as follows: based on the data annotation completeness index and the data annotation completeness threshold of the dataset to be segmented, the data annotation completeness deviation value of the dataset to be segmented is obtained, and the newly added number of word segmentation rules is matched from the training database. The word segmentation rules are added to the multilingual pre-training model CINO based on the newly added number of word segmentation rules, and word segmentation conversion is performed on the dataset to be segmented; after the word segmentation conversion is completed, Tibetan sentences with a length greater than a defined length are screened out from the dataset to be segmented after the secondary optimization is completed, and marked as a dataset to be warned, and the dataset to be segmented is refilled into the dataset to be trained, and data warning is performed on the dataset to be warned at the same time; at the warning limit time point, it is determined whether a data update instruction is received. If the data update instruction is not received, it is determined to directly perform data division on the dataset to be trained; if the data update instruction is received, the data in the data update instruction is filled into the dataset to be trained, and data division is performed on the dataset to be trained.
[0010] As a further method, the multilingual pre-training model CINO is initialized. The specific initialization process is: dividing the training data set into several batches of training data sets, and dividing the verification data set into several batches of verification data sets; training the multilingual pre-training model CINO with the first batch of training data sets; collecting and analyzing the first batch of training process parameters to obtain the training anomaly index of the first batch of training data sets, and comparing and analyzing the training anomaly threshold to obtain the training anomaly deviation value of the first batch of training data sets, and matching the verification accuracy threshold correction coefficient from the training database to correct the verification accuracy threshold; after the training of the first batch of training data sets is completed, the trained multilingual pre-training model CINO is verified with the first batch of verification data sets, collecting and analyzing the first batch of verification process parameters to obtain the verification accuracy coefficient of the first batch of verification data sets, and comparing it with the verification accuracy threshold after correction to determine whether to adjust the training process of the second batch of training data sets.
[0011] As a further method, it is determined whether to adjust the training process of the second batch of training data sets. The specific determination process is: if the verification accuracy coefficient of the first batch of verification data sets is greater than or equal to the verification accuracy threshold, it is determined that the training process of the second batch of training data sets is not adjusted; if the verification accuracy coefficient of the first batch of verification data sets is less than the verification accuracy threshold, it is determined that the training process of the second batch of training data sets is adjusted. The specific adjustment process is: based on the verification accuracy coefficient and the verification accuracy threshold of the first batch of verification data sets, the verification accuracy deviation value of the first batch of verification data sets is obtained, and the learning rate correction coefficient and the training round number correction coefficient are matched from the training database, so as to correct and adjust the learning rate and the number of training rounds of the training process of the second batch of training data sets; if the correction adjustment If the third condition exists after the holiday, a correction warning will be issued. The third condition refers to the learning rate being greater than the defined learning rate or the number of training rounds being greater than the defined number of training rounds. If the third condition does not exist after the correction and adjustment, no correction warning will be issued, and the training process parameters of the next adjacent batch will be continuously collected and analyzed, as well as the verification process parameters of the next adjacent batch will be collected. If the training process of a certain batch of training data sets is adjusted, and the verification accuracy coefficient of the next adjacent batch of training data sets corresponding to the batch of training data sets is still less than the verification accuracy coefficient, the next adjacent batch of training data sets corresponding to the batch of training data sets will be marked as the data set to be trained, so that the multilingual pre-training model CINO will be trained using the data set to be trained. If the data in the training data set has been used up, the initialization of the multilingual pre-training model CINO is completed.
[0012] A second aspect of the present invention provides a Tibetan word segmentation system based on a multilingual pre-training model CINO, comprising: a data partitioning module, for collecting a data set to be labeled, performing word segmentation conversion on the data set to be labeled, thereby obtaining a data set to be trained, acquiring and analyzing attribute parameters of the data set to be trained, thereby determining whether to perform data segmentation on the data set to be trained; a model initialization module, for obtaining a training data set and a verification data set through data partitioning, training the multilingual pre-training model CINO using the training data set and the verification data set, collecting and analyzing training process parameters, thereby initializing the multilingual pre-training model CINO; and a word segmentation visualization module, for obtaining a test data set, and performing Tibetan word segmentation on the test data set using the initialized multilingual pre-training model CINO, thereby performing a visual display of the Tibetan word segmentation.
[0013] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0014] (1) The present invention provides a Tibetan word segmentation method and system based on a multilingual pre-training model CINO. First, a data set to be annotated is collected to provide a massive text resource for subsequent research. Then, word segmentation conversion is performed to obtain a data set to be trained, and the text is converted into a word sequence suitable for model processing, which is convenient for the model to learn the structure. Then, the attribute parameters of the data set to be trained are analyzed to determine whether the data should be divided. Reasonable division can ensure the representativeness of the training set and the verification set, avoid data distribution deviation, and improve the generalization ability of the model. Subsequently, training and verification data sets are obtained through division and used to train the multilingual pre-training model CINO. The training process parameters are collected and analyzed, which can provide insight into the model training status, timely adjust the strategy and hyperparameters, complete the model initialization, and enable the model to fully utilize the training information, quickly obtain a good performance foundation, and improve the training efficiency. Finally, a test data set is obtained, and the initialized model is used to perform Tibetan word segmentation, and the results are then visualized. The test data set can objectively evaluate the model performance, and the visualization can intuitively present the word segmentation effect, promote the improvement of Tibetan word segmentation accuracy and reliability, and assist in the application of multilingual processing technology in the Tibetan field.
[0015] (2) The present invention obtains the completeness index of annotation by analyzing the attribute parameters and compares it with the threshold, which can scientifically determine whether to divide the data set, improve the pertinence of training, optimize the data set that does not meet the conditions, screen long sentences for secondary word segmentation and correct the threshold before re-comparison, which can improve the quality of data annotation. The whole process helps to ensure the completeness and rationality of the annotation of the training data set, so that the subsequent Tibetan word segmentation work based on the multilingual pre-training model CINO can be carried out based on better quality data, improve the accuracy of word segmentation and model performance, and enhance the Tibetan processing effect.
[0016] (3) The present invention processes the data set in batches, which can monitor the training and verification process in more detail. By analyzing the training parameters to obtain the abnormal index and correct the verification threshold, the training status can be accurately evaluated, potential problems can be discovered in time, and whether to adjust the subsequent training process can be determined based on the verification accuracy coefficient. The learning rate and the number of training rounds can be dynamically optimized to improve the training efficiency and effect. After the correction and adjustment, the early warning conditions can be set to avoid over-adjustment. Batches that still do not meet the standards after multiple adjustments are marked as training data sets, which can strengthen the training in a targeted manner. Finally, when the training data is used up, the model is initialized, which ensures that the model is built based on high-quality data and scientific processes, and improves the Tibetan word segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.
[0018] Figure 1 Schematic diagram of the method steps of the present invention.
[0019] Figure 2 This is a schematic diagram of system module connections of the present invention.
[0020] Figure 3 Schematic diagram of the data set partitioning process of the present invention.
[0021] Figure 4 Schematic diagram of the training process adjustment flow of the present invention.
[0022] Figure 5 This is a structural model diagram of the multilingual pre-training model CINO based on the present invention.
[0023] Figure 6 This is an input example diagram of the present invention.
[0024] Figure 7 These are examples of Tibetan word segmentation data in different formats according to the present invention.
[0025] Figure 8 This is the Tibetan word segmentation and visualization interface diagram of the present invention.
[0026] Figure 9 This is a diagram of the training set details interface of the present invention.
[0027] Figure 10 This is a diagram of the model parameter management interface of the present invention. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0029] Reference Figure 1 As shown, the first aspect of the present invention provides a Tibetan word segmentation method based on a multilingual pre-training model CINO, including: step 1, collecting a data set to be labeled, and performing word segmentation conversion on the data set to be labeled, thereby obtaining a data set to be trained, obtaining and analyzing attribute parameters of the data set to be trained, thereby determining whether to perform data division on the data set to be trained; step 2, obtaining a training data set and a verification data set through data division, using the training data set and the verification data set to train the multilingual pre-training model CINO, collecting and analyzing training process parameters, thereby initializing the multilingual pre-training model CINO; step 3, obtaining a test data set, and using the initialized multilingual pre-training model CINO to perform Tibetan word segmentation on the test data set, thereby performing a visual display of Tibetan word segmentation.
[0030] The aforementioned dataset to be labeled refers to a dataset that has not been manually or automatically labeled; the aforementioned dataset to be trained refers to a dataset that has undergone a series of preprocessing operations (such as word segmentation conversion) and already contains feature information that can be used for model training.
[0031] In the following text, the multilingual pre-training model CINO is referred to as the model.
[0032] In a specific embodiment, the present invention analyzes attribute parameters to obtain a labeling completeness index and compares it with a threshold value, which can scientifically determine whether to divide the data set, improve the targeted training, optimize the data set that does not meet the conditions, screen long sentences for secondary word segmentation, and correct the threshold before comparing again, which can improve the quality of data labeling. The entire process helps to ensure the labeling integrity and rationality of the data set to be trained, so that subsequent Tibetan word segmentation work based on the multilingual pre-training model CINO can be carried out based on higher-quality data, improve word segmentation accuracy and model performance, and enhance the Tibetan processing effect.
[0033] Figure 5 This is a structural model diagram based on the multilingual pre-training model CINO of the present invention, which consists of an input layer, a CINO layer, and an output layer. The main function of the input layer is to convert the original Tibetan sentence to be segmented into a syllable sequence, and the syllable sequence into a word sequence. The syllable sequence is obtained by segmenting the Tibetan syllable separator, and the word sequence is obtained by segmenting the word segmenter provided by the pre-training model CINO. It means sharpening the sword before the battle. Figure 6 This is an input example diagram of the present invention, which shows the original Tibetan word sequence to be segmented, Tibetan syllable sequence and word unit sequence in the input layer. After the syllable sequence in the input layer is converted into a word unit sequence, the length of the word unit sequence changes, and a new start mark " <s> "、End tag"< / s> " and connection mark " ", as well as multiple tokens generated after some syllable characters are further segmented. For classification tasks, the original classification labels are still applicable to the token sequence, but for sequence labeling tasks, the labels corresponding to each syllable character in the syllable character sequence cannot correspond one-to-one to each token in the token sequence. For this problem, the present invention uniformly assigns a "-100" label to the newly generated tokens, which effectively solves the problem of mismatch between the token sequence and the original label sequence. Specifically, in the token sequence, for un-tokenized syllable characters and when the syllable characters are segmented into multiple tokens after tokenization, the first token uses its original corresponding label, and the newly added start mark, end mark and connection mark, as well as the non-first token when the syllable character is segmented into multiple tokens, are assigned a "-100" label. The "-100" label will be automatically ignored in the process of calculating the loss function and will not affect the training of the model. Figure 6 The data to be segmented in is It means all living things.
[0034] The core of this Tibetan word segmentation model is the CINO layer. Its main function is to assign an initial vector to each word based on the word sequence provided by the input layer and output a label for each word using the CINO pre-trained model. Each word vector assigned by the CINO layer incorporates rich general Tibetan language knowledge, as well as rich knowledge of languages such as English and Chinese (acquired through the language transfer capabilities of the multilingual pre-trained model). To examine the impact of different numbers of Tibetan word segmentation labels on the Tibetan word segmentation model, a comparison was conducted between 2-tag and 4-tag Tibetan word segmentation data. When the model uses 2-tag Tibetan word segmentation data, the CINO layer outputs a label set of {"B", "M"}; when the model uses 4-tag Tibetan word segmentation data, the CINO layer outputs a label set of {"B", "M", "E", "S"}. The output layer's main function is to segment the Tibetan text to be segmented based on the sequence labels output by the CINO layer and output the segmentation results. Specifically, the output layer identifies word boundaries based on the output labels of the CINO layer, i.e., labels such as "S" and "B", and adds a word segmentation symbol " / " to the right boundary of the word. It should be noted that the output label of the CINO layer contains the sequence start marker " <s>", etc. These labels have no effect on Tibetan word segmentation, but affect the annotation of syllable word sequences. Therefore, these redundant labels are cleaned in the output layer.
[0035] Figure 7 This is an example diagram of Tibetan word segmentation data in different formats of the present invention, where Figure 7 The manual word segmentation data is converted into a 2-tag annotation sequence. The specific steps are as follows: First, the input manual word segmentation sequence (space-delimited) into a list of individual words, where It means to sharpen your sword before the battle; then, define the word segmentation tags as "B" (for the initial syllable) and "M" (for non-initial syllable); initialize an empty list result to store the final annotation sequence; then, traverse each word in the manual word segmentation sequence: for the current word, use the syllable separator "·" to split it into a syllable list; if the current word contains only one syllable (that is, the length of the syllable list is 1), directly mark the syllable as "B" and add it to the result list; if the current word contains two or more syllables (that is, the length of the syllable list is greater than or equal to 2), mark the initial syllable as "B" and add it to the result list; then, loop through the remaining syllables (the number of loops is the syllable list length minus 1), mark all non-initial syllables as "M", and add them to the result list in turn; finally, the result list is the converted 2-tag annotation sequence, in which each syllable is correctly marked as "B" or "M".
[0036] in Figure 7 The 4-tag tagging sequence in
[15] is constructed as follows: first, the input manual word segmentation sequence (each word is separated by a space) is split into a list of individual words; then, the word segmentation tags are defined, including "B" (for the initial syllable), "E" (for the final syllable), "M" (for non-initial and non-final syllables in a word), and "S" (for monosyllabic words); then, an empty list result is initialized to store the final tagging sequence. Then traverse each word in the manual word segmentation sequence: for the current word, use the syllable separator "·" to split it into a syllable list; if the current word is a monosyllabic word (that is, the syllable list length is 1), add the label "S" to the result list; if the current word consists of two syllables, add the labels "B" and "E" to the result list in turn; if the current word has three syllables, add the labels "B", "M", and "E" to the result list in turn; if the current word has four or more syllables, first add the label "B" to the result list, and then loop through the syllables except the first and last (the number of loops is the syllable list length minus 2), mark them all as "M" and add them to the result list, and finally add the label "E"; finally, the result list is the converted 4-tag annotation sequence, in which each syllable is correctly labeled as "B", "E", "M", or "S".
[0037] Specifically, it is determined whether to divide the data set to be trained. The specific determination process is: by analyzing the attribute parameters of the data set to be trained, the data labeling completeness index of the data set to be trained is obtained; the data labeling completeness threshold is extracted from the training database and compared with the data labeling completeness index of the data set to be trained. If the first condition exists in the data set to be trained, it is determined that the data set to be trained is divided; the above-mentioned data labeling completeness threshold represents the minimum value allowed by the data labeling completeness index.
[0038] If the second condition exists in the training data set, it is determined that the training data set will not be divided and data optimization will be performed on the training data set; the first condition means that the data labeling completeness index is greater than or equal to the data labeling completeness threshold; the second condition means that the data labeling completeness index is less than the data labeling completeness threshold.
[0039] The attribute parameters of the dataset to be trained include the ratio of unlabeled words in the dataset to be trained, the extreme difference in the labeling density of the dataset to be trained, and the labeling coverage of the dataset to be trained; the above-mentioned ratio of unlabeled words refers to the ratio of the number of unlabeled words in the dataset to be trained to the total number of words; the above-mentioned extreme difference in the labeling density refers to the difference between the maximum and minimum values of the labeling density in the dataset. The labeling density represents the ratio of labels within a unit length (such as a sentence or paragraph), and the extreme difference reflects the fluctuation range of the labeling density in the dataset; the above-mentioned labeling coverage refers to the ratio of the number of labeled words or entities in the dataset to the total number of words or entities that should be labeled. The ratio of unlabeled words, the extreme difference in the labeling density, and the labeling coverage can all be analyzed using programming languages (such as Python).
[0040] Metrics are introduced to quantify the degree of influence of the proportional relationship between the ratio of unlabeled words and the ratio of defined unlabeled words, the proportional relationship between the extreme difference of annotation density and the extreme difference of defined annotation density, and the proportional relationship between the annotation coverage and the defined annotation coverage on the data annotation completeness index. The various influence degrees are aggregated to obtain the data annotation completeness index.
[0041] It should be explained that the unlabeled word ratio reflects the proportion of unlabeled words in the dataset. When the unlabeled word ratio is high, it means that a large number of words have not been labeled, which will directly lead to a decrease in the labeling density (because the labeling density reflects the proportion of labeled words in the whole. If there are many unlabeled words, the proportion of labeled words will naturally decrease). At the same time, the labeling coverage (that is, the extent to which the labeled words cover the entire vocabulary range) will also decrease accordingly, because many words are not labeled, making the labeling unable to fully cover the vocabulary content of the dataset. The labeling density range describes the degree of difference in the labeling density of different parts of the dataset. If the labeling density range is large, it means that the dataset The labeling of each part is uneven, with some areas densely labeled and some areas sparsely labeled. This imbalance will further affect the labeling coverage, because sparsely labeled areas will lead to incomplete overall labeling coverage. At the same time, the ratio of unlabeled words will appear relatively high due to insufficient labeling in some areas. Overall, the increase in the ratio of unlabeled words and the increase in the extreme difference in labeling density will lead to a decrease in the labeling coverage, and the changes in these three parameters will eventually have a comprehensive effect on the data labeling completeness index. A too high ratio of unlabeled words, a too large extreme difference in labeling density, and a too low labeling coverage will all reduce the data labeling completeness index, indicating that the labeling work of the dataset is not perfect and comprehensive.
[0042] The data annotation completeness index of the training dataset represents the completeness of the data annotation of the training dataset. The specific expression is:
[0043]
[0044] Where DA is the data annotation completeness index of the training dataset, LB is the ratio of unlabeled words in the training dataset, JC is the range of annotation density in the training dataset, FG is the annotation coverage of the training dataset, J_LB is the ratio of defined unlabeled words preset in the training database, J_JC is the range of defined annotation density preset in the training database, J_FG is the defined annotation coverage preset in the training database, ro1 is the unlabeled word ratio measurement value preset in the training database, ro2 is the range of annotation density measurement value preset in the training database, and ro3 is the annotation coverage measurement value preset in the training database.
[0045] The above-mentioned unlabeled word ratio measurement value is used to quantify the influence of the unlabeled word ratio unit value on the data labeling completeness index; the above-mentioned labeling density range measurement value is used to quantify the influence of the labeling density range unit value on the data labeling completeness index; the above-mentioned labeling coverage measurement value is used to quantify the influence of the labeling coverage unit value on the data labeling completeness index; the training database stores the mapping relationship between the unlabeled word ratio, labeling density range and labeling coverage and their corresponding measurement values. For example, the unlabeled word ratio, labeling density range and labeling coverage are input into the training database, and the training database can retrieve the unlabeled word ratio measurement value, labeling density range measurement value and labeling coverage measurement value. The value range of the unlabeled word ratio measurement value, labeling density range measurement value and labeling coverage measurement value is between 0 and 1.
[0046] The above definition of the unlabeled word ratio indicates the maximum value allowed for the unlabeled word ratio; the above definition of the annotation density range indicates the maximum value allowed for the annotation density range; the above definition of the annotation coverage rate indicates the minimum value allowed for the annotation coverage rate.
[0047] Specifically, the training dataset is optimized. The specific optimization process is as follows: Tibetan sentences with a length greater than a defined length are screened from the training dataset, aggregated and marked as a dataset to be segmented, a secondary word segmentation conversion is performed on the segmented dataset, and the data annotation completeness index of the dataset after the secondary word segmentation conversion is completed is obtained. The aforementioned defined length represents the maximum allowable length and is extracted from the training database. The dataset to be segmented is a set of Tibetan sentences selected from the training dataset according to a specific rule (i.e., sentence length greater than the defined length). After the initial word segmentation conversion and other preprocessing, these sentences may exceed the preset defined length and may not meet the specific requirements of subsequent model training or processing. Therefore, they need to be separated for further processing.
[0048] After the secondary word segmentation conversion is completed, the lengths of the Tibetan sentences in the dataset to be segmented are averaged to obtain the average length of the Tibetan sentences in the dataset to be segmented, and the threshold correction coefficient is matched from the training database to correct and update the data annotation completeness threshold. The training database stores an average length-threshold correction coefficient mapping table. The average length of the Tibetan sentences in the dataset to be segmented is directly queried in the training database to obtain the threshold correction coefficient corresponding to the average length of the Tibetan sentences in the dataset to be segmented. The threshold correction coefficient represents the proportional coefficient for correcting the data annotation completeness threshold. The threshold correction coefficient is multiplied by the data annotation completeness threshold to obtain the corrected and updated data annotation completeness threshold.
[0049] The data annotation completeness index of the data set to be segmented after the secondary word segmentation conversion is completed is compared with the corrected and updated data annotation completeness threshold. If the data set to be segmented after the secondary word segmentation conversion is completed meets the first condition, the data set to be segmented is refilled into the data set to be trained, and the data division of the training data set is determined; if the data set to be segmented after the secondary word segmentation conversion is completed meets the second condition, the data set to be segmented is optimized twice.
[0050] Furthermore, the data set to be segmented is optimized twice. The specific optimization process is as follows: based on the data labeling completeness index and the data labeling completeness threshold of the data set to be segmented, the data labeling completeness deviation value of the data set to be segmented is obtained, and the number of new word segmentation rules is matched from the training database. The word segmentation rules are added to the multilingual pre-training model CINO based on the number of new word segmentation rules, and the data set to be segmented is converted into words; the above-mentioned data labeling completeness deviation value is used to quantify the degree of deviation between the data labeling completeness index and the data labeling completeness threshold. Specifically, the data labeling completeness threshold is subtracted from the data labeling completeness index, and the result is the data labeling completeness deviation value; the above-mentioned number of new word segmentation rules refers to the number of new word segmentation rules that need to be added to the model. The specific matching process is: the data labeling completeness deviation value-new number of word segmentation rules mapping table is stored in the training database, and the data labeling completeness deviation value is directly queried in the training database to obtain the number of new word segmentation rules corresponding to the data labeling completeness deviation value.
[0051] After the word segmentation conversion is completed, Tibetan sentences with a length greater than the defined length are screened out from the dataset to be segmented after the secondary optimization is completed, and marked as a dataset to be warned. The dataset to be segmented is refilled into the dataset to be trained, and at the same time, a data warning is issued for the dataset to be warned. The specific warning refers to displaying the dataset to be warned with visual instructions, prompting relevant personnel to manually perform word segmentation conversion. The above-mentioned dataset to be warned is a collection of Tibetan sentences further screened out from the dataset to be segmented after the data processing step of word segmentation conversion is completed. These Tibetan sentences have a significant feature, that is, their length exceeds the pre-set defined length. Since this extra-long feature may bring potential problems or risks to subsequent data processing, model training and other links, they need to be paid special attention to and warned, and are therefore marked as datasets to be warned.
[0052] At the warning definition time point, it is determined whether the data update instruction is received. If the data update instruction is not received, it is determined to directly divide the data into the training data set; the above-mentioned warning definition time point refers to the latest time point for receiving the data update instruction after the data warning operation is completed for the warning data set; the above-mentioned data update instruction includes the results generated by relevant personnel after dividing the Tibetan sentences in the warning data set.
[0053] If a data update instruction is received, the data in the data update instruction is filled into the data set to be trained, and the data set to be trained is divided.
[0054] Step 2: Obtain a training dataset and a validation dataset through data partitioning, use the training dataset and the validation dataset to train the multilingual pre-training model CINO, collect and analyze the training process parameters, and thus initialize the multilingual pre-training model CINO.
[0055] The training dataset is the basic data set for model learning and parameter adjustment. The validation dataset is the data set used to evaluate model performance, adjust model hyperparameters, and prevent model overfitting during model training.
[0056] The above-mentioned data division to obtain the training data set and the verification data set refers to dividing the data set to be trained according to the data volume ratio preset by relevant personnel.
[0057] In a specific embodiment, the present invention processes the data set in batches, which can monitor the training and verification process more carefully. By analyzing the training parameters to obtain the abnormality index and correct the verification threshold, the training status can be accurately evaluated, potential problems can be discovered in a timely manner, and whether to adjust the subsequent training process can be determined based on the verification accuracy coefficient. The learning rate and the number of training rounds can be dynamically optimized to improve the training efficiency and effect. After the correction and adjustment, the early warning conditions can be set to avoid excessive adjustment. Batches that still do not meet the standards after multiple adjustments are marked as data sets that need training, and targeted training can be strengthened. Finally, when the training data is used, the model is initialized, ensuring that the model is built based on high-quality data and scientific processes, thereby improving the Tibetan word segmentation performance.
[0058] Specifically, the multilingual pre-training model CINO is initialized. The specific initialization process is: divide the training dataset into several batches of training datasets, and divide the verification dataset into several batches of verification datasets; relevant experimenters divide the training dataset into several batches of training datasets based on factors such as the total amount of data, computing resource carrying capacity, and model training strategy. Each batch contains a certain amount of sample data; at the same time, according to the same division logic and scale standards, the verification dataset is divided into a corresponding number of batches of verification datasets. Under this division system, the one-to-one correspondence rule is followed, that is, a batch of training datasets forms a corresponding relationship with a batch of verification datasets. During subsequent model training, each batch of data can be used in turn for training and verification according to this correspondence to ensure the scientific nature of the training process and the reliability of the evaluation results.
[0059] The multilingual pre-training model CINO is trained using the first batch of training data sets; the first batch of training process parameters are collected and analyzed to obtain the training anomaly index of the first batch of training data sets, and compared and analyzed with the training anomaly threshold to obtain the training anomaly deviation value of the first batch of training data sets, and the verification accuracy threshold correction coefficient is matched from the training database to correct the verification accuracy threshold; the above-mentioned training anomaly threshold represents the maximum value allowed by the training anomaly index and is extracted from the training database; the above-mentioned training anomaly deviation value refers to the result of subtracting the training anomaly threshold from the training anomaly index; the above-mentioned verification accuracy threshold correction coefficient refers to the proportional value for correcting the verification accuracy threshold, and the verification accuracy threshold correction coefficient is multiplied by the verification accuracy threshold, and the result is the corrected verification accuracy threshold, where the verification accuracy threshold correction coefficient has the following specific matching rules: the training anomaly deviation value-verification accuracy threshold correction coefficient mapping table is stored in the training database, and the training anomaly deviation value is directly queried in the training database to obtain the verification accuracy threshold correction coefficient corresponding to the training anomaly deviation value.
[0060] After the training of the first batch of training data sets is completed, the trained multilingual pre-training model CINO is verified using the first batch of verification data sets. The first batch of verification process parameters are collected and analyzed to obtain the verification accuracy coefficient of the first batch of verification data sets. The coefficient is then compared with the verification accuracy threshold after correction to determine whether to adjust the training process of the second batch of training data sets.
[0061] Specifically, the training anomaly index of the first batch of training data sets, the specific analysis process is as follows: the first batch of training process parameters, including: the loss function value of the first batch of training data sets, the parameter update amplitude of the first batch of training data sets and the training time of the first batch of training data sets; the above-mentioned loss function value is an indicator to measure the difference between the prediction results of the model on the first batch of training data sets and the true labels; the above-mentioned parameter update amplitude refers to the amplitude of the update of the model parameters after backpropagation and optimization algorithm (such as gradient descent) on the first batch of training data sets; the above-mentioned training time refers to the time required for the model to process the first batch of training data sets; the loss function value, parameter update amplitude and training time can all be obtained from the training log of the model.
[0062] Obtain the final value of the data labeling complete index of the data set to be trained, and match the correction coefficient from the training database; the above-mentioned final value of the data labeling complete index refers to the data labeling complete index of the data set to be trained that is finally input into the model for data training, which can be obtained by analysis in a programming language (such as Python); the above-mentioned correction coefficient refers to the proportional coefficient for correcting the training anomaly index and the verification accuracy coefficient. The specific matching process is: the data labeling complete index final value-correction coefficient mapping table is stored in the training database, and the data labeling complete index final value is directly queried in the training database to obtain the correction coefficient corresponding to the data labeling complete index final value.
[0063] A measurement factor is introduced to quantify the influence of the proportional relationship between the loss function value and the defined loss function value, the proportional relationship between the parameter update amplitude and the defined parameter update amplitude, and the proportional relationship between the training time and the defined training time on the training anomaly index. The influence degrees are summarized and the summary results are corrected using the correction coefficient to obtain the training anomaly index of the first batch of training data sets.
[0064] It needs to be explained that the loss function value is closely related to the parameter update amplitude. The parameter update amplitude is often calculated based on the gradient of the loss function value. The larger the loss function value, the worse the fit between the model and the data, and the larger the gradient, which in turn leads to a larger parameter update amplitude and a drastic adjustment of the model parameters. Conversely, when the loss function value is small, the parameter update amplitude is also relatively small, and the model parameters are adjusted smoothly. The training time is related to the parameter update amplitude. When the parameter update amplitude is too large, the model parameters change rapidly, resulting in an unstable and oscillating training process, which may take longer to converge, indicating that there are abnormalities in the batch training process, and the model training is unstable or has not been carried out effectively.
[0065] The training anomaly index of the first batch of training data sets represents the degree of training anomaly of the first batch of training data sets. The specific expression is:
[0066]
[0067] Where HUI is the training anomaly index of the first batch of training data sets, Z_DA is the correction coefficient, BM is the loss function value of the first batch of training data sets, PN is the parameter update amplitude of the first batch of training data sets, TO is the training time of the first batch of training data sets, J_BM is the bounded loss function value preset in the training database, J_PN is the bounded parameter update amplitude preset in the training database, J_TO is the bounded training time preset in the training database, do1 is the loss function value measurement factor preset in the training database, do2 is the parameter update amplitude measurement factor preset in the training database, and do3 is the training time measurement factor preset in the training database.
[0068] The above-mentioned definition of the loss function value indicates the maximum value allowed by the loss function value; the above-mentioned definition of the parameter update amplitude indicates the maximum value allowed by the parameter update amplitude; the above-mentioned definition of the training duration indicates the maximum value allowed by the training duration.
[0069] The above-mentioned loss function value measurement factor is used to quantify the degree of influence of the unit value of the loss function value on the training anomaly index; the above-mentioned parameter update amplitude measurement factor is used to quantify the degree of influence of the unit value of the parameter update amplitude on the training anomaly index; the above-mentioned training duration measurement factor is used to quantify the degree of influence of the unit value of the training duration on the training anomaly index; the training database stores the mapping relationship between the loss function value, parameter update amplitude and training duration and their corresponding measurement factors. For example, the loss function value, parameter update amplitude and training duration are input into the training database, and the training database can retrieve the loss function value measurement factor, parameter update amplitude measurement factor and training duration measurement factor. The value range of the loss function value measurement factor, parameter update amplitude measurement factor and training duration measurement factor is between 0 and 1.
[0070] Furthermore, the verification accuracy coefficient of the first batch of verification data sets, the specific analysis process is as follows: the first batch of verification process parameters, including: the accuracy of the first batch of verification data sets, the recall rate of the first batch of verification data sets and the F1 value of the first batch of verification data sets; the above-mentioned accuracy rate refers to the proportion of the number of samples correctly predicted by the model on the verification data set to the total number of samples; the above-mentioned recall rate refers to the proportion of the number of samples correctly predicted as positive by the model on the verification data set to the actual number of positive samples; the above-mentioned F1 value is the harmonic mean of the accuracy rate and the recall rate; the accuracy rate, recall rate and F1 value can all be obtained from the training log of the model.
[0071] Metric factors are introduced to quantify the influence of the training anomaly index, the proportional relationship between the accuracy and the defined accuracy, the proportional relationship between the recall rate and the defined recall rate, and the ratio relationship between the F1 value and the defined F1 value on the verification accuracy coefficient. The influence degrees are aggregated and the aggregated results are corrected using the correction coefficient to obtain the verification accuracy coefficient.
[0072] It should be explained that the training anomaly index reflects the degree of anomaly of the model during the training process. If the training anomaly index is too high, it means that there may be data problems and algorithm instability during the training process. These anomalies will directly interfere with the model's normal learning of the data, making it difficult for the model to accurately capture the features and rules in the data, and then causing the model to make more errors in predictions and reduce the accuracy. Because the accuracy measures the proportion of samples predicted correctly by the model to the total samples, training anomalies reduce the proportion of correct predictions. At the same time, training anomalies will also lead to a decrease in the recall rate. The recall rate focuses on the proportion of positive samples correctly identified by the model to the actual positive samples. Training anomalies may cause the model to miss many samples that are actually positive. The F1 value is the harmonic mean of the precision and recall rates. The reduction of the precision and recall rates will inevitably lead to a decrease in the F1 value, because the F1 value combines the evaluation of the model performance by the two. When both perform poorly, the F1 value will also be at a low level. Accuracy, recall and F1 values are key indicators of model performance. Their performance will have a comprehensive effect on the verification accuracy coefficient, which is used to measure the overall performance of the model on the verification set. The reduction of accuracy, recall and F1 values means that the prediction effect of the model on the verification set has deteriorated, thereby reducing the verification accuracy coefficient, indicating that the performance of the model in the verification stage is not ideal, and there may be problems such as overfitting, underfitting or training anomalies.
[0073] The verification accuracy coefficient of the first batch of verification data sets represents the accuracy of the verification of the first batch of verification data sets. The specific expression is:
[0074]
[0075] Where MKO is the verification accuracy coefficient of the batch training dataset, Z_DA is the correction coefficient, HUI is the training anomaly index of the batch training dataset, VA is the accuracy of the first batch verification dataset, VR is the recall of the first batch verification dataset, BS is the F1 value of the first batch verification dataset, J_VA is the bounded accuracy preset in the training database, J_VR is the bounded recall preset in the training database, J_BS is the bounded F1 value preset in the training database, sc1 is the training anomaly index measurement factor preset in the training database, sc2 is the accuracy measurement factor preset in the training database, sc3 is the recall measurement factor preset in the training database, and sc4 is the F1 value measurement factor preset in the training database.
[0076] The above definition of accuracy indicates the minimum value allowed for accuracy; the above definition of recall indicates the minimum value allowed for recall; the above definition of F1 value indicates the minimum value allowed for F1 value.
[0077] The above-mentioned training anomaly index measurement factor is used to quantify the influence of the unit value of the training anomaly index on the verification accuracy coefficient; the above-mentioned accuracy measurement factor is used to quantify the influence of the unit value of the accuracy on the verification accuracy coefficient; the above-mentioned recall rate measurement factor is used to quantify the influence of the unit value of the recall rate on the verification accuracy coefficient; the above-mentioned F1 value measurement factor is used to quantify the influence of the unit value of the F1 value on the verification accuracy coefficient; the training database stores the mapping relationship between the training anomaly index, accuracy, recall rate and F1 value and their corresponding measurement factors. For example, the training anomaly index, accuracy, recall rate and F1 value are input into the training database, and the training database can retrieve the training anomaly index measurement factor, accuracy measurement factor, recall measurement factor and F1 value measurement factor. The value range of the training anomaly index measurement factor, accuracy measurement factor, recall measurement factor and F1 value measurement factor is between 0 and 1.
[0078] Furthermore, it is determined whether to adjust the training process of the second batch of training data sets. The specific determination process is: if the verification accuracy coefficient of the first batch of verification data sets is greater than or equal to the verification accuracy threshold, it is determined not to adjust the training process of the second batch of training data sets; the above-mentioned verification accuracy threshold represents the minimum value allowed by the verification accuracy coefficient, which is extracted from the training database.
[0079] If the verification accuracy coefficient of the first batch of verification data sets is less than the verification accuracy threshold, it is determined that the training process of the second batch of training data sets is adjusted. The specific adjustment process is: based on the verification accuracy coefficient and verification accuracy threshold of the first batch of verification data sets, the verification accuracy deviation value of the first batch of verification data sets is obtained, and the learning rate correction coefficient and the number of training rounds correction coefficient are matched from the training database, so as to correct and adjust the learning rate and the number of training rounds of the training process of the second batch of training data sets; the above-mentioned verification accuracy deviation value refers to the result obtained by subtracting the verification accuracy coefficient from the verification accuracy threshold; the above-mentioned learning rate correction coefficient refers to the correction of the learning rate The above-mentioned training round correction coefficient refers to the proportional value for correcting the number of training rounds; the learning rate correction coefficient is multiplied by the learning rate, and the training round correction coefficient is multiplied by the number of training rounds. The result is the corrected learning rate and number of training rounds, among which the learning rate correction coefficient and the training round correction coefficient are specifically matched as follows: a mapping table of verified accurate deviation value-learning rate correction coefficient and a mapping table of verified accurate deviation value-training round correction coefficient are stored in the training database. The verified accurate deviation value is directly queried in the training database to obtain the learning rate correction coefficient and the training round correction coefficient corresponding to the verified accurate deviation value.
[0080] If the third condition exists after the correction and adjustment, a correction warning will be issued. The third condition means that the learning rate is greater than the defined learning rate or the number of training rounds is greater than the defined number of training rounds; the above-mentioned defined learning rate refers to the maximum value allowed for the learning rate; the above-mentioned defined number of training rounds refers to the maximum value allowed for the number of training rounds; the defined learning rate and the defined number of training rounds are both extracted from the training database; the above-mentioned correction warning means notifying relevant experimental personnel of abnormalities in the training process through a visual warning method.
[0081] If the third condition does not exist after the correction and adjustment, no correction warning will be issued, and the training process parameters of the next adjacent batch will continue to be collected and analyzed, and the verification process parameters of the next adjacent batch will be collected; if the training process of a batch of training data sets is adjusted, and the verification accuracy coefficient of the next adjacent batch training data set corresponding to the batch training data set is still less than the verification accuracy coefficient, then the next adjacent batch training data set corresponding to the batch training data set will be marked as the data set to be trained, so that the multilingual pre-training model CINO will be trained using the data set to be trained; if the data in the training data set has been used up, the initialization of the multilingual pre-training model CINO is completed.
[0082] Step 3: Obtain a test dataset and use the initialized multilingual pre-trained model CINO to perform Tibetan word segmentation on the test dataset, thereby visually displaying the Tibetan word segmentation.
[0083] Specifically, a visual display of Tibetan word segmentation is performed. The specific analysis process is as follows: the test data set is input into the initialized multilingual pre-training model CINO, and the result output by the multilingual pre-training model CINO is marked as the result data set; the reference result data set of the test data set is obtained, and the result data set is overlapped and compared with the result data set, and the result accuracy of the test data set is obtained; the above test data set is formulated by relevant experimental personnel, and the test set data content includes fields such as religion, culture, law and politics, and the styles include gatha and long lines. Because the Tibetan word segmentation training data has foreign punctuation marks such as "(", as well as Punctuation marks in the Tibetan language, as well as non-Tibetan characters such as English and Chinese, are segmented from Tibetan syllables. To objectively evaluate model performance, syllable separators are added after these symbols and non-Tibetan characters in the test set to prevent them from cluttering with Tibetan syllables and affecting model prediction and evaluation. The reference result dataset for the aforementioned test dataset was annotated by relevant experimenters. The accuracy of the test dataset results is a metric that measures the degree of consistency between the result dataset and the reference result dataset. It reflects the overall performance of the model on the entire test dataset. For example, based on overlap comparison using data processing software (such as Matrix Lab), the number of samples whose model output results are consistent with the reference results is counted across all samples. For example, if the test dataset has 100 samples and, after a one-by-one comparison, the model output for 80 samples is consistent with the reference results, the accuracy can be calculated based on this statistical result. The reference result dataset for the aforementioned test dataset represents the reference results of the result dataset and is established by relevant experimenters.
[0084] Obtain the comparison process of the test dataset, and based on the test dataset's accuracy and the comparison process, visualize the Tibetan word segmentation of the test dataset. The comparison process of the test dataset can be obtained from the model training log. The visualization refers to the use of intuitive visual elements such as graphics, charts, and images to clearly and vividly present the Tibetan word segmentation in the test dataset.
[0085] In a specific embodiment, the present invention provides a Tibetan word segmentation method and system based on a multilingual pre-training model CINO. First, a data set to be annotated is collected to provide massive text resources for subsequent research. Then, word segmentation conversion is performed to obtain a data set to be trained, and the text is converted into a word sequence suitable for model processing to facilitate model learning structure. Then, the attribute parameters of the data set to be trained are analyzed to determine whether data partitioning is required. Reasonable partitioning can ensure the representativeness of the training set and the validation set, avoid data distribution bias, and improve the generalization ability of the model. Subsequently, training and validation data sets are obtained through partitioning and used to train the multilingual pre-training model CINO. The training process parameters are collected and analyzed, which can provide insight into the model training status, timely adjust the strategy and hyperparameters, complete model initialization, and enable the model to fully utilize the training information, quickly obtain a good performance foundation, and improve training efficiency. Finally, a test data set is obtained, and Tibetan word segmentation is performed using the initialized model. The results are then visualized. The test data set can objectively evaluate the model performance, and the visual display can intuitively present the word segmentation effect, promote the improvement of Tibetan word segmentation accuracy and reliability, and facilitate the application of multilingual processing technology in the Tibetan field.
[0086] Reference Figure 2 As shown, the second aspect of the present invention provides a Tibetan word segmentation system based on the multilingual pre-training model CINO, including: a data partitioning module, a model initialization module, a word segmentation visualization module and a training database.
[0087] The training database is used to store parameters involved in a Tibetan word segmentation system based on the multilingual pre-training model CINO.
[0088] The data partitioning module is connected to the model initialization module, the model initialization module is connected to the word segmentation visualization module, and the data partitioning module, the model initialization module and the word segmentation visualization module are all connected to the training database.
[0089] The data partitioning module is used to collect the data set to be labeled, and perform word segmentation conversion on the data set to be labeled to obtain the data set to be trained, obtain and analyze the attribute parameters of the data set to be trained, and thus determine whether to perform data partitioning on the data set to be trained.
[0090] The model initialization module is used to obtain a training dataset and a validation dataset through data partitioning, use the training dataset and the validation dataset to train the multilingual pre-training model CINO, collect and analyze the training process parameters, and thus initialize the multilingual pre-training model CINO.
[0091] The word segmentation visualization module is used to obtain a test dataset and use the initialized multilingual pre-trained model CINO to perform Tibetan word segmentation on the test dataset, thereby visually displaying the Tibetan word segmentation.
[0092] Figure 3 The data set division process diagram of the present invention is as follows. After the process starts, the data set is first collected and word segmentation conversion is performed, that is, the Tibetan data set to be annotated is collected and word segmentation conversion is performed to obtain the data set to be trained; then the attribute parameters of the data set are analyzed to understand its characteristics, and based on the analysis results, it is determined whether the data set needs to be divided into a training set and a validation set. If the division conditions are met, the data set is divided; if not, the data set is optimized, including processing long sentences and re-evaluating the division conditions. If the conditions are still not met after optimization, further optimization is required; after the division or optimization is completed, the multilingual pre-training model CINO is trained using the training set and the validation set. Figure 4 This is a schematic diagram of the training process adjustment process of the present invention. The training set and validation set are trained and verified in batches. During the process, relevant parameters are collected and analyzed to evaluate the model training effect and verification accuracy. Based on the analysis results, it is determined whether the training process needs to be adjusted (such as adjusting the learning rate, the number of training rounds, etc.). If no adjustment is required, the next batch of training is continued. If adjustment is required, the corresponding adjustment is made. Subsequently, it is determined whether the model has completed all batches of training. If not, the training is continued. If so, the model is initialized. After that, a Tibetan data set for testing is collected, and Tibetan word segmentation is performed on it using the initialized model. The word segmentation results are visualized for easy observation and analysis. Finally, the process ends.
[0093] Figure 8 Tibetan word segmentation and visualization interface of the present invention Figure 2 From a functional perspective, the interface has a Tibetan text input area where users can enter Tibetan sentences or paragraphs to be segmented. The system will process the input text based on the built-in Tibetan word segmentation algorithm, dividing the continuous Tibetan character sequence into semantically meaningful words. The interface will display the segmentation results in a clear and intuitive manner. Figure 9 This is the training set details interface of the present invention. It lists basic information about the training set, such as the name of the training set, to help users understand the source and timeliness of the training set and clearly understand the amount of data used to train the model. At the same time, it may also display the distribution of training set samples, helping users evaluate the coverage of the training set for various Tibetan texts and determine whether the model can learn comprehensive Tibetan language features. Figure 10 This is the model parameter management interface diagram of the present invention, which clearly displays various parameter information of the model. These parameters may include but are not limited to learning rate, number of iterations, etc. Users can view the current setting values of various parameters on this interface.
[0094] The above content is merely an example and explanation of the structure of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the scope of protection of the present invention.< / s>
Claims
1. A Tibetan word segmentation method based on the multilingual pre-training model CINO, characterized by: include: Step 1: Collect the dataset to be labeled and perform word segmentation conversion on the dataset to be labeled, thereby obtaining the dataset to be trained, obtaining and analyzing the attribute parameters of the dataset to be trained, and determining whether to perform data segmentation on the dataset to be trained; Step 2: Determine the training dataset and validation dataset through data partitioning, use the training dataset and validation dataset to train the multilingual pre-training model CINO, collect and analyze the training process parameters, and initialize the multilingual pre-training model CINO; Step 3: Obtain a test dataset and use the initialized multilingual pre-trained model CINO to perform Tibetan word segmentation on the test dataset, thereby visually displaying the Tibetan word segmentation.
2. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 1, characterized in that: The specific process of determining whether to perform data division on the training data set is as follows: By analyzing the attribute parameters of the training data set, the data annotation completeness index of the training data set is obtained; Extract the data annotation completeness threshold from the training database and compare it with the data annotation completeness index of the training data set. If the training data set meets the first condition, the training data set is determined to be divided into data; If the training data set meets the second condition, it is determined that the training data set is not to be divided, and data optimization is performed on the training data set; The first condition refers to that the data annotation completeness index is greater than or equal to the data annotation completeness threshold; The second condition refers to the data annotation completeness index being less than the data annotation completeness threshold; The attribute parameters of the training dataset include the ratio of unlabeled words in the training dataset, the extreme difference in labeling density of the training dataset, and the labeling coverage of the training dataset; The metrics are introduced to quantify the impact of the ratio between the unlabeled word ratio and the defined unlabeled word ratio, the ratio between the extreme difference in annotation density and the extreme difference in defined annotation density, and the ratio between the annotation coverage and the defined annotation coverage on the data annotation completeness index. The impact of each factor is aggregated to derive the data annotation completeness index. The data annotation completeness index of the dataset to be trained represents the completeness of the data annotation of the dataset to be trained.
3. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 2, characterized in that: The data optimization for the training dataset is performed as follows: Screen out Tibetan sentences with a length greater than a defined length from the training dataset, aggregate and mark them as a dataset to be segmented, perform a secondary word segmentation conversion on the dataset to be segmented, and obtain the data annotation completeness index of the dataset to be segmented after the secondary word segmentation conversion is completed; After the secondary word segmentation conversion is completed, the length of the Tibetan sentences in the dataset to be segmented is averaged to obtain the average length of the Tibetan sentences in the dataset to be segmented. The threshold correction coefficient is then matched from the training database to correct and update the data annotation completeness threshold. The data annotation completeness index of the data set to be segmented after the second word segmentation conversion is completed is compared with the corrected and updated data annotation completeness threshold. If the data set to be segmented after the second word segmentation conversion meets the first condition, the data set to be segmented is refilled into the training data set, and the data partitioning of the training data set is determined; If the data set to be segmented meets the second condition after the secondary word segmentation conversion is completed, the data set to be segmented is optimized twice.
4. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 3, characterized in that: The secondary optimization of the data set to be segmented is performed, and the specific optimization process is as follows: Based on the data annotation completeness index and data annotation completeness threshold of the dataset to be segmented, the data annotation completeness deviation value of the dataset to be segmented is obtained, and the number of newly added word segmentation rules is matched from the training database. The newly added word segmentation rules are added to the multilingual pre-training model CINO based on the number of newly added word segmentation rules, and the segmentation conversion is performed on the dataset to be segmented. After the word segmentation conversion is completed, Tibetan sentences with a length greater than the defined length are screened out from the dataset to be segmented after the secondary optimization is completed, and marked as a dataset to be warned. The dataset to be segmented is then re-filled into the dataset to be trained, and data warnings are issued for the dataset to be warned. At the warning definition time point, it is determined whether the data update instruction is received. If the data update instruction is not received, it is determined to directly divide the training data set; If a data update instruction is received, the data in the data update instruction is filled into the data set to be trained, and the data set to be trained is divided.
5. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 1, characterized in that: The specific initialization process of the multilingual pre-training model CINO is as follows: Divide the training dataset into several batches of training datasets, and divide the validation dataset into several batches of validation datasets; Train the multilingual pre-training model CINO using the first batch of training datasets; Collect and analyze the first batch of training process parameters to obtain the training anomaly index of the first batch of training data sets, and compare and analyze them with the training anomaly threshold to obtain the training anomaly deviation value of the first batch of training data sets, and match the verification accuracy threshold correction coefficient from the training database to correct the verification accuracy threshold; After the training of the first batch of training data sets is completed, the trained multilingual pre-training model CINO is verified using the first batch of verification data sets. The first batch of verification process parameters are collected and analyzed to obtain the verification accuracy coefficient of the first batch of verification data sets. The coefficient is then compared with the verification accuracy threshold after correction to determine whether to adjust the training process of the second batch of training data sets.
6. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 5, characterized in that: The determination of whether to adjust the training process of the second batch of training data sets is specifically as follows: If the verification accuracy coefficient of the first batch of verification data sets is greater than or equal to the verification accuracy threshold, it is determined that the training process of the second batch of training data sets will not be adjusted; If the verification accuracy coefficient of the first batch of verification data sets is less than the verification accuracy threshold, it is determined that the training process of the second batch of training data sets is adjusted. The specific adjustment process is: based on the verification accuracy coefficient and the verification accuracy threshold of the first batch of verification data sets, the verification accuracy deviation value of the first batch of verification data sets is obtained, and the learning rate correction coefficient and the training round number correction coefficient are matched from the training database, so as to correct and adjust the learning rate and the number of training rounds of the training process of the second batch of training data sets; If the third condition exists after the correction adjustment, a correction warning is issued, wherein the third condition refers to the learning rate being greater than the defined learning rate or the number of training rounds being greater than the defined number of training rounds; If the third condition does not exist after the correction and adjustment, no correction warning is issued, and the training process parameters of the next adjacent batch are continuously collected and analyzed, and the verification process parameters of the next adjacent batch are collected; If the training process of a certain batch of training data sets is adjusted, and the verification accuracy coefficient of the next adjacent batch of training data sets corresponding to the batch of training data sets is still smaller than the verification accuracy coefficient, the next adjacent batch of training data sets corresponding to the batch of training data sets will be marked as the training data set to be trained, so that the training data set to be trained is used to train the multilingual pre-training model CINO; If the data in the training dataset is used up, the multilingual pre-training model CINO is initialized.
7. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 5, characterized in that: The specific analysis process of the training anomaly index of the first batch of training data sets is as follows: The first batch of training process parameters include: the loss function value of the first batch of training data sets, the parameter update amplitude of the first batch of training data sets, and the training time of the first batch of training data sets; Obtain the final value of the data annotation complete index of the training data set and match the correction coefficient from the training database; The measurement factors are introduced to quantify the influence of the proportional relationship between the loss function value and the defined loss function value, the proportional relationship between the parameter update amplitude and the defined parameter update amplitude, and the proportional relationship between the training time and the defined training time on the training anomaly index. The influence degrees are summarized and the summary results are corrected using the correction coefficient to obtain the training anomaly index of the first batch of training data sets. The training anomaly index of the first batch of training data sets represents the degree of training anomaly of the first batch of training data sets.
8. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 5, characterized in that: The verification accuracy coefficient of the first batch of verification data sets, the specific analysis process is as follows: The first batch verification process parameters include: the accuracy of the first batch verification data set, the recall rate of the first batch verification data set, and the F1 value of the first batch verification data set; Metric factors are introduced to quantify the influence of the training anomaly index, the proportional relationship between the accuracy rate and the defined accuracy rate, the proportional relationship between the recall rate and the defined recall rate, and the ratio between the F1 value and the defined F1 value on the verification accuracy coefficient. The influence degrees are aggregated and the aggregated results are corrected using the correction coefficient to obtain the verification accuracy coefficient. The verification accuracy coefficient of the first batch of verification data sets represents the accuracy of the verification of the first batch of verification data sets.
9. The Tibetan word segmentation method based on the multilingual pre-training model CINO according to claim 1, characterized in that: The above-mentioned visualization display of Tibetan word segmentation is as follows: The test dataset is input into the initialized multilingual pre-training model CINO. The output of the multilingual pre-training model CINO is marked as the result dataset. Obtain the reference result dataset of the test dataset and compare it with the result dataset to obtain the accuracy of the test dataset; Obtain the comparison process of the test dataset, and visualize the Tibetan word segmentation of the test dataset based on the result accuracy of the test dataset and the comparison process of the test dataset.
10. A system using the Tibetan word segmentation method based on the multilingual pre-training model CINO according to any one of claims 1 to 9, characterized in that: include: The data partitioning module is used to collect the data set to be labeled, perform word segmentation conversion on the data set to be labeled, thereby obtaining the data set to be trained, obtain and analyze the attribute parameters of the data set to be trained, and determine whether to perform data partitioning on the data set to be trained; The model initialization module is used to obtain a training dataset and a validation dataset through data partitioning, train the multilingual pre-training model CINO using the training dataset and the validation dataset, collect and analyze the training process parameters, and thus initialize the multilingual pre-training model CINO; The word segmentation visualization module is used to obtain a test dataset and use the initialized multilingual pre-trained model CINO to perform Tibetan word segmentation on the test dataset, thereby visually displaying the Tibetan word segmentation.
Citation Information
Patent Citations
Tibetan word segmentation method based on Transformer-CRF
CN114330328B
A Tibetan word segmentation and part-of-speech tagging integrated method and system
CN117556814B
Cited By
Low-resource language word segmentation model training and cross-model word list migration injection method
CN122197877A