Natural language type labeling method and device for large model training data, equipment and medium

The combination of Spark and FastText method to perform natural language type annotation on large-scale data, solving the problems of inefficiency and insufficient accuracy in the existing technology, and achieving efficient and accurate labeling results.

CN120372359APending Publication Date: 2025-07-25SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510518975.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing natural language type annotation methods are inefficient and inaccurate when processing large-scale data.

Method used

The target application programming interface of Spark is used to preprocess the annotation training data, and the data is sliced to nodes in the Spark cluster, language detection is performed in parallel through the FastText model, and the preliminary annotation results are optimized using Spark.

Benefits of technology

It significantly improves the efficiency and accuracy of natural language type annotation, reduces time costs, ensures workload balancing in distributed processing, avoids data skew problems, and improves the reliability of the annotation results through correction rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372359A_ABST
    Figure CN120372359A_ABST
Patent Text Reader

Abstract

The invention discloses a natural language type labeling method and device for large model training data, equipment and a medium, and relates to the technical field of natural language processing, and the method comprises the steps: carrying out the preprocessing of to-be-labeled training data through a target application programming interface of Spark, and storing the obtained processed data to the local; reading the processed data from the local based on the target application programming interface, and fragmenting the processed data, so as to distribute the obtained fragmented data to each node in a Spark cluster; and performing language detection on the fragmented data in parallel through a FastText model on each node to obtain a corresponding preliminary labeling result, and optimizing the preliminary labeling result by using the Spark to obtain an optimized target labeling result. Therefore, the problems of low efficiency and insufficient accuracy in the natural language type labeling process when large-scale data is processed can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method, device, equipment and medium for natural language type annotation of large model training data. Background Technique

[0002] With the advent of the big data era, the amount of data faced by natural language processing tasks is increasing day by day. Existing language type annotation methods are carried out by using some language recognition tools and libraries, but they often rely on single-machine processing, resulting in obvious performance bottlenecks and poor performance when dealing with large-scale data sets. That is, there are problems such as low efficiency and insufficient accuracy when dealing with large-scale data.

[0003] In summary, it can be seen that how to solve the problems of low efficiency and insufficient accuracy in the process of natural language type annotation when dealing with large-scale data is an urgent problem to be solved at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for natural language type annotation of large model training data, which can solve the problems of low efficiency and insufficient accuracy in the process of natural language type annotation when dealing with large-scale data. The specific solutions are as follows:

[0005] In the first aspect, the present application provides a method for natural language type annotation of large model training data, including:

[0006] Preprocess the training data to be annotated by using the target application programming interface of Spark, and save the processed data obtained to the local;

[0007] Read the processed data from the local based on the target application programming interface, and slice the processed data to distribute the sliced data to each node in the Spark cluster;

[0008] Perform language detection on the sliced data in parallel through the FastText model on each node to obtain corresponding preliminary annotation results, and use the Spark to optimize the preliminary annotation results to obtain optimized target annotation results.

[0009] Optionally, before preprocessing the training data to be annotated by using the target application programming interface of Spark, it further includes:

[0010] Determine the scale of the Spark cluster according to the data volume of the training data to be annotated;

[0011] After determining the scale, deploy the FastText language detection library to each node of the Spark cluster;

[0012] Train the FastText language detection library based on a pre-collected multilingual corpus and adjust the parameters of the FastText language detection library to obtain the FastText model.

[0013] Optionally, the preprocessing of the data to be labeled and trained using the target application programming interface of Spark to save the processed data obtained to the local includes:

[0014] Use the target application programming interface of Spark to scan the data to be labeled and trained to clean the data to be labeled and trained, and filter the data to be labeled and trained based on preset data quality rules and a preset function to obtain filtered data;

[0015] Convert the filtered data into a preset encoding format and perform normalization processing on the filtered data according to preset normalization rules to obtain normalized data;

[0016] According to a preset stop word list, use Spark to process the normalized data to remove the stop words in the normalized data, and perform text enhancement processing on the normalized data based on a preset enhancement technique to obtain enhanced data;

[0017] Judge whether the data effect of the enhanced data meets the preset preprocessing effect;

[0018] If the data effect meets the preset preprocessing effect, partition the enhanced data according to the language directory of the enhanced data to obtain corresponding processed data, and save the processed data to the local.

[0019] Optionally, the judging whether the data effect of the enhanced data meets the preset preprocessing effect includes:

[0020] Perform hierarchical processing on the enhanced data according to the characteristics of the enhanced data, and randomly sample the data in each obtained layer to display the sampled data through a preset interface, so as to obtain corresponding data inspection feedback through the preset interface;

[0021] If the data inspection feedback indicates that the data inspection passes, obtain the judgment result that the data effect of the enhanced data meets the preset preprocessing effect;

[0022] If the data check feedback indicates that the data check fails, a judgment result that the data effect of the enhanced data does not meet the preset preprocessing effect is obtained.

[0023] Optionally, the parallel language detection of the sharded data by the FastText models on each node to obtain corresponding preliminary annotation results includes:

[0024] Parallelly execute respective corresponding language detection tasks by the FastText models on each node to perform language detection on the sharded data, and obtain each language detection result output by each node;

[0025] Based on Spark, integrate the language detection results to obtain a language type annotation result set aggregated from the language detection results, and use the language type annotation result set as the preliminary annotation result;

[0026] Among them, in the process of parallelly executing respective corresponding language detection tasks by the FastText models on each node, it includes:

[0027] Continuously monitor the running status of each node, and when a target node with an abnormal situation is detected, process the language detection task on the target node according to the abnormal situation.

[0028] Optionally, the optimization of the preliminary annotation result by using Spark includes:

[0029] Use Spark to extract features from the text data in the preliminary annotation result to obtain corresponding data features;

[0030] Construct corresponding feature vectors according to the extracted data features;

[0031] Based on several preset machine learning algorithms in Spark, use the feature vectors to train a preset initial model respectively to obtain several trained models corresponding to the several preset machine learning algorithms;

[0032] Evaluate the several trained models respectively through preset evaluation indicators, and determine the annotation result optimization model with the best performance from the several trained models according to the obtained evaluation results;

[0033] Apply the annotation result optimization model to the preliminary annotation result to optimize the preliminary annotation result.

[0034] Optionally, after optimizing the preliminary annotation result by using Spark to obtain an optimized target annotation result, it further includes:

[0035] Based on a preset correction rule, correct the target annotation result to correct the incorrect text data in the target annotation result into correct text data;

[0036] According to the analysis result of analyzing relevant cases of the incorrect text data, adjust the feature extraction project and the parameters of the annotation result optimization model, so as to obtain the target annotation result whose annotation accuracy meets the preset accuracy condition.

[0037] In a second aspect, the present application provides a natural language type annotation device for large model training data, including:

[0038] A data preprocessing module, configured to preprocess the training data to be annotated by using a target application programming interface of Spark, and save the obtained processed data locally;

[0039] A data sharding module, configured to read the processed data from local based on the target application programming interface, and shard the processed data, so as to distribute the obtained sharded data to each node in the Spark cluster;

[0040] A result optimization module, configured to perform language detection on the sharded data in parallel through the FastText model on each node to obtain corresponding preliminary annotation results, and use the Spark to optimize the preliminary annotation results to obtain an optimized target annotation result.

[0041] In a third aspect, the present application provides an electronic device, including:

[0042] A memory, configured to store a computer program;

[0043] A processor, configured to execute the computer program to implement the natural language type annotation method for large model training data described above.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the natural language type annotation method for large model training data described above is implemented.

[0045] In this embodiment, the target application programming interface of Spark is used to preprocess the training data to be labeled, and the processed data obtained is saved locally. Based on the target application programming interface, the processed data is read from local, and the processed data is sharded to distribute the sharded data to each node in the Spark cluster. The FastText model on each node is used to parallelly perform language detection on the sharded data to obtain corresponding preliminary annotation results, and Spark is used to optimize the preliminary annotation results to obtain optimized target annotation results. As can be seen from the above, in this application, the target application programming interface of Spark is first used to preprocess the training data to be labeled to obtain processed data, and the processed data is saved locally. Subsequently, the processed data is read from local through the target application programming interface, and the processed data is sharded to distribute the sharded data to each node in the Spark cluster. The FastText model on each node is used to parallelly perform language detection on the sharded data, and Spark is used to optimize the obtained preliminary annotation results to obtain optimized target annotation results. In this way, through the above process of this application, using the target application programming interface of Spark to preprocess the training data to be labeled can significantly improve the data processing speed and reduce the time cost. The form of combining FastText with Spark is used to perform language detection on the sharded data and optimize the preliminary annotation results, improving the accuracy of natural language type annotation. Through the data sharding strategy, the workload balance in the distributed processing process is ensured, avoiding the data skew problem, and thus solving the problems of low efficiency and insufficient accuracy in the natural language type annotation process when dealing with large-scale data. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a method for natural language type annotation of large model training data disclosed in this application;

[0048] Figure 2 It is a timing flowchart of a method for natural language type annotation of large model training data disclosed in this application;

[0049] Figure 3 It is a schematic structural diagram of a device for natural language type annotation of large model training data disclosed in this application;

[0050] Figure 4 This is a structural diagram of an electronic device disclosed in the present application. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] With the advent of the big data era, the amount of data faced by natural language processing tasks is increasing day by day. The existing language type annotation methods are carried out by using some language recognition tools and libraries, but they often rely on single-machine processing, resulting in obvious performance bottlenecks and poor performance when processing large-scale data sets. That is, there are problems such as low efficiency and insufficient accuracy when processing large-scale data.

[0053] To overcome the above technical problems, the present application provides a natural language type annotation method for large model training data to solve the problems of low efficiency and insufficient accuracy in the natural language type annotation process when processing large-scale data.

[0054] See Figure 1 As shown, the embodiments of the present invention disclose a natural language type annotation method for large model training data, including:

[0055] Step S11: Preprocess the training data to be annotated by using the target application programming interface of Spark, and save the obtained processed data locally.

[0056] In this embodiment, the training data to be annotated is preprocessed by using the target application programming interface of Spark (an open-source big data processing engine) to obtain the corresponding processed data, and the processed data is saved locally. Among them, the target application programming interface may be the DataFrame (a core data structure in Spark) application programming interface; the training data to be annotated is data with natural language type annotation requirements.

[0057] It should be noted that before preprocessing the to-be-annotated training data using the target application programming interface, it is necessary to deploy the FastText (a fast text classification algorithm) language detection library to the Spark to support the function of natural language type annotation. The processing flow is as follows: Determine the scale of the Spark cluster according to the data volume of the to-be-annotated training data; After determining the scale, deploy the FastText language detection library to each node of the Spark cluster; Train the FastText language detection library based on the pre-collected multilingual corpus and adjust the parameters of the FastText language detection library to obtain the FastText model. Among them, the parameters include but are not limited to the learning rate, the number of iterations, the vector dimension, etc. That is to say, in this embodiment, the scale of the Spark cluster can be dynamically adjusted according to the data volume of the to-be-annotated training data, and the number of nodes of the corresponding Spark cluster can be determined according to the scale to adapt to different-scale data processing requirements. After determining the scale and the number of nodes, deploy the FastText language detection library to each node of the Spark cluster and ensure that the versions and configurations of the FastText language detection libraries are consistent so as to run seamlessly in a distributed environment. Subsequently, train the FastText language detection library based on the pre-collected multilingual corpus and optimize the model performance by adjusting the parameters of the FastText language detection library to make it more suitable for language detection tasks in a distributed environment.

[0058] It should be further noted that in this embodiment, the target application programming interface can be used to perform preprocessing operations such as data cleaning, format unification, noise removal, and text normalization on the to-be-annotated training data, such as Figure 2The following is a time-sequence flowchart of a natural language type annotation method for large model training data provided by this application. Among them, the processing flow of the preprocessing is as follows: Use the target application programming interface of Spark to scan the training data to be annotated, so as to clean the data of the training data to be annotated, and filter the training data to be annotated based on preset data quality rules and a preset function to obtain filtered data; Convert the filtered data into a preset encoding format, and perform normalization processing on the filtered data according to preset normalization rules to obtain normalized data; According to the preset stop word list, use Spark to process the normalized data to remove the stop words in the normalized data, and perform text enhancement processing on the normalized data based on a preset enhancement technique to obtain enhanced data; Judge whether the data effect of the enhanced data meets the preset preprocessing effect; If the data effect meets the preset preprocessing effect, partition the enhanced data according to the language directory of the enhanced data to obtain corresponding processed data, and save the processed data to the local. Among them, the data quality rules include but are not limited to the text length range, field content type, etc.; The preset function is a built-in function of Spark or a user-defined UDF (i.e., User Defined Function); The stop word list is used to store high-frequency but less informative words, such as "de", "he", "shi", etc.; The preset enhancement techniques include synonym replacement technique, stemming technique.That is, use the target application programming interface of Spark to scan the training data to be labeled, identify and delete data records containing null values, null values (representing a null object), or incomplete fields to ensure that the dataset for subsequent processing is complete and meaningful, so as to perform data cleaning on the training data to be labeled; and define data quality rules, and filter the training data to be labeled based on the preset data quality rules and preset functions, so as to filter out data records that do not conform to the rules through the data quality rules and remove them from the dataset, and identify and remove HTML (i.e., HyperText Markup Language) tags, script codes, and non-text characters in the text, such as control characters, invisible characters, etc. through the preset function, so as to obtain the filtered data; then convert the filtered data into a preset encoding format; perform normalization processing on the filtered data according to the preset normalization rules to obtain normalized data; construct corresponding stop word lists for different languages, so as to process the normalized data according to the preset stop word lists to remove the stop words therein, reduce data noise, and thus improve the efficiency and accuracy of language detection; use a preset enhancement technique to perform text enhancement processing on the normalized data to obtain enhanced data; then determine whether the data effect of the enhanced data meets the preset preprocessing effect. If it meets, partition the enhanced data according to the language directory of the enhanced data to facilitate more efficient subsequent distributed processing, so as to obtain the corresponding processed data and save it locally.

[0059] Specifically, for converting the filtered data into a preset encoding format, in this embodiment, the encoding format can be unified through the configuration options of Spark or the parameter settings during data reading, such as UTF-8 (a data encoding format), so as to avoid encoding errors during the processing. In addition, data from different sources and in different formats can be converted into a consistent structure. For example, all data can be converted into the DataFrame format for subsequent distributed processing. Subsequently, the text processing library of Spark can be used to normalize the filtered data. For example, to improve the accuracy of language detection, all text can be converted into a unified case format through the string processing functions of Spark, such as converting all to lowercase. According to the characteristics of different languages, the punctuation marks in the text can be standardized, such as removing unnecessary punctuation marks and unifying the punctuation marks of different languages. For text containing numbers, it can be converted into text form or standardized as needed to reduce interference with language detection. The long text can be segmented into smaller units, such as sentences or paragraphs, for more accurate detection of the language type. For the text enhancement processing of the normalized data based on the preset enhancement technology, in order to enhance the diversity of the data, in this embodiment, the synonym replacement technology can be used to replace some words in the text and extract the stems of the text to restore them to the basic form to improve the generalization ability of language detection. At the same time, when partitioning the enhanced data, in order to ensure the workload balance of each partition, the data can also be evenly partitioned to avoid the data skew problem.

[0060] It should be noted that to determine whether the data effect of the enhanced data meets the preset preprocessing effect, the processing flow is as follows: perform hierarchical processing on the enhanced data according to the characteristics of the enhanced data, and randomly sample the data in each obtained layer, so as to display the sampled data through a preset interface, so as to obtain corresponding data inspection feedback through the preset interface; if the data inspection feedback indicates that the data inspection passes, obtain the judgment result that the data effect of the enhanced data meets the preset preprocessing effect; if the data inspection feedback indicates that the data inspection fails, obtain the judgment result that the data effect of the enhanced data does not meet the preset preprocessing effect. Among them, the characteristics include but are not limited to data source, data theme, etc. That is, in order to verify the preprocessing effect, in this embodiment, the enhanced data can be hierarchically processed according to the characteristics of the enhanced data, and the data in each obtained layer can be randomly sampled to ensure that the data in each layer is representative in the sampling, and the sampled data is displayed through a preset interface for manual inspection of the data, and then corresponding data inspection feedback is obtained through the preset interface, and it is judged whether the preset preprocessing effect is met according to the data inspection feedback. For example, if the data inspection feedback indicates that the data inspection passes, obtain the judgment result that the data effect of the enhanced data meets the preset preprocessing effect; if the data inspection feedback indicates that the data inspection fails, obtain the judgment result that the data effect of the enhanced data does not meet the preset preprocessing effect. In this way, before preprocessing the data to be labeled and trained, the FastText language detection library is deployed to the Spark in this embodiment to support the function of natural language type annotation and support horizontal expansion, and the cluster scale can be dynamically adjusted according to the size of the data volume to adapt to the data processing requirements of different scales and maintain high processing capabilities. At the same time, preprocessing operations including data cleaning, format unification, noise removal, text normalization, etc. are performed on the data to be labeled and trained, which can provide a high-quality, uniformly formatted, and minimized-noise data set for distributed language detection, provide a reliable basis for subsequent annotation, thus significantly improving the efficiency and accuracy of natural language type annotation, and through data partitioning and balanced partitioning strategies, ensure workload balance during the distributed processing process and avoid data skew problems.

[0061] Step S12: Read the processed data from the local based on the target application programming interface, and fragment the processed data to distribute the fragmented data to each node in the Spark cluster.

[0062] In this embodiment, the processed data is read locally using the target application programming interface, loaded into the distributed memory, and the processed data is divided into multiple shards for parallel processing, so as to distribute the sharded data to each node in the Spark cluster. Each of the shards contains a certain amount of text data. It should be noted that after distributing the sharded data to each node, it is also necessary to allocate the language detection task to each node through the task scheduler of Spark, so that each node can execute the detection task in parallel, thereby realizing the distributed processing of data. It can be understood that due to the in-memory computing characteristics of Spark, it allows data to be iterated quickly in memory, thereby reducing disk I / O (i.e., Input / Output, input / output) operations and further improving the speed of language detection. In this way, in this embodiment, the preprocessed processed data is sharded, which is convenient for parallel processing, and the data is distributed to each node of the spark cluster. Using the distributed computing ability of Spark enables a large amount of data to be processed in parallel, significantly improving the processing speed and reducing the time cost; at the same time, using the in-memory computing characteristics of Spark, reducing disk I / O operations, improving the computing performance and the utilization rate of computing resources in a distributed environment, reducing the processing cost, and further improving the speed of language detection.

[0063] Step S13: Parallelly perform language detection on the sharded data through the FastText model on each node to obtain corresponding preliminary annotation results, and use Spark to optimize the preliminary annotation results to obtain optimized target annotation results.

[0064] In this embodiment, the FastText model on each node is used to perform language detection on the text data assigned to itself in the sharded data in a parallel manner to obtain the preliminary annotation results output by the FastText model, and then Spark is used to optimize the preliminary annotation results to obtain optimized target annotation results. The preliminary annotation results include the output language type and its corresponding probability distribution.

[0065] It should be noted that the processing flow for obtaining the preliminary annotation results is as follows: The FastText models on the respective nodes execute their corresponding language detection tasks in parallel to perform language detection on the sharded data, obtaining the language detection results output by each node; based on Spark, the language detection results are integrated to obtain a language type annotation result set aggregated from the language detection results, and this language type annotation result set is used as the preliminary annotation result; among them, in the process of the FastText models on the respective nodes executing their corresponding language detection tasks in parallel, it includes: continuously monitoring the running status of the respective nodes, and when a target node with an abnormal situation is detected, processing the language detection task on the target node according to the abnormal situation. That is to say, the FastText models on the respective nodes execute the language detection tasks assigned to them in parallel to perform language detection on the sharded data, obtaining the language detection results output by each node, that is, the corresponding language types and their probabilities. Subsequently, Spark collects the language detection results on the respective nodes and integrates the language detection results to aggregate them into a complete language type annotation result set, and this language type annotation result is used as the preliminary annotation result. It can be understood that during the distributed detection process, that is, during the process of the respective nodes executing their corresponding language detection tasks, continuously monitor the running status of the respective nodes, and when a target node with an abnormal situation is detected, handle possible errors or abnormalities according to the abnormal situation to ensure the stability of the detection process.

[0066] It should be further pointed out that the processing flow for optimizing the preliminary annotation results using the Spark is as follows: extracting the feature of the text data in the preliminary annotation results using the Spark to obtain the corresponding data features; constructing the corresponding feature vectors according to the extracted data features; based on a number of preset machine learning algorithms in the Spark, using the feature vectors to train the preset initial models respectively to obtain a number of trained models corresponding to the number of preset machine learning algorithms; evaluating the number of trained models respectively through preset evaluation metrics to determine the annotation result optimization model with the best performance from the number of trained models according to the obtained evaluation results; applying the annotation result optimization model to the preliminary annotation results to optimize the preliminary annotation results. Among them, the preset evaluation metrics include but are not limited to accuracy, recall rate, F1 score (a metric for evaluating the performance of classification models), etc.; the preset machine learning algorithms include but are not limited to algorithms such as logistic regression and support vector machine. That is, importing the preliminary annotation results into the Spark environment, extracting the data features of the text data of the preliminary annotation results through the Spark, constructing the corresponding feature vectors using the data features, based on a number of preset machine learning algorithms in the Spark, using the feature vectors to train the preset initial models respectively to obtain the corresponding number of trained models, and evaluating the number of trained models using the preset evaluation metrics, selecting the annotation result optimization model with the best performance from the number of trained models according to the obtained evaluation results, and applying the annotation result optimization model to the preliminary annotation results to further optimize the annotation of the preliminary annotation results.

[0067] It should be noted that, in order to ensure the accuracy and reliability of natural language type annotation, after optimizing the preliminary annotation results, this embodiment can further correct and refine the obtained target annotation results, and the processing flow is as follows: Correct the target annotation results based on a preset correction rule to correct the incorrect text data in the target annotation results into correct text data; According to the analysis results of analyzing relevant cases of the incorrect text data, adjust the feature extraction project and the parameters of the annotation result optimization model, so as to obtain the target annotation results whose annotation accuracy meets the preset accuracy conditions. Among them, the correction rule is a series of rules formulated according to prior linguistic knowledge and specific business rules, including but not limited to the language attribution of specific words, the judgment of context, etc. That is, based on the prediction of the annotation result optimization model, correct the target annotation results according to the preset correction rule to correct the incorrect text data in the target annotation results into correct text data. For example, if the annotation result optimization model mislabels a text containing French words as English, the correction rule can correctly label it as French; At the same time, it is possible to deeply analyze the cases where the annotation result optimization model mislabels, that is, the relevant cases of the incorrect text data, find out the reasons for the errors, and adjust the feature extraction project and the parameters of the annotation result optimization model according to the obtained analysis results. Through multiple iterations and continuous adjustments, until the target annotation results whose annotation accuracy meets the preset accuracy conditions are obtained, thereby achieving satisfactory annotation accuracy. Among them, for some data with relatively fuzzy language types, this embodiment can further perform refined annotation, such as distinguishing dialects or regional language variants. In this way, this embodiment uses the distributed computing power of Spark combined with the FastText language detection library to perform natural language type annotation. The FastText language detection library can identify multiple languages and has high accuracy and generalization ability, realizing efficient and accurate natural language type annotation processing for large-scale data sets. At the same time, it is optimized on the basis of preliminary annotation, further improving the accuracy of language type annotation, and introducing correction rules and the correction and refinement of annotation results, realizing precise annotation of data with relatively fuzzy language types, meeting the requirements in different scenarios, effectively reducing incorrect annotations, and improving the reliability of annotation results.

[0068] As can be seen from the above, in the embodiment of the present application, the target application programming interface of Spark is first used to preprocess the training data to be labeled to obtain the processed data, and the processed data is saved locally. Subsequently, the processed data is read from the local through the target application programming interface, and the processed data is sharded to distribute the sharded data to each node in the Spark cluster. The FastText model on each node is used to perform language detection on the sharded data in parallel, and Spark is used to optimize the obtained preliminary annotation results to obtain the optimized target annotation results. In this way, through the above process of the embodiment of the present application, on the one hand, before preprocessing the training data to be labeled, the FastText language detection library is first deployed to Spark to support the function of natural language type annotation; on the other hand, horizontal expansion is supported, and the cluster scale can be dynamically adjusted according to the size of the data volume to adapt to different-scale data processing requirements and maintain high processing capabilities; on the one hand, preprocessing operations including data cleaning, format unification, noise removal, text normalization, etc. are performed on the training data to be labeled, which can provide a high-quality, uniformly formatted, and minimized-noise data set for distributed language detection and provide a reliable basis for subsequent annotation, thus significantly improving the efficiency and accuracy of natural language type annotation; on the one hand, through data partitioning and balanced partitioning strategies, workload balance in the distributed processing process is ensured, and data skew problems are avoided; on the other hand, the processed data after preprocessing is sharded, which is convenient for parallel processing, and the data is distributed to each node of the spark cluster. The distributed computing power of Spark enables parallel processing of a large amount of data, significantly improving the processing speed and reducing the time cost; on the one hand, the memory computing characteristics of Spark are utilized to reduce disk I / O operations, improve the computing performance and utilization rate of computing resources in a distributed environment, reduce the processing cost, and further improve the speed of language detection; on the one hand, the distributed computing power of Spark is combined with the FastText language detection library to perform natural language type annotation. The FastText language detection library can identify multiple languages and has high accuracy and generalization ability, realizing efficient and accurate natural language type annotation processing for large-scale data sets; on the one hand, optimization is performed on the basis of preliminary annotation, further improving the accuracy of language type annotation; on the other hand, correction rules and correction and refinement of annotation results are introduced to achieve precise annotation of data with relatively fuzzy language types, which can meet the requirements in different scenarios, effectively reduce misannotations, improve the reliability of annotation results, and thus solve the problems of low efficiency and insufficient accuracy in the natural language type annotation process when processing large-scale data.

[0069] Correspondingly, referring to Figure 3As shown in the figure, an embodiment of the present application further provides a natural language type annotation device for large model training data, including:

[0070] A data preprocessing module 11, configured to preprocess the training data to be annotated by using the target application programming interface of Spark, and save the obtained processed data locally;

[0071] A data sharding module 12, configured to read the processed data from local based on the target application programming interface, and shard the processed data, so as to distribute the sharded data to each node in the Spark cluster;

[0072] A result optimization module 13, configured to perform language detection on the sharded data in parallel through the FastText model on each node to obtain corresponding preliminary annotation results, and use Spark to optimize the preliminary annotation results to obtain optimized target annotation results.

[0073] As can be seen from the above, in the embodiment of the present application, the training data to be annotated is first preprocessed by using the target application programming interface of Spark to obtain processed data, and the processed data is saved locally. Subsequently, the processed data is read from local through the target application programming interface, and the processed data is sharded to distribute the sharded data to each node in the Spark cluster. The sharded data is subjected to language detection in parallel through the FastText model on each node, and Spark is used to optimize the obtained preliminary annotation results to obtain optimized target annotation results. In this way, through the above process of the embodiment of the present application, preprocessing the training data to be annotated by using the target application programming interface of Spark can significantly improve the data processing speed and reduce the time cost; adopting the form of combining FastText and Spark to perform language detection on the sharded data and optimize the preliminary annotation results improves the accuracy of natural language type annotation; through the data sharding strategy, it ensures workload balance during the distributed processing process and avoids data skew problems, thereby solving the problems of low efficiency and insufficient accuracy in the natural language type annotation process when dealing with large-scale data.

[0074] In some specific embodiments, the natural language type annotation device for large model training data may further include:

[0075] A size determination unit, configured to determine the scale of the Spark cluster according to the data volume of the training data to be annotated;

[0076] A node deployment unit, configured to deploy the FastText language detection library to each node of the Spark cluster after determining the scale.

[0077] The first parameter adjustment unit is used to train the FastText language detection library based on a pre-collected multilingual corpus and adjust the parameters of the FastText language detection library to obtain the FastText model.

[0078] In some specific embodiments, the data preprocessing module 11 may specifically include:

[0079] The data screening unit is used to scan the to-be-annotated training data by using the target application programming interface of Spark to clean the to-be-annotated training data, and screen the to-be-annotated training data based on a preset data quality rule and a preset function to obtain the screened data;

[0080] The normalization processing unit is used to convert the screened data into a preset encoding format and perform normalization processing on the screened data according to a preset normalization rule to obtain the normalized data;

[0081] The data enhancement unit is used to process the normalized data by using the Spark according to a preset stop word list to remove the stop words in the normalized data, and perform text enhancement processing on the normalized data based on a preset enhancement technique to obtain the enhanced data;

[0082] The condition judgment sub-module is used to judge whether the data effect of the enhanced data meets a preset preprocessing effect;

[0083] The data saving unit is used to, if the data effect meets the preset preprocessing effect, partition the enhanced data according to the language directory of the enhanced data to obtain the corresponding processed data, and save the processed data locally.

[0084] In some specific embodiments, the condition judgment sub-module may specifically include:

[0085] The data sampling unit is used to perform hierarchical processing on the enhanced data according to the characteristics of the enhanced data, and randomly sample the data in each obtained layer to display the sampled data through a preset interface, so as to obtain corresponding data inspection feedback through the preset interface;

[0086] The first result determination unit is used to, if the data inspection feedback indicates that the data inspection passes, obtain a judgment result that the data effect of the enhanced data meets the preset preprocessing effect;

[0087] A second result determination unit, configured to obtain a determination result that the data effect of the enhanced data does not meet a preset preprocessing effect if the data check feedback indicates that the data check fails.

[0088] In some specific embodiments, the result optimization module 13 may specifically include:

[0089] A task execution sub-module, configured to parallelly execute respective corresponding language detection tasks through the FastText models on the respective nodes, so as to perform language detection on the sharded data and obtain respective language detection results output by the respective nodes;

[0090] A result integration unit, configured to integrate the respective language detection results based on the Spark to obtain a language type annotation result set aggregated from the respective language detection results, and use the language type annotation result set as a preliminary annotation result;

[0091] Wherein, the task execution sub-module may specifically include:

[0092] A task processing unit, configured to continuously monitor the running states of the respective nodes, and when a target node with an abnormal situation is detected, process the language detection task on the target node according to the abnormal situation.

[0093] In some specific embodiments, the result optimization module 13 may specifically include:

[0094] A feature extraction unit, configured to extract features of text data in the preliminary annotation result by using the Spark to obtain corresponding data features;

[0095] A vector construction unit, configured to construct corresponding feature vectors according to the extracted data features;

[0096] A model training unit, configured to respectively train a preset initial model by using the feature vectors based on several preset machine learning algorithms in the Spark to obtain several trained models corresponding to the several preset machine learning algorithms;

[0097] A model evaluation unit, configured to evaluate the several trained models respectively through preset evaluation metrics, so as to determine an annotation result optimization model with the optimal performance from the several trained models according to the obtained evaluation results;

[0098] A model application unit, configured to apply the annotation result optimization model to the preliminary annotation result to optimize the preliminary annotation result.

[0099] In some specific embodiments, the natural language type annotation device for large model training data may further include:

[0100] A result correction unit for correcting the target annotation result based on a preset correction rule to correct the incorrect text data in the target annotation result into correct text data;

[0101] A second parameter adjustment unit for adjusting the parameters of the feature extraction project and the annotation result optimization model according to the analysis result of analyzing relevant cases of the incorrect text data, so as to obtain the target annotation result whose annotation accuracy meets the preset accuracy condition.

[0102] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the natural language type annotation method for large model training data disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0103] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0104] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0105] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the natural language type annotation method of the large model training data executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0106] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the natural language type annotation method of the large model training data disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0107] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0108] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0109] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0110] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0111] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A natural language type annotation method for large model training data, characterized in that, Including: Using the target application programming interface of Spark to preprocess the training data to be annotated, and saving the obtained processed data locally; Reading the processed data from local based on the target application programming interface, and sharding the processed data to distribute the obtained sharded data to each node in the Spark cluster; Parallelly performing language detection on the sharded data through the FastText model on each node to obtain corresponding preliminary annotation results, and using Spark to optimize the preliminary annotation results to obtain optimized target annotation results.

2. The method for natural language type annotation of large model training data according to claim 1, characterized in that Before using the target application programming interface of Spark to preprocess the training data to be annotated, it further includes: Determining the scale of the Spark cluster according to the data volume of the training data to be annotated; After determining the scale, deploying the FastText language detection library to each node of the Spark cluster; Training the FastText language detection library based on a pre-collected multilingual corpus and adjusting the parameters of the FastText language detection library to obtain the FastText model.

3. The method for natural language type annotation of large model training data according to claim 1, wherein, Using the target application programming interface of Spark to preprocess the training data to be annotated and saving the obtained processed data locally includes: Using the target application programming interface of Spark to scan the training data to be annotated, performing data cleaning on the training data to be annotated, and screening the training data to be annotated based on preset data quality rules and a preset function to obtain screened data; Converting the screened data into a preset encoding format and performing normalization processing on the screened data according to preset normalization rules to obtain normalized data; According to a preset stop word list, using Spark to process the normalized data to remove stop words in the normalized data, and performing text enhancement processing on the normalized data based on a preset enhancement technique to obtain enhanced data; Judging whether the data effect of the enhanced data meets a preset preprocessing effect; If the data effect meets the preset preprocessing effect, partitioning the enhanced data according to the language directory of the enhanced data to obtain corresponding processed data, and saving the processed data locally.

4. The natural language type annotation method for large model training data according to claim 3, wherein Judging whether the data effect of the enhanced data meets a preset preprocessing effect includes: Performing hierarchical processing on the enhanced data according to the characteristics of the enhanced data, randomly sampling the data in each obtained layer, and displaying the sampled data through a preset interface so as to obtain corresponding data inspection feedback through the preset interface; If the data inspection feedback indicates that the data inspection passes, obtaining a judgment result that the data effect of the enhanced data meets the preset preprocessing effect; If the data inspection feedback indicates that the data inspection fails, obtaining a judgment result that the data effect of the enhanced data does not meet the preset preprocessing effect.

5. The method for natural language type annotation of large model training data according to claim 1, wherein Performing language detection on the sharded data in parallel through the FastText models on each node to obtain corresponding preliminary annotation results, including: Performing respective corresponding language detection tasks in parallel through the FastText models on each node to perform language detection on the sharded data, and obtaining language detection results output by each node; Integrating the language detection results based on Spark to obtain a language type annotation result set aggregated from the language detection results, and using the language type annotation result set as the preliminary annotation result; Among them, in the process of performing respective corresponding language detection tasks in parallel through the FastText models on each node, it includes: Continuously monitoring the running status of each node, and when a target node with an abnormal situation is detected, processing the language detection task on the target node according to the abnormal situation.

6. The method for natural language type annotation of large model training data according to any one of claims 1 to 5, characterized in that, Optimizing the preliminary annotation result by using Spark, including: Extracting features from the text data in the preliminary annotation result by using Spark to obtain corresponding data features; Constructing corresponding feature vectors according to the extracted data features; Based on several preset machine learning algorithms in Spark, using the feature vectors to train a preset initial model respectively to obtain several trained models corresponding to the several preset machine learning algorithms; Evaluating the several trained models respectively through preset evaluation indicators, and determining the best-performing annotation result optimization model from the several trained models according to the obtained evaluation results; Applying the annotation result optimization model to the preliminary annotation result to optimize the preliminary annotation result.

7. The method for natural language type annotation of large model training data according to claim 6, wherein After optimizing the preliminary annotation result by using Spark to obtain an optimized target annotation result, it further includes: Correcting the target annotation result based on a preset correction rule to correct incorrect text data in the target annotation result into correct text data; Adjusting the feature extraction project and the parameters of the annotation result optimization model according to the analysis result of analyzing relevant cases of the incorrect text data, so as to obtain the target annotation result whose annotation accuracy meets the preset accuracy condition.

8. A natural language type annotation device for large model training data, characterized in that, Including: A data preprocessing module for preprocessing the training data to be annotated by using the target application programming interface of Spark and saving the obtained processed data locally; A data sharding module for reading the processed data from local based on the target application programming interface and sharding the processed data to distribute the obtained sharded data to each node in the Spark cluster; A result optimization module for performing language detection on the sharded data in parallel through the FastText models on each node to obtain corresponding preliminary annotation results, and optimizing the preliminary annotation results by using Spark to obtain optimized target annotation results.

9. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the natural language type annotation method for large model training data according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the natural language type annotation method for large model training data according to any one of claims 1 to 7 is implemented.