Method, system and equipment for automatically labeling data labels in real time and storage medium
Through the real-time automated data labeling method, the AI model is used to automatically generate data labels, which solves the lag, high cost and flexibility of manual labeling, and realizes efficient and real-time data labeling and management.
Patent Information
- Application Number
- CN202510132634.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
The existing manual annotation methods have lag, high cost, flexibility and poor practicality, making it difficult to quickly adapt to the dynamic changes of data.
It provides a real-time automated data labeling method, including defining data acquisition tasks and configuring data fields and task scheduling, judging training models based on outliers, and using intelligent annotation to perform model iterative optimization.
Automatically generate data labels through AI models, reducing the time and cost of manual labeling, improving the efficiency of data processing, ensuring the consistency and standardization of labels, and real-time update of labels, reducing operation and maintenance costs.
Smart Images

Figure CN120067856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data processing and artificial intelligence, and particularly to a method, system, device, and storage medium for real-time automated annotation of data labels. Background Art
[0002] In the era of big data, the application of data labels plays a crucial role in all walks of life. Data labels are not only an important cornerstone for data analysis and mining, but also the core means to improve data management and quality control. By effectively using data labels, enterprises can significantly improve data processing efficiency, optimize model performance, and thus realize a more intelligent and automated data management and analysis process. Data labels are particularly critical in the training of artificial intelligence models, especially in machine learning algorithms, where the quality of labeled data directly determines the accuracy and performance of the model. For this reason, many fields increasingly rely on the quality and generation efficiency of data labels, especially in the processing and annotation tasks of large-scale data sets. With the continuous expansion of data scale, manual annotation can no longer meet the requirements of efficient and accurate data processing.
[0003] Traditional manual annotation methods are time-consuming and laborious, and are prone to annotation errors. This method has high requirements for personnel and is also limited by the scale of the annotation task. Especially when facing continuously updated data, the lag and high cost of manual annotation become more obvious. In recent years, with the rapid development of artificial intelligence technology, automated annotation technology has gradually emerged and become an important trend in the field of data annotation. Existing automated annotation methods mainly rely on pre-trained models. Although these models can reduce some manual annotation work, they have significant limitations: First, pre-trained models require a large amount of labeled data sets as the training basis, which is difficult to implement for application scenarios with less initial data; Second, existing pre-trained models are usually static and difficult to quickly adapt to the dynamic changes of data, and cannot meet the needs of real-time data processing. In addition, once the annotation model has deviations or errors, the correction process is complex and difficult to quickly iterate, which limits the flexibility and practicality of the automated annotation system. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: the existing manual annotation method has lag, high cost, poor flexibility and practicality, and how it is difficult to quickly adapt to the dynamic changes of data.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: A real-time automated annotation data tagging method, including defining a data collection task and configuring data fields and task scheduling; training a model based on outlier judgment; and iteratively optimizing the model through intelligent annotation.
[0007] As a preferred solution of the real-time automated annotation data tagging method described in the present invention, wherein: the defining a data collection task and configuring data fields and task scheduling includes defining a data collection task, and data source information and data table information corresponding to each collection task;
[0008] Defining the primary key, data text, and classification label fields corresponding to the data table that needs to be automatically annotated, and the text field supports configuring the combination and sorting of more than one physical table field;
[0009] If the text field is configured as more than one field, then when fetching the data set, more than one text field will be concatenated in sequence;
[0010] For the data collection task, configure the scheduling period, and the collection periods of training sample data and automatically annotated data need to be configured separately.
[0011] As a preferred solution of the real-time automated annotation data tagging method described in the present invention, wherein: the training the model based on outlier judgment includes automatically collecting a data set containing name text and corresponding classification labels according to the basic configuration information;
[0012] Cleaning the data set for missing values, outliers, and duplicate data;
[0013] Dividing the cleaned data set into a training set and a test set;
[0014] The training set accounts for 80% of the total data set, and the test set accounts for 20% of the total data set;
[0015] Outliers are divided into two categories: text type and numerical type. Text type outliers are verified through a reference data set;
[0016] For numerical data, it is discriminated by outputting the Z-score of each data point, and the points exceeding the threshold are considered outliers;
[0017] Adopt a threshold of 2 as the judgment benchmark, and |Z-score|>2 is abnormal data.
[0018] As a preferred solution of the real-time automated annotation data tagging method described in the present invention, wherein: the training the model based on outlier judgment includes reducing the noise in the text by removing punctuation marks, special characters, and space special characters in the text, and reducing the complexity that the tokenizer needs to process;
[0019] The text is decomposed into words or phrases using the tokenizer of the NLTK library based on the Punkt algorithm. It learns sentence segmentation rules and the usage of punctuation marks through training data, without the need for predefined rules;
[0020] The Punkt algorithm is based on unsupervised learning, punctuation and abbreviation processing, delimiter rules, text preprocessing - word vector conversion;
[0021] Words are converted into word vectors through the Word2Vec model. The CBOW model is adopted to predict the central word through context words.
[0022] As a preferred solution of the real-time automatic annotation data tagging method described in the present invention, wherein: the training model based on outlier judgment includes using the Naive Bayes machine learning algorithm and preprocessed data to train the model, inputting text data, and outputting classification label data;
[0023] The model is used to predict the test set. By comparing the prediction results of the model with the true labels, evaluation metrics are output, including accuracy, precision, recall, and F1 score, to evaluate the performance of the model;
[0024] The cross-validation technique is used to adjust the hyperparameters of the model to improve performance and automatically find the best combination of hyperparameters.
[0025] As a preferred solution of the real-time automatic annotation data tagging method described in the present invention, wherein: the performance evaluation of the training model based on outlier judgment includes K-fold cross-validation for each hyperparameter combination θ, including:
[0026] The dataset D is divided into K subsets D1, D2,..., DK;
[0027] For each k ∈ {1, 2,..., K}, D-k = D\Dk is used as the training set to train the model fθ, and Dk is used as the validation set to calculate the validation error Ek;
[0028] All validation errors are averaged to obtain the average validation error of the hyperparameter combination θ:
[0029]
[0030] Select the hyperparameter combination with the minimum average validation error:
[0031]
[0032] where argmin represents the value range of θ when E θ is the minimum value.
[0033] As a preferred solution of the real-time automated annotation data tagging method described in the present invention, wherein: the model iteration optimization through intelligent annotation includes automatically collecting incremental data sets according to basic configuration information;
[0034] Data with empty data tags, including text data, primary keys, and data date field information.
[0035] As a preferred solution of the real-time automated annotation data tagging method described in the present invention, wherein: the AI model is called, the input is text data, and the output is the classification tags automatically recommended by the AI model. According to the primary key value, the classification tag field of the data table is updated, and the data status is marked as "00";
[0036] The data status data is uniformly displayed, and a special person reviews the data tags, makes manual adjustments, and updates the recorded data status to "01" after manual confirmation;
[0037] According to the basic configuration program, data with a status of "01" and a data date within the last month is regularly obtained from the data table, and an incremental training data set is constructed through the data cleaning and preprocessing processes;
[0038] Through online incremental training of the model, the parameters of the model are gradually updated;
[0039] Through online incremental training of the model, the AI model is continuously iteratively optimized.
[0040] Another object of the present invention is to provide a real-time automated annotation data tagging system, which aims to utilize artificial intelligence and machine learning algorithms to achieve the automated generation and real-time update of data tags through the real-time automated annotation data tagging technology based on AI, and solves the problem of high cost in the current manual annotation method.
[0041] As a preferred solution of the real-time automated annotation data tagging system described in the present invention, it includes a basic configuration module (100), a model online training module (200), and an intelligent annotation module (300); the basic configuration module (100) is used to configure the scheduling period of data collection tasks, data table information, and data columns; the model online training module (200) is used to preprocess the collected data, including data cleaning and format conversion steps, to form training sample data, and uses the Naive Bayes algorithm to train the model based on the collected labeled data to generate an AI model for intelligent recommendation of data tags. The model training process is based on an online machine learning platform and supports dynamic iteration and update of the model; the intelligent annotation module (300) is used to collect newly generated business data in real time, call the trained AI model, generate and update data tags, manually verify the automatic annotation results, re-annotate the data with unreasonable automatic annotations, and use the data for further training and optimization of the model to continuously improve the accuracy of the model.
[0042] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the real-time automated annotation data tagging method are implemented.
[0043] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the real-time automated annotation data tagging method are implemented.
[0044] Advantages of the present invention: The real-time automated annotation data tagging method provided by the present invention automatically generates data tags through an AI model, reducing the time and cost of manual annotation and improving the efficiency of data processing. Intelligent data tags can ensure the consistency and standardization of tags, avoiding subjective biases that may occur during the manual annotation process. Intelligent tags can classify and annotate data more accurately, improving the accuracy and usability of data. The tags can be updated in real time according to changes in the data to ensure that the data tags always reflect the latest status and content. Through intelligent tag generation and management, manual intervention can be greatly reduced, the operation and maintenance costs can be reduced, and data tags can be better managed and monitored, reducing system failures and risks caused by tag errors. The present invention achieves better results in terms of data management efficiency, data quality and accuracy, and operation and maintenance costs. Description of the Drawings
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 This is the overall flowchart of a method for real-time automated annotation data tags provided by the first embodiment of the present invention.
[0047] Figure 2 This is the data automated tag scheme diagram of a method for real-time automated annotation data tags provided by the first embodiment of the present invention.
[0048] Figure 3 This is the training specimen diagram of a method for real-time automated annotation data tags provided by the second embodiment of the present invention.
[0049] Figure 4 This is the training sample diagram of the material label model of a method for real-time automated annotation data tags provided by the second embodiment of the present invention. Detailed implementation manners
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Embodiment 1, referring to Figure 1 - Figure 2 , an embodiment of the present invention provides a method for real-time automated annotation data tags, including:
[0052] S1: Define a data collection task and configure data fields and task scheduling.
[0053] Furthermore, data collection configuration: Define a data collection task, as well as the data source information and data table information corresponding to each collection task. The data relationship configuration is shown in Table 1:
[0054] Table 1 Data collection configuration table
[0055]
[0056] Data field configuration: Define fields such as the primary key, data text, and classification label corresponding to the data table that needs to be automatically annotated. Among them, the text field supports configuring the combination and sorting of multiple physical table fields. (Note: If the text field is configured as multiple fields, when fetching the data set, multiple text fields will be concatenated in order), and the data relationship configuration is shown in Table 2:
[0057] Table 2 Data field configuration table
[0058]
[0059] Task scheduling configuration: For data collection tasks, configure the scheduling period, and the collection periods for training sample data and automatically labeled data need to be configured separately. The data relationship configuration is shown in Table 3:
[0060] Table 3 Task Scheduling Configuration Table
[0061] Scheduling encoding Collection task Sample collection period Annotation period DP20240801001 TASK2024080101 00?*MON 0001*? DP20240801001 TASK2024080102 00?*MON 0001*?
[0062] S2: Judge the training model based on outliers.
[0063] Furthermore, according to the basic configuration information, automatically collect a dataset containing name texts and their corresponding classification labels. And clean the dataset for missing values, outliers, and duplicate data. Divide the cleaned dataset into a training set (80%) and a test set (20%).
[0064] Outlier judgment: Outliers are divided into two categories: text type and numerical type. Usually, text type outliers are verified through a reference dataset. For example, for the "gender" field, the reference dataset is (male (encoded as M); female (encoded as F)). If there is a data record in the record that is not "M" or "F", then this data is abnormal data; for numerical data, it is judged by the Z-score method (calculate the Z-score of each data point, and points exceeding a certain threshold are considered outliers).
[0065] It should be noted that the Z-score is expressed as:
[0066]
[0067] Among them, Z represents the Z-score, X represents the value of the data point, μ represents the average value of the dataset, and σ represents the standard deviation of the dataset;
[0068] Adopt the threshold = 2 as the judgment benchmark, and |Z-score| > 2 is abnormal data.
[0069] It should also be noted that by removing special characters such as punctuation marks, special characters, and spaces in the text, the noise in the text is reduced, thereby reducing the complexity that the tokenizer needs to process. Use the tokenizer of the NLTK library based on the Punkt algorithm to decompose the text into words or phrases. The Punkt tokenizer uses unsupervised learning technology to learn sentence segmentation rules and the usage of punctuation marks through training data, without the need for predefined rules. Train the Punkt model by sorting out text data in the industrial field containing professional terms, professional descriptions, etc., to improve the accuracy of the Punkt model for tokenizing industrial field descriptions. The training specimens are as follows Figure 3 shown.
[0070] Furthermore, the Punkt algorithm is based on the following key mathematical principles:
[0071] Unsupervised learning: The Punkt algorithm uses unsupervised learning techniques to automatically infer sentence segmentation rules from training texts. It learns patterns of punctuation and sentence boundaries through statistical analysis.
[0072] Punctuation and abbreviation handling: The Punkt algorithm can recognize and handle punctuation, abbreviations, and other special cases that may affect segmentation. It uses statistical models to distinguish true sentence boundaries from non-boundary markers (such as dots in abbreviations).
[0073] Separator rules: The Punkt algorithm determines sentence boundaries by analyzing punctuation in the training data. It learns which punctuation marks (such as full stops, question marks) usually indicate the end of a sentence and differentiates these markers from other cases (such as abbreviations).
[0074] It should be noted that for text preprocessing - word vector conversion: Words are converted into word vectors through the Word2Vec model. Word2Vec learns the distributed representation of words by training a language model, captures the semantic relationships between words, and adopts the CBOW model to predict the target word through context words.
[0075] The algorithm principle of the CBOW model is as follows:
[0076] 1. Calculation from the input layer to the hidden layer:
[0077]
[0078] Among them, is the one-hot encoding of the context words, which is multiplied by the input layer weight matrix W to obtain the word vector, and m represents the context window size;
[0079] 2. Calculation from the hidden layer to the output layer:
[0080] u = W'h
[0081] Among them, W' is the weight matrix of the output layer, and u is the unnormalized probability distribution of the output layer;
[0082] 3. Softmax function of the output layer:
[0083]
[0084] Among them, w t represents the target word, V represents the vocabulary, w t-m , …, w t-1 , w t+1 , …, w t+mThe context words representing the central word, m represents the context window size, P(w t ∣w t-m ,…,w t-1 ,w t+1 ,…,w t+m ) represents maximizing the probability of the central word.
[0085] It should also be noted that the CBOW model optimization method is as follows:
[0086] The CBOW model is optimized by using the Negative Sampling method to improve the training efficiency and prediction effect. Negative Sampling is a technique to reduce the computational complexity, used to replace the traditional Softmax layer, thus accelerating the model training. The specific implementation method is as follows:
[0087] 1. In the traditional Softmax, the loss function is:
[0088]
[0089] where, is the prediction score of the central word w t
[0090] 2. Negative Sampling reduces the computational cost by introducing positive and negative samples and only optimizing the probabilities of a part of the words.
[0091]
[0092] where, is the prediction score of the central word w t , P n (w) is the negative sample distribution sampled from the noise distribution, k is the number of negative samples, is the sigmoid function.
[0093] S3: The model is iteratively optimized through intelligent annotation.
[0094] Furthermore, the Naive Bayes machine learning algorithm and the preprocessed data are used to train the model (input: text data, output: classification label data). The Naive Bayes algorithm is a simple but effective text classification method. The Naive Bayes classifier is based on Bayes' theorem and the feature conditional independence assumption. It assumes that the features are independent given a certain class. Although this assumption may not hold completely in reality, in many cases, the Naive Bayes classifier still performs well. Bayes' theorem describes the posterior probability (the probability of an event occurring given some information), and its formula is:
[0095]
[0096] Among them, P(C∣X) represents the posterior probability of class C given feature X, P(X∣C) represents the likelihood probability of feature X under class C, P(C) represents the prior probability of class C, and P(X) represents the marginal probability of feature X;
[0097] In the Naive Bayes classifier, the conditional independence assumption of features makes:
[0098]
[0099] where x i is the i-th feature in feature X;
[0100] Use the trained model to predict the test set. By comparing the prediction results of the model with the true labels, calculate evaluation metrics (such as accuracy, precision, recall, F1-score, etc.) to evaluate the performance of the model;
[0101] Accuracy is the proportion of samples correctly predicted by the model in the total samples.
[0102] Accuracy is expressed as:
[0103]
[0104] where Accuracy represents accuracy, TP represents the number of samples with the true label being the positive class and the model predicting as the positive class, TN represents the number of samples with the true label being the negative class and the model predicting as the negative class, FP represents the number of samples with the true label being the negative class and the model predicting as the positive class, and FN represents the number of samples with the true label being the positive class and the model predicting as the negative class;
[0105] Precision is the proportion of samples actually being the positive class among the samples predicted as the positive class by the model.
[0106] Precision is expressed as:
[0107]
[0108] where Precision represents precision;
[0109] Recall is the proportion of samples actually being the positive class that are correctly predicted as the positive class by the model.
[0110] Recall is expressed as:
[0111]
[0112] where Recall represents recall
[0113] F1-score is the harmonic mean of precision and recall, used to balance the influence of precision and recall.
[0114] The F1 score is expressed as:
[0115]
[0116] where F1Score represents the F1 score;
[0117] The hyperparameters of the model are adjusted using cross - validation techniques to improve performance, and the best combination of hyperparameters is automatically found.
[0118] Hyperparameter tuning: The hyperparameters of the model are adjusted using cross - validation techniques to improve performance, and the best combination of hyperparameters is automatically found.
[0119] Performance evaluation of K - fold cross - validation
[0120] For each combination of hyperparameters θ:
[0121] The dataset D is divided into K subsets D1, D2, …, DK.
[0122] For each k ∈ {1, 2, …, K}:
[0123] Use D - k = D\Dk as the training set to train the model fθ.
[0124] Use Dk as the validation set to calculate the validation error Ek.
[0125] Average all the validation errors to obtain the average validation error of the hyperparameter combination θ:
[0126]
[0127] Select the combination of hyperparameters with the minimum average validation error:
[0128]
[0129] where argmin represents the value of θ when E θ is the minimum value.
[0130] It should be noted that intelligent annotation includes:
[0131] Incremental data collection: According to the basic configuration information, incrementally collect the dataset (data with empty data labels, including field information such as text data, primary keys, and data dates).
[0132] Automatic annotation: Call the trained AI model, with the input being text data and the output being the classification labels automatically recommended by the AI model. Update the classification label field of the data table according to the primary key value, and mark the data status as "00" (AI recommended status value).
[0133] Manual review: Uniformly display data with a data status of "00", and have a dedicated person review the data labels. For unreasonable label recommendations, manual adjustment is made. After manual confirmation, the recorded data status is updated to "01" (the completion status value of manual review). When the model accuracy reaches a certain level, the manual review link can be considered omitted.
[0134] Model iteration and optimization include:
[0135] According to the basic configuration program, regularly obtain data with a status of "01" and a data date within the last month from the data table. Through processes such as data cleaning and preprocessing, an incremental training data set is constructed. Through online incremental training of the model, the model parameters are gradually updated to enable it to adapt to new data in real time. Through online incremental training of the model, the AI model is continuously iteratively optimized to make the model recommendations more accurate.
[0136] Embodiment 2, an embodiment of the present invention, provides a real-time automated annotation data label system, including:
[0137] A basic configuration module (100), a model online training module (200), and an intelligent annotation module (300).
[0138] The basic configuration module (100) is used to configure the scheduling cycle of the data collection task, data table information, and data columns;
[0139] The model online training module (200) is used to preprocess the collected data, including data cleaning and format conversion steps, to form training sample data. Using the Naive Bayes algorithm, model training is performed based on the collected labeled data to generate an AI model for intelligent recommendation of data labels. The model training process is based on an online machine learning platform, supporting dynamic iteration and update of the model;
[0140] The intelligent annotation module (300) is used to collect newly generated business data in real time, call the trained AI model, generate and update data labels, manually verify the automatic annotation results, re-annotate data with unreasonable automatic annotations, and use the data for further training and optimization of the model to continuously improve the model accuracy.
[0141] Embodiment 3, an embodiment of the present invention, is different from the previous two embodiments in that:
[0142] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.
[0143] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0144] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0145] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
[0146] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
[0147] Example 4, referring to Figure 3 - Figure 4 , which is an embodiment of the present invention, provides a method for real-time automated annotation of data tags. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0148] First, configure the data source, data fields, collection period, etc. of the material information table, and the configured data is shown in Tables 4, 5, and 6 below.
[0149] Table 4 Data Source Configuration Table
[0150]
[0151] Table 5 Data Table Information Configuration
[0152]
[0153] Table 6 Collection Task Configuration
[0154]
[0155] After the configuration is completed, according to the scheduled task, the data of the material information table is collected regularly. After data cleaning, 80% of it is used as training data and 20% is used as test data. The training specimen data is as Figure 4 shown.
[0156] Text preprocessing: Through Word2Vec for text preprocessing, the original text data can be converted into word vectors with semantic information.
[0157] Use the Naive Bayes machine learning algorithm to train with the preprocessed data to generate a material label model (input: material description, output: material label).
[0158] Collect the data in the material information table with empty labels, call the material label model to generate recommended labels, and after manual verification, the recommended labels are processed for storage.
[0159] In summary, the present invention achieves better effects in terms of data management efficiency, data quality and accuracy, and operation and maintenance costs.
Claims
1. A real-time automatic data labeling method, characterized in that: include: Define data collection tasks and configure data fields and task scheduling; Training model based on outlier judgment; Iteratively optimize the model through intelligent annotation.
2. The real-time automatic data labeling method according to claim 1, characterized in that: Defining data collection tasks and configuring data fields and task scheduling includes defining data collection tasks, as well as data source information and data table information corresponding to each collection task; Define the primary key, data text, and classification label fields corresponding to the data table that needs to be automatically labeled. The text field supports the combination and sorting of more than one physical table field. If the text field is configured as more than one field, more than one text field will be concatenated in order when taking the data set; To configure the scheduling cycle for data collection tasks, you need to configure the collection cycle for training sample data and automatic annotation data separately.
3. The real-time automatic data labeling method according to claim 2, characterized in that: The training model based on outlier judgment includes automatically collecting a data set including name text and corresponding classification labels according to basic configuration information; Clean the data set for missing values, outliers, and duplicate data; Divide the cleaned data set into training set and test set; The training set accounts for 80% of the total data set, and the test set accounts for 20% of the total data set; Outliers are divided into two categories: text and numerical. Text outliers are verified using a reference dataset. For numerical data, the Z-score of each data point is output, and the points exceeding the threshold are considered as outliers for judgment; The threshold value = 2 is used as the judgment standard, and |Z-score|>2 is considered abnormal data.
4. The real-time automatic data labeling method according to claim 3, characterized in that: The training model includes reducing the noise in the text by removing punctuation marks, special characters, and space special characters in the text, thereby reducing the complexity that the word segmenter needs to process; The NLTK library uses the Punkt algorithm-based tokenizer to decompose the text into words or phrases, and learns sentence segmentation rules and punctuation usage through training data, without the need for predefined rules; The Punkt algorithm is based on unsupervised learning, punctuation and abbreviation processing, separator rules, text preprocessing-word embedding conversion; The words are converted into word vectors through the Word2Vec model, and the CBOW model is used to predict the central word through the context words.
5. The real-time automatic data labeling method according to claim 4, characterized in that: The training model includes using a naive Bayesian machine learning algorithm and a preprocessed data training model, inputting text data, and outputting classification label data; The model is used to predict the test set, and the evaluation indicators, including accuracy, precision, recall, and F1 score, are output by comparing the model's prediction results with the true labels to evaluate the model performance. Use cross-validation techniques to adjust model hyperparameters to improve performance and automatically find the best hyperparameter combination.
6. The real-time automatic data labeling method according to claim 5, characterized in that: The outlier judgment-based training model includes a K-fold cross-validation performance evaluation for each hyperparameter combination θ, including: Divide the data set D into K subsets D1, D2, …, DK; For each k∈{1, 2, …, K}, use Dk=D\Dk as the training set, train the model fθ, use Dk as the validation set, and calculate the validation error Ek; Average all validation errors to get the average validation error for the hyperparameter combination θ: Select the hyperparameter combination that minimizes the average validation error: Among them, argmin refers to E θ The value range of θ when is the minimum value.
7. The real-time automatic data labeling method according to claim 6, characterized in that: The iterative optimization of the model through intelligent annotation includes automatically collecting incremental data sets according to basic configuration information; Data with empty data labels, including text data, primary keys, and data date field information.
8. The real-time automatic data labeling method according to claim 7, characterized in that: The iterative model optimization includes calling the AI model, inputting text data, outputting classification labels automatically recommended by the AI model, updating the classification label field of the data table according to the primary key value, and marking the data status as "00"; The data status data is displayed uniformly, and a dedicated person reviews the data label and makes manual adjustments. After manual confirmation, the record data status is updated to 01; According to the basic configuration program, data with status 01 and data date of the past month are regularly obtained from the data table, and an incremental training data set is constructed through data cleaning and preprocessing. Through online incremental training of the model, the model parameters are gradually updated; Through online incremental model training, the AI model is continuously iterated and optimized.
9. A system for real-time automatic data labeling method, characterized in that: include, A basic configuration module (100), used to configure the scheduling period, data table information, and data columns of the data collection task; The model online training module (200) is used to pre-process the collected data, including data cleaning and format conversion steps, to form training sample data, and to use the naive Bayes algorithm to perform model training based on the collected labeled data to generate an AI model for intelligent recommendation of data labels. The model training process is based on an online machine learning platform to support dynamic iteration and updating of the model. The intelligent labeling module (300) is used to collect newly generated business data in real time, call the trained AI model, generate and update data labels, manually verify the automatic labeling results, re-label the data that is not reasonably labeled automatically, and use the data for further training and optimization of the model, so as to continuously improve the accuracy of the model.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the real-time automatic data labeling method are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the real-time automatic data labeling method are implemented.
Citation Information
Patent Citations
Automatic text labeling method and device and terminal
CN113312476A
Government affair data processing method based on artificial intelligence, electronic equipment and storage medium
CN115098671A
Data labeling method based on rule, model and manual combination
CN117493903A
Method and system for intelligently realizing data cleaning by using AI model
CN118568423A
Medical data cleaning method, system and equipment based on machine learning optimization and medium
CN119003999A
Cited By
Intelligent equipment visual data management and AI model development platform and method
CN120356038A