Data marking method and device and machine readable storage medium

By using text clustering, entity recognition and entity clustering methods in the data labeling process in the field of machine learning, multi-angle labeling data is generated, which solves the problem of time-consuming and costly acquisition of high-quality labeling data, and improves the accuracy and efficiency of data labeling.

CN120217023APending Publication Date: 2025-06-27HUNAN ZOOMLION CONCRETE MASCH STATION EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202544.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the field of machine learning, especially in supervised learning tasks, obtaining high-quality labeled data is a time-consuming and costly task, especially in the field of engineering machinery. Due to the complex working environment and a wide variety of equipment, engineers with deep expertise need to accurately label, resulting in high labor costs and inconsistencies and errors in labeled data.

Method used

By clustering text vectors of labeled data based on the preset number of clusters, combining entity recognition and entity clustering, multi-angle labeled data is generated to improve the accuracy and efficiency of data labeling.

Benefits of technology

By combining the advantages of text clustering, entity recognition and entity clustering, the labeled data can be comprehensively and in-depth tagged from multiple angles, improving the accuracy and efficiency of data labeling, reducing labor costs, and reducing inconsistency and error of labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217023A_ABST
    Figure CN120217023A_ABST
Patent Text Reader

Abstract

The invention discloses a data marking method and device and a machine readable storage medium, and relates to the technical field of computers. The method comprises the steps of performing text clustering on a text vector corresponding to data to be marked based on a preset clustering number to obtain a text clustering result; when the text clustering result meets a preset clustering condition, performing text marking on the to-be-marked data based on the text clustering result to obtain a first mark; inputting the to-be-marked data into a preset entity recognition model to obtain an entity; performing entity marking on the to-be-marked data based on the entity to obtain a second mark; combining the entities, and determining a corresponding entity combination vector; performing entity clustering on the entity combination vector, and performing entity combination marking on the to-be-marked data based on an entity clustering result to obtain a third mark; and based on the first mark, the second mark and the third mark, generating a mark of the to-be-marked data. The advantages of text clustering, entity recognition and entity clustering are combined, and the accuracy and efficiency of data marking are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a data marking method, apparatus, and machine-readable storage medium. Background Art

[0002] In the field of machine learning, especially in supervised learning tasks, the accuracy and richness of training data are crucial for the performance of the model. Training data usually consists of input features and corresponding labels, which are used to guide the model to learn how to map the input features to the correct output. For many application scenarios, such as image recognition, speech recognition, and natural language processing, obtaining high-quality labeled data is a time-consuming and costly task. In the field of construction machinery, this challenge is particularly prominent. The working environment of construction machinery is complex and variable, involving a wide variety of equipment with different operating states. Therefore, accurately marking the images, videos, or sensor data of construction machinery for training machine learning models requires engineers with profound professional knowledge and rich experience. This not only increases the labor cost but also may lead to inconsistencies and errors in the labeled data due to human factors. Summary of the Invention

[0003] Aiming at the above deficiencies in the prior art, the purpose of the embodiments of this application is to provide a data marking method, apparatus, and machine-readable storage medium.

[0004] To achieve the above purpose, the first aspect of this application provides a data marking method, including:

[0005] Performing text clustering on the text vectors corresponding to the data to be marked based on a preset number of clusters to obtain a text clustering result;

[0006] When the text clustering result meets the preset clustering condition, performing text marking on the data to be marked based on the text clustering result to obtain a first mark;

[0007] Inputting the data to be marked into a preset entity recognition model to obtain an entity;

[0008] Performing entity marking on the data to be marked based on the entity to obtain a second mark;

[0009] Combining the entities and determining the corresponding entity combination vector;

[0010] Performing entity clustering on the entity combination vector and performing entity combination marking on the data to be marked based on the entity clustering result to obtain a third mark;

[0011] Generating a mark for the data to be marked based on the first mark, the second mark, and the third mark.

[0012] In an embodiment of the present application, entity clustering is performed on the entity combination vectors, and entity combination labeling is performed on the data to be labeled based on the entity clustering result to obtain a third label, including:

[0013] Determine the target number of clusters corresponding to the text clustering result that meets the preset clustering conditions;

[0014] Perform entity clustering on the entity combination vectors based on the target number of clusters, and perform entity combination labeling on the data to be labeled based on the entity clustering result of the entity clustering to obtain a third label.

[0015] In an embodiment of the present application, entity clustering is performed on the entity combination vectors based on the target number of clusters, and entity combination labeling is performed on the data to be labeled based on the entity clustering result of the entity clustering to obtain a third label, including:

[0016] Perform entity clustering on the entity combination vectors based on the target number of clusters, and for each entity combination category in the entity clustering result, obtain the entity clustering center of the entity combination category;

[0017] Determine the first target data to be labeled among all the data to be labeled included in the entity combination category, where the first target data to be labeled is the data to be labeled closest to the entity clustering center among all the data to be labeled included in the entity combination category;

[0018] Perform entity combination labeling on the first target data to be labeled to obtain a third label.

[0019] In an embodiment of the present application, text clustering is performed on the text vectors corresponding to the data to be labeled based on the preset number of clusters to obtain a text clustering result, including:

[0020] Perform text clustering on the text vectors corresponding to the data to be labeled based on the preset number of clusters to obtain an initial text clustering result;

[0021] Determine whether there is a target category in the initial text clustering result, where the target category is a category in all the categories of the initial text clustering result that includes a number of data to be labeled less than the preset sample number;

[0022] In the case where there is a target category, update the preset number of clusters based on the preset number of clusters and the number of the target category;

[0023] Perform text clustering on the text vectors corresponding to the data to be labeled again based on the updated preset number of clusters until there is no target category in the initial text clustering result, and use the initial text clustering result as the text clustering result.

[0024] In an embodiment of the present application, in the case where there is a target category, updating the preset number of clusters based on the preset number of clusters and the number of the target category includes:

[0025] In the case where there is a target category, determine the difference between the preset number of clusters and the number of target categories;

[0026] Update the preset number of clusters based on the difference.

[0027] In the embodiments of the present application, in the case where the text clustering result meets the preset clustering condition, text marking is performed on the data to be marked based on the text clustering result to obtain a first mark, including:

[0028] In the case where the text clustering result meets the preset clustering condition, for each text category in the text clustering result, obtain the text clustering center of the text category;

[0029] Determine the second target data to be marked among all the data to be marked included in the text category, where the second target data to be marked is the data to be marked closest to the text clustering center among all the data to be marked included in the text category;

[0030] Perform text marking on the second target data to be marked to obtain a first mark.

[0031] In the embodiments of the present application, based on the preset number of clusters, text clustering is performed on the text vectors corresponding to the data to be marked to obtain a text clustering result, including:

[0032] Convert the data to be marked into text vectors based on a preset text embedding model;

[0033] Perform text clustering on the text vectors based on the preset number of clusters to obtain a text clustering result.

[0034] In the embodiments of the present application, before the step of performing text clustering on the text vectors corresponding to the data to be marked based on the preset number of clusters to obtain a text clustering result, it further includes:

[0035] Obtain the data to be marked;

[0036] Perform preprocessing on the data to be marked;

[0037] Convert the preprocessed data to be marked into a preset text format.

[0038] A second aspect of the present application provides a computing device, including:

[0039] A memory configured to store instructions;

[0040] A processor configured to call instructions from the memory and capable of implementing the data marking method as described in the above embodiments when executing the instructions.

[0041] A third aspect of the present application provides a machine-readable storage medium, on which instructions are stored for causing a machine to execute the data marking method as described in the above embodiments.

[0042] Through the above technical solution, text clustering is performed on the text vectors corresponding to the data to be marked based on a preset number of clusters to obtain a text clustering result; when the text clustering result meets the preset clustering condition, text marking is performed on the data to be marked based on the text clustering result to obtain a first mark; the data to be marked is input into a preset entity recognition model to obtain an entity; entity marking is performed on the data to be marked based on the entity to obtain a second mark; the entities are combined, and the corresponding entity combination vector is determined; entity clustering is performed on the entity combination vector, and entity combination marking is performed on the data to be marked based on the entity clustering result to obtain a third mark; based on the first mark, the second mark, and the third mark, a mark for the data to be marked is generated. By combining the advantages of text clustering, entity recognition, and entity clustering, it is possible to comprehensively and deeply mark the data to be marked from multiple perspectives, improving the accuracy and efficiency of data marking.

[0043] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. Description of the Drawings

[0044] The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. They are used together with the following specific implementation to explain the embodiments of the present application, but do not limit the embodiments of the present application. In the drawings:

[0045] Figure 1 Schematically shows a flowchart of a data marking method according to an embodiment of the present application;

[0046] Figure 2 Schematically shows a flowchart of a data marking method according to another embodiment of the present application. Detailed Description of the Invention

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiments of the present application and is not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0048] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain industry-existing solutions such as software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0049] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of this application, then such directional indications are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the attached drawings). If this specific posture changes, then the directional indications will also change accordingly.

[0050] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0051] Figure 1 A schematic flowchart of a data marking method according to an embodiment of this application is schematically shown. As Figure 1 shown, the embodiments of this application provide a data marking method, and this method may include the following steps:

[0052] Step 100, perform text clustering on the text vectors corresponding to the data to be marked based on a preset number of clusters to obtain a text clustering result;

[0053] In this embodiment, it should be noted that the data to be marked refers to data that has not been given label or classification information. These data need to be processed and marked for use in the training, testing, or evaluation of machine learning models. The data to be marked can come from various fields, such as text, images, audio, video, etc., specifically depending on the application scenario and requirements. In the field of machine learning, the importance of the data to be marked is self-evident. The data to be marked is the basis for training the model. Without a sufficient quantity and quality of the data to be marked, it is impossible to train a high-performance model. Therefore, effectively processing and marking the data to be marked is one of the key steps in a machine learning project.

[0054] Refer to Figure 2, specifically, in one embodiment, before the step of performing text clustering on the text vectors corresponding to the data to be labeled based on a preset number of clusters to obtain a text clustering result, the following steps are further included:

[0055] Obtain the data to be labeled;

[0056] Preprocess the data to be labeled;

[0057] Convert the preprocessed data to be labeled into a preset text format.

[0058] It should be noted that the sources of the data to be labeled may include data extracted from a database, data crawled from the Internet, or data collected by other means. According to the data source, use the corresponding tools or methods to collect the data. Ensure that the collected data to be labeled meets the requirements of the actual application and is sufficient in quantity for subsequent preprocessing and labeling work. In one embodiment, the collected data to be labeled will also be preliminarily verified to ensure the integrity and accuracy of the data to be labeled. Check for missing values, outliers, or duplicate data and perform necessary processing on them.

[0059] The preprocessing may include data cleaning, data transformation, feature extraction, etc. Among them, data cleaning may include removing noise, outliers, and duplicate data from the data. For text data, operations such as spelling checking, removing stop words, and punctuation marks may also be required. Data transformation includes transforming the data as needed to improve the quality and adaptability of the data. For example, for numerical data, normalization or standardization processing may be required; for text data, stemming or lemmatization operations may be required. Feature extraction includes extracting useful feature information for model training from the original data. This may require using specific algorithms or tools to extract features and convert them into a format suitable for model training. The preset text format is a specific file format, such as CSV (Comma Separated Values), TXT (plain text file), JSON (JavaScript Object Notation), etc., as well as the arrangement of data in the file and field names. According to the preset text format, convert the preprocessed data into the corresponding format. Verify the converted data to ensure the integrity and accuracy of the data. By obtaining the data to be labeled, preprocessing it, and converting it into a preset text format, preparations are made for subsequent data labeling and model training work.

[0060] It should be noted that the text vectors of the data to be labeled are grouped based on a preset number of clusters to form a text clustering result. Clustering algorithms such as K-means (K-means clustering), Hierarchical Clustering, and DBSCAN (Density-Based Spatial Clustering of Applications with Noise) can be used, and the specific selection depends on the characteristics and requirements of the data. The text clustering result is output, that is, the data is divided into several different clusters.

[0061] Reference Figure 2 , in one embodiment, text clustering is performed on the text vectors corresponding to the data to be labeled based on a preset number of clusters, and the text clustering result is obtained, including:

[0062] Convert the data to be labeled into text vectors based on a preset text embedding model;

[0063] Perform text clustering on the text vectors based on a preset number of clusters to obtain a text clustering result.

[0064] It should be noted that the preset text embedding model can include m3e (Moka Massive Mixed Embedding), m3e-large, Word2Vec (Word to Vector), GloVe (Global Vectors for Word Representation), BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc. Each text in the data to be labeled is input into the text embedding model to obtain the corresponding text vector. These vectors are usually represented as points in a high-dimensional space and can capture the semantic relationships between texts. For example, after passing the text "There is a problem with the aggregate bin" through the preset text embedding model, the resulting three-dimensional vector can be represented as: [[0.001, 0.05, 0.3], [0.1, 0.02, 0.85], [0.33, 0.71, 0.92]]. The preset number of clusters refers to the number of categories into which the data to be labeled is expected to be divided. In this embodiment, in order to discover new categories, the preset number of clusters can be set to a relatively large value, such as 100. This allows the algorithm to search for more potential structures or patterns in the data, even if these structures or patterns are not obvious initially. Through subsequent analysis and possible convergence processes, more appropriate and meaningful clustering results can be obtained.

[0065] Step 200, when the text clustering result meets the preset clustering condition, perform text labeling on the data to be labeled based on the text clustering result to obtain the first label;

[0066] It should be noted that during the text clustering process, if the clustering result meets the preset clustering condition, corresponding labels can be assigned to the data to be labeled based on these clustering results, and this step is called text labeling. The preset clustering condition can be determined based on the number of clusters, the similarity of data within the clusters, the difference between clusters, etc. In one embodiment, it can be determined based on the number of data to be labeled included in all categories of the clustering result. When the number of data to be labeled included in all categories of the clustering result is greater than the preset value, it is considered that the clustering result meets the preset clustering condition. The preset value can be adjusted according to actual application requirements, for example, 30. Assign a unique label or category name to each cluster, and this label can be any name that can reflect the characteristics of the text clustering data. Each data item to be labeled is assigned the corresponding label according to the cluster it belongs to, that is, the first label.

[0067] Step 300: Input the data to be labeled into a preset entity recognition model to obtain entities.

[0068] It should be noted that the preset entity recognition model refers to a model that has learned entity recognition knowledge from a large amount of text data during the training phase. The preset entity recognition model is based on deep learning technology and is pre-trained with a large corpus, so as to accurately identify named entities in the text, such as person names, place names, organization names, etc. The preset entity recognition model can include BERT (Bidirectional Encoder Representations from Transformers), LSTM (Long Short-Term Memory), CRF (Conditional Random Fields), and the Bert+BiLSTM (Bi-directional Long Short-Term Memory)+CRF model, etc. Input the data to be labeled into the preset entity recognition model, and the preset entity recognition model outputs the recognized entity information, which usually includes the text representation of the entity, the type label, and the position in the original text.

[0069] Step 400: Perform entity labeling on the data to be labeled based on the entities to obtain a second label.

[0070] It should be noted that performing entity labeling on the data to be labeled based on the recognized entities aims to mark specific entities in the text and assign a corresponding label to each entity. Specifically, for each recognized entity, a corresponding label is assigned according to its type. The labels are predefined and correspond to the entity types supported by the entity recognition model. Integrating the labeled entity information into the data to be labeled can be to add specific marking symbols, such as XML tags, JSON key-value pairs, etc., before and after each entity to represent the boundaries and types of the entities. The original data to be labeled is converted into a formatted data containing entity labels, that is, the second label.

[0071] Step 500: Combine the entities and determine the corresponding entity combination vector.

[0072] Step 600: Cluster the entity combination vectors, and perform entity combination labeling on the data to be labeled based on the entity clustering result to obtain a third label.

[0073] It should be noted that the process of combining entities, determining the corresponding entity combination vectors, and subsequent entity clustering and labeling aims to explore the associations and semantic structures among entities. According to predefined rules or algorithms, the identified entities are combined. The combination can be based on entity types, such as combining entities of the same type; it can also be based on the relationships between entities, such as the combination of "person name - organization name"; the recognition results of each entity contain the sequential relationship between entities, and the combination can also be based on the recognition order of entities and their positions in the original text. A vector representation is generated for each entity combination. This can be achieved by concatenating, averaging the entity vectors in the combination, or using more complex fusion methods. It can also be directly vectorizing the entity combination. The vectorization of entities or entity combinations is usually learned from text data through deep learning models such as text embedding models. The entity combination vectors are input into a clustering algorithm to perform clustering operations. This will divide the entity combinations into different clusters, and each cluster represents a group of entity combinations with similar characteristics. According to the clustering results, a unique label or category is assigned to each cluster. This can be determined by analyzing the common features or semantic content of the entity combinations in the cluster. The information of the labeled entity combinations is integrated into the data to be labeled to generate the third label containing the entity combination labels.

[0074] Step 700, generate the label of the data to be labeled based on the first label, the second label, and the third label.

[0075] It should be noted that generating the final label of the data to be labeled based on the first label, the second label, and the third label is a process of integrating multiple information sources to form a more comprehensive and accurate annotation. For example, for the data to be labeled "There is aggregate in the aggregate bin that cannot be loaded or unloaded continuously", the first label is cls: "The aggregate bin is full", the second label is ne: "aggregate bin", "aggregate", and the third label is nec: "The aggregate bin cannot be loaded or unloaded"; finally, the label of the data to be labeled is: "There is aggregate in the aggregate bin that cannot be loaded or unloaded continuously, cls: The aggregate bin is full, ne: aggregate bin, aggregate, nec: The aggregate bin cannot be loaded or unloaded".

[0076] In this embodiment, text clustering is performed on the text vectors corresponding to the data to be labeled based on a preset number of clusters to obtain a text clustering result. When the text clustering result meets the preset clustering condition, text labeling is performed on the data to be labeled based on the text clustering result to obtain a first label. The data to be labeled is input into a preset entity recognition model to obtain entities. Entity labeling is performed on the data to be labeled based on the entities to obtain a second label. The entities are combined, and the corresponding entity combination vector is determined. Entity clustering is performed on the entity combination vector, and entity combination labeling is performed on the data to be labeled based on the entity clustering result to obtain a third label. Based on the first label, the second label, and the third label, a label for the data to be labeled is generated. Furthermore, the labeled training data can be obtained. By combining the advantages of text clustering, entity recognition, and entity clustering, the data to be labeled can be comprehensively and deeply labeled from multiple perspectives, improving the accuracy and efficiency of training data labeling. Subsequently, the model is more likely to converge to obtain better training results.

[0077] In one embodiment, performing entity clustering on the entity combination vector and performing entity combination labeling on the data to be labeled based on the entity clustering result to obtain a third label includes:

[0078] Determine the target number of clusters corresponding to the text clustering result that meets the preset clustering condition;

[0079] Perform entity clustering on the entity combination vector based on the target number of clusters, and perform entity combination labeling on the data to be labeled based on the entity clustering result of the entity clustering to obtain a third label.

[0080] It should be noted that in this embodiment, it is determined based on the number of data to be labeled included in all categories of the clustering result. When the number of data to be labeled included in each category of the clustering result is greater than a preset value, it is considered that the clustering result meets the preset clustering condition. The text clustering result that meets the preset clustering condition, that is, when performing text clustering, the number of data to be labeled included in each category of the final text clustering result is greater than the preset value. At this time, the number of cluster centers corresponding to the current text clustering result is determined as the target number of clusters. It can be understood that when the text clustering result does not meet the preset clustering condition, the number of cluster centers of the text clustering algorithm can be updated, that is, by updating the preset number of clusters to re-execute the text clustering algorithm until the text clustering result of the text clustering algorithm meets the preset clustering condition. Based on this target number of clusters, clustering operations are performed on the entity combination vector. The process of gathering similar entity combinations together to form clusters is entity clustering. The clustering algorithm can be K-means, hierarchical clustering, DBSCAN, etc. According to the result of entity clustering, the entity combinations in the data to be labeled are labeled with the labels of their respective clusters. The new label or tag obtained through the entity clustering result is the third label.

[0081] In this embodiment, by determining the target clustering number corresponding to the text clustering result and performing entity combination marking on the data to be labeled, it helps to better understand the internal structure of the text data and improve the efficiency of data analysis and processing.

[0082] Reference Figure 2 , in one embodiment, entity clustering is performed on the entity combination vectors based on the target clustering number, and entity combination marking is performed on the data to be labeled based on the entity clustering result of the entity clustering, to obtain a third marking, including:

[0083] Entity clustering is performed on the entity combination vectors based on the target clustering number, and for each entity combination category in the entity clustering result, the entity clustering center of the entity combination category is obtained;

[0084] The first target data to be labeled among all the data to be labeled included in the entity combination category is determined, where the first target data to be labeled is the data to be labeled closest to the entity clustering center among all the data to be labeled included in the entity combination category;

[0085] Entity combination marking is performed on the first target data to be labeled to obtain a third marking.

[0086] In this embodiment, it should be noted that the entity combination category is each cluster or category output by the clustering algorithm during entity clustering. For each entity combination category, the entity clustering center can be understood as the mean value or a certain representative vector of all the vectors in this category. Specifically, after entity clustering is performed on the entity combination vectors based on the target clustering number, the clustering center of each category is extracted from the output of the clustering algorithm. For example, if the K-means algorithm is used, the clustering center is usually the mean vector of each cluster. For each entity combination category, the distance between the clustering center of this entity combination category and the entity combination vectors among all the data to be labeled is calculated. The data to be labeled with the smallest distance is selected as the first target data to be labeled. Entity combination marking is performed on the selected first target data to be labeled to obtain a third marking.

[0087] In this embodiment, finding the data to be labeled closest to the clustering center for each entity combination category for marking helps to provide more effective data for subsequent analysis, classification, or retrieval tasks.

[0088] Reference Figure 2 , in one embodiment, text clustering is performed on the text vectors corresponding to the data to be labeled based on a preset clustering number to obtain a text clustering result, including:

[0089] Text clustering is performed on the text vectors corresponding to the data to be labeled based on the preset clustering number to obtain an initial text clustering result;

[0090] Determine whether there is a target category in the initial text clustering result, where the target category is a category in all categories of the initial text clustering result that includes a number of unlabeled data less than a preset sample quantity;

[0091] In the case where there is a target category, update the preset clustering quantity based on the preset clustering quantity and the quantity of the target category;

[0092] Based on the updated preset clustering quantity, re - perform text clustering on the text vectors corresponding to the unlabeled data until there is no target category in the initial text clustering result, and use the initial text clustering result as the text clustering result.

[0093] It should be noted that for label marking of data related to the construction machinery field, since construction machinery texts contain many unique components, a certain number of samples are required to ensure model training. Insufficient samples will lead to inaccurate models obtained from the training samples. That is to say, after performing text clustering on the text vectors corresponding to the unlabeled data, it is necessary to ensure that each category of the text clustering result includes a certain number or more of unlabeled data.

[0094] It should be noted that the preset sample quantity refers to the minimum number of unlabeled data that a valid category should contain. The target category refers to those categories in the initial text clustering result that contain a number of unlabeled data less than the preset sample quantity. In this embodiment, the result obtained by performing text clustering on the text vectors corresponding to the unlabeled data based on the preset clustering quantity is determined as the initial text clustering result. If there is a target category in the initial text clustering result, it indicates that the current preset clustering quantity may be too large or the algorithm fails to correctly distinguish the data. According to the quantity of the target category, methods such as increasing or decreasing the preset clustering quantity can be selected, such as reducing a certain proportion or a fixed quantity, to achieve the update of the preset clustering quantity. Use the updated preset clustering quantity to perform iterative clustering on the text vectors of the unlabeled data again. When all categories in the initial text clustering result contain at least the preset sample quantity of unlabeled data, stop the iterative process. Output the initial text clustering result at this time as the final text clustering result. In the case where there is no target category, directly use the initial text clustering result as the text clustering result.

[0095] Specifically, in one embodiment, in the case where there is a target category, updating the preset clustering quantity based on the preset clustering quantity and the quantity of the target category includes:

[0096] In the case where there is a target category, determine the difference between the preset clustering quantity and the quantity of the target category;

[0097] Update the preset clustering quantity based on the difference.

[0098] In this embodiment, it should be noted that the target category refers to those categories in the initial text clustering result that contain fewer unlabeled data than the preset sample quantity. The appearance of the target category indicates that the clustering quantity is excessive at this time, and the preset clustering quantity can be reduced. Specifically, the difference between the preset clustering quantity and the quantity of the target category can be calculated, and this difference can be used as the new preset clustering quantity; alternatively, the difference plus a small positive integer can be used as the new preset clustering quantity, and a certain redundancy can be reserved through this small positive integer to cope with possible category subdivision or noise data.

[0099] In this embodiment, the result of clustering analysis is optimized by updating the preset clustering quantity.

[0100] In one embodiment, when the text clustering result meets the preset clustering condition, text labeling is performed on the unlabeled data based on the text clustering result to obtain the first label, including:

[0101] When the text clustering result meets the preset clustering condition, for each text category in the text clustering result, obtain the text clustering center of the text category;

[0102] Determine the second target unlabeled data among all the unlabeled data included in the text category, where the second target unlabeled data is the unlabeled data closest to the text clustering center among all the unlabeled data included in the text category;

[0103] Perform text labeling on the second target unlabeled data to obtain the first label.

[0104] In this embodiment, it should be noted that the text category is each cluster or category output by the clustering algorithm during text clustering. For each text category, the text clustering center can be understood as the mean value or a certain representative vector of all vectors in this category. Specifically, after text clustering is performed on text vectors based on the preset clustering quantity, the clustering center of each category is extracted from the output of the clustering algorithm. For example, if the K-means algorithm is used, the clustering center is usually the mean vector of each cluster. For each text category, calculate the distance between the clustering center of this text category and the text vectors among all the unlabeled data. Select the unlabeled data with the smallest distance as the second target unlabeled data. Perform text labeling on the selected second target unlabeled data to obtain the first label.

[0105] In this embodiment, finding the unlabeled data closest to the clustering center for each text category for labeling helps to provide more effective data for subsequent analysis, classification, or retrieval tasks.

[0106] The embodiment of the present application also provides a computing device, which may include:

[0107] A memory configured to store instructions;

[0108] A processor configured to call instructions from the memory and capable of implementing the above data marking method when executing the instructions.

[0109] An embodiment of the present application also provides a machine-readable storage medium having instructions stored thereon for causing a machine to execute the above data marking method.

[0110] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0112] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0114] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0115] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0116] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0117] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0118] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A data labeling method, characterized in that: include: Performing text clustering on the text vectors corresponding to the unlabeled data based on a preset number of clusters to obtain a text clustering result; In the case where the text clustering result satisfies a preset clustering condition, text marking is performed on the data to be marked based on the text clustering result to obtain a first mark; Inputting the data to be labeled into a preset entity recognition model to obtain an entity; Performing entity labeling on the data to be labeled based on the entity to obtain a second label; Combining the entities and determining a corresponding entity combination vector; Performing entity clustering on the entity combination vectors, and performing entity combination labeling on the to-be-labeled data based on the entity clustering result to obtain a third label; A label for the data to be labeled is generated based on the first label, the second label, and the third label.

2. The data labeling method according to claim 1, characterized in that: The entity combination vectors are entity clustered, and entity combination labeling is performed on the data to be labeled based on the entity clustering result to obtain a third label, including: Determine the target clustering quantity corresponding to the text clustering results that meet the preset clustering condition; The entity combination vectors are entity clustered based on the target cluster quantity, and the to-be-labeled data are entity-combined labeled based on the entity clustering result of the entity clustering to obtain a third label.

3. The data labeling method according to claim 2, characterized in that: The performing entity clustering on the entity combination vector based on the target cluster quantity, and performing entity combination labeling on the to-be-labeled data based on the entity clustering result of the entity clustering to obtain a third label, comprises: Performing entity clustering on the entity combination vector based on the target number of clusters, and obtaining the entity cluster center of each entity combination category in the entity clustering result; Determine a first target to-be-labeled data among all to-be-labeled data included in the entity combination category, wherein the first target to-be-labeled data is the to-be-labeled data that is closest to the entity cluster center among all to-be-labeled data included in the entity combination category; Perform entity combination labeling on the first target data to be labeled to obtain a third label.

4. The data labeling method according to claim 1, characterized in that: The text clustering is performed on the text vectors corresponding to the to-be-labeled data based on the preset number of clusters to obtain the text clustering result, including: Performing text clustering on the text vectors corresponding to the unlabeled data based on a preset number of clusters to obtain an initial text clustering result; Determine whether there is a target category in the initial text clustering result, wherein the target category is a category in which the number of to-be-labeled data is less than a preset number of samples among all categories of the initial text clustering result; In the case where the target category exists, updating the preset number of clusters based on the preset number of clusters and the number of the target categories; Based on the updated preset clustering number, the text vectors corresponding to the to-be-labeled data are re-clustered until the target category does not exist in the initial text clustering result, and the initial text clustering result is used as the text clustering result.

5. The data labeling method according to claim 4, characterized in that: In the case where the target category exists, updating the preset number of clusters based on the preset number of clusters and the number of the target categories includes: In the case where the target category exists, determining a difference between the preset number of clusters and the number of the target category; The preset number of clusters is updated based on the difference.

6. The data labeling method according to claim 1, characterized in that: When the text clustering result satisfies a preset clustering condition, text marking is performed on the to-be-marked data based on the text clustering result to obtain a first mark, including: When the text clustering result satisfies a preset clustering condition, for each text category in the text clustering result, obtaining a text clustering center of the text category; Determine second target data to be labeled among all the data to be labeled included in the text category, wherein the second target data to be labeled is the data to be labeled that is closest to the text cluster center among all the data to be labeled included in the text category; Text labeling is performed on the second target data to be labeled to obtain a first label.

7. The data labeling method according to claim 1, characterized in that: The text clustering is performed on the text vectors corresponding to the to-be-labeled data based on the preset number of clusters to obtain the text clustering result, including: Convert the to-be-labeled data into a text vector based on a preset text embedding model; The text vector is clustered based on a preset number of clusters to obtain a text clustering result.

8. The data labeling method according to claim 1, characterized in that: Before the step of performing text clustering on the text vectors corresponding to the to-be-labeled data based on the preset number of clusters to obtain the text clustering result, the step further includes: Obtain the data to be labeled; Preprocessing the data to be labeled; The preprocessed data to be marked is converted into a preset text format.

9. A computing device, characterized in that include: a memory configured to store instructions; A processor is configured to call the instructions from the memory and implement the data marking method according to any one of claims 1 to 8 when executing the instructions.

10. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores instructions for causing a machine to execute the data marking method according to any one of claims 1 to 8.