Data label automatic generation method and system

By combining data label generation library and model, priority is given to using the library to generate tags, and after reaching the threshold, the model is used to generate tags and supplement the library, the problems of waste of resources and inefficiency in the existing technology are solved, and efficient automatic generation of data labels and dynamic improvement of libraries are achieved.

CN120508830APending Publication Date: 2025-08-19DATANG SOFTWARE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511006435.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art requires the processing of content feature information through a fusion encoder when generating data tags, resulting in waste of resources and inefficient efficiency, and the inability to effectively utilize the existing generated data tags.

Method used

Using a method of combining data label generation library and model, we give priority to using the data label generation library to generate data labels. If the library cannot be generated or reaches the threshold, the data label generation model is used to generate data labels, and the final label is determined through feature information similarity and weight calculation, and the label generated by the model is supplemented and improved.

Benefits of technology

It improves the efficiency of automatic generation of data tags, realizes effective utilization of resources and dynamic improvement of tag libraries, and reduces resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508830A_ABST
    Figure CN120508830A_ABST
Patent Text Reader

Abstract

The invention provides a data label automatic generation method and system. The method comprises the steps that a data label generation library and a data label generation model are preset; and obtaining the data quantity of the to-be-generated data labels, preferentially generating the data labels based on the data label generation library, and if the data labels cannot be generated based on the data label generation library or the data quantity of the data labels generated based on the data label generation library reaches a first preset threshold value, generating the data labels based on the data label generation model. According to the method, the data tag generation library and the generation model are preset, the data tag generation library is preferentially used for generating the data tag based on the feature information similarity, and if the data tag generation library cannot generate the data tag, the data tag generation model is used for understanding and generating the data tag based on the feature information; and meanwhile, shunting can be realized when the data volume is large to synchronously generate the data label, compared with the prior art, resource utilization can be performed on the generated data label, and shunting can improve the automatic generation efficiency of the data label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for automatically generating data labels. Background Art

[0002] Data labels refer to labels or tags used to identify, classify, describe or explain each data point or data set in a data set, providing contextual information about the data to help users understand the meaning and purpose of the data.

[0003] Currently, existing technologies generally use an encoder to process the content feature information and fusion information of published content to obtain data labels. For example, the "Method, Device, Electronic Device, and Readable Storage Medium for Generating Data Labels" disclosed in Chinese patent literature, with publication number CN111125177A, includes the following steps: obtaining published content, inputting the data in the published content into encoders corresponding to the types of data contained in the published content to obtain corresponding content feature information, inputting all the content feature information of the published content into a fusion encoder to obtain fusion information, and inputting the content feature information and fusion information corresponding to the published content into a preset decoder to obtain a data label corresponding to the published content. Although this solution takes into account multiple types of data and the relationships between the data, thereby improving the accuracy of data label generation, the generation of the above-mentioned data labels requires each time the content feature information of the published content is processed by the fusion encoder to obtain fusion information, and then the content feature information and fusion information of the published content are processed by the encoder to obtain the data label. This makes it impossible to utilize the existing generated data labels, resulting in a waste of resources. Summary of the Invention

[0004] To this end, one purpose of the present invention is to propose a method and system for automatically generating data tags to solve the problems mentioned in the background technology and overcome the shortcomings of the existing technology.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a method for automatically generating data labels, comprising:

[0007] Preset data label generation library;

[0008] Preset data label generation model;

[0009] Obtain the amount of data to be generated for data labels, and give priority to generating data labels based on the data label generation library. If data labels cannot be generated based on the data label generation library or the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generate data labels based on the data label generation model.

[0010] Preferably, the data label generation library is composed of multiple data labels and multiple feature information, each data label is associated with multiple related feature information, and each feature information is set with a weight; the data label generation model is associated with the data label generation library.

[0011] Preferably, generating data labels based on a data label generation library includes:

[0012] A1: Obtains the data type of the data label to be generated;

[0013] A2: Select a specific model based on the data type to extract data feature information. If the extracted data feature information is a single one, directly execute A3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information exceeds a second preset threshold, and execute A3.

[0014] A3: Calculate the similarity between the feature information retained in the data and the feature information in the data label generation library, and make a judgment: If the similarity between the feature information of the data and the feature information in the data label generation library is all higher than a third preset threshold, execute A4; if the similarity between a certain feature information of the data and the feature information in the data label generation library is lower than the third preset threshold, execute A5;

[0015] A4: The feature information in the data label generation library is used as the feature information of the data. The weight value of the associated data label is calculated based on the weight value of the feature information in the data label generation library. The associated data label with the highest weight value is used as the final data label generated for the data.

[0016] A5: First generate a data label according to A4, use the data label as the first data label, then calculate the similarity between a certain feature information and the first data label, and make a judgment: if the similarity between a certain feature information and the first data label is higher than the fourth preset threshold, use the first data label as the final data label for data generation; if the similarity between a certain feature information and the first data label is lower than the fourth preset threshold, generate the final data label based on the data label generation model.

[0017] Preferably, if the data label generation library cannot generate the data label, generating the data label based on the data label generation model includes:

[0018] A certain feature information and a first data label are input into a data label generation model. The data label generation model extracts the certain feature information and the feature information of the first data label and fuses them to generate a second data label. The second data label is used as the final data label for data generation.

[0019] Preferably, if the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generating data labels based on the data label generation model includes:

[0020] B1: Obtain the data type of the data label to be generated;

[0021] B2: Select a specific model based on the data type to extract feature information of the data. If the extracted data feature information is a single one, execute B3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and execute B3.

[0022] B3: The data label generation model generates data labels based on the fusion of feature information retained by the data, and uses the data labels as the final data labels generated by the data.

[0023] Preferably, the step of obtaining the data type of the data tag to be generated includes obtaining a file header and / or a suffix.

[0024] Preferably, the method of selecting a specific model based on the data type to extract the feature information of the data includes: if the data type of the data label to be generated is a file type, extracting the feature information of the data based on a large language model; if the data type of the data label to be generated is an image or video, extracting the feature information of the data based on a convolutional neural network.

[0025] As an advantage, the method further includes: adding the data labels generated by the data label generation model into the data label generation library to improve the data label generation library;

[0026] The data labels generated by the data label generation model are added to the data label generation library. The data label generation library is improved to include:

[0027] Calculating the similarity between the second data label generated by the data label generation model and the data labels in the data label generation library;

[0028] If the similarity between the second data tag and the data tag in the data tag generation library is higher than a fifth preset threshold, a certain feature information of the data is associated with the data tag in the data tag generation library;

[0029] If the similarity between the second data tag and the data tag in the data tag generation library is lower than the fifth preset threshold, the second data tag is retained as a new data tag in the data tag generation library, and a certain feature information of the data is associated with the newly generated data tag in the data tag library.

[0030] Preferably, the calculation of the similarity includes: converting the feature information or data label into a vector, calculating the distance between the two vectors according to a calculation formula, and judging the similarity between the two feature information or data labels based on the distance between the two vectors. If the distance is smaller, the similarity is higher, and vice versa, the similarity is lower.

[0031] In a second aspect, the present invention provides a data label automatic generation system, comprising:

[0032] A data label generation library preset module is used to preset the data label generation library;

[0033] A data label generation model preset module is used to preset a data label generation model;

[0034] The data diversion module is used to obtain the amount of data to be generated for data labels, and preferentially divert the data to the preset data label generation library. If the preset data label generation library cannot generate data labels or the amount of data generated by the data label generation library reaches a preset threshold, the data is diverted to the preset data label generation model;

[0035] The data label generation module is used to generate data labels based on a preset data label generation library and a preset data label generation model.

[0036] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for automatically generating data labels as described above are implemented.

[0037] In a fourth aspect, the present invention provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for automatically generating data tags.

[0038] Therefore, the present invention has the following beneficial effects:

[0039] The present invention provides a method and system for automatically generating data labels, which presets a data label generation library and a data label generation model. The data label generation library is used to generate data labels based on the similarity of feature information. If the data label generation library cannot generate data labels, the data label generation model is used to generate data labels based on the understanding of feature information. At the same time, the data labels generated by the data label generation model can supplement and improve the data label generation library based on label similarity. At the same time, when the amount of data is large, the two can be diverted to generate data labels synchronously. Compared with the existing technology, the generated data labels can be utilized as resources, and the diversion can improve the efficiency of automatic data label generation.

[0040] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0042] Figure 1 is a flow chart of the method of the present invention;

[0043] Figure 2 is an overall flow chart of the method according to an embodiment of the present invention;

[0044] Figure 3 This is a flow chart of generating data labels based on a data label generation library according to an embodiment of the present invention;

[0045] Figure 4 is a flow chart of generating data labels based on a data label generation model according to an embodiment of the present invention;

[0046] Figure 5 This is a structural block diagram of a data tag automatic generation system according to an embodiment of the present invention;

[0047] Figure 6 It is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0049] The terms "first" and "second" in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to the process, method, product, or device.

[0050] The present invention provides a method for automatically generating data labels. Figure 1 Shown, including:

[0051] Preset data label generation library;

[0052] Preset data label generation model;

[0053] Obtain the amount of data to be generated for data labels, and give priority to generating data labels based on the data label generation library. If data labels cannot be generated based on the data label generation library or the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generate data labels based on the data label generation model.

[0054] The present invention presets a data label generation library and a data label generation model, and preferentially uses the data label generation library to generate data labels based on feature information similarity. If the data label generation library cannot generate data labels, the data label generation model is used to generate data labels based on feature information understanding. At the same time, the data labels generated by the data label generation model can supplement and improve the data label generation library based on label similarity, and can also realize the synchronous generation of data labels by diverting the two when the amount of data is large. Compared with the existing technology, the generated data labels can be utilized as resources, and the diversion can improve the efficiency of automatic data label generation.

[0055] Furthermore, the data label generation library is composed of multiple data labels and multiple feature information, each data label is associated with multiple related feature information, and a weight is set for each feature information; the data label generation model is associated with the data label generation library.

[0056] Further, such as Figure 3 As shown, generating data labels based on the data label generation library includes:

[0057] A1: Obtains the data type of the data label to be generated;

[0058] A2: Select a specific model based on the data type to extract data feature information. If the extracted data feature information is a single one, directly execute A3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information exceeds a second preset threshold, and execute A3.

[0059] A3: Calculate the similarity between the feature information retained in the data and the feature information in the data label generation library, and make a judgment: If the similarity between the feature information of the data and the feature information in the data label generation library is all higher than a third preset threshold, execute A4; if the similarity between a certain feature information of the data and the feature information in the data label generation library is lower than the third preset threshold, execute A5;

[0060] A4: The feature information in the data label generation library is used as the feature information of the data. The weight value of the associated data label is calculated based on the weight value of the feature information in the data label generation library. The associated data label with the highest weight value is used as the final data label generated for the data.

[0061] A5: First generate a data label according to A4, use the data label as the first data label, then calculate the similarity between a certain feature information and the first data label, and make a judgment: if the similarity between a certain feature information and the first data label is higher than the fourth preset threshold, use the first data label as the final data label for data generation; if the similarity between a certain feature information and the first data label is lower than the fourth preset threshold, generate the final data label based on the data label generation model.

[0062] Furthermore, if the data label generation library cannot generate data labels, generating data labels based on the data label generation model includes:

[0063] A certain feature information and a first data label are input into a data label generation model. The data label generation model extracts the certain feature information and the feature information of the first data label and fuses them to generate a second data label. The second data label is used as the final data label for data generation.

[0064] Further, such as Figure 4 As shown, if the amount of data generated by the data label generation library reaches a first preset threshold, generating data labels based on the data label generation model includes:

[0065] B1: Obtain the data type of the data label to be generated;

[0066] B2: Select a specific model based on the data type to extract feature information of the data. If the extracted data feature information is a single one, execute B3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and execute B3.

[0067] B3: The data label generation model generates data labels based on the fusion of feature information retained by the data, and uses the data labels as the final data labels generated by the data.

[0068] Furthermore, obtaining the data type of the data tag to be generated includes obtaining a file header and / or a suffix.

[0069] Furthermore, selecting a specific model based on the data type to extract feature information of the data includes: if the data type of the data label to be generated is a file type, extracting feature information of the data based on a large language model; if the data type of the data label to be generated is an image or video, extracting feature information of the data based on a convolutional neural network.

[0070] Furthermore, the method further includes: adding the data labels generated by the data label generation model into the data label generation library to improve the data label generation library;

[0071] The data labels generated by the data label generation model are added to the data label generation library. The data label generation library is improved to include:

[0072] Calculating the similarity between the second data label generated by the data label generation model and the data labels in the data label generation library;

[0073] If the similarity between the second data tag and the data tag in the data tag generation library is higher than a fifth preset threshold, a certain feature information of the data is associated with the data tag in the data tag generation library;

[0074] If the similarity between the second data tag and the data tag in the data tag generation library is lower than the fifth preset threshold, the second data tag is retained as a new data tag in the data tag generation library, and a certain feature information of the data is associated with the newly generated data tag in the data tag library.

[0075] Furthermore, the calculation of similarity includes: converting feature information or data labels into vectors, calculating the distance between two vectors according to a calculation formula, and judging the similarity between two feature information or data labels based on the distance between the two vectors. If the distance is smaller, the similarity is higher, and vice versa, the similarity is lower.

[0076] Next, the data tag automatic generation method of the present invention is described in detail in the form of a step flow and in combination with the accompanying drawings.

[0077] like Figure 2 As shown, a method for automatically generating data labels includes steps S1-S4.

[0078] S1: A preset data label generation library is composed of multiple data labels and multiple feature information. Each data label is associated with multiple related feature information, and a weight is set for each feature information.

[0079] In one embodiment, it is assumed that the data tag generation library has data labels, respectively , each data label Associate multiple feature information .

[0080] Feature Information The weight values set include 0.1, 0.2, and 0.3, and each data label Multiple feature information under The sum of the weights is 1.

[0081] S2: A data label generation model based on a large language model is preset, and the data label generation model is associated with the data label generation library.

[0082] S3: Obtain the amount of data to generate data labels, and give priority to generating data labels based on a preset data label generation library. If the preset data label generation library cannot generate data labels or the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generate data labels based on a preset data label generation model.

[0083] Further, such as Figure 3 As shown, in step S3, generating data labels based on the preset data label generation library specifically includes steps A1-A5:

[0084] A1: Obtains the data type of the data label to be generated;

[0085] A2: Select a specific model based on the data type to extract data feature information. If the extracted data feature information is a single one, proceed to the next step. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and proceed to the next step.

[0086] In this embodiment, it is assumed that there are multiple pieces of data feature information, specifically , then compare and 、 and 、…、 and ,…,as well as and The similarity of and If the similarity exceeds the second preset threshold, or , and so on.

[0087] A3: Calculate the similarity between the feature information retained in the data and the feature information in the data label generation library. If the similarity between the feature information of the data and the feature information in the data label generation library is all higher than a third preset threshold, execute A4. If the similarity between a certain feature information of the data and the feature information in the data label generation library is lower than the third preset threshold, execute A5.

[0088] A4: The feature information in the data label generation library is used as the feature information of the data. The weight value of the associated data label is calculated based on the weight value of the feature information in the data label generation library. The associated data label with the highest weight value is used as the final data label generated for the data.

[0089] A5: First, determine the first data label for data generation according to A4 through feature information whose similarity is higher than that of feature information in the data label generation library, and then calculate the similarity between a certain feature information and the first data label. If the similarity between a certain feature information and the first data label is higher than a fourth preset threshold, then use the data label as the final data label for data generation. If the similarity between a certain feature information and the first data label is lower than the fourth preset threshold, then generate a data label based on the data label generation model.

[0090] In one embodiment, the similarity between the feature information of the data and the feature information in the data tag generation library is all higher than the third preset threshold value.

[0091] Assume that the characteristic information retained by the data is , then calculate and Feature information under data labels similarities between

[0092] If the data feature information and the feature information under the data label If the similarities between the three features are higher than the third preset threshold, the three feature information As characteristic information of data;

[0093] Determine the three characteristic information Data label , such as feature information Belongs to data label , feature information Belongs to data label , feature information Belongs to data label , then determine the feature information The weight value of , then the data label As data labels generated by the data, and so on.

[0094] In one embodiment, the similarity between a certain feature information of the data and the feature information in the data tag generation library is lower than the third preset threshold as an example for explanation:

[0095] Assume that the characteristic information retained by the data is 、 , then calculate 、 and Feature information under data labels similarities between

[0096] If the data feature information and the feature information under the data label If the similarities between the three are higher than the third preset threshold, the data feature information and the feature information under the data label If the similarity between the three features is lower than the third preset threshold, the three feature information As characteristic information of data;

[0097] Determine the three characteristic information Data label , such as feature information Belongs to data label , feature information Belongs to data label , feature information Belongs to data label , then determine the feature information The weight value of , then the data label As the first data label;

[0098] Calculating feature information With the first data label The similarity between them is determined by: With the first data label If the similarity is higher than the fourth preset threshold, the first data label is used as the final data label generated by the data. If the similarity with the first data label is lower than a fourth preset threshold, a final data label is generated based on the data label generation model.

[0099] It is understood that a certain characteristic information in the present invention is characteristic information Any item in can also be feature information More specifically, a certain feature information refers to one or more feature information in the feature information retained by the data whose similarity with the feature information in the data label generation library is lower than a third preset threshold.

[0100] Furthermore, when the preset data label generation library cannot generate a data label, that is, the similarity between a certain feature information and the first data label is lower than a fourth preset threshold, the specific steps of generating a data label based on the preset data label generation model include:

[0101] A certain feature information and a first data label are input into a data label generation model. The data label generation model extracts the certain feature information and the feature information of the first data label and fuses them to generate a second data label, and the second data label is used as the final data label for data generation.

[0102] Further, in step S3, if the amount of data generated based on the data tag generation library reaches a first preset threshold, then generating data tags based on the data tag generation model includes steps B1-B3, such as Figure 4 As shown:

[0103] B1: Obtain the data type of the data label to be generated;

[0104] B2: Select a specific model based on the data type to extract feature information of the data. If the extracted data feature information is a single one, execute B3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and execute B3.

[0105] B3: The data label generation model generates data labels based on the fusion of feature information retained by the data, and uses the data labels as the final data labels generated by the data.

[0106] Furthermore, in step S3, obtaining the data type of the data tag to be generated includes one or more ways of obtaining the file header or suffix.

[0107] Specifically, in step S3, if the data type is a file type, feature information of the data is extracted based on a large language model; if the data type is an image or video, feature information of the data is extracted based on a convolutional neural network.

[0108] It is understandable that extracting feature information of data based on a large language model and extracting feature information of data based on a convolutional neural network are both existing mature technologies and will not be described in detail in the present invention.

[0109] S4: The data labels generated by the data label generation model are added to the data label generation library to improve the data label generation library.

[0110] As an implementation method, step S4 includes:

[0111] Calculating the similarity between the second data label generated by the data label generation model and the data labels in the data label generation library;

[0112] If the similarity between the second data tag and the data tag in the data tag generation library is higher than a fifth preset threshold, a certain feature information of the data is associated with the data tag in the data tag generation library;

[0113] If the similarity between the second data tag and the data tag in the data tag generation library is lower than the fifth preset threshold, the second data tag is retained as a new data tag in the data tag generation library, and a certain feature information of the data is associated with the newly generated data tag in the data tag library.

[0114] Furthermore, the specific steps of similarity calculation in the above steps include:

[0115] Convert feature information or data labels into vectors;

[0116] Calculate the distance between two vectors. The expression is:

[0117]

[0118] Where: and Represented as a vector of two feature information or data labels;

[0119] The similarity between two feature information or data labels is determined based on the distance between the two vectors. The smaller the distance, the higher the similarity, and vice versa.

[0120] It can be understood that the first preset threshold, the second preset threshold, the third preset threshold, the fourth preset threshold, and the fifth preset threshold in the present invention are set by those skilled in the art according to actual circumstances. The method of the present invention does not specifically limit the above preset thresholds. For example, the first preset threshold can be 10, or 100, or 1000, etc., or it can be a percentage of the number of data. For example, if the number of data is N, the first preset threshold can be N×1%, N×10%, or N×20%, etc. The second to fifth preset thresholds can be any number in the similarity value range of 0-1. The closer to 1, the higher the similarity. The values that can be taken are 0.5 and 0.7. The second to fifth preset thresholds can also be represented by the distance between two vectors, because the calculation of similarity / similarity is to convert feature information or data labels into vectors, calculate the distance between the two vectors according to the calculation formula, and judge the similarity between the two feature information or data labels based on the distance between the two vectors. If the distance is smaller, the similarity is higher, and vice versa, the similarity is lower.

[0121] Furthermore, the above-mentioned data label generation model is a data label generation model based on a large language model, which is an existing mature technology and will not be described in detail here.

[0122] As an implementation method, the data label generation model based on the large language model is a machine learning technology based on deep neural networks. Therefore, the method of the present invention is also a data label automatic generation method based on deep learning.

[0123] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0124] In another embodiment, Figure 5 As shown, the present invention provides a data label automatic generation system, comprising:

[0125] A data label generation library preset module is used to preset the data label generation library;

[0126] A data label generation model preset module is used to preset a data label generation model;

[0127] The data diversion module is used to obtain the amount of data to be generated for data labels, and preferentially divert the data to the preset data label generation library. If the preset data label generation library cannot generate data labels or the amount of data generated by the data label generation library reaches a preset threshold, the data is diverted to the preset data label generation model;

[0128] The data label generation module is used to obtain the amount of data to be generated for data labels, and to generate data labels based on the data label generation library first. If the data label cannot be generated based on the data label generation library or the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, the data label is generated based on the data label generation model.

[0129] Furthermore, the data label generation library is composed of multiple data labels and multiple feature information, each data label is associated with multiple related feature information, and a weight is set for each feature information; the data label generation model is associated with the data label generation library.

[0130] Furthermore, generating data labels based on the data label generation library includes:

[0131] A1: Obtains the data type of the data label to be generated;

[0132] A2: Select a specific model based on the data type to extract data feature information. If the extracted data feature information is a single one, directly execute A3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information exceeds a second preset threshold, and execute A3.

[0133] A3: Calculate the similarity between the feature information retained in the data and the feature information in the data label generation library, and make a judgment: If the similarity between the feature information of the data and the feature information in the data label generation library is all higher than a third preset threshold, execute A4; if the similarity between a certain feature information of the data and the feature information in the data label generation library is lower than the third preset threshold, execute A5;

[0134] A4: The feature information in the data label generation library is used as the feature information of the data. The weight value of the associated data label is calculated based on the weight value of the feature information in the data label generation library. The associated data label with the highest weight value is used as the final data label generated for the data.

[0135] A5: First generate a data label according to A4, use the data label as the first data label, then calculate the similarity between a certain feature information and the first data label, and make a judgment: if the similarity between a certain feature information and the first data label is higher than the fourth preset threshold, use the first data label as the final data label for data generation; if the similarity between a certain feature information and the first data label is lower than the fourth preset threshold, generate the final data label based on the data label generation model.

[0136] Furthermore, if the data label generation library cannot generate data labels, generating data labels based on the data label generation model includes:

[0137] A certain feature information and a first data label are input into a data label generation model. The data label generation model extracts the certain feature information and the feature information of the first data label and fuses them to generate a second data label. The second data label is used as the final data label for data generation.

[0138] Furthermore, if the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generating data labels based on the data label generation model includes:

[0139] B1: Obtain the data type of the data label to be generated;

[0140] B2: Select a specific model based on the data type to extract feature information of the data. If the extracted data feature information is a single one, execute B3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and execute B3.

[0141] B3: The data label generation model generates data labels based on the fusion of feature information retained by the data, and uses the data labels as the final data labels generated by the data.

[0142] Furthermore, obtaining the data type of the data tag to be generated includes obtaining a file header and / or a suffix.

[0143] Furthermore, selecting a specific model based on the data type to extract feature information of the data includes: if the data type of the data label to be generated is a file type, extracting feature information of the data based on a large language model; if the data type of the data label to be generated is an image or video, extracting feature information of the data based on a convolutional neural network.

[0144] Furthermore, the method further includes: adding the data labels generated by the data label generation model into the data label generation library to improve the data label generation library;

[0145] The data labels generated by the data label generation model are added to the data label generation library. The data label generation library is improved to include:

[0146] Calculating the similarity between the second data label generated by the data label generation model and the data labels in the data label generation library;

[0147] If the similarity between the second data tag and the data tag in the data tag generation library is higher than a fifth preset threshold, a certain feature information of the data is associated with the data tag in the data tag generation library;

[0148] If the similarity between the second data tag and the data tag in the data tag generation library is lower than the fifth preset threshold, the second data tag is retained as a new data tag in the data tag generation library, and a certain feature information of the data is associated with the newly generated data tag in the data tag library.

[0149] Furthermore, the calculation of similarity includes: converting feature information or data labels into vectors, calculating the distance between two vectors according to a calculation formula, and judging the similarity between two feature information or data labels based on the distance between the two vectors. If the distance is smaller, the similarity is higher, and vice versa, the similarity is lower.

[0150] Furthermore, it should be understood that since the configuration of each module is merely to illustrate the functional units of the system disclosed herein, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.

[0151] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] To solve the above technical problems, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned method for automatically generating data labels are implemented.

[0153] like Figure 6 As shown, the computer / electronic device includes a memory, a processor, and a network interface that are interconnected and communicated through a system bus. It should be noted that the figure only shows a computer device with component memory, a processor, a network interface, and an operating system, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer / electronic device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0154] Computers / electronic devices can be desktop computers, laptops, PDAs, cloud servers, and other computing devices. Computers / electronic devices can interact with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0155] The memory may be one or more than one, and may include at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disks, optical disks, and the like. In some embodiments, the memory may be an internal storage unit of a computer device, such as the computer device's hard disk or internal memory. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, and the like. Of course, the memory may also include both the internal storage unit and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device, such as the program code for a method for automatically generating data tags. Furthermore, the memory may also be used to temporarily store various types of data that have been output or are about to be output.

[0156] In some embodiments, the processor can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of a computer device. In this embodiment, the processor is used to execute program code stored in a memory or process data, such as executing program code for a method for automatically generating data tags.

[0157] The network interface may include a wireless network interface and / or a wired network interface, which is generally used to establish a communication connection between a computer device and other electronic devices.

[0158] The present invention also provides another embodiment, namely, providing a readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned method for automatically generating data tags when executed by a processor.

[0159] Explanation of terms:

[0160] Data labeling: refers to labels or tags used to identify, classify, describe or explain each data point or data set in a data set, providing contextual information about the data to help users understand the meaning and purpose of the data.

[0161] Deep learning: A machine learning technology based on deep neural networks. It has developed on the basis of statistical machine learning and artificial neural networks, combined with the development of big data and powerful computing power. By building and training deep neural network models, it can learn and automatically extract features from large amounts of data.

[0162] Feature information: This article mainly refers to the field of machine learning. Feature engineering involves extracting useful feature information from raw data to help the model learn and predict better.

[0163] Data diversion: This refers to distributing data to different processing paths based on certain rules. This processing method can help achieve parallel data processing. When large-scale data needs to be processed in real time, diversion can significantly improve system performance.

[0164] Data label generation model: It is a machine learning model that aims to generate new data samples by learning the distribution characteristics of data. By analyzing and learning the data, it automatically assigns or modifies labels to the data, thereby improving the efficiency and accuracy of data processing.

[0165] Data features refer to the parts or indicators in the data that have specific meanings or special attributes. They are used to describe and analyze data, thereby helping to better understand the nature and structure of the data.

[0166] Data feature information is a quantifiable attribute or variable in a dataset that describes the characteristics of a data point. Feature information can be continuous values, such as height and weight, or discrete categories, such as gender.

[0167] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0168] Those skilled in the art will readily understand that the present invention encompasses any combination of the components described in the Summary and Detailed Description of the Invention and the accompanying drawings. Due to space limitations and for the sake of clarity, not all of the various solutions resulting from these combinations are described. Any modifications, equivalent substitutions, and improvements within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0169] Although the embodiments of the present invention have been shown and described above, it should be understood that the above embodiments are illustrative and are not to be construed as limiting the present invention. Those skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments without departing from the principles and intent of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatically generating data labels, characterized in that: include: Preset data label generation library; Preset data label generation model; Obtain the amount of data to be generated for data labels, and give priority to generating data labels based on the data label generation library. If data labels cannot be generated based on the data label generation library or the amount of data for generating data labels based on the data label generation library reaches a first preset threshold, generate data labels based on the data label generation model.

2. The method for automatically generating data labels according to claim 1, characterized in that: The data tag generation library consists of multiple data tags and multiple feature information. Each data tag is associated with multiple related feature information, and each feature information is weighted. The data label generation model is associated with the data label generation library.

3. The method for automatically generating data labels according to claim 1, characterized in that: Generating data labels based on the data label generation library includes: A1: Obtains the data type of the data label to be generated; A2: Select a specific model based on the data type to extract data feature information. If the extracted data feature information is a single one, directly execute A3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information exceeds a second preset threshold, and execute A3. A3: Calculate the similarity between the feature information retained in the data and the feature information in the data label generation library, and make a judgment: If the similarity between the feature information of the data and the feature information in the data label generation library is all higher than a third preset threshold, execute A4; if the similarity between a certain feature information of the data and the feature information in the data label generation library is lower than the third preset threshold, execute A5; A4: The feature information in the data label generation library is used as the feature information of the data. The weight value of the associated data label is calculated based on the weight value of the feature information in the data label generation library. The associated data label with the highest weight value is used as the final data label generated for the data. A5: First generate a data label according to A4, use the data label as the first data label, then calculate the similarity between a certain feature information and the first data label, and make a judgment: if the similarity between a certain feature information and the first data label is higher than the fourth preset threshold, use the first data label as the final data label for data generation; if the similarity between a certain feature information and the first data label is lower than the fourth preset threshold, generate the final data label based on the data label generation model.

4. The method for automatically generating data labels according to claim 3, characterized in that: If data labels cannot be generated based on the data label generation library, data labels are generated based on the data label generation model, including: A certain feature information and a first data label are input into a data label generation model. The data label generation model extracts the certain feature information and the feature information of the first data label and fuses them to generate a second data label. The second data label is used as the final data label for data generation.

5. The method for automatically generating data labels according to claim 1, characterized in that: If the amount of data generated based on the data label generation library reaches a first preset threshold, generating data labels based on the data label generation model includes: B1: Obtain the data type of the data label to be generated; B2: Select a specific model based on the data type to extract feature information of the data. If the extracted data feature information is a single one, execute B3. If the extracted data feature information is multiple, calculate the similarity between the multiple feature information, eliminate any feature information whose similarity between the two feature information is higher than a second preset threshold, and execute B3. B3: The data label generation model generates data labels based on the fusion of feature information retained by the data, and uses the data labels as the final data labels generated by the data.

6. A method for automatically generating data tags according to claim 3 or 5, characterized in that: Acquiring the data type of the data tag to be generated includes acquiring a file header and / or a suffix.

7. A method for automatically generating data tags according to claim 3 or 5, characterized in that: The method of extracting feature information of data by selecting a specific model based on the data type includes: if the data type of the data label to be generated is a file type, extracting feature information of the data based on a large language model; if the data type of the data label to be generated is an image or video, extracting feature information of the data based on a convolutional neural network.

8. The method for automatically generating data labels according to claim 4, characterized in that: Also includes: Add the data labels generated by the data label generation model to the data label generation library to improve the data label generation library; The data labels generated by the data label generation model are added to the data label generation library. The data label generation library is improved to include: Calculating the similarity between the second data label generated by the data label generation model and the data labels in the data label generation library; If the similarity between the second data tag and the data tag in the data tag generation library is higher than a fifth preset threshold, a certain feature information of the data is associated with the data tag in the data tag generation library; If the similarity between the second data tag and the data tag in the data tag generation library is lower than the fifth preset threshold, the second data tag will be retained as a new data tag in the data tag generation library, and a certain feature information of the data will be associated with the newly generated data tag in the data tag library.

9. The method for automatically generating data labels according to claim 3, characterized in that: The calculation of the similarity includes: converting the feature information or data label into a vector, calculating the distance between the two vectors according to a calculation formula, and judging the similarity between the two feature information or data labels based on the distance between the two vectors. If the distance is smaller, the similarity is higher, and vice versa, the similarity is lower.

10. A data label automatic generation system, characterized in that: include: A data label generation library preset module is used to preset the data label generation library; A data label generation model preset module is used to preset a data label generation model; The data diversion module is used to obtain the amount of data to be generated for data labels, and preferentially divert the data to the preset data label generation library. If the preset data label generation library cannot generate data labels or the amount of data generated by the data label generation library reaches a preset threshold, the data is diverted to the preset data label generation model; The data label generation module is used to generate data labels based on a preset data label generation library and a preset data label generation model.

Citation Information

Patent Citations

  • Data label generation method and device, electronic equipment and readable storage medium

    CN111125177A

  • Intelligent event marking method and device and storage medium

    CN116992034A

  • Image label generation method, system, equipment and medium

    CN120220149A

  • Method and apparatus for tagging text based on adversarial learning

    US20200342172A1