Data labeling method and device, electronic equipment, storage medium and program product
By combining the labeling strategies of large models and professional models and adopting different labeling methods for data with different attributes, the problems of high cost and long cycle of manual labeling are solved, fast and accurate data labeling is achieved, and the quality and efficiency of data labeling are improved.
Patent Information
- Application Number
- CN202510725446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-23
AI Technical Summary
Manual labeling is costly and time-consuming, making it difficult to adapt to the model development needs of rapid iteration and precise services.
A labeling method that combines large models with professional models is adopted. The appropriate labeling model is selected according to data attributes and preset strategies to label business data. The labeling results are integrated through weighted averaging, voting calculation or inductive summary to improve the labeling quality.
It achieves fast and accurate data labeling, reduces the cost and cycle of manual labeling, and improves the efficiency and quality of data labeling.
Smart Images

Figure CN120687830A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a data annotation method, device, electronic device, storage medium, and program product. Background Art
[0002] With the rapid development of artificial intelligence technology, machine learning and deep learning models are widely used in many key fields, such as medical diagnosis, autonomous driving, financial risk control, and intelligent customer service. To adapt to different fields, effective training of machine learning and deep learning models is required.
[0003] Traditional model training requires manual labeling of training data to generate and conduct training. Manual labeling involves human experts classifying, labeling, and annotating raw data such as images, text, and speech to create structured training data that the model can understand.
[0004] Manual labeling is costly and time-consuming, making it difficult to adapt to the model development needs of rapid iteration and precise services.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0006] The present disclosure provides a data labeling method, device, electronic device, storage medium and program product, which at least to some extent overcome the problems of high cost and long cycle of manual labeling in related technologies.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a data labeling method is provided, including: acquiring business data; determining data attributes of the business data; selecting a large model and / or a professional model as a labeling model based on the data attributes and a preset labeling strategy; calling the labeling model to label the business data to obtain a labeling result of the business data.
[0009] In some possible embodiments, data attributes include data source and / or data type; based on data attributes and preset labeling strategies, selecting a large model and / or a professional model as the labeling model includes: when the data source is a set source and / or the data type is a set type, selecting a professional model as the labeling model; when the data source is a non-set source and the data type is a non-set type, selecting a large model and a professional model as the labeling model.
[0010] In some possible embodiments, when the annotation model includes a large model and a professional model, calling the annotation model to annotate the business data to obtain the annotation results of the business data includes: calling the large model to annotate the business data, and receiving a first annotation result output by the large model; calling the professional model to annotate the business data, and receiving a second annotation result output by the professional model; determining the annotation results of the business data based on the first annotation result and the second annotation result.
[0011] In some possible embodiments, determining the annotation result of business data based on the first annotation result and the second annotation result includes: if the similarity between the first annotation result and the second annotation result is greater than or equal to the set similarity threshold, performing a weighted average calculation on the first annotation result and the second annotation result to obtain the annotation result of the business data; if the similarity between the first annotation result and the second annotation result is less than the set similarity threshold, performing a voting calculation on the first annotation result and the second annotation result to obtain the annotation result of the business data.
[0012] In some possible embodiments, determining the annotation result of the business data based on the first annotation result and the second annotation result includes: calling a large model to summarize the first annotation result and the second annotation result; and receiving the annotation result of the business data output by the large model.
[0013] In some possible embodiments, it also includes: extracting data features of the target model; determining the format requirements of the labeled data based on the data features; accordingly, calling the labeling model to label the business data, and obtaining the labeling results of the business data includes: calling the labeling model to label the business data based on the format requirements, and obtaining the labeling results of the business data.
[0014] In some possible embodiments, extracting data features of the target model includes: reading training sample data of the target model using a data extraction script; and extracting data features of the training sample data using a cluster analysis algorithm.
[0015] In some possible embodiments, the method further includes: if a set keyword is extracted from the business data, searching the standard rule library for annotation results that match the set keyword; and using the annotation information that matches the set keyword as the annotation result of the business data.
[0016] According to another aspect of the present disclosure, a data labeling device is also provided, including: a data acquisition module for acquiring business data; an attribute determination module for determining data attributes of business data; a model selection module for selecting a large model and / or a professional model as a labeling model based on data attributes and preset labeling strategies; a data labeling module for calling the labeling model to label the business data and obtain the labeling results of the business data.
[0017] According to another aspect of the present disclosure, an electronic device is also provided, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above-mentioned data labeling methods by executing the executable instructions.
[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned data labeling methods is implemented.
[0019] According to another aspect of the present disclosure, a computer program product is further provided, including: a computer program or instructions, which implements any of the above-mentioned data labeling methods when the computer program or instructions are executed by a processor.
[0020] The data annotation method provided in the embodiments of the present disclosure obtains business data; determines the data attributes of the business data; selects a large model and / or a professional model as an annotation model based on the data attributes and a preset annotation strategy; calls the annotation model to annotate the business data, and obtains the annotation results of the business data. In this embodiment, an annotation method that combines a large model with a professional model is adopted, and different annotation strategies are adopted for data with different attributes. By leveraging the powerful learning and generalization capabilities of the large model and the professional model, data can be annotated quickly and accurately, improving the quality of data annotation and avoiding the high cost and long cycle of manual annotation in related technologies.
[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0023] Figure 1 A schematic diagram of an exemplary application system architecture to which the data annotation method in the embodiments of the present disclosure can be applied is shown;
[0024] Figure 2 A flow chart of a data annotation method according to an embodiment of the present disclosure is shown;
[0025] Figure 3 A flow chart of another data annotation method according to an embodiment of the present disclosure is shown;
[0026] Figure 4A flowchart of a method for fusing annotation results according to an embodiment of the present disclosure is shown;
[0027] Figure 5 A flowchart showing another method for fusing annotation results according to an embodiment of the present disclosure is shown;
[0028] Figure 6 A flow chart of another data annotation method according to an embodiment of the present disclosure is shown;
[0029] Figure 7 A flow chart of another data annotation method according to an embodiment of the present disclosure is shown;
[0030] Figure 8 A schematic diagram showing a data annotation architecture in an embodiment of the present disclosure is shown;
[0031] Figure 9 A schematic diagram showing a data labeling and model training method process in an embodiment of the present disclosure;
[0032] Figure 10 A schematic diagram showing a data classification in an embodiment of the present disclosure;
[0033] Figure 11 A schematic diagram showing a data annotation example in an embodiment of the present disclosure;
[0034] Figure 12 A schematic diagram of a data annotation device according to an embodiment of the present disclosure is shown;
[0035] Figure 13 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0037] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:
[0039] Call center platform: A channel platform for interactive communication between enterprises and customers. Customers use this platform to consult, make complaints, handle business, and other operations with enterprises through voice, text, video, etc. It is the source of customer interaction data.
[0040] AI (Artificial Intelligence) Gateway: A key device that connects the call center platform with the back-end AI platform and data analysis and processing systems. It is responsible for receiving data backflow from the call center platform and transmitting it to the corresponding modules for processing. It also plays an important role in data collection, caching, and preprocessing processes.
[0041] K-means clustering algorithm: A clustering algorithm used for text data feature analysis. After preprocessing text data (such as word segmentation and stop word removal), it extracts feature vectors, such as bag-of-words models, TF-IDF (term frequency-inverse document frequency) vectors, or word embedding vectors. By adjusting the clustering parameter (the number of clusters K), the feature vectors are clustered to discover potential data patterns and feature groupings, helping to understand the distribution and inherent structure of text data.
[0042] MFCC (Mel-Frequency Cepstral Coefficients) feature algorithm: This algorithm is used for cluster analysis of media (speech) data. It first calculates the Mel-Frequency Cepstral Coefficient (MFCC) features of the speech data, which effectively capture the spectral characteristics of speech. MFCC features are then clustered using a clustering algorithm (such as K-Means or hierarchical clustering) to identify groupings of different speech patterns (such as voice intonation, speaking rate, and speaker characteristics), providing information for model optimization.
[0043] Conflict Resolution - Voting: A solution used by the conflict resolution unit. The historical accuracy of the two methods on similar data is counted and weighted accordingly. The votes for the current data are calculated based on the weights, and the annotation result with the most votes is selected as the final result.
[0044] Conflict resolution - weighted average method: A solution method used by the conflict resolution unit. When the annotation results are in the form of probabilities, the final result is obtained by weighted averaging the two annotation results according to their weights.
[0045] Rule-based conflict resolution: Based on pre-configured rules, match the rules that meet the characteristics and determine the annotation results based on the rules.
[0046] Data fusion algorithm: The algorithms stored in the labeling strategy management unit, such as the random forest algorithm and the weighted average algorithm, are used for data fusion after multi-model labeling. The labeling results of different models are comprehensively processed to obtain more accurate and reliable final labeling results.
[0047] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0048] Figure 1 FIG. 1 shows an exemplary application system architecture diagram to which the data annotation method in the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture may include a terminal device 101 , a network 102 and a server 103 .
[0049] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103 , and can be a wired network or a wireless network.
[0050] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0051] The terminal device 101 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.
[0052] Optionally, the client of the application installed in different terminal devices 101 is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile phone client, a PC client, etc.
[0053] The server 103 may be a server that provides various services, such as a background management server that provides support for the devices operated by the user using the terminal device 101. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.
[0054] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0055] Those skilled in the art will know that Figure 1 The number of terminal devices, networks, and servers in the embodiment is merely illustrative, and any number of terminal devices, networks, and servers may be provided based on actual needs. This embodiment of the present disclosure does not limit this.
[0056] Under the above system architecture, an embodiment of the present disclosure provides a data labeling method, which can be executed by any electronic device with computing and processing capabilities.
[0057] In some embodiments, the data labeling method provided in the embodiments of the present disclosure can be executed by the terminal device of the above-mentioned system architecture; in other embodiments, the data labeling method provided in the embodiments of the present disclosure can be executed by the server in the above-mentioned system architecture; in other embodiments, the data labeling method provided in the embodiments of the present disclosure can be implemented by the terminal device and the server in the above-mentioned system architecture through interaction.
[0058] Figure 2 A flow chart of a data annotation method according to an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the data labeling method provided in the embodiment of the present disclosure includes the following steps: S202-S208.
[0059] S202: Obtain business data.
[0060] Business data refers to the structured or unstructured data generated, collected or used by an enterprise in the course of daily operations around various business activities. For example: order data, logistics data, complaint data, etc. In this embodiment, the business data includes the return data fed back by the call platform as an example. Return data is a special type of data in the field of business data, which is related to the "reflow" of user behavior and business processes. For example, return data refers to the interaction data generated when a user returns after leaving a business scenario or product, or the data that is reconnected after leaving the original path in the business process.
[0061] In one possible implementation scenario, the device that implements the data labeling method of this embodiment is called a data labeling device. This data labeling device is deployed on the AI gateway to effectively connect the call platform with the back-end AI platform and data analysis and processing system. When the service data of the call platform flows back to the AI gateway, the data labeling device provided in this embodiment is activated to collect, cache, and pre-process the data; it also intelligently labels the service data; and then, during the model training process, the model parameters and structure are continuously optimized based on the labeled service data to enable it to better adapt to actual service needs.
[0062] Collect business data from the call center. The AI gateway captures the return data fed back by the call center as business data. For example, the call center pushes the return data to the intelligent platform via the AI gateway. The data annotation device is activated, and the data collection module in the data annotation device collects the return data in real time.
[0063] In one possible implementation, each piece of business data carries a tag that indicates how the data was generated. Exemplary tags include, but are not limited to, intelligent interaction, manual service, and traditional keystrokes. Based on the tags, the business data is initially categorized, grouping business data with the same tag into the same category. Categorizing business data by tag effectively improves the annotation efficiency of business flow data and allows for faster identification of data sources, data types, and scenario types.
[0064] In one possible implementation, data cleaning is performed on the classified business data, including but not limited to: format conversion, data deduplication, data noise reduction, etc.
[0065] Format conversion refers to converting business data of different types and formats into training samples; for example, voice data in formats such as MP3, WAV, and AMR, and text data in different encoding formats and text formats are converted into the format required for training samples.
[0066] Data deduplication includes: using real-time data cleaning and incremental data cleaning methods to clean data duplication based on call flow ID, and performing data labeling on the cleaned data.
[0067] Data noise reduction, including: deleting unreadable business data and business data with error codes.
[0068] In one possible implementation, the business data is analyzed and identified by type, category, and style. Based on the sample data's ID, the archive database is checked to see if the data exists. If so, the annotation results are extracted and written into auxiliary annotation fields. Feature recognition is used to write the sample data into a cache using the annotation style's wide table format. The sample data is traversed to determine if there are duplicates based on the unique identifier, service number, and call summary tag, and any duplicates are deleted. During the traversal process, key fields of the data to be annotated, such as the service number, call summary (content summary), call summary tag, and detailed conversation content, are determined to be missing. If so, the data is deleted as invalid.
[0069] S204: Determine the data attributes of the business data.
[0070] Data attributes are labels that describe data characteristics and are used to clarify the basic characteristics, data type, data source, data type, etc. In business scenarios, the determination of data attributes is the basis for data management and data labeling.
[0071] Exemplarily, data attributes include at least one or more of the following: data source and data type. Data source can be understood as the collection channel, generation method, or acquisition method of business data, answering the question of "where does the data come from?" Data type refers to whether business data involves sensitive information based on the nature of the business and data content, answering the question of "whether the data requires special protection."
[0072] Determining the data source of business data includes: adding channel identifiers to business data through technical means when the data is generated or collected, and after collecting the business data, determining the data source of the business data based on the channel identifiers in the business data.
[0073] Determining the data type of the business data includes: after collecting the business data, determining whether the business data includes sensitive information to determine the data type of the business data.
[0074] S206: Based on data attributes and a preset annotation strategy, a large model and / or a professional model is selected as an annotation model.
[0075] The preset labeling strategy can be understood as the rules set in advance before business data labeling, which are used to guide the selection of labeling models and the execution of labeling tasks.
[0076] A large model refers to a large-scale pre-trained model based on deep learning. It has strong generalization and multi-task processing capabilities, a large number of parameters, and can complete multiple tasks with a small number of samples or prompts. For example: a common natural language processing model. A professional model refers to a model optimized for a specific field, task, or data attribute, and usually has higher accuracy or efficiency on a single task. For example, a professional model can refer to a professional model built for a specific field or problem in the call service. Professional models include but are not limited to: decision trees, vector professional models, etc. A labeling model can be understood as a model used to label business data.
[0077] Based on the data attributes of the business data and combined with the pre-set labeling strategy, decide whether to use a general large model for data labeling, a professional model for a specific field to represent the data, or a combination of large models and professional models to perform data labeling.
[0078] S208: Call the annotation model to annotate the business data and obtain the annotation results of the business data.
[0079] In this embodiment, calling refers to starting the annotation model and transferring business data to the annotation model through an interface, SDK, or local function, so that the annotation model can annotate the business data. Annotation refers to the process of adding labels to business data. Annotation results are labels added to business data, which may include but are not limited to: classification results, such as "positive review" and "negative review"; entity recognition, such as "XX" is a place name and "**" is a person's name; sentiment analysis, such as "angry" and "happy"; and relationship extraction, such as "**-works-at-company-A".
[0080] The business data is input into a selected annotation model, which analyzes and labels the business data, and receives the annotation results of the business data output by the annotation model.
[0081] After obtaining the annotation results of the business data, the model training module is used to perform model training. After the model training task is completed, the annotated data is stored in the archive library for traceability.
[0082] The data annotation method provided in the embodiments of the present disclosure obtains business data; determines the data attributes of the business data; selects a large model and / or a professional model as an annotation model based on the data attributes and a preset annotation strategy; calls the annotation model to annotate the business data, and obtains the annotation results of the business data. In this embodiment, an annotation method that combines a large model with a professional model is adopted, and different annotation strategies are adopted for data with different attributes. By leveraging the powerful learning and generalization capabilities of the large model and the professional model, data can be annotated quickly and accurately, improving the quality of data annotation and avoiding the high cost and long cycle of manual annotation in related technologies.
[0083] Based on the above embodiment, the embodiment of the present disclosure further optimizes the data labeling method. The optimized data labeling method includes the following steps:
[0084] S302: Obtain business data.
[0085] S304: Determine data attributes of the business data, where the data attributes include data source and / or data type.
[0086] Data sources can be understood as the collection channels, generation methods, or acquisition methods of business data, answering the question of "where does the data come from?" Data types refer to whether business data involves sensitive information based on the nature of the business and data content, answering the question of "whether the data requires special protection."
[0087] S306: When the data source is a set source and / or the data type is a set type, select a professional model as the annotation model.
[0088] A defined source is a predefined, focused data source, typically with specific content. Data from a defined source often requires stricter processing policies, such as using specialized models or prohibiting uploads to the cloud.
[0089] The set type means that the business data is sensitive data or includes sensitive data and requires a more secure processing method.
[0090] A rule base is preset, which includes several annotation strategies. Each annotation strategy describes the type of annotation model that should be used for a specific data source or data type.
[0091] After acquiring business data, the system analyzes its source and content type, such as whether it contains sensitive information. If the business data is found to come from a pre-defined source or is identified as a sensitive data type, a specialized model is automatically selected for annotation.
[0092] S308: Call the professional model to annotate the business data and obtain the annotated results of the business data.
[0093] The business data is input into the professional model so that the professional model labels the business data, and the labeling results of the business data output by the professional model are received.
[0094] In one possible implementation, the annotation results of the business data are encrypted twice to improve the security of the business data.
[0095] S310. When the data source is a non-set source and the data type is a non-set type, a large model and a professional model are selected as annotation models.
[0096] When the data source is non-set and the data type is non-set, it indicates that the business data is regular data and does not require special processing. At the same time, select the large model and the professional model as the annotation model to annotate the business data.
[0097] The generalization capabilities of large models and the domain precision of specialized models complement each other. Large models address the question of "can the data be processed?" while specialized models address the question of "is the processing accurate?" This allows for efficient and accurate labeling in standardized data scenarios.
[0098] S312: Call the big model to annotate the business data, and receive the first annotation result output by the big model.
[0099] The first labeling result refers to the output of the large model after labeling the business data. It has a certain generalization ability, but may have errors in details or professional fields.
[0100] Preprocess the business data into a format acceptable to the big model, such as a text string, call the big model through the API (Application Programming Interface) interface or the locally deployed model service, and receive the first annotation result returned by the big model, for example, the text sentiment is labeled as "positive" and the image object is labeled as "truck".
[0101] S314: Call the professional model to annotate the business data, and receive the second annotation result output by the professional model.
[0102] The second labeling result refers to the output of a professional model after labeling business data. It is more accurate in a specific field, but may not be able to process data beyond its training range.
[0103] Preprocess the business data into a format acceptable to the professional model, such as a text string, call the professional model through the API interface or a locally deployed model service, and receive the second annotation result returned by the professional model.
[0104] S316: Determine the annotation result of the business data according to the first annotation result and the second annotation result.
[0105] The first annotation result output by the large model and the second annotation result output by the professional model are integrated to determine the final annotation result of the business data, balancing generalization ability and professional accuracy.
[0106] When using large models and professional models to label data simultaneously, the final labeling results are determined through voting or weighted average methods. If, during the result merging process, there are large differences in the labeling results and the final results cannot be determined through voting or weighted average methods, the conflict resolution mechanism is activated.
[0107] When automatic labeling fails or there is a labeling conflict, the labeling data will be recorded and pushed to the manual assisted processing module, allowing operators to quickly intervene and resolve the labeling problem through manual processing.
[0108] If the aforementioned algorithms and strategies fail to resolve the labeling result fusion problem, the business data, along with the first and second labeling results, are pushed to the manual processing module, along with a prompt notification for manual intervention. This notification should include key data information, labeling conflicts, and resolution requirements, allowing manual reviewers to quickly identify the issue and accurately label the data. After the manual review is complete, the results are fed back to the system for updating the labeling results and optimizing the model strategy.
[0109] In this embodiment, the large model and the professional model are used in collaboration. With their powerful learning and generalization capabilities, they can quickly and accurately label the data and improve the quality of data labeling.
[0110] On the basis of the above embodiment, the specific implementation method of "determining the annotation result of the business data according to the first annotation result and the second annotation result" is optimized in this embodiment. The optimized annotation result determination process includes steps S402-S408.
[0111] S402: Calculate the similarity between the first annotation result and the second annotation result.
[0112] Similarity measures the consistency between the first and second annotation results, and is used to determine whether the annotations of the large model and the specialized model for the same business data are converging. A higher similarity indicates a smaller difference in annotations between the two models.
[0113] When the annotation result is text data, the semantic similarity between the first annotation result and the second annotation result is calculated. Optionally, the cosine similarity between the first annotation result and the second annotation result is calculated using a natural language processing model, or the overlap between the first annotation result and the second annotation result is calculated by keyword matching.
[0114] When the annotation results are image annotations, the overlap of the object bounding boxes and the consistency of the classification labels are compared. For example, if the labels are both "car," the similarity is 100%; otherwise, the similarity is 0%. For numerical or enumerated value differences, for example, if the label output by the large model is "medium" and the label output by the specialized model is "medium-high," the level difference can be set to ≤ 1; the similarity is 80%.
[0115] S404: Determine whether the similarity between the first annotation result and the second annotation result is greater than or equal to a set similarity threshold. If so, execute S406; if not, execute S408.
[0116] Setting a similarity threshold refers to a pre-set judgment standard. For example, the similarity threshold is set to 80%. The similarity threshold is used to determine how to fuse the first annotation result and the second annotation result.
[0117] If the similarity between the first annotation result and the second annotation result is greater than or equal to the set similarity threshold, it is considered that the annotation differences between the large model and the professional model are relatively consistent; if the similarity between the first annotation result and the second annotation result is lower than the threshold, the annotation differences between the first annotation result and the second annotation result are large.
[0118] In a possible implementation, the similarity threshold may be dynamically adjusted according to actual application scenarios.
[0119] S406: If the similarity between the first annotation result and the second annotation result is greater than or equal to the set similarity threshold, a weighted average calculation is performed on the first annotation result and the second annotation result to obtain an annotation result of the business data.
[0120] Weighted average calculation refers to assigning weights to the annotation results of large models and professional models according to their credibility, and generating the final annotation results through weighted summation.
[0121] The weights are determined based on the performance of the large and specialized models on the validation set. For example, if the large model's validation set accuracy is 80% and the specialized model's is 60%, the large model's weight is 0.8 and the specialized model's weight is 0.6. For probabilistic annotation results, the final result is obtained by taking the weighted average of the two annotation results.
[0122] The model weight represents the credibility or importance of the model's annotation results. A higher value indicates a more reliable model. For example, the weight of the large model is 80%, and the weight of the specialized model is 60%, indicating that the large model has a higher overall credibility than the specialized model in this scenario.
[0123] In one possible implementation, the annotation results of the larger model with the higher weight are prioritized as the basis. Based on the first annotation result output by the larger model, the second annotation result of the specialized model is referenced to correct the divergent parts. However, the degree of correction is affected by the weight difference. The larger the weight difference, the smaller the influence of the specialized model.
[0124] In one possible implementation, the business data is: ["I", "love", "nature", "language", "processing"]; the first labeling result output by the large model is: ["pronoun", "verb", "noun", "noun", "noun"]; the second labeling result output by the specialized model is: ["pronoun", "verb", "adjective", "noun", "verb"]. The labeling results for "I", "love", and "language" are identical, so the first labeling result output by the large model is directly used as the final labeling result. The labeling results of the "nature" and "processing" models are different, and the large model's weight is 20% higher than that of the specialized model. Therefore, the opinions of the specialized models have limited influence on the final labeling result and are only partially integrated.
[0125] Specifically, the large-scale model considers the "natural" label a "noun," while the specialized model considers it an "adjective." Using the large-scale model's "noun" as the core and referencing the specialized model's "adjective," we might generate "noun (can be used as an adjective)." This means retaining the core label and supplementing other possible labels. The final weighted average result is ["pronoun," "verb," "noun (can be used as an adjective)," "noun," and "noun / verb"].
[0126] In this embodiment, weighted calculations focus on semantic fusion. The core goal is to retain the subjective judgment of high-weight models while selectively absorbing the reasonable disagreements of low-weight models. By assigning different credibility weights to different models, we maximize the reasonable opinions of multiple models while ensuring annotation efficiency, improving the efficiency and quality of data annotation.
[0127] S408: If the similarity between the first annotation result and the second annotation result is less than the set similarity threshold, the first annotation result and the second annotation result are voted to obtain an annotation result of the business data.
[0128] If the similarity between the first and second annotation results is lower than the threshold, then the annotation differences between the first and second annotation results are large. In this case, a voting mechanism is used to determine the final annotation result. The essence of this mechanism is to reduce the risk of misjudgment of a single model through the principle of "majority rule" or "authority priority".
[0129] When the labeling results of the large model and the professional model conflict, the credibility is measured by the historical accuracy of the large model and the professional model on similar data, and dynamic weights are assigned to the large model and the professional model. The higher the accuracy, the greater the weight. Ultimately, the more reliable labeling result is selected through weighted voting.
[0130] In one possible implementation, historically annotated data of the same data type and similar domain as the current business data is selected. For example, if the current business data is "call text," the annotation performance of the two models on this historical call text is then tallied. The accuracy of the two models on this historical call text is calculated, with the accuracy being the ratio of the number of correct annotations to the total number of annotations. For example, if the large model correctly annotates 70 out of 100 call texts, the accuracy is 70%, while the specialized model correctly annotates 60 times, giving it an accuracy of 60%. Each time an annotation is completed, the accuracy is automatically recorded, and the accuracy is recalculated after N cumulative annotations.
[0131] The weight of the model is determined based on the historical accuracy of the model. For example, if the historical accuracy of the large model is 70%, the weight of the large model is 0.7; if the historical accuracy of the professional model is 60%, the weight of the professional model is 0.6.
[0132] The voting calculation process involves comparing the weights of the two models and selecting the labeling result with the higher weight as the final labeling result. For example, the first labeling result output by the large model is "Label A," and the second labeling result output by the specialized model is "Label B." Since the weight of the large model is 0.7 and the weight of the specialized model is 0.6, 0.7 > 0.6, the first labeling result output by the large model, "Label A," is selected as the final labeling result for the business data.
[0133] When the annotation results of large models and specialized models are highly similar, a weighted average method is used to combine the strengths of both. Large models excel at generalized understanding, while specialized models are intensive in domain detail. Weighted averaging balances comprehensiveness and expertise. When the annotation results of large models and specialized models are less similar, a voting method is used, using historical accuracy as a weight. This gives models with more reliable historical performance greater weight, avoiding ambiguous conclusions caused by simple compromises.
[0134] On the basis of the above embodiment, the specific implementation method of “determining the annotation result of the business data according to the first annotation result and the second annotation result” is optimized in this embodiment, and the optimized annotation result determination process includes steps S502-S504.
[0135] S502: Call the large model to summarize the first annotation result and the second annotation result.
[0136] Summarization involves analyzing and refining the first and second annotation results, extracting commonalities or integrating differences to form a more comprehensive and unified annotation conclusion. For example, "positive sentiment" and "good product reviews" can be summarized as "positive reviews."
[0137] A large model refers to a large model used for induction and summarization, which can include a large language model or a multimodal model. It has the ability to process complex semantics and generate structured outputs, and can understand the annotation logic of different models and perform high-level abstraction.
[0138] The input data for the large model includes the first and second annotation results. The large model's processing logic includes identifying differences between the first and second annotation results and retaining consistent items. For divergent items, the model analyzes semantic associations, merges similar labels, and selects a more reasonable single label based on the large model's general knowledge or prompts.
[0139] In one possible implementation, the input data of the large model includes: a first annotation result, a second annotation result, and a prompt word. For example, the input data of the large model is: "You need to integrate the annotation results of the two models. The consistent parts are directly retained, and the inconsistent parts are correctly labeled according to grammatical rules. The first annotation result: [XXX], the second annotation result: XXX]. ".
[0140] In one possible implementation, for annotation results with more complex or ambiguous semantics, more accurate annotation results can be given by re-summarizing the content of the large model and combining it with the powerful semantic understanding ability of the large model.
[0141] S504: Receive the annotation results of the business data output by the large model.
[0142] Receive the inference results output by the large model and directly use them as the final annotation results of the business data without manual intervention or additional model voting.
[0143] The large model summarizes the first and second annotation results, leveraging its general intelligence to resolve semantic divergences or redundancies between models. This eliminates the need for manual design of complex fusion rules, as the large model automatically performs semantic summarization, improving processing efficiency. This approach is suitable for scenarios with complex annotation logic that require high-level semantic understanding.
[0144] On the basis of the above embodiment, the embodiment of the present disclosure further optimizes the data labeling method. The optimized data labeling method includes steps: S602-S612.
[0145] S602: Extract data features of the target model.
[0146] The target model is the model that needs to be trained using labeled business data. Data features are extracted from the target model's existing training sample data and are information dimensions that reflect the essential patterns of the business data or are relevant to the target model.
[0147] Extract the data features of the target model for identification, so that the data format and data arrangement of the training data are more compatible with the trained model, thereby improving the availability of the training data.
[0148] In one possible implementation, extracting data features of the target model includes steps S6021 - S6022 .
[0149] S6021. Use the data extraction script to read the training sample data of the target model.
[0150] Use a data extraction script to read sample data from the target model and perform appropriate parsing and processing based on the data format. A data extraction script is a pre-written code snippet or a script extraction tool. The sample data is training data or validation data for optimizing machine learning or deep learning models.
[0151] S6022. Use a cluster analysis algorithm to extract data features of the training sample data.
[0152] When the training sample data is text data, the K-Means clustering algorithm is used to conduct an in-depth analysis of the text data's features. First, the text data is preprocessed, including operations such as word segmentation and stop word removal. Feature vectors (such as bag-of-words models, TF-IDF vectors, or word embedding vectors) are then extracted. Next, the K-Means algorithm is used to cluster the feature vectors. By adjusting clustering parameters (such as the number of clusters, K), potential data patterns (data patterns refer to the structured definition of data, describing its organization, field (column) names, data types, constraints, and relationships) and feature groupings are discovered to better understand the distribution and inherent structure of the text data. Data patterns refer to the structured definition of data, describing its organization, field (column) names, data types, constraints, and relationships. Feature grouping involves categorizing features (fields) in the raw data according to certain logic or business rules to form feature sets with specific meanings.
[0153] When the training sample data is media data, the MFCC feature algorithm is combined with cluster analysis. First, the MFCC features of the speech data are calculated. These features effectively capture the spectral characteristics of speech. Then, clustering algorithms (such as K-Means or hierarchical clustering) are used to cluster the MFCC features. This allows the identification of groupings of different speech patterns, such as different intonations, speaking rates, and speaker characteristics, providing valuable information for subsequent model optimization.
[0154] When training sample data is image data, images with similar content are clustered together by extracting visual features such as SIFT (Scale-invariant feature transform) local features, HOG (Histogram of Oriented Gradients) shape features, or CNN (Convolutional Neural Networks) high-level semantic features. Images are preprocessed, normalized, and denoised. Feature vectors are then extracted using traditional hand-crafted features or deep learning models. Finally, clustering is performed using algorithms such as K-Means, hierarchical clustering, or DBSCAN (Density-Based Spatial Clustering of Applications with). This approach is suitable for scenarios such as image retrieval and automatic annotation, helping to discover image themes, object types, or scene patterns.
[0155] When the training sample data is video data, clustering is performed by combining the spatial (image frame) and temporal (motion) information of the video. The video is first broken down into key frames and spatiotemporal features, such as optical flow motion features and 3D (3-dimensional) CNN spatiotemporal features, are extracted. The sequence data is then processed using Dynamic Time Warping (DTW), Hidden Markov Models (HMM), or hierarchical clustering to group action patterns or scene types. This is commonly used for video recommendation, action recognition preprocessing, and video content analysis, solving the problems of spatiotemporal alignment and semantic modeling of long sequences.
[0156] Data extraction scripts can quickly read batches of data from complex data sources, avoiding errors caused by manual operations and significantly reducing data preprocessing time. Clustering algorithms can be used to discover implicit features in sample training data, providing more suitable training data for the target model.
[0157] S604: Determine the format requirements of the labeled data based on the data characteristics.
[0158] Among them, format requirements include annotation samples, annotation examples, annotation rules, etc.
[0159] Each model has its own data model characteristics, including its table header, wide table fields, and data format. Through the initial training samples of the target model and the execution results of the model, the format requirements of the annotation data can be simulated to extract the annotation style.
[0160] By extracting the features of the target model, the format requirements of the input data of the target model are constructed, which is used to perform auxiliary model optimization training tasks after sample annotation.
[0161] S606: Obtain business data.
[0162] S608: Determine the data attributes of the business data.
[0163] S610: Based on data attributes and a preset annotation strategy, a large model and / or a professional model is selected as an annotation model.
[0164] S612: Call the annotation model to annotate the business data based on the format requirements to obtain the annotation results of the business data.
[0165] The connected big model is called according to the configured preset annotation strategy, and the annotation samples, annotation examples, and annotation rules are pushed to the big model to perform data annotation. After the annotation is completed, the first annotation result output by the big model is obtained, including data labels, data summary content, and data analysis content;
[0166] According to the configured strategy and data type, the corresponding professional model is called to perform data annotation, and the first annotation result output by the professional model is obtained, including data labels, data summary content and data analysis content.
[0167] In this embodiment, the target model features are automatically identified and the training sample style is automatically compiled according to the target model requirements, ensuring that the annotation results can meet the training requirements of the target model.
[0168] On the basis of the above embodiment, the embodiment of the present disclosure further optimizes the data labeling method. The optimized data labeling method includes steps: S702-S706.
[0169] S702: Obtain business data.
[0170] S704: If the set keyword is extracted from the business data, query the standard rule library for the annotation results that match the set keyword.
[0171] Preset keywords are predefined words or phrases with specific business meanings that trigger rule matching. For example, in e-commerce scenarios, these keywords are often directly related to business needs or annotation objectives.
[0172] A standard rule base is a database or knowledge base that stores the mapping between keywords and annotation results. It includes keywords, annotation results, matching rules, and more. Rules are developed based on business knowledge and experience. If business data contains "quality issues" and mentions specific product models, it will be labeled as a "product quality complaint" first.
[0173] Business data is segmented to obtain multiple keywords within the data. Using these keywords as indexes, corresponding annotation results are retrieved from the standard rule library. For example, the keyword "quality issue" corresponds to the annotation result "product quality complaint," with a "high" processing priority.
[0174] S706: The annotation information matching the set keyword is used as the annotation result of the business data.
[0175] The labeling result refers to the label or conclusion generated based on the keyword matching rules. For example, business data containing "quality issues" is labeled as "product quality complaints-refund related".
[0176] The matching annotation information in the rule base is directly assigned to the business data to complete the annotation.
[0177] For data with specific keywords or clear semantic patterns, we prioritize annotation results that conform to the rules rather than simply relying on model voting or weighting. When model annotation results conflict with the rules, the rules will prevail in determining the annotation results.
[0178] In this embodiment, the standard rule library predefines unique labeling results corresponding to keywords to avoid labeling differences caused by human subjective judgment or model fluctuations, and ensure consistent labeling results for similar data.
[0179] In this embodiment, the data annotation method is applied to the traffic data scenario as an example. Figure 8 As shown, this embodiment proposes a device for traffic platform data annotation and model optimization training, such as Figure 8As shown, data annotation device 811 is deployed in AI gateway 810, and AI gateway 10 is deployed in gateway server 800. Gateway server 800 also includes CPU 820, GPU 830, memory 840, and storage device 850. Data annotation device 811 mainly includes data acquisition module 8111, data preprocessing module 8112, intelligent annotation module 8113, manual auxiliary processing module 8114, data storage module 8115 and model-assisted training module 8116.
[0180] The data annotation device 811 can effectively connect the call center platform 860, the customer service operation and management system 861, the big data lake 862, and other connections. In addition, it can also interact with the big model 870, the professional model 880, and other professional algorithms 890. When the business data of the call center platform 860 flows back to the AI gateway 810, the data annotation device 811 is started to collect, cache, and pre-process the business data. The data is intelligently annotated through the intelligent annotation module; then, during the model training process, the data annotation device 811 continuously optimizes the parameters and structure of the model based on the annotated data, so that it can better adapt to actual business needs.
[0181] The data collection module 8111 collects interactive data by connecting to the traffic platform 860 and the data analysis system. The data collection module 8111 mainly includes two control units: interface management and data storage.
[0182] The data preprocessing module 8112 is a common module for all data processing devices and is mainly used to preprocess the collected data, including data noise reduction, data cleaning and format conversion.
[0183] The intelligent annotation module 8113 is the core module and includes: an annotation strategy management unit, a model docking unit, and a conflict resolution unit. The specific units are described below.
[0184] The annotation strategy management unit is used to store annotation strategies and data fusion algorithms. Initialization involves manually entering annotation strategies and importing annotation tools. Stored data fusion algorithms include random forests and weighted averages, which are used to fuse data after multi-model annotation.
[0185] The model docking unit is used to connect to large models and schedule specialized algorithms. Based on the configured annotation strategy, it schedules large models or specialized algorithms to perform annotation.
[0186] Different business types and data sources use specific labeling strategies. The operator configures the labeling strategy. The system uses business characteristics, data characteristics, and the source of the returned data to identify the scenario type / data type / scenario+data type, and then selects the specified strategy.
[0187] The conflict resolution unit, when using multi-model collaborative annotation, if there are large differences in the annotation results during the result merging process and the final result cannot be determined by voting or weighted average method, the conflict resolution mechanism is activated.
[0188] The manual assisted processing module records the annotation data when the device fails to automatically label or there is a labeling conflict, and pushes it to the manual assisted processing module, allowing operators to quickly intervene and resolve the labeling problem through manual processing.
[0189] The data storage module 8115 is used to store the collected data and the labeled data, and store the pre-processed and labeled data in the database according to a predetermined storage strategy.
[0190] The model-assisted training module 8116 is mainly composed of a model training unit and a model evaluation unit. The main function of the module is to assist the training of large models and professional models, provide high-quality training sample data, execute training scripts and evaluate training results.
[0191] Based on the above embodiment, a process for processing business data is provided, such as Figure 9 As shown, it mainly includes the following steps.
[0192] S902: Reflux data collection.
[0193] The call center platform pushes the return data to the intelligent platform through the AI gateway, the data annotation device is activated, and the data collection module collects the return data in real time.
[0194] S904: Data analysis and classification.
[0195] The data preprocessing module preliminarily classifies the returned data according to the label attributes (intelligent interaction, manual service, traditional button) of the returned data.
[0196] like Figure 10 As described, the reflow data is classified into text records, voice records and video records according to the data classification algorithm, and the text records are divided into intelligent interaction and manual service according to the labels carried by the data. The voice records are converted into text, and the converted text is divided into intelligent interaction and manual service. The video records are converted into text, and the converted text is divided into intelligent interaction and manual service.
[0197] De-duplication and noise reduction are performed on the classified business data.
[0198] Preliminary classification according to labels can effectively improve the labeling efficiency of returned data and more quickly identify data sources, data types, and scenario types.
[0199] S906: Data cleaning.
[0200] The data preprocessing module performs data cleaning on the classified data, including format conversion, data deduplication and data noise reduction.
[0201] Format conversion: Traffic data of different types and formats, such as voice data (MP3, WAV, AMR, etc.) and text data (different encoding formats, text formats, etc.), are converted into training samples.
[0202] Data deduplication: Through real-time data cleaning and incremental data cleaning, data duplication is removed based on the call flow ID; the cleaned data is pushed to the annotation module for data annotation.
[0203] Data noise reduction: Delete unreadable data and data with error codes.
[0204] S908: Call the large model.
[0205] Use a model such as the Transformer architecture to perform annotation of the reflow data and cache the annotation results in the server memory.
[0206] S910. Call the professional model.
[0207] Build professional models for specific areas or problems in the call traffic business, such as decision trees and vector professional models, to annotate the return data;
[0208] S912. Result integration and conflict resolution.
[0209] After the large model and the professional model have labeled the same data, the labeling results of the two are fused. When the content is judged to be similar, the fusion is performed through weighted averaging. When the labeling results of the two are different, the voting algorithm, preset weight judgment, etc. are used to perform content fusion.
[0210] When both the large model and the professional model cannot recognize the same data or the annotation result cannot be selected during conflict resolution, the annotation content is written into the auxiliary processing module and a notification is pushed to allow operators to intervene and handle the matter.
[0211] S914. Record and push notification.
[0212] S916, model optimization training.
[0213] After labeling the sample data, model training is performed through the model-assisted training module.
[0214] S918. Archive the marked data.
[0215] After completing the model training task, the labeled data is stored in the archive library for traceability.
[0216] In a possible implementation, a specific implementation of a data annotation module is provided, such as Figure 11 As shown in the figure, the data annotation module task execution steps are as follows:
[0217] S1102: Feature extraction of target model.
[0218] Clarify the interface specifications between the call center platform and the AI platform, including the interface type, transmission protocol, and interface calling method and permission settings. At the same time, determine the connection method between media streams such as voice streams and video streams and data. Among them, interface types include but are not limited to: RESTful (Representational State Transfer) API, RPC (Remote Procedure Call), etc., and transmission protocols include but are not limited to: HTTP (Hypertext Transfer Protocol), HTTPS (Hypertext Transfer Protocol Secure), etc.
[0219] Using data extraction scripts, we can read the sample data of the model and perform corresponding analysis and processing according to the data format. Using cluster analysis algorithms, we can extract the data features of the sample data of the target model training.
[0220] S1104: Obtain pre-processed reflux data.
[0221] Data deduplication and noise reduction. The sample data is traversed, and duplicates are determined based on unique identifiers, service numbers, and call summary tags. Any duplicates are then deleted. During the traversal process, key fields of the data to be annotated, such as service numbers, call summaries (content summary), call summary tags, and detailed conversation content, are determined to be missing. If missing, the data is deleted as invalid.
[0222] S1106. Data source analysis.
[0223] Analyze source data and identify data types, categories, patterns, etc.
[0224] S1108. Assessment of historical annotation status.
[0225] Based on the ID (Identity document) of the sample data, check whether the archived data exists in the archive library; if so, extract the annotation results and write them into the auxiliary annotation field.
[0226] S1110. Labeling model selection and weight configuration.
[0227] Based on the sample data, match the strategy of the annotation strategy management unit and determine the compilation model combination to perform the annotation task.
[0228] For general data, a combination of large and specialized models is used for labeling. For data from a specific source or type, a specific specialized model is used for labeling, and the information is then re-encrypted.
[0229] Inject the selected model, the model weight configuration, and the processed business data into the annotation module.
[0230] like Figure 11 As shown, the annotation module first calls the large model annotation unit. The large model annotation unit calls the connected large model according to the configured strategy, pushes the annotation samples, annotation examples, and annotation rules to the large model, and performs data annotation. After the annotation is completed, the first annotation result is obtained, including data labels, data summary content, and data analysis content;
[0231] The annotation model is then annotated by a professional model annotation unit; the professional model annotation unit calls the corresponding algorithm to perform data annotation based on the configured strategy and data type, and obtains the second annotation result, including data labels, data summary content and data analysis content.
[0232] The first annotation result and the second annotation result are pushed to the conflict resolution unit. The conflict resolution unit performs content merging according to the weights of the configured large model and professional model and the conflict resolution strategy. When a content conflict is found, the conflict resolution is performed.
[0233] If the conflict resolution unit is unable to resolve the labeling result fusion using the aforementioned algorithms and strategies, the data record and existing labeling results are pushed to the manual processing module, along with a prompt notification for manual intervention. This notification should include key data information, labeling conflict details, and resolution requirements, allowing manual reviewers to quickly understand the issue and accurately label. After the manual review is complete, the results are fed back to the system for updating the labeling results and optimizing the model strategy.
[0234] This embodiment achieves a complete closed loop from data generation, data application, and data feedback. The vast amount of customer interaction data generated by the call center platform can be efficiently utilized, eliminating the reliance on traditional, inefficient manual labeling methods and improving the efficiency of data application. By synergizing large and specialized models, with their powerful learning and generalization capabilities, data can be quickly and accurately labeled, improving data labeling quality.
[0235] It should be noted that the acquisition, storage, use, and processing of data in the technical solution disclosed herein are in compliance with the relevant provisions of national laws and regulations. Various types of data such as personal identity data, operation data, behavioral data, etc. related to individuals, customers, and groups obtained in the embodiments of the present disclosure have been authorized.
[0236] Based on the same inventive concept, the present disclosure also provides a data tagging device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0237] Figure 12 A schematic diagram of a data annotation device according to an embodiment of the present disclosure is shown in FIG. Figure 12 As shown, the device includes: a data acquisition module 1210, an attribute determination module 1220, a model selection module 1230 and a data annotation module 1240.
[0238] Among them, the data acquisition module 1210 is used to acquire business data; the attribute determination module 1220 is used to determine the data attributes of the business data; the model selection module 1230 is used to select large models and / or professional models as annotation models based on data attributes and preset annotation strategies; the data annotation module 1240 is used to call the annotation model to annotate the business data and obtain the annotation results of the business data.
[0239] In some possible embodiments, data attributes include data source and / or data type; the model selection module 1230 is specifically used to select a professional model as the annotation model when the data source is a set source and / or the data type is a set type; when the data source is a non-set source and the data type is a non-set type, select a large model and a professional model as the annotation model.
[0240] In some possible embodiments, when the annotation model includes a large model and a professional model, the data annotation module 1240 includes: a model calling unit, used to call the large model to annotate the business data, and receive a first annotation result output by the large model; calling the professional model to annotate the business data, and receive a second annotation result output by the professional model; and an annotation result processing unit, used to determine the annotation result of the business data based on the first annotation result and the second annotation result.
[0241] In some possible embodiments, the annotation result processing unit is specifically used to perform a weighted average calculation on the first annotation result and the second annotation result to obtain the annotation result of the business data if the similarity between the first annotation result and the second annotation result is greater than or equal to the set similarity threshold; if the similarity between the first annotation result and the second annotation result is less than the set similarity threshold, perform a voting calculation on the first annotation result and the second annotation result to obtain the annotation result of the business data.
[0242] In some possible embodiments, the annotation result processing unit is specifically used to call the large model to summarize the first annotation result and the second annotation result; and receive the annotation results of the business data output by the large model.
[0243] In some possible embodiments, it also includes: a format requirement determination module for extracting data features of the target model; determining the format requirements of the labeled data based on the data features; a data labeling module 1240, specifically for calling the labeling model to label the business data based on the format requirements to obtain the labeling results of the business data.
[0244] In some possible embodiments, the format requirement determination module is specifically configured to read the training sample data of the target model using a data extraction script; and extract data features of the training sample data using a clustering analysis algorithm.
[0245] In some possible embodiments, the data annotation module 1240 is also used to query the annotation results that match the set keywords in the standard rule library if the set keywords are extracted from the business data; and use the annotation information that matches the set keywords as the annotation results of the business data.
[0246] It should be noted that the examples and application scenarios implemented by the modules in the above-mentioned apparatus embodiment are the same as those implemented by the corresponding steps in the method embodiment, but are not limited to the contents disclosed in the above-mentioned method embodiment. It should be noted that the above-mentioned modules, as part of the apparatus, can be executed in a computer system, such as a set of computer-executable instructions.
[0247] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0248] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned data labeling methods by executing the executable instructions. Since the principles for solving the problem in this electronic device embodiment are similar to those in the above-mentioned method embodiment, the implementation of this electronic device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0249] Refer to the following Figure 13 1300 according to this embodiment of the present disclosure will be described. Figure 13 The electronic device 1300 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0250] like Figure 13 As shown, electronic device 1300 is implemented as a general-purpose computing device. Components of electronic device 1300 may include, but are not limited to, the aforementioned at least one processing unit 1310, the aforementioned at least one storage unit 1320, and a bus 1330 connecting various system components (including storage unit 1320 and processing unit 1310).
[0251] The storage unit stores program code, which can be executed by the processing unit 1310, causing the processing unit 1310 to perform the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of the present disclosure. For example, the processing unit 1310 can perform the following steps of the aforementioned method embodiment: obtaining business data; determining data attributes of the business data; selecting a large model and / or a specialized model as a labeling model based on the data attributes and a preset labeling strategy; and calling the labeling model to label the business data to obtain a labeling result for the business data.
[0252] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 13201 and / or a cache 13202 , and may further include a read-only memory unit (ROM) 13203 .
[0253] The storage unit 1320 may also include a program / utility 13204 having a set (at least one) of program modules 13205, such program modules 13205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0254] Bus 1330 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0255] Electronic device 1300 may also communicate with one or more external devices 1340 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1300, and / or any device that enables electronic device 1300 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication may occur via input / output (I / O) interface 1350. Furthermore, electronic device 1300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1360. As shown, network adapter 1360 communicates with other modules of electronic device 1300 via bus 1330. It should be understood that, although not shown, other hardware and / or software modules may be used in conjunction with electronic device 1300, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0256] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0257] Based on the same inventive concept, embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the aforementioned data labeling methods. Because the principles underlying the problem solved by this computer-readable storage medium embodiment are similar to those of the aforementioned method embodiment, the implementation of this computer-readable storage medium embodiment can be referenced to the implementation of the aforementioned method embodiment, and any repetitions will not be repeated.
[0258] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0259] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0260] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0261] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0262] Based on the same inventive concept, the present disclosure also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the data labeling method of any one of the above-mentioned method embodiments. Since the principles for solving the problems in this computer program product embodiment are similar to those in the above-mentioned method embodiment, the implementation of this computer program product embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated here.
[0263] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0264] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0265] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0266] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A data annotation method, characterized in that: include: Obtain business data; Determining data attributes of the business data; Based on the data attributes and the preset annotation strategy, a large model and / or a specialized model is selected as the annotation model; The annotation model is called to annotate the business data to obtain an annotation result of the business data.
2. The method according to claim 1, characterized in that The data attributes include data source and / or data type; based on the data attributes and the preset annotation strategy, selecting a large model and / or a specialized model as the annotation model includes: When the data source is a set source and / or the data type is a set type, selecting the professional model as the annotation model; When the data source is a non-set source and the data type is a non-set type, the large model and the professional model are selected as annotation models.
3. The data annotation method according to claim 2, characterized in that: When the annotation model includes the large model and the specialized model, the calling of the annotation model to annotate the business data to obtain an annotation result of the business data includes: Calling the large model to annotate the business data, and receiving a first annotation result output by the large model; calling the professional model to annotate the business data, and receiving a second annotation result output by the professional model; Determine a labeling result for the business data according to the first labeling result and the second labeling result.
4. The data annotation method according to claim 3, characterized in that: Determining the annotation result of the business data according to the first annotation result and the second annotation result includes: If the similarity between the first annotation result and the second annotation result is greater than or equal to a set similarity threshold, performing a weighted average calculation on the first annotation result and the second annotation result to obtain an annotation result for the business data; If the similarity between the first annotation result and the second annotation result is less than the set similarity threshold, the first annotation result and the second annotation result are voted to obtain the annotation result of the business data.
5. The data annotation method according to claim 3, characterized in that: Determining the annotation result of the business data according to the first annotation result and the second annotation result includes: Calling the large model to summarize the first annotation result and the second annotation result; Receive the annotation results of the business data output by the large model.
6. The data labeling method according to any one of claims 1 to 5, characterized in that: Also includes: Extract data features of the target model; Determining format requirements for the labeled data based on the data characteristics; Accordingly, calling the annotation model to annotate the business data, and obtaining the annotation results of the business data includes: The annotation model is called to annotate the business data based on the format requirement to obtain an annotation result of the business data.
7. The data annotation method according to claim 6, characterized in that: The data features of the target model are extracted as follows: Using a data extraction script to read the training sample data of the target model; A cluster analysis algorithm is used to extract data features of the training sample data.
8. The data annotation method according to claim 1, wherein: Also includes: If a set keyword is extracted from the business data, query the standard rule library for a marking result matching the set keyword; The annotation information matching the set keyword is used as the annotation result of the business data.
9. A data labeling device, characterized in that: include: Data acquisition module, used to obtain business data; An attribute determination module, configured to determine the data attributes of the business data; A model selection module, configured to select a large model and / or a specialized model as a labeling model based on the data attributes and a preset labeling strategy; The data annotation module is used to call the annotation model to annotate the business data and obtain the annotation results of the business data.
10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the data labeling method described in any one of claims 1 to 8 by executing the executable instructions.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data labeling method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the data labeling method described in any one of claims 1 to 8.
Citation Information
Cited By
Heterogeneous data semantic metadata discovery method and system based on big and small model collaboration
CN120873264A