Data processing method, device and equipment
By using a large language model to identify and split the annotation requirements of the annotation data, combined with an appropriate data processing model, the problem of low manual annotation efficiency and accuracy is solved, and efficient and accurate processing of data annotation is achieved.
Patent Information
- Application Number
- CN202510097248.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
With the increase in the amount of data to be marked and the complexity of the data structure, the efficiency and accuracy of manual labeling are reduced, making it difficult to effectively handle complex data labeling tasks.
By receiving the annotation request for the data to be marked, using the large language model to identify the annotation demand information and split it into subtasks, and determining the appropriate data processing model, processing the annotation data, and finally determining the target annotation result.
The efficiency and accuracy of data labeling are improved, especially when the data to be labeled and the requirements for labeling are complex, intent identification, demand splitting and model determination can be carried out quickly and accurately, thereby improving the overall quality of data labeling.
Smart Images

Figure CN119988974A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a data processing method, device and equipment. Background Art
[0002] In order to protect user privacy and ensure data security, business data, user data and other data can be detected through detection models. Therefore, the data label can be determined through data annotation, so that the detection model can be trained with the annotated data. Among them, data annotation refers to the process of adding structured information to the original data. The added annotated data can be used to assist in understanding the data characteristics of the original data, such as adding annotated data to the original data through manual annotation.
[0003] However, as the amount of data to be annotated increases and the data structure becomes more complex, the efficiency and accuracy of manual annotation are low. To this end, the embodiments of this specification provide a better technical solution for improving data annotation efficiency and annotation accuracy. Summary of the invention
[0004] The purpose of the embodiments of this specification is to provide a better technical solution to improve data annotation efficiency and annotation accuracy.
[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided by an embodiment of the present specification includes: receiving a labeling request for data to be labeled, the labeling request including labeling requirement information corresponding to the data to be labeled; in response to the labeling request, using a large language model to perform intent recognition processing on the labeling requirement information to obtain an intent recognition result; using the large language model, according to the intent recognition result, splitting the labeling requirement information into multiple subtasks, and determining a data processing model for processing each of the subtasks; using the data processing model, processing the data to be labeled respectively to obtain a processing result corresponding to each of the subtasks; and determining a target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
[0006] A data processing device provided in an embodiment of the present specification includes: a request receiving module, which is used to receive a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; an intention recognition module, which is used to respond to the labeling request and use a large language model to perform intent recognition processing on the labeling requirement information to obtain an intention recognition result; a requirement splitting module, which is used to use the large language model to split the labeling requirement information into multiple subtasks according to the intention recognition result, and determine a data processing model for processing each of the subtasks; a data processing module, which is used to use the data processing model to process the data to be labeled respectively to obtain a processing result corresponding to each of the subtasks; and a result determination module, which is used to determine a target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
[0007] A data processing device provided in an embodiment of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor: receives a labeling request for data to be labeled, wherein the labeling request comprises labeling requirement information corresponding to the data to be labeled; in response to the labeling request, performs intent recognition processing on the labeling requirement information using a large language model to obtain an intent recognition result; uses the large language model to split the labeling requirement information into a plurality of subtasks according to the intent recognition result, and determines a data processing model for processing each of the subtasks; uses the data processing model to process the data to be labeled respectively to obtain a processing result corresponding to each of the subtasks; and determines a target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
[0008] An embodiment of the present specification also provides a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, they implement the following process: receiving a labeling request for data to be labeled, the labeling request including labeling requirement information corresponding to the data to be labeled; in response to the labeling request, using a large language model, performing intent recognition processing on the labeling requirement information to obtain an intent recognition result; using the large language model, according to the intent recognition result, splitting the labeling requirement information into multiple subtasks, and determining a data processing model for processing each of the subtasks; using the data processing model, processing the data to be labeled separately to obtain a processing result corresponding to each of the subtasks; and determining a target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
[0009] The embodiments of the present specification also provide a computer program product, including a computer program, which implements the following process when executed by a processor: receiving a labeling request for data to be labeled, the labeling request including labeling requirement information corresponding to the data to be labeled; in response to the labeling request, using a large language model to perform intent recognition processing on the labeling requirement information to obtain an intent recognition result; using the large language model, according to the intent recognition result, splitting the labeling requirement information into multiple subtasks, and determining a data processing model for processing each of the subtasks; using the data processing model, processing the data to be labeled separately to obtain a processing result corresponding to each of the subtasks; and determining a target labeling result for the data to be labeled according to the processing result corresponding to the subtask. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is an embodiment of a data processing method of this specification; Figure 2 This is another data processing method embodiment of the present specification; Figure 3 A schematic diagram of a display interface for annotated data in this manual; Figure 4 This is a schematic diagram of a feedback content receiving interface of this manual; Figure 5 A schematic diagram of a data annotation system for this specification; Figure 6 A schematic diagram of a data processing process of this specification; Figure 7 This is an embodiment of a data processing device of the present specification; Figure 8 This is an embodiment of a data processing device in this specification. DETAILED DESCRIPTION
[0011] The embodiments of this specification provide a data processing method, device and equipment.
[0012] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0013] The embodiment of this specification provides a better technical solution to improve data annotation efficiency and annotation accuracy. Data annotation is the process of adding structured information to the original data. The added annotation data can be used to assist in understanding the data features of the original data to improve the subsequent data processing efficiency of the original data, such as adding annotation data to the original data by manual annotation. However, as the amount of data to be annotated increases and the data structure becomes more and more complex, the annotation efficiency and annotation accuracy of manual annotation are low. To this end, the embodiment of this specification provides a better technical solution to improve data annotation efficiency and annotation accuracy. In this solution, by receiving an annotation request for the data to be annotated, wherein the annotation request includes the annotation requirement information corresponding to the data to be annotated, in response to the annotation request, using a large language model, the annotation requirement information is processed for intent recognition, and the intent recognition result is obtained. Using the large language model, according to the intent recognition result, the annotation requirement information is split into multiple subtasks, and a data processing model for processing each subtask is determined. Using the data processing model, the data to be annotated is processed separately to obtain the processing result corresponding to each subtask, and according to the processing result corresponding to the subtask, the target annotation result for the data to be annotated is determined. In this way, when the data to be annotated and the annotation requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately identify intentions, split requirements and determine models (i.e., determine the data processing model corresponding to the subtask), and then process the data to be annotated separately by calling the data processing model to obtain the processing results corresponding to each subtask. Finally, the target annotation results for the data to be annotated can be quickly and accurately determined through the processing results corresponding to the subtasks, thereby improving the annotation efficiency and standard accuracy of data annotation. For specific processing, please refer to the specific content in the following embodiments.
[0014] like Figure 1As shown, the embodiment of this specification provides a data processing method, the execution subject of the method can be a server, wherein the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server for financial services or online shopping services, or a background server for an application, etc. In this embodiment, the execution subject is taken as an example for detailed description, and the method can specifically include the following steps: In step S102, a labeling request for the data to be labeled is received.
[0015] Among them, the annotation request may include annotation requirement information corresponding to the data to be annotated. The data to be annotated may be any data, for example, the data to be annotated may be image data, text data, voice data, video data, point cloud data, etc. The annotation requirement information may include annotation specifications, annotation rules, annotation strategies, etc. for the data to be annotated.
[0016] In implementation, the server may receive a labeling request triggered by a user for one or more data to be labeled, or the server may also trigger a labeling request for data to be labeled obtained within a preset labeling period based on a preset labeling period, or the server may also receive a labeling request for sample data used to train a business processing model, where the sample data may serve as the data to be labeled, and the business processing model may be a model built based on a preset machine learning algorithm for processing preset businesses.
[0017] The above-mentioned method for obtaining the annotation request of the data to be annotated is an optional and feasible acquisition method. In actual application scenarios, there may be a variety of different acquisition methods. Different acquisition methods can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.
[0018] In addition, the labeling requirement information of different types of data to be labeled may be different. For example, taking the data to be labeled as image data, text data, voice data, video data, and point cloud data as an example, the labeling requirement information for different data to be labeled may be as shown in Table 1 below.
[0019] Table 1
[0020] Among them, the target detection requirement can be to identify and mark the location of a specific object in an image, usually using a bounding box. The semantic segmentation requirement can be to classify each pixel of the image and mark different regions or objects. The instance segmentation requirement can be to distinguish different instances in the same category. The key point detection requirement can be to mark specific key points on an object, such as the key parts of a face. The sentiment analysis requirement can be to mark the sentiment tendency of a text, such as positive, negative, or neutral. The named entity recognition requirement can be to mark entities such as names of people, places, and organizations in the text. The relationship extraction requirement can be to mark the relationship between entities in the text. The text classification requirement can be to classify the text into predefined categories. The speech-to-text requirement can be to convert speech data into text. The audio event annotation requirement can be to mark specific events or sound types in the audio. The frame-level annotation requirement can be to annotate images for each frame of the video. The behavior recognition requirement can be to mark the behavior or actions of people in the video. The time series annotation requirement can be to mark the time point when the event occurs. The point cloud target detection requirement can be to identify points with significant geometric features from unordered point cloud data to describe the local or global features of the object. The requirement for key point annotation of point cloud can be that key points usually contain key information about the shape, structure or function of an object. The requirement for semantic segmentation of point cloud can be that in the application scenario of self-driving cars, points in the point cloud can be classified as roads, vehicles, pedestrians, buildings, vegetation, etc. The requirement for multi-dimensional fusion annotation can be to annotate the image data collected by sensors in 2D and 3D dimensions at the same time and establish connections.
[0021] As shown in Table 1 above, the annotation requirement information corresponding to different types of data to be labeled may be different. At the same time, the annotation requirement information corresponding to the same type of data to be labeled may be the same or different. For example, the annotation requirement information for image data 1 may include classification requirements and target detection requirements, and the annotation requirement information for image data 2 may include component requirements, speech segmentation requirements, and instance segmentation requirements.
[0022] The annotation requirement information corresponding to the data to be annotated can select different annotation requirements according to different actual application scenarios, and the embodiments of this specification do not make specific limitations on this.
[0023] In step S104, in response to the annotation request, the large language model is used to perform intent recognition processing on the annotation requirement information to obtain an intent recognition result.
[0024] Among them, large language models (LLMs) can be advanced natural language processing (NLP) models built on the basis of deep learning technology. Their underlying converters are a set of neural networks, which can be composed of encoders and decoders with self-attention functions. They can understand and generate human language by processing large amounts of text data, and can process and generate high-dimensional data to perform various natural language processing tasks. Since large language models have a large number of parameters and a wide range of training data sets, they can capture the richness and subtle differences of language expressions, and can display the captured information through powerful computing power of billions to trillions of parameters. They have powerful prediction and analysis capabilities in multiple fields such as language understanding and image recognition.
[0025] In implementation, since the data to be labeled may contain data of multiple different data types, and the labeling requirement information corresponding to different data to be labeled may also contain multiple different labeling requirements, the server can use the understanding and analysis capabilities of the large language model to improve the processing efficiency and processing effect of intent recognition processing of labeling requirement information.
[0026] For example, if the data to be annotated is image data, the corresponding annotation requirement information may include target detection requirements and semantic segmentation requirements, that is, the annotation requirement information corresponding to the data to be annotated obtained by the server may be: "Recognize and use a bounding box to mark the location of a specific object in the image, and classify each pixel of the image to mark different areas or objects". The server may input the first prompt information (propot) and the above-mentioned annotation requirement information into the large language model to obtain the intent recognition result for the annotation requirement information, and the obtained intent recognition result may be "target detection intent" and "semantic segmentation intent".
[0027] In step S106, the large language model is used to split the labeling requirement information into multiple subtasks according to the intention recognition result, and a data processing model for processing each subtask is determined.
[0028] Among them, the data processing model can be a model for data processing built based on a preset machine learning algorithm, and the data processing models corresponding to different subtasks may be different.
[0029] In implementation, since there are a large number of available data processing models, a large language model can be used to achieve reasonable resource planning. That is, the powerful computing power of the large language model can be used to determine the data processing model used to process each subtask, so as to assign the subtask to the appropriate data processing model for processing and improve resource allocation efficiency.
[0030] For example, taking the above-mentioned intent recognition results including "target detection intent" and "semantic segmentation intent" as an example, a large language model can be used to split the labeling requirement information into multiple subtasks based on the intent recognition results. The multiple subtasks obtained by splitting can be: Subtask 1 "target detection task (i.e., identifying and using bounding boxes to mark the location of specific objects in the image)" and Subtask 2 "semantic segmentation task (i.e., classifying each pixel of the image and marking different areas or objects)".
[0031] The server can use the large language model to determine the data processing model used to process each subtask. For example, the data processing model 1 determined to process subtask 1 can be a model for object recognition built based on a machine learning algorithm, and the data processing model 2 used to process subtask 2 can be a model for region or object classification built based on a deep learning algorithm.
[0032] In step S108, the data to be labeled are processed respectively using the data processing model to obtain processing results corresponding to each subtask.
[0033] In step S110, a target labeling result for the data to be labeled is determined according to the processing result corresponding to the subtask.
[0034] In implementation, the server can determine the target labeling result for the data to be labeled according to the processing result corresponding to each subtask. For example, taking the above subtask 1 and subtask 2 as an example, the server can determine the target labeling result according to the target detection result and the semantic segmentation result, such as marking the target detection result and the semantic segmentation result in the data to be labeled (i.e., image data).
[0035] After annotating the data to be annotated (i.e., obtaining the target annotation results for the annotated data), the service can use the data to be annotated as sample data to train the business processing model, so as to process the preset business through the trained business processing model.
[0036] The embodiment of the present specification provides a data processing method, by receiving a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responding to the labeling request, using a large language model, performing intent recognition processing on the labeling requirement information, obtaining an intent recognition result, using the large language model, according to the intent recognition result, splitting the labeling requirement information into multiple subtasks, and determining a data processing model for processing each subtask, using the data processing model, processing the data to be labeled respectively, obtaining a processing result corresponding to each subtask, and determining a target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask, and finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0037] In practical applications, in step S110, there are many ways to determine the specific processing method of the target labeling result for the data to be labeled according to the processing result corresponding to the subtask. The following is an optional processing method, such as Figure 2 As shown, the process may specifically include the following step S1102.
[0038] In step S1102, the large language model is used to perform information fusion processing on the processing results corresponding to the subtasks to obtain the target labeling results for the data to be labeled.
[0039] In implementation, the server can use the powerful understanding ability of the large language model to perform information fusion processing on the processing results corresponding to the subtasks to obtain the target labeling results for the data to be labeled. That is, the server can input the second prompt information and the processing results corresponding to each subtask into the large language model for information fusion processing to obtain the target labeling results.
[0040] In practical applications, in the above step S1102, a large language model is used to perform information fusion processing on the processing results corresponding to the subtasks, and there are many specific processing methods to obtain the target labeling results for the data to be labeled. The following is an optional processing method, which may specifically include the following steps A1 to A3.
[0041] In step A1, the large language model is used to perform information fusion processing on the processing results corresponding to the subtasks to obtain a first labeling result for the data to be labeled.
[0042] In step A2, keyword screening is performed on the data to be annotated to obtain target keywords contained in the data to be annotated, and target knowledge matching the target keywords is obtained based on a preset knowledge base.
[0043] In implementation, the server can perform keyword screening on the data to be annotated based on a pre-trained keyword extraction model to obtain target keywords contained in the data to be annotated, wherein the keyword extraction model can be a model built based on a preset machine learning algorithm.
[0044] In addition, when the data to be annotated is multimodal data, text conversion processing can be performed on the non-text data included in the data to be annotated in advance, and keyword screening processing can be performed on the data to be annotated obtained by the text conversion processing.
[0045] In addition, the above-mentioned keyword screening processing method is an optional and feasible processing method. In actual application scenarios, there may be a variety of different processing methods. Different processing methods can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.
[0046] After filtering out the target keyword, the server can obtain a preset knowledge base corresponding to the target keyword, and then determine the matching degree between each knowledge in the preset knowledge base and the target keyword, and then obtain the target knowledge matching the target keyword from the preset knowledge base based on the matching degree.
[0047] For example, assuming that the target keyword is "salicylic acid", the server can determine the knowledge base corresponding to the target keyword according to the knowledge field corresponding to the target keyword, such as a medical knowledge base. Then, the server can obtain the target knowledge matching "salicylic acid" in the medical knowledge base (such as a detailed introduction to "salicylic acid", etc.).
[0048] In step A3, the first annotation result and the target knowledge matched by the target keyword are sent to a preset processing party, and a target annotation result determined by the preset processing party according to the first annotation result and the target knowledge matched by the target keyword is received.
[0049] In implementation, in order to improve the accuracy of data annotation, manual review can be introduced to review the first annotation results output by the large language model. In addition, since the data to be annotated may contain words that are difficult for humans to understand (such as highly professional keywords), the first annotation results and the target knowledge matched by the target keywords can be sent to the preset processing party, so that the preset processing party can deeply understand the data to be annotated through the target knowledge, thereby improving the accuracy of manual review.
[0050] In addition, on the side of the preset processing party, the data to be annotated, the target knowledge and the first annotation result can be presented to the preset processing party through a variety of data presentation modes. Figure 3 As shown, the data to be annotated can be displayed on the left side of the display page, and the target knowledge matched by the target key can be displayed in the display area of the data to be annotated by means of word marking. The target annotation result can be displayed on the right side of the display page. The middle area between the target annotation result and the data to be annotated can be an input area for the target annotation result. The preset processing party can input the target annotation result in the input area according to the target knowledge and the first annotation result. The above display mode is an optional interaction mode. In addition to this interaction mode, there can be a variety of different interaction modes, such as adding sidebar options, etc. Different interaction modes can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.
[0051] In addition, an AI intelligent assistant can be introduced in the display area of the first annotation result, that is, the preset processor can also input query information for the first annotation result in the display area. The AI intelligent assistant can obtain feedback information of the query information by calling knowledge bases, data processing algorithms and other tools, and interact with the preset processor through feedback information to help the preset processor determine the target annotation result more accurately.
[0052] In practical applications, data processing models include, but are not limited to, object recognition models, sentiment analysis models, risk recognition models, data prediction models, and data classification models. For example, an object recognition model may be a model built based on a preset machine learning algorithm for identifying specific objects contained in the data to be detected. Specifically, an object recognition model may be a model built based on a convolutional neural network algorithm for identifying people and vehicles contained in image data.
[0053] The sentiment analysis model can be a model built based on a preset machine learning algorithm for identifying the sentiment tendency type of the data to be detected. Specifically, the sentiment analysis model can be a model built based on a neural network algorithm for identifying the sentiment tendency type of users in human-computer interaction data, wherein the sentiment tendency type can include positive sentiment type, negative sentiment type, and the like.
[0054] The risk identification model can be a model built based on a preset machine learning algorithm to identify the risk type of the data to be detected. Specifically, the risk identification model can be a model built based on a neural network algorithm to identify the risk type of business data, where the risk type can include high risk, medium risk, low risk and other types.
[0055] The data prediction model may be a model for performing data prediction based on a machine learning algorithm. Specifically, the data prediction model may be a model for performing data prediction on time series data based on a long short-term memory network (LSTM) algorithm.
[0056] The data classification model may be a model for data classification built based on a preset classification algorithm (such as a random forest algorithm, a k-means algorithm, etc.).
[0057] In addition, the same subtask can correspond to one or more data processing models.
[0058] In practical applications, it is also possible to determine whether the large language model needs to be fine-tuned. There are many specific ways to fine-tune the model. The following is an optional way to do this: Figure 2 As shown, the processing may specifically include the following steps S302 to S306.
[0059] In step S302, the calling time and calling success rate of the data processing model are obtained.
[0060] In step S304, it is determined whether to perform fine-tuning on the large language model according to the calling time and the calling success rate.
[0061] During implementation, the server can obtain the calling time and calling success rate of the data processing model corresponding to each subtask, and determine whether it is necessary to fine-tune the large language model based on the calling time and calling success rate of the data processing model corresponding to each subtask.
[0062] For example, the server can determine whether it is necessary to fine-tune the large language model based on the mean (or maximum value, weighted summary value, minimum value, etc.) of the call time of the data processing model corresponding to each subtask and the mean (or maximum value, weighted summary value, minimum value, etc.) of the call success rate, according to the preset time threshold and the preset success rate threshold. Specifically, the server can determine whether it is necessary to fine-tune the large language model based on whether the mean of the call time is greater than the preset time threshold and / or whether the mean of the call success rate is less than the preset success rate threshold.
[0063] In addition, in addition to the call time and the call success rate, the server can also obtain other indicator data to determine whether to fine-tune the large language model. Different indicator data can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.
[0064] In step S306, when it is determined that the large language model is to be fine-tuned, the large language model is fine-tuned according to the target labeling result.
[0065] In implementation, when it is determined to fine-tune the large language model, the server can use the data to be labeled as sample data, and fine-tune the large language model through a preset fine-tuning method according to the data to be labeled and the corresponding target labeling results.
[0066] Among them, there can be multiple preset fine-tuning methods, for example, the preset fine-tuning method can be LoRA fine-tuning, QLoRA fine-tuning, LongLoRA fine-tuning, etc. Different fine-tuning methods can be selected according to different actual application scenarios. This specification does not make specific limitations on this.
[0067] In practical applications, it is also possible to determine whether the large language model needs to be fine-tuned. There are many specific ways to fine-tune the model. The following is an optional way to do this: Figure 2 As shown, the processing may specifically include the following steps S308 to S310.
[0068] In step S308, it is determined whether to perform fine-tuning processing on the large language model according to the target labeling result and the first labeling result.
[0069] In implementation, the server may determine whether to perform fine-tuning processing on the large language model according to the similarity between the target annotation result and the first annotation result, and a preset similarity threshold.
[0070] For example, the server may determine the similarity between the target annotation result and the first annotation result according to a preset similarity algorithm, and determine to perform fine-tuning on the large language model when the similarity is not greater than a preset similarity threshold.
[0071] In step S310, when it is determined that the large language model is to be fine-tuned, the large language model is fine-tuned according to the target labeling result.
[0072] The specific processing process of the above S310 can refer to the relevant content of S306 in the above embodiment, which will not be repeated here.
[0073] In practical applications, it is also possible to determine whether the large language model needs to be fine-tuned. There are many specific ways to fine-tune the model. The following is an optional way to do this: Figure 2 As shown, the processing may specifically include the following steps S312 to S314.
[0074] In step S312, feedback information of a preset processing method on the first annotation result is received, and it is determined whether to perform fine-tuning processing on the large language model according to the feedback information.
[0075] In implementation, a feedback information input interface may be provided on the preset processing side, and the preset processing method may input feedback information for the first annotation result, wherein the feedback information may include information input by the preset processing side on whether the first annotation result is useful, and the feedback content input for the first annotation result, etc. Figure 3 In the display page shown, the preset input party may be provided with option labels (i.e., useful and useless labels) for whether the first annotation result is useful, as well as a feedback label. The preset processing party may click the feedback label and, in the following example, Figure 4 Enter the feedback content for the first annotation result in the page shown.
[0076] The server may determine whether to fine-tune the large language model based on the feedback information. For example, if a click operation on a "useless" label is received from a preset feedback party, the server may determine to fine-tune the large language model.
[0077] In step S314, when it is determined that the large language model is to be fine-tuned, the large language model is fine-tuned according to the target labeling result.
[0078] In implementation, the specific processing process of the above S314 can refer to the relevant content of S306 in the above embodiment, which will not be repeated here.
[0079] In addition, when it is determined to perform fine-tuning on the large language model, the server may also perform fine-tuning on the large language model according to the feedback content input by the preset processing party and the target annotation result.
[0080] In practical applications, the data processing model is used in step S108 to process the data to be labeled respectively, and the specific processing methods to obtain the processing results corresponding to each subtask can be varied. The following is an optional processing method, such as Figure 2 As shown, the processing may specifically include the following step S1082.
[0081] In step S1082, according to the intention recognition result, the data to be labeled is split and processed to obtain the sub-data corresponding to each subtask in the data to be labeled, and the sub-data corresponding to each subtask is processed separately using the data processing model to obtain the processing result corresponding to each subtask.
[0082] In implementation, the server can split the data to be annotated using a large language model based on the intent recognition results to obtain sub-data corresponding to each subtask in the data to be annotated. Then, the data processing model corresponding to each subtask is called to process the sub-data corresponding to each subtask and obtain the processing result corresponding to each subtask.
[0083] In addition, the server can also use intelligent agent technology to build a data annotation system. The framework of the data annotation system can be as follows: Figure 5 As shown in the figure, in computer science, an agent can refer to a class of independent software entities that can run autonomously in a certain environment, respond to external events, and manage themselves. It has a certain degree of intelligence and can complete tasks according to designed strategies. It is often used in fields such as automated operations and information retrieval.
[0084] The data annotation system can include a large language model agent and an annotation agent, wherein the large language model agent can be an agent built on the basis of the large language model, and has the ability to automatically perform specific tasks. It can use the deep learning advantages of the large language model to conduct self-learning and adjustment to complete complex decision-making, prediction and operation functions, and can improve the efficiency and intelligence level of automated processing.
[0085] An annotation agent can be a software system or tool that integrates artificial intelligence technology, especially machine learning and deep learning algorithms, and can be used to complete data annotation tasks automatically or semi-automatically. This system can understand the context of the data, identify and mark specific information or attributes in the data, such as entity recognition in text, object detection in images, etc. An annotation agent can support the rapid processing of large-scale data sets by improving the efficiency and accuracy of the annotation process, thereby accelerating the training and deployment of machine learning models. In addition, it can continuously learn and optimize based on feedback to improve the quality of annotation and processing speed.
[0086] The data labeling system is important. Task consultation / planning refers to the analysis and capability matching of corresponding tasks in an interactive form. The implementation of planning refers to capability recommendation (i.e. matching the corresponding data processing model to the subtask), parameter recommendation and automatic deployment.
[0087] Among them, mission planning can include: 1. Confirm data; 2. Create a task flow (quickly create a task); 3. Select UDF / intelligent capabilities; 4. Configure input parameters; 5. Configure output parameters; 6. Configure / modify the template to adapt to the output results of UDF (after intelligence); 7. Configure annotation requirements; 8. Deploy.
[0088] The data tagging system based on the large language model is a "new management method" for large language models. The tagging agent (Tag Agent) can use LLM as the overall controller. Humans stand at a high point to describe a "task goal" and hand over the work of completing this tagging task to the Tag Agent. Figure 6 As shown in the figure, Tag Agent can receive the target and gradually execute a series of tasks-oriented actions, including "task planning", "decision execution", "tool call", "labeling assistant interaction", and "human / system feedback". The actions of each process are as follows: (1) Task planning: Before the labeling task begins, the preset processor can input the data to be labeled and the labeling requirement information to the Tag Agent. The large language model can act as a thinking brain, realize the recognition of human intentions based on the labeling requirement information, and parse the user request (i.e., the labeling requirement information) into multiple subtasks based on the intention recognition results. For specific subtasks, the large language model can be used to analyze and select an appropriate data processing model, that is, to select an appropriate intelligent tool for organizational planning. The plan can include a recommended capability list, a specific calling method / system, and corresponding parameter configurations.
[0089] (2) Decision execution: According to the calling method, calling capability object and corresponding parameter configuration in the task plan, the corresponding data processing model is called step by step to process the subtasks to obtain the processing results corresponding to each subtask. According to the processing results corresponding to each subtask, the target labeling results of the data to be labeled are obtained through merging and processing. At the same time, the labeling assistant can also be used to adjust the format of the target labeling results to the target format.
[0090] (3) Tool calling: All capability tools can be integrated and coordinated into a knowledge base at the bottom of the system. The knowledge base is then packaged into a large capability store (the capability store can also include a variety of tools such as image and text retrieval, image security detection, and document knowledge base). The capability is called by step (2) through the above-mentioned calling method, filling in the corresponding calling object and parameters, and the calling action is triggered by step (2), and the calling result is returned to step (2).
[0091] (4) Annotation assistant: The target data provided by step (2) is transmitted to the annotation task corresponding to the annotation platform in the data annotation system through the interface between platforms. The annotation agent can present the first annotation result to the preset processing party, wherein the annotation agent can present the label of the first annotation result (such as can be determined according to the intention recognition result) to the preset processing party to improve the annotation efficiency of the preset processing party. In addition, the annotation intelligence can also provide human-computer assistance functions. For example, during the annotation process, the preset processing party can call up the annotation agent through word marking and chat box (Chat UI). The annotation agent will recommend tools that can be used, so that the preset processing party can confirm the call by clicking to select or communicating through dialogue.
[0092] (5) Manual / system feedback: In order to continuously improve or correct the large language model, a manual feedback button can be set on the annotation assistant side. Each time it is manually called up through word marking or a dialogue chat box, a corresponding feedback window can be provided to collect feedback information for optimizing or correcting the model's return results. In addition, the data annotation system can also recycle call duration, call success rate, agent thinking chain, call effect and other indicator data to fine-tune the large language model to discover the missing points of the large language model and improve it.
[0093] In this way, since the labeling agent is a system that integrates several intelligent algorithms and can automatically perform data labeling tasks, the labeling agent has the following advantages: (1) Integration: The labeling agent integrates a variety of technologies (such as pre-labeling, algorithm-assisted tools, knowledge base, etc.) to provide a one-stop data processing solution, and has the ability to flexibly choose and adapt to complex tasks.
[0094] (2) Automation: After running, the intelligent agent can perform labeling tasks relatively independently and automatically plan and analyze labeling tasks, reducing human intervention and saving time and labor costs.
[0095] (3) Human-machine collaborative labeling and feedback links: The labeling agents are designed to continuously review and correct large language models through human feedback and system feedback links. They can continuously evolve based on newly added labeling data, thereby improving the accuracy of the labeling results.
[0096] In addition, using LLM as the thinking brain and utilizing Agent technology to connect the underlying intelligent tools (Tool Set) to automatically analyze task portraits and match the most appropriate tools can reduce the cost of manual intervention and generate a set of intelligent solutions for labeling tasks without relying on manual experience.
[0097] In addition, the real-time interactive function provided by the annotation system can realize human-machine collaboration to improve annotation efficiency. And the feedback interaction link and perfect system indicator system revealed by the annotation assistant can ensure that the intelligent agent's ability is always at a high level for sustainable use.
[0098] The embodiment of the present specification provides a data processing method, by receiving a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responding to the labeling request, using a large language model, performing intent recognition processing on the labeling requirement information, obtaining an intent recognition result, using the large language model, according to the intent recognition result, splitting the labeling requirement information into multiple subtasks, and determining a data processing model for processing each subtask, using the data processing model, processing the data to be labeled respectively, obtaining a processing result corresponding to each subtask, and determining a target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask, and finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0099] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 7 shown.
[0100] The data processing device includes: a request receiving module 701, an intention recognition module 702, a demand splitting module 703, a data processing module 704 and a result trustworthy module 705, wherein: A request receiving module 701 is used to receive a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; The intention recognition module 702 is used to respond to the annotation request and use a large language model to perform intention recognition processing on the annotation requirement information to obtain an intention recognition result; A demand splitting module 703 is used to use the large language model to split the annotation demand information into multiple subtasks according to the intention recognition result, and determine a data processing model for processing each of the subtasks; The data processing module 704 is used to process the data to be labeled respectively by using the data processing model to obtain the processing result corresponding to each subtask; The result determination module 705 is used to determine a target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
[0101] In the embodiment of this specification, the result determination module 705 is used to: The large language model is used to perform information fusion processing on the processing results corresponding to the subtasks to obtain target labeling results for the data to be labeled.
[0102] In the embodiment of this specification, the result determination module 705 is used to: Using the large language model, performing information fusion processing on the processing results corresponding to the subtasks to obtain a first labeling result for the data to be labeled; Perform keyword screening processing on the data to be annotated to obtain target keywords contained in the data to be annotated, and obtain target knowledge matching the target keywords according to a preset knowledge base; The first annotation result and the target knowledge matched by the target keyword are sent to a preset processing party, and the target annotation result determined by the preset processing party according to the first annotation result and the target knowledge matched by the target keyword is received.
[0103] In the embodiments of this specification, the data processing model includes but is not limited to an object recognition model, a sentiment analysis model, a risk recognition model, a data prediction model, and a data classification model.
[0104] In the embodiment of this specification, the device further includes: An information acquisition module, used to obtain the calling time and calling success rate of the data processing model; A first judgment module is used to determine whether to perform fine-tuning processing on the large language model according to the call time and the call success rate; The first fine-tuning module is used to fine-tune the language model according to the target labeling result when it is determined that the large language model is to be fine-tuned.
[0105] In the embodiment of this specification, the device further includes: A second judgment module, used for determining whether to perform fine-tuning processing on the large language model according to the target labeling result and the first labeling result; The second fine-tuning module is used to fine-tune the language model according to the target labeling result when it is determined that the large language model is to be fine-tuned.
[0106] In the embodiment of this specification, the device further includes: A third judgment module is used to receive feedback information from the preset processing party regarding the first annotation result, and determine whether to perform fine-tuning processing on the large language model according to the feedback information; The third fine-tuning module is used to fine-tune the language model according to the target labeling result when it is determined that the large language model is to be fine-tuned.
[0107] In the embodiment of this specification, the data processing module is used to: According to the intention recognition result, the data to be labeled is split and processed to obtain sub-data corresponding to each subtask in the data to be labeled, and the data processing model is used to process the sub-data corresponding to each subtask separately to obtain the processing result corresponding to each subtask.
[0108] The embodiment of the present specification provides a data processing device, which receives a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responds to the labeling request, uses a large language model to perform intent recognition processing on the labeling requirement information, obtains the intent recognition result, uses the large language model, and splits the labeling requirement information into multiple subtasks according to the intent recognition result, and determines the data processing model used to process each subtask, uses the data processing model to process the data to be labeled respectively, obtains the processing result corresponding to each subtask, and determines the target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask. Finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0109] The above is a data processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 8 shown.
[0110] The data processing device may provide a terminal device or a server, etc. for the above embodiments.
[0111] The data processing device may have relatively large differences due to different configurations or performances, and may include one or more processors 801 and memory 802, and the memory 802 may store one or more storage applications or data. Among them, the memory 802 may be a short-term storage or a persistent storage. The application stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the data processing device. Furthermore, the processor 801 may be configured to communicate with the memory 802 and execute a series of computer executable instructions in the memory 802 on the data processing device. The data processing device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input and output interfaces 805, and one or more keyboards 806.
[0112] Specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs are configured to be executed by one or more processors, including computer executable instructions for performing the following: Receiving a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; In response to the annotation request, using a large language model, performing intent recognition processing on the annotation requirement information to obtain an intent recognition result; Using the large language model, according to the intention recognition result, the labeling requirement information is split into a plurality of subtasks, and a data processing model for processing each of the subtasks is determined; Using the data processing model, the data to be labeled are processed respectively to obtain processing results corresponding to each of the subtasks; According to the processing result corresponding to the subtask, a target labeling result for the data to be labeled is determined.
[0113] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0114] The embodiment of the present specification provides a data processing device, which receives a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responds to the labeling request, uses a large language model to perform intent recognition processing on the labeling requirement information, obtains the intent recognition result, uses the large language model, and splits the labeling requirement information into multiple subtasks according to the intent recognition result, and determines the data processing model used to process each subtask, uses the data processing model to process the data to be labeled respectively, obtains the processing result corresponding to each subtask, and determines the target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask. Finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0115] Furthermore, based on the above Figures 1 to 6 In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Receiving a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; In response to the annotation request, using a large language model, performing intent recognition processing on the annotation requirement information to obtain an intent recognition result; Using the large language model, according to the intention recognition result, the labeling requirement information is split into a plurality of subtasks, and a data processing model for processing each of the subtasks is determined; Using the data processing model, the data to be labeled are processed respectively to obtain processing results corresponding to each of the subtasks; According to the processing result corresponding to the subtask, a target labeling result for the data to be labeled is determined.
[0116] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0117] The embodiment of the present specification provides a storage medium, which receives a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responds to the labeling request, uses a large language model to perform intent recognition processing on the labeling requirement information, obtains the intent recognition result, uses the large language model, and splits the labeling requirement information into multiple subtasks according to the intent recognition result, and determines the data processing model used to process each subtask, uses the data processing model to process the data to be labeled respectively, obtains the processing result corresponding to each subtask, and determines the target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask. Finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0118] Furthermore, based on the above Figures 1 to 6 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Receiving a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; In response to the annotation request, using a large language model, performing intent recognition processing on the annotation requirement information to obtain an intent recognition result; Using the large language model, according to the intention recognition result, the labeling requirement information is split into a plurality of subtasks, and a data processing model for processing each of the subtasks is determined; Using the data processing model, the data to be labeled are processed respectively to obtain processing results corresponding to each of the subtasks; According to the processing result corresponding to the subtask, a target labeling result for the data to be labeled is determined.
[0119] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0120] The embodiment of the present specification provides a computer program product, which receives a labeling request for data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled, responds to the labeling request, uses a large language model to perform intent recognition processing on the labeling requirement information, obtains the intent recognition result, uses the large language model, and splits the labeling requirement information into multiple subtasks according to the intent recognition result, and determines the data processing model used to process each subtask, uses the data processing model to process the data to be labeled respectively, obtains the processing result corresponding to each subtask, and determines the target labeling result for the data to be labeled according to the processing result corresponding to the subtask. In this way, in the case where the data to be labeled and the labeling requirements are relatively complex, the understanding, summarization and reasoning capabilities of the large language model can be used to quickly and accurately perform intent recognition, requirement splitting and model determination (i.e., determine the data processing model corresponding to the subtask), and then by calling the data processing model, the data to be labeled is processed respectively to obtain the processing result corresponding to each subtask. Finally, the target labeling result for the data to be labeled can be quickly and accurately determined according to the processing result corresponding to the subtask, thereby improving the labeling efficiency and standard accuracy of data labeling.
[0121] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0122] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0123] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0124] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0125] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0126] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0127] The embodiments of this specification are described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0129] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0130] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0131] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0132] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0133] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0134] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0136] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0137] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.
Claims
1. A data processing method, comprising: Receiving a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; In response to the annotation request, using a large language model, performing intent recognition processing on the annotation requirement information to obtain an intent recognition result; Using the large language model, according to the intention recognition result, the labeling requirement information is split into a plurality of subtasks, and a data processing model for processing each of the subtasks is determined; Using the data processing model, the data to be labeled are processed respectively to obtain processing results corresponding to each of the subtasks; According to the processing result corresponding to the subtask, a target labeling result for the data to be labeled is determined.
2. According to the method of claim 1, determining the target labeling result for the data to be labeled according to the processing result corresponding to the subtask comprises: The large language model is used to perform information fusion processing on the processing results corresponding to the subtasks to obtain target labeling results for the data to be labeled.
3. According to the method of claim 2, the step of using the large language model to perform information fusion processing on the processing results corresponding to the subtasks to obtain target labeling results for the data to be labeled includes: Using the large language model, performing information fusion processing on the processing results corresponding to the subtasks to obtain a first labeling result for the data to be labeled; Perform keyword screening processing on the data to be annotated to obtain target keywords contained in the data to be annotated, and obtain target knowledge matching the target keywords according to a preset knowledge base; The first annotation result and the target knowledge matched by the target keyword are sent to a preset processing party, and the target annotation result determined by the preset processing party according to the first annotation result and the target knowledge matched by the target keyword is received.
4. According to the method of claim 3, the data processing model includes but is not limited to an object recognition model, a sentiment analysis model, a risk recognition model, a data prediction model, and a data classification model.
5. The method according to claim 4, further comprising: Obtaining the calling time and calling success rate of the data processing model; Determining whether to perform fine-tuning processing on the large language model according to the calling time and the calling success rate; In the case where it is determined to perform fine-tuning on the large language model, fine-tuning is performed on the large language model according to the target labeling result.
6. The method according to claim 4, further comprising: Determining whether to perform fine-tuning processing on the large language model according to the target labeling result and the first labeling result; In the case where it is determined to perform fine-tuning on the large language model, fine-tuning is performed on the large language model according to the target labeling result.
7. The method according to claim 4, further comprising: receiving feedback information from the preset processing party regarding the first annotation result, and determining whether to perform fine-tuning processing on the large language model according to the feedback information; In the case where it is determined to perform fine-tuning on the large language model, fine-tuning is performed on the large language model according to the target labeling result.
8. The method according to claim 1, wherein the data to be labeled is processed using the data processing model to obtain processing results corresponding to each subtask, including: According to the intention recognition result, the data to be labeled is split and processed to obtain sub-data corresponding to each subtask in the data to be labeled, and the data processing model is used to process the sub-data corresponding to each subtask separately to obtain the processing result corresponding to each subtask.
9. A data processing device, comprising: A request receiving module, configured to receive a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; An intention recognition module, used to respond to the annotation request and use a large language model to perform intention recognition processing on the annotation requirement information to obtain an intention recognition result; A demand splitting module, used to use the large language model to split the labeling demand information into multiple subtasks according to the intention recognition result, and determine a data processing model for processing each of the subtasks; A data processing module, used to process the data to be labeled respectively using the data processing model to obtain a processing result corresponding to each subtask; The result determination module is used to determine the target labeling result for the data to be labeled according to the processing result corresponding to the subtask.
10. A data processing device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Receiving a labeling request for the data to be labeled, wherein the labeling request includes labeling requirement information corresponding to the data to be labeled; In response to the annotation request, using a large language model, performing intent recognition processing on the annotation requirement information to obtain an intent recognition result; Using the large language model, according to the intention recognition result, the labeling requirement information is split into a plurality of subtasks, and a data processing model for processing each of the subtasks is determined; Using the data processing model, the data to be labeled are processed respectively to obtain processing results corresponding to each of the subtasks; According to the processing result corresponding to the subtask, a target labeling result for the data to be labeled is determined.