Data processing method and device, electronic equipment and computer readable storage medium
By acquiring and classifying the initial annotation items of business data, determining candidate annotation items and performing data classification and annotation, the problem of low efficiency and accuracy of business data processing in the intelligent service system is solved, and more efficient and accurate data annotation is achieved.
Patent Information
- Application Number
- CN202311776506.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when processing business data, intelligent service systems are affected by manual factors, resulting in low processing efficiency and accuracy.
By obtaining business data and its initial annotation items, candidate annotation items are determined based on the business type and initial annotation items, data classification and labeling are carried out, and the efficiency and accuracy of the annotation are improved.
This method can improve the efficiency and accuracy of business data processing, and improve the consistency and accuracy of labeling by focusing on the annotation of the same type of business data.
Smart Images

Figure CN120197797A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a data processing method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence, intelligent services have been widely applied to various industries. For example, in the financial service industry, intelligent services replace manual services, reducing the labor intensity and saving a large amount of labor costs. A large amount of business data is generated during the operation of intelligent services, and the business data includes, but is not limited to, voice, text, and pictures. In order to improve the service level, it is necessary to classify and label this business data, and then apply the processed business data to the operation department. However, currently, the business data is mainly processed manually, and this processing method is affected by subjective factors, with low accuracy and efficiency. Summary of the Invention
[0003] The present disclosure provides a data processing method and apparatus, an electronic device, and a computer-readable storage medium, which can improve the processing efficiency and accuracy of business data.
[0004] In a first aspect, the present disclosure provides a data processing method, which includes:
[0005] Obtaining a plurality of business data and initial annotation items of the business data, where the initial annotation items are preliminarily determined based on the intent of the business data;
[0006] Determining candidate annotation items based on the business type of the business data and the initial annotation items;
[0007] Classifying the business data based on the candidate annotation items and the initial annotation items of the business data to obtain a first classification result;
[0008] Annotating the business data in each first classification result to obtain an annotation result, where the annotation result is used to identify the intent of the business data.
[0009] In a second aspect, the present disclosure provides a data processing apparatus, which includes:
[0010] An obtaining module, configured to obtain a plurality of business data and initial annotation items of the business data, where the initial annotation items are preliminarily determined based on the intent of the business data;
[0011] A determining module, configured to determine candidate annotation items based on the business type of the business data and the initial annotation items;
[0012] A classification module, configured to classify the service data based on the candidate annotation items and the initial annotation items of the service data, so as to obtain a first classification result;
[0013] An annotation module, configured to annotate the service data in each first classification result to obtain an annotation result, where the annotation result is used to identify the intention of the service data.
[0014] In a third aspect, the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor, so that the at least one processor can execute the above data processing method.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program implements the above data processing method when executed by a processor / processing core.
[0016] In the data processing method provided by the embodiments of the present disclosure, after obtaining a plurality of service data and the initial annotation items of the service data, candidate annotation items are determined based on the service type of the service data and the initial annotation items, and then, the service data is classified based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result, so that the service data with the same or similar candidate annotation items is classified into one classification result. Finally, the service data in each first classification result is annotated. Since the service data with the same or similar candidate annotation items is annotated, compared with the disordered service data to be processed, focusing on annotating the same type of service data can improve the annotation efficiency, and annotating the same type of service data together has good annotation consistency, and the differences between the same type of service data are small, thereby improving the annotation accuracy.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0018] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, and do not constitute a limitation to the present disclosure. By describing the detailed exemplary embodiments with reference to the drawings, the above and other features and advantages will become more obvious to those skilled in the art. In the drawings:
[0019] Figure 1Flowchart of a data processing method provided by an embodiment of the present disclosure;
[0020] Figure 2 Application scenario of a data annotation processing process provided by an embodiment of the present application;
[0021] Figure 3 Flowchart of a data processing method provided by an embodiment of the present application;
[0022] Figure 4 Block diagram of a data processing device provided by an embodiment of the present disclosure;
[0023] Figure 5 Block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0024] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0025] Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0026] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0027] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "include" and / or "consist of" are used in this specification, the specified features, wholes, steps, operations, elements, and / or components are present, but one or more other features, wholes, steps, operations, elements, components, and / or their groups are not excluded. "Connection" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0029] During the operation of the intelligent service system, a large amount of business data will be generated. Processing these business data and accurately determining the true intention of each piece of business data is beneficial to improving the service ability of the intelligent service system. Currently, it is mainly manually annotated by annotators. When an annotator selects the intention of marking business data, the annotated data and the data to be annotated are compared. Among them, the background provides candidate annotation items for the data to be annotated, and the annotator can select the annotation item of the data to be annotated from the candidate annotation items according to the comparison result. This annotation method is inefficient and affected by human factors, resulting in low annotation accuracy.
[0030] Embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and computer-readable storage medium, which can improve the efficiency of business data processing to improve the accuracy of annotation.
[0031] The data processing method according to the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of performing cloud computing. This method can be implemented by a processor calling computer-readable program instructions stored in a memory.
[0032] Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present disclosure. Refer to Figure 1 , this method includes:
[0033] Step S101, obtain multiple business data and initial annotation items of the business data.
[0034] The business data in the embodiments of the present disclosure can be the original business data generated during the operation of the intelligent service system and the data determined after screening.
[0035] The type of the original business data can be voice, text, and picture information. Among them, the voice can be the conversation generated during the conversation between the intelligent service system and the customer, the text can be the chat information generated during the chat between the intelligent service system and the customer, and the picture information can be the picture uploaded by the customer to the intelligent service system, and the voice or text generated after picture analysis.
[0036] In some embodiments, obtaining multiple business data includes: obtaining the original business data and the business type of the original business data; screening the original business data according to the business type to obtain the business data to be processed.
[0037] Among them, the business type can be the business type of the service provided by the intelligent service system. For example, if the intelligent service system provides financial services, then the business type is finance. If the intelligent service system provides communication services, then the business type is communication. If the intelligent service system provides insurance services, then the business type is insurance.
[0038] The initial annotation item is initially determined based on the intention of the business data. Since the initial annotation item is the initially determined intention of the business data, it may not be consistent with the true intention. However, the embodiments of the present disclosure provide multiple initial annotation items, and subsequent data processing can be performed according to the multiple initial annotation items.
[0039] In some embodiments, obtaining the initial annotation item of the business data includes: inputting the business data into a pre-trained model; determining the initial annotation item based on the model annotation result output by the pre-trained model.
[0040] Among them, the pre-trained model is a pre-trained model for initially annotating business data. The pre-trained model can be a BERT model trained using the initial business data. The embodiments of the present disclosure do not limit the structure of the BERT model. The pre-trained model can initially annotate the business data and output an initial annotation result.
[0041] In some embodiments, the model annotation result output by the pre-trained model includes a model annotation item and a probability value of the model annotation item. Among them, the probability value (confidence) of the model annotation item refers to the probability that the model annotation item becomes the annotation item of the business data. The larger the probability value, the greater the probability of becoming the initial annotation item, and vice versa.
[0042] In some embodiments, based on the model annotation result output by the pre-trained model, determining the initial annotation item includes: determining the model annotation item whose probability value of the initial annotation item is greater than or equal to a preset probability threshold as the initial annotation item.
[0043] Among them, the probability threshold is preset, and the embodiments of the present disclosure do not limit the specific value of the probability threshold.
[0044] In the embodiments of the present disclosure, the model annotation items with probability values greater than or equal to the probability threshold are determined as the initial annotation items, and the model annotation items with probability values less than the probability threshold are discarded because the probability of becoming an initial annotation item is small. This can reduce the number of initial annotation items and is beneficial to improving the efficiency of subsequent processing.
[0045] Step S102: Determine candidate annotation items based on the service type of the service data and the initial annotation items.
[0046] The service type refers to the service type of the services provided by the intelligent service system. For details, please refer to the above text and will not be elaborated here.
[0047] In the embodiments of the present disclosure, for different service data belonging to the same service type, the initial annotation items are not necessarily exactly the same. When there is a large amount of service data, the number of initial annotation items is large. To improve the processing efficiency of subsequent steps, the initial annotation items can be streamlined, and candidate annotation items can be selected from the initial annotation items.
[0048] For example, the initial annotation items determined by service data a include A, B, C, D, and E, and the initial annotation items determined by service data b include A, B, C, D, and F. The candidate annotation items of service data a and service data b can be determined as A, B, C, and D.
[0049] In some embodiments, the candidate annotation items can be determined according to the statistical probability of the service type of the service data and the initial annotation items. The larger the statistical probability value, the greater the probability that the initial annotation item becomes a candidate annotation item.
[0050] For example, there are 1000 pieces of service data in total, and these 1000 pieces of service data include 10 initial annotation items. Among them, the first initial annotation item A appears 700 times, and the statistical probability value is 70%; the second initial annotation item B appears 900 times, and the statistical probability value is 90%; the third initial annotation item C appears 800 times, and the statistical probability value is 80%; the fourth initial annotation item D appears 850 times, and the statistical probability value is 85%; the fifth initial annotation item E appears 1000 times, and the statistical probability value is 100%;..., the tenth initial annotation item J appears 100 times, and the statistical probability value is 10%. Then, the initial annotation items including A, B, C, D, and E can be determined as candidate annotation items.
[0051] In some embodiments, a probability threshold can be set, and the initial annotation items with probability values greater than or equal to the probability threshold are determined as candidate annotation items. The specific data of the probability threshold can be set by the user himself / herself, and the embodiments of the present disclosure do not limit this.
[0052] Step S103: Classify the service data based on the candidate annotation items and the initial annotation items of the service data to obtain the first classification result.
[0053] In some embodiments, the first classification result includes business data. That is to say, the first classification result is the corresponding business data itself. For example, by matching candidate annotation items and initial annotation items in the business data, the first classification result (business data) is obtained. The matching methods include: the candidate annotation items in the business data are the same as, different from, or partially the same as the initial annotation items, etc.
[0054] In some embodiments, the first classification result includes a classification label and business data. That is to say, the first classification result includes not only the business data itself, but also the classification label corresponding to the business data. Similarly, by matching candidate annotation items and initial annotation items in the business data, the first classification result is obtained. The first classification result includes the corresponding business data and the corresponding preset classification label for the matching method.
[0055] In some embodiments, multiple business data are classified based on candidate annotation items and initial annotation items of the business data to obtain a first classification result, including: determining a reference annotation item based on the candidate annotation item and the business type of the business; classifying multiple businesses based on the candidate annotation item, the reference annotation item, and the initial annotation items of the business data to obtain a first classification result. Similarly, in this embodiment, the first classification result may include business data, or the first classification result includes business data and a classification label. Specifically, the first classification result is obtained by matching candidate annotation items, initial annotation items, and reference annotation items in the business data.
[0056] When classifying a business to obtain a first classification result, first determine a reference annotation item according to the candidate annotation item and the business type of the business. Among them, the reference annotation item can be determined according to statistical probability. For example, the initial annotation item E with a statistical probability value of 100% is determined as the reference annotation item.
[0057] In some embodiments, determining a reference annotation item based on the candidate annotation item and the business type of the business includes: randomly extracting any business data from multiple business data as reference business data; determining the matching degree between the reference business data and the candidate annotation item based on the business type of the reference business data; and determining the candidate annotation item with the highest matching degree as the reference annotation item.
[0058] For example, the business data is "Why is the interest rate so high? Can it be lowered?" The candidate annotation items determined in step S102 include: A. Potential customer, B. Apply for after-sales service, C. Appeal, D. Customer satisfaction, E. Business handling problem, F. Business consultation problem. If the type of this business data is finance, based on the type of the business data "Why is the interest rate so high? Can it be lowered?", match it with the candidate annotation items A, B, C, D, E, and F respectively according to the intention. If the matching degrees are 40%, 50%, 35%, 60%, 90%, and 20% respectively, then determine the candidate annotation item E as the reference annotation item.
[0059] The embodiment of the present disclosure can also determine the reference annotation item through manual judgment, that is, the annotator judges the initial annotation item with the highest matching degree with the business, and determines the initial annotation item with the highest matching degree as the reference annotation item.
[0060] After determining the reference annotation item, classify multiple services based on the candidate annotation item and the reference annotation item to obtain the first classification result.
[0061] In some embodiments, the business data of the first classification result includes one or more of the following:
[0062] The first type of data. The first type of data is that the initial annotation item is exactly the same as the candidate annotation item, that is, the business data with the initial annotation item exactly the same as the candidate annotation item belongs to the first type of data. Among them, exactly the same means that the initial annotation item and the candidate annotation item are not only the same in number, but also the annotation items are the same.
[0063] For example, if the candidate annotation items include A, B, C, D, E, and the reference annotation item is E. If the initial annotation item of the business data is A, B, C, D, E, the number of the candidate annotation item and the initial annotation item is 5, and the annotation items are all A, B, C, D, E, that is, the initial annotation item of the business data is exactly the same as the candidate annotation item, then this business data belongs to the first type of data.
[0064] The second type of data. The second type of data is that the initial annotation item includes the reference annotation item and some candidate annotation items, and moreover, some candidate annotation items do not include other than the reference annotation item, that is, the initial annotation item is the same as a part of the candidate annotation item, and the part that is the same as the candidate annotation item not only includes the reference annotation item, but also includes other candidate annotation items. This type of business data belongs to the second type of data.
[0065] For example, if the candidate annotation items include A, B, C, D, and E, the reference annotation item is E. If the initial annotation items of the business data are A, B, C, and E, that is, the initial annotation items of the business data include E, and also include the candidate annotation items A, B, and C, then the business data belongs to the second category of data. It should be noted that, in addition to the reference annotation item E, the initial annotation items of the business data may have one item that is the same as the candidate annotation item, or may have two or three items that are the same as the candidate annotation item, and this embodiment of the application does not limit this.
[0066] In addition to E, the initial annotation items of the business data are the same as the candidate annotation items, and different numbers of the initial annotation items are the same as the candidate annotation items. At this time, they can be sorted according to the number of the same candidate annotation items, and the more the same number, the higher the ranking.
[0067] For example, if the initial marking items of the first business data include A, B, and E, and the initial marking items of the second business data include A, B, D, and E, the initial marking items of the first business data and the candidate marking items, except for E, also include two initial marking items A and B, which are the same as the candidate marking items, and the initial marking items of the second business data and the candidate marking items, except for E, also include three initial marking items A, B, and D, which are the same as the candidate marking items. Therefore, the second business data is ranked before the first business data.
[0068] The business data is sorted according to the same number of initial annotation items and candidate annotation items. In subsequent annotation, the annotation can be performed according to the sorting result, which is conducive to improving the efficiency of annotation.
[0069] The third type of data: The third type of data is that the initial annotation items only include reference annotation items, and other initial annotation items are different from the candidate annotation items. This type of business data belongs to the third type of data.
[0070] For example, if the candidate annotation items include A, B, C, D, and E, the reference annotation item is E. If the initial annotation items of the business data include E, F, G, and H, that is, only E is the same, and the other initial annotation items are different from the candidate annotation items, therefore, the business data belongs to the third category of data.
[0071] The fourth category of data is that the initial marking items include some candidate marking items, and the reference marking items do not belong to the partial candidate marking items, that is, the initial marking items are the same as some candidate marking items, but the same part does not include the reference marking items. This type of business data belongs to the fourth category of data.
[0072] For example, if the candidate annotation items include A, B, C, D, and E, the reference annotation item is E. If the initial annotation items of the business data include A, B, C, and D, although the initial annotation items of the business data are the same as A, B, C, and D in the candidate annotation items, they are different from the reference annotation item E. Therefore, the business data belongs to the fourth category of data.
[0073] The fifth type of data is data in which the initial annotation items are completely different from the candidate annotation items, that is, the business data in which the initial annotation items and the candidate annotation items do not have any common items belongs to the fifth type of data.
[0074] For example, if the candidate annotation items include A, B, C, D, and E, the reference annotation item is E. If the initial annotation items of the business data include F, G, H, and I, the initial annotation items of the business data are completely different from the candidate annotation items, so the business data belongs to the fifth category of data.
[0075] It should be noted that, in this embodiment, if the first classification result includes business data and classification labels, the first category data, the second category data, the third category data, the fourth category data and the fifth category data are classification labels corresponding to the business data.
[0076] Step S104: annotate the business data in each first classification result to obtain an annotated result, where the annotated result is used to identify the intent of the business data.
[0077] Since the candidate annotation items of the business data in each first classification result are the same or similar, the annotation efficiency can be improved when annotating the business data. Moreover, the same type of business data can be annotated together, the annotation consistency is good, and the difference between the same type of business data is small, thereby improving the annotation accuracy.
[0078] Figure 2 This is an application scenario of a data annotation process provided by an embodiment of the present application. Figure 2 As shown in Figure 1, the data annotation process includes:
[0079] Step S201, obtaining a model to be trained and training samples.
[0080] The model to be trained may be a BERT model, and the training sample may be original business data, that is, the model to be trained is trained using original business data, which can increase the robustness of the model.
[0081] Step S202: train the model to be trained using the training samples.
[0082] When the output of the model to be trained meets the preset conditions, a pre-trained model is obtained.
[0083] The model to be trained can not only output model annotation items, but also output the probability values of model annotation items.
[0084] Step S203: input the business data into the pre-trained model, and the pre-trained model outputs the model annotation items and the probability values of the model annotation items, where the probability values are also called confidence levels.
[0085] Step S204: determining initial annotation items based on the model annotation items.
[0086] Model annotation items whose probability values in the model annotation results are greater than or equal to a preset probability threshold are determined as initial annotation items, wherein the probability threshold can be set by the user.
[0087] Step S205: data classification.
[0088] Determine candidate annotation items based on the business type of the business data and the initial annotation items; classify the business data based on the candidate annotation items and the initial annotation items of the business data to obtain a first classification result.
[0089] The method for determining the candidate annotation items may refer to step S102 , and the method for determining the first classification result may refer to step S103 , which will not be described in detail here.
[0090] Step S206: annotate the business data in each first classification result to obtain an annotation result. The annotation method can refer to step S104, which will not be described in detail here.
[0091] Step S207, uploading the annotated results to the operation department, and the operation department operates the business data according to the annotated results. The specific operations include but are not limited to archiving, querying, extracting, issuing, reviewing and exporting.
[0092] After step S207, the verified labeled data to be processed may also be used for model training to improve the accuracy of the pre-trained model.
[0093] Figure 3 A flowchart of a data processing method provided in an embodiment of the present application. Figure 3 As shown, the data processing method includes:
[0094] Step S301: Perform preliminary annotation on business data using an algorithm to obtain preliminary annotation items.
[0095] The business data is annotated using a pre-trained model, which outputs model annotation items and probability values of model annotation items. Preliminary annotation items are determined based on the model annotation items and probability values of the model annotation items.
[0096] Step S302: annotate the business data based on the preliminary annotated items.
[0097] Step S321, classify the business data to obtain data of the same type.
[0098] Determine candidate annotation items based on the business type of the business data and the initial annotation items; and classify multiple business data based on the candidate annotation items and the initial annotation items of the business data to obtain a first classification result.
[0099] As described above, the service data can be classified into the first type of data, the second type of data, the third type of data, the fourth type of data, and the fifth type of data. The specific definition methods of these five types of data can be found in the previous text. The specific classification method can refer to step S103 and will not be elaborated here again.
[0100] Step S322: Sort the data of the same type.
[0101] When the service data is the second type of data, the service data can be sorted according to the same quantity as the candidate annotation items, and the more the same quantity, the more forward.
[0102] In step S302, the service data in each first classification result is annotated to obtain an annotation result, which is more accurate than the preliminary annotation items in step S301.
[0103] Step S303: Upload the service data and the annotation result to the operation department.
[0104] Step S304: Annotate the service data to be processed in each first classification result to obtain an annotation result.
[0105] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0106] In the data processing method provided by the embodiment of the present disclosure, after obtaining multiple service data and the initial annotation items of the service data, candidate annotation items are determined based on the service type of the service data and the initial annotation items. Then, the service data is classified based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result, so that the service data with the same or similar candidate annotation items is classified into one classification result. Finally, the service data in each first classification result is annotated. Since the service data with the same or similar candidate annotation items is annotated, compared with the disordered service data, focusing on annotating the service data of the same type can improve the annotation efficiency, and annotating the service data of the same type together has good annotation consistency, and the differences between the service data of the same type are small, thus improving the annotation accuracy.
[0107] Figure 4 It is a block diagram of a data processing device provided by an embodiment of the present disclosure.
[0108] Referring to Figure 4 , the embodiment of the present disclosure provides a data processing device. The data processing device 400 includes:
[0109] An acquisition module 401, configured to acquire multiple pieces of service data and initial annotation items of the service data, where the initial annotation items are preliminarily determined based on the intent of the service data.
[0110] A determination module 402, configured to determine candidate annotation items based on the service type of the service data and the initial annotation items.
[0111] A classification module 403, configured to classify the service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result.
[0112] An annotation module 404, configured to annotate the service data in each first classification result to obtain an annotation result, where the annotation result is used to identify the intent of the service data.
[0113] In some embodiments, the annotation module 404 is further configured to determine a reference annotation item based on the candidate annotation item and the service type of the service; classify multiple services based on the candidate annotation item, the reference annotation item, and the initial annotation item of the service data to obtain a first classification result.
[0114] In some embodiments, the annotation module 404 is further configured to randomly extract any piece of service data from multiple pieces of service data as a reference service data; determine the matching degree between the reference service data and the candidate annotation item based on the service type of the reference service data; and determine the candidate annotation item with the highest matching degree as the reference annotation item.
[0115] In some embodiments, the first classification result includes one or more of the following classification data:
[0116] First type of data, where the first type of data is service data whose initial annotation items are completely consistent with the candidate annotation items;
[0117] Second type of data, where the second type of data is service data whose initial annotation items include the reference annotation item, and moreover, the initial annotation items further include at least some of the candidate annotation items other than the reference annotation item;
[0118] Third type of data, where the third type of data is service data whose initial annotation items include the reference annotation item but do not include other candidate annotation items other than the reference annotation item;
[0119] Fourth type of data, where the fourth type of data is service data whose initial annotation items include other candidate annotation items other than the reference annotation item; and
[0120] Fifth type of data, where the fifth type of data is service data whose initial annotation items are completely different from the candidate annotation items.
[0121] In some embodiments, the obtaining module 401 is further configured to obtain the original service data and the service type of the original service data; screen the original service data according to the service type to obtain the service data to be processed.
[0122] In some embodiments, the obtaining module 401 is further configured to input the service data into a pre-trained model, where the pre-trained model is a model that has been pre-trained and is used to perform preliminary annotation on the service data; determine the initial annotation items based on the model annotation results output by the pre-trained model.
[0123] In some embodiments, the model annotation results include model annotation items and probability values of the model annotation items. The obtaining module 401 is further configured to determine the model annotation items with probability values greater than or equal to a preset probability threshold in the model annotation results as the initial annotation items.
[0124] After the obtaining module of the data processing device provided by the embodiments of the present disclosure obtains multiple service data and the initial annotation items of the service data, the determining module determines candidate annotation items based on the service type and the initial annotation items of the service data. Then, the classification module classifies the service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result, so as to classify the service data with the same or similar candidate annotation items into one classification result. Finally, the annotation module annotates the service data in each first classification result. Since the annotation is performed on the service data with the same or similar candidate annotation items, compared with the disordered data to be processed, focusing on annotating the same type of service data can improve the annotation efficiency, and the annotation consistency is good when annotating the same type of service data together, and the differences of the same type of service data are small, thereby improving the annotation accuracy.
[0125] Figure 5 It is a block diagram of an electronic device provided by the embodiments of the present disclosure.
[0126] Referring to Figure 5 , the embodiments of the present disclosure provide an electronic device, which includes: at least one processor 501; at least one memory 502, and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein, the memory 502 stores one or more computer programs executable by at least one processor 501, and the one or more computer programs are executed by at least one processor 501 so that at least one processor 501 can execute the above data processing method.
[0127] In some embodiments, the processor 501 is configured to obtain a plurality of service data and initial annotation items of the service data, where the initial annotation items are preliminarily determined based on the intent of the service data; determine candidate annotation items based on the service type of the service data and the initial annotation items; classify the service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result; annotate the service data in each first classification result to obtain an annotation result, where the annotation result is used to identify the intent of the service data.
[0128] In some embodiments, the processor 501 is further configured to determine reference annotation items based on the candidate annotation items and the service type of the service; classify a plurality of services based on the candidate annotation items and the reference annotation items to obtain a first classification result.
[0129] In some embodiments, the processor 501 is further configured to randomly extract any service data from the plurality of service data as reference service data; determine the matching degree between the reference service data and the candidate annotation items based on the service type of the reference service data; and determine the candidate annotation item with the highest matching degree as the reference annotation item.
[0130] In some embodiments, the processor 501 is further configured to obtain original service data and the service type of the original service data; screen the original service data according to the service type to obtain the service data to be processed.
[0131] In some embodiments, the processor 501 is further configured to input the service data into a pre-trained model, where the pre-trained model is a model that is pre-trained and used to perform preliminary annotation on the service data; determine the initial annotation items based on the model annotation results output by the pre-trained model.
[0132] In some embodiments, the model annotation results include model annotation items and probability values of the model annotation items. The processor 501 is further configured to determine the model annotation items with probability values greater than or equal to a preset probability threshold in the model annotation results as the initial annotation items.
[0133] Each module in the above electronic device can be implemented in whole or in part by software, hardware, and combinations thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0134] The embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, where the computer program, when executed by a processor / processing core, implements the above data processing method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0135] Embodiments of the present disclosure also provide a computer program product, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above data processing method.
[0136] Those of ordinary skill in the art can understand that all or some of the steps in the above-disclosed methods, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium).
[0137] As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable program instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.
[0138] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0139] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0140] The computer program product described herein may be implemented specifically by hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0141] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0142] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, including instructions that implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0143] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0144] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0145] Example embodiments have been disclosed herein, and although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in connection with a particular embodiment may be used singly or in combination with other embodiments. Accordingly, those skilled in the art will appreciate that various forms and details may be changed without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A data processing method, characterized in that, including: obtaining a plurality of service data and initial annotation items of the service data, where the initial annotation items are determined based on the intent of the service data; determining candidate annotation items based on the service type of the service data and the initial annotation items; classifying the service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result; annotating the service data in each of the first classification results to obtain an annotation result, where the annotation result is used to identify the intent of the service data.
2. The method according to claim 1, wherein The classifying the service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result includes: determining a reference annotation item based on the candidate annotation items and the service type of the service; classifying the service based on the candidate annotation items, the reference annotation item, and the initial annotation items of the service data to obtain the first classification result.
3. The method according to claim 2, wherein The determining a reference annotation item based on the candidate annotation items and the service type of the service includes: randomly extracting any service data from a plurality of service data as reference service data; determining the matching degree between the reference service data and the candidate annotation items based on the service type of the reference service data; determining the candidate annotation item with the highest matching degree as the reference annotation item.
4. The method according to claim 3, wherein The service data of the first classification result includes one or more of the following: type I data, where the type I data is that the initial annotation item is exactly the same as the candidate annotation item; type II data, where the type II data is that the initial annotation item includes the reference annotation item and some of the candidate annotation items, and the some candidate annotation items do not include the reference annotation item; type III data, where the type III data is that the initial annotation item only includes the reference annotation item; type IV data, where the type IV data is that the initial annotation item includes some candidate annotation items, and the reference annotation item does not belong to the some candidate annotation items; type V data, where the type V data is that the initial annotation item is completely different from the candidate annotation item.
5. The method according to claim 1, wherein The obtaining a plurality of service data includes: obtaining original service data and the service type of the original service data; screening the original service data according to the service type to obtain the service data to be processed.
6. The method according to claim 1, wherein The obtaining the initial annotation items of the service data includes: inputting the service data into a pre-trained model, where the pre-trained model is a model that is pre-trained and used to perform preliminary annotation on the service data; determining the initial annotation items based on the model annotation results output by the pre-trained model.
7. The method according to claim 6, wherein The model annotation results include the model annotation items and the probability values of the model annotation items; The determining the initial annotation items based on the model annotation results output by the pre-trained model includes: determining the model annotation items with probability values greater than or equal to a preset probability threshold in the model annotation results as the initial annotation items.
8. A data processing device, characterized in that, including: an obtaining module, configured to obtain a plurality of service data and initial annotation items of the service data, where the initial annotation items are preliminarily determined based on the intent of the service data; A determination module, configured to determine candidate annotation items of the service data based on the service type of the service data and the initial annotation items; A classification module, configured to classify the multiple service data based on the candidate annotation items and the initial annotation items of the service data to obtain a first classification result; An annotation module, configured to annotate the service data in each first classification result to obtain an annotation result, where the annotation result is used to identify the intention of the service data.
9. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the data processing method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the data processing method according to any one of claims 1-7.