Data labeling method and device based on natural language processing, equipment and medium
Through the data annotation method based on natural language processing, the annotation model is used to label data according to the preset label classification system, which solves the problem of inconsistent data annotation in the existing technology, and realizes the efficiency and accuracy of data processing, and adapts to the diversity of natural language.
Patent Information
- Application Number
- CN202510571673.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
The lack of a unified and systematic data annotation system in the prior art has led to uneven accuracy of classification results, making it difficult to adapt to the diversity and complexity of natural language, and affecting the quality of model training and data processing efficiency.
The data labeling method based on natural language processing is adopted, and the data is marked using the labeling model, and the preset label classification system is marked, including classification subsystem and hierarchical description to ensure that the data of different information sources are marked according to the same rule.
It improves the consistency and accuracy of data processing, reduces manual annotation dependence, improves data processing efficiency, adapts to the diversity of natural language, and enhances the applicability of the model in different contexts.
Smart Images

Figure CN120493098A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of data processing technology, artificial intelligence technology, large model technology, and large language model technology, and specifically to data annotation methods, devices, equipment, and media based on natural language processing. Background Art
[0002] Understanding the true needs and problem tendencies of large-scale models is crucial for training. Furthermore, the acquisition and processing of large-scale corpora form the foundation for large-scale model training. Data generated by various information sources is a valuable resource. Labeling this data allows for more efficient management and optimization of data resources. Therefore, labeling acquired data is a pressing issue. Summary of the Invention
[0003] In view of this, the present application provides a data annotation method, apparatus, device and medium based on natural language processing to solve the problem of data annotation.
[0004] In a first aspect, the present application provides a data annotation method based on natural language processing, comprising:
[0005] Acquire first data, where the first data comes from one or more preset information sources;
[0006] The first data is labeled using a labeling model to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes a labeling level and the labeling content under the labeling level.
[0007] In a second aspect, the present application provides a data annotation device based on natural language processing, comprising:
[0008] A first data acquisition module is used to acquire first data, where the first data comes from one or more preset information sources;
[0009] A labeling module is used to label the first data using a labeling model to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes a labeling level and the labeling content under the labeling level.
[0010] In a third aspect, the present application provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data labeling method based on natural language processing of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0011] In a fourth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data labeling method based on natural language processing of the above-mentioned first aspect or any corresponding embodiment thereof.
[0012] In a fifth aspect, the present application provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the data labeling method based on natural language processing of the above-mentioned first aspect or any corresponding embodiment thereof.
[0013] The data labeling method based on natural language processing provided in an embodiment of the present application, after obtaining the first data from one or more preset information sources, uses the labeling model to label the first data to obtain labeling information. Among them, the labeling model is configured to label the input data based on a preset label classification system, and the preset label classification system includes a classification subsystem, and under the classification subsystem, there is a hierarchical description corresponding to the classification category, and the labeling information of the first data includes the labeling level and the labeling content under the labeling level. The labeling model in this method labels the first data from different information sources, that is, the data from different information sources can be labeled according to the same preset label classification system to obtain labeling information with unified rules. At the same time, the preset label classification system includes a classification subsystem, and under the classification subsystem, there is a hierarchical description corresponding to the classification category. Under this labeling rule system, accurate and clear labeling information can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the specific implementation methods of this application or the technical solutions in related technologies, the following is a brief introduction to the drawings required for use in the specific implementation methods or related technical descriptions. Obviously, the drawings described below are some implementation methods of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present application;
[0016] Figure 2 is a flowchart of a data annotation method based on natural language processing according to an embodiment of the present application;
[0017] Figure 3 is a schematic diagram of a preset label classification system according to an embodiment of the present application;
[0018] Figure 4 is a schematic diagram of a hierarchical description of scene content according to an embodiment of the present application;
[0019] Figure 5-Figure 6 is a schematic diagram of a hierarchical description of data content according to an embodiment of the present application;
[0020] Figure 7 is a schematic diagram of a hierarchical description of a knowledge base according to an embodiment of the present application;
[0021] Figure 8 is a flowchart of another data annotation method based on natural language processing according to an embodiment of the present application;
[0022] Figure 9 is a structural block diagram of a data annotation device based on natural language processing according to an embodiment of the present application;
[0023] Figure 10 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0029] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0030] In related technologies, different teams use their own standards for data labeling, such as rule-based or keyword-based labeling. This approach can lead to inconsistent classification accuracy and can also cause conflicts during data integration and sharing. Therefore, the solutions in related technologies primarily lack a unified and systematic classification system. This lack of a classification system makes it difficult to ensure consistency and reliability in the classification process when processing and analyzing large-scale data.
[0031] In addition, since text data is represented by natural language, which itself is highly diverse and complex, the way language is expressed can be due to cultural, regional, contextual and other factors, which makes it difficult for data annotation in related technologies to fully cover all possible language variants.
[0032] As described above, related technologies typically use rule-based or keyword-based tagging. This approach lacks flexibility and is difficult to adapt to the ever-changing language environment. This uncertainty and complexity further exacerbates the difficulty of classification and affects the overall effectiveness of data processing.
[0033] When subsequent training tasks are heavy, inaccurate or inconsistent labeling results will directly affect the quality of model training, resulting in poor performance in practical applications. In addition, low labeling efficiency may also lead to delays in data processing, affecting the timeliness and effectiveness of business decisions. Therefore, there is an urgent need to provide a data labeling method to address these challenges.
[0034] Based on this, an embodiment of the present application provides a data labeling method based on natural language processing, which, after obtaining first data from one or more preset information sources, uses a labeling model to label the first data to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system, and the preset label classification system includes a classification subsystem, under which there is a hierarchical description corresponding to the classification category, and the labeling information of the first data includes a labeling level and the labeling content under the labeling level. The labeling model in this method labels the first data from different information sources, that is, the data from different information sources can be labeled according to the same preset label classification system to obtain labeling information with unified rules. At the same time, the preset label classification system includes a classification subsystem, and under the classification subsystem there is a hierarchical description corresponding to the classification category. Under this labeling rule system, accurate and clear labeling information can be obtained.
[0035] In the embodiments of the present application, the presence of a preset label classification system ensures consistency and standardization of data processing, thereby improving the accuracy of data labeling. Secondly, data labeling is automatically completed based on the labeling model, that is, automated labeling. This automated labeling function greatly improves the efficiency of data processing, reduces reliance on manual labeling, and saves time and human resources. In addition, this method can better adapt to the diversity and complexity of natural language and enhance the ability to respond to user needs.
[0036] As used in the embodiments of the present application, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology, etc., taking deep learning as an example, deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In the embodiments of the present application, the model can also be referred to as a machine learning model, a machine learning network or a network, and these terms can be used interchangeably in this article. Among them, a model can also include different types of processing units or networks.
[0037] As an optional application scenario of the embodiment of the present application, Figure 1 As shown, it includes: server 101, log 102 and data pool 103. The data of the preset information source is stored in the corresponding log, that is, the log is used to store data from the corresponding preset information source, including but not limited to text, etc., and there is no limitation on the data type of the data in the log.
[0038] Server 101 can be an edge node, a data center, or the like, with no specific limitation. The data annotation method based on natural language processing described in the embodiments of this application is provided in server 101. Server 101 annotates the data in log 102 to obtain corresponding annotation information, and associates the data and annotation information and stores them in data pool 103. A annotation model is deployed in server 101, and the annotation model is used to label the data in log 102 based on a preset label classification system.
[0039] The log 102 and the data pool 103 may be deployed in the server 101, or in other servers, or on other terminals, etc., and no limitation is made here.
[0040] Exemplarily, server 101 is used to deploy a data annotation platform, which is in communication with a preset information source. The data annotation platform labels the data produced by the preset information source and stores it in a data pool 103. The data produced by the data annotation platform can be stored in a corresponding log 102, from which the data annotation platform obtains data for labeling.
[0041] In some embodiments, the data pool 103 may provide a data access interface, allowing authorized business access requests to obtain labeled data from the data pool 103 for business processing. For example, before training a business model, a business access request may be used to obtain sample data of the currently trained business model from the data pool 103 to form a sample dataset, which is used to train the business model.
[0042] Of course, there are other applications for the marking data in the data pool 103, which are not limited here and can be set according to actual needs.
[0043] According to an embodiment of the present application, an embodiment of a data labeling method based on natural language processing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] In this embodiment, a data annotation method based on natural language processing is provided, which can be used in the above-mentioned server. Figure 2 is a flow chart of a data annotation method based on natural language processing according to an embodiment of the present application. Figure 2 As shown, the process includes the following steps:
[0045] Step S201: Acquire first data.
[0046] The first data comes from one or more preset information sources.
[0047] The preset information source can be a preset data processing platform, such as a resource recommendation platform, a resource interaction platform, a video playback platform, etc., which is not limited here. The data annotation method based on natural language processing in this embodiment can be executed by a data annotation platform deployed on a server, which accesses one or more of the preset information sources to implement labeling of the data generated by these preset information sources.
[0048] The data generated by the preset information source includes, but is not limited to, data entered by users after interacting with the corresponding page of the preset information source, and may also be data obtained after the preset information source performs data processing, etc. There is no limitation on the data generated by the preset information source, and it can be set according to actual needs.
[0049] The type of the first data may be text. For example, the data type produced by the preset information source may be a variety of media types. The data of these media types may be converted into text described in natural language, and the converted text is referred to as the first data.
[0050] Step S202: label the first data using the labeling model to obtain labeling information.
[0051] Among them, the labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes the labeling level and the labeling content under the labeling level.
[0052] The labeling model receives input including first data and outputs labeling information for the first data. The labeling model is configured to label data based on a preset label classification system, which is applicable to all preset data sources connected to the data labeling platform. Alternatively, it can be understood that the labeling model labels data generated by all preset data sources connected to the data labeling platform using the same preset label classification system.
[0053] The preset label classification system represents the labeling architecture under at least one classification subsystem. The classification subsystems represent different division dimensions. For example, the preset label classification system includes three classification subsystems, namely, classification subsystem 1 to classification subsystem 3. During the labeling process, you can first label according to classification subsystem 1, then according to classification subsystem 2, and finally according to classification subsystem 3.
[0054] The classification subsystem has a hierarchical description corresponding to the classification category. The classification category represents different classification dimensions. The hierarchical description includes the division of the label hierarchy under the classification category and the label content under each label hierarchy. For example, classification subsystem 1 includes two classification categories, classification category 1 and classification category 2. The label hierarchy under classification category 1 includes first-level labels, second-level labels, and third-level labels. The label hierarchy under classification category 2 includes first-level labels, second-level labels, and each label hierarchy has corresponding label content.
[0055] It should be noted that the classification subsystem under the preset label classification system and the hierarchical description under the classification subsystem are set according to actual needs and are not limited here.
[0056] The preset label classification system can be represented by a configuration file, and the annotation model labels the input data by reading the configuration file. In addition, when the preset label classification system needs to be updated, the configuration file can be updated synchronously.
[0057] In some optional implementations, the classification subsystem includes one or more of scene content, data content, and a knowledge base.
[0058] The scenario content represents the scenarios involved in the data to be labeled, including but not limited to business-specific analysis, enhanced analysis, and advanced analysis; the data content represents the specific content in the data to be labeled, including but not limited to time description and visualization.
[0059] The hierarchical description includes the hierarchical levels under the classification category and the hierarchical content under the hierarchical level. The hierarchical level represents the tag level, for example, the number of levels of tags included; the hierarchical content represents the tag content under each level of tags.
[0060] For example, Figure 3 As shown, the preset tag classification system includes three classification subsystems 301, namely scene content, data content and knowledge base. Under the scene content classification subsystem, there is one classification category 302, and the label level under this classification category is 2 levels, and there are multiple second-level labels corresponding to the first-level label, and the corresponding labels have corresponding label content. Under the data content classification subsystem, there are two classification categories, and the label level under each classification category is 2 levels, and there are multiple second-level labels corresponding to the first-level label, and the corresponding labels have corresponding label content. Under the knowledge base classification subsystem, there is one classification category, and the label level under this classification category is 2 levels, and there are multiple second-level labels corresponding to the first-level label, and the corresponding labels have corresponding label content.
[0061] It is understandable that Figure 3 This is only an example and does not limit the scope of protection of this application.
[0062] Scene content and text content are labeled for data obtained from the data source. In addition, the data in the knowledge base also has annotation value. Therefore, including these three in the classification subsystem can ensure the richness of the data source and provide rich sample data for downstream model training after labeling.
[0063] In some optional embodiments, the classification categories under each classification subsystem in the preset tag classification system include primary tags and secondary tags. For example, the classification categories under scenario content include business scenarios, and the classification categories under data content include time descriptions and visualization descriptions.
[0064] For example, Figure 4 As shown, the classification subsystem is scenario content, and the classification category is business scenario. Within the business scenario, there are primary and secondary tags. The primary tags include business-specific analysis, enhanced analysis, advanced analysis, and basic interpretation, among others. For the primary tag "Business-specific analysis," the secondary tags include path analysis, strategy analysis, and so on. For the primary tag "Enhanced analysis," the secondary tags include attribution analysis, data interpretation, and so on.
[0065] Exemplarily, in a business scenario, the annotation information provided by the annotation model to the first data may be business topic analysis-path analysis / strategy analysis, or enhanced analysis-attribution analysis.
[0066] like Figure 5 As shown in the figure, the classification subsystem is data content, the classification category is time description, and under the time description, there are primary tags and secondary tags. The tag content of the primary tag includes time description, etc. The primary tag: time description. Correspondingly, the tag content of the secondary tag includes relative date and absolute date, etc.
[0067] Exemplarily, under the time description, the annotation information given by the annotation model to the first data may be time description-relative date, or time description-absolute date.
[0068] like Figure 6 As shown, the classification subsystem is data content, and the classification category is visualization description. Under visualization description, there are primary tags and secondary tags. The primary tags include numerical visualization and chart visualization, etc. The primary tag: numerical visualization, correspondingly, the secondary tags include the specified numerical display format and the specified numerical display style, etc. The primary tag: chart visualization, correspondingly, the secondary tags include the bar chart and line chart, etc.
[0069] Exemplarily, under the visualization description, the annotation information given by the annotation model to the first data may be numerical visualization - a specified numerical display style, or chart visualization - a bar chart.
[0070] like Figure 7 As shown, the classification subsystem is the knowledge base, and the classification category is others, including first-level labels and second-level labels. The label content of the first-level labels includes others, and correspondingly, the label content of the second-level labels includes general knowledge and high-base dimensions, etc.
[0071] Exemplarily, under the knowledge base, the annotation information given by the annotation model to the first data may be other-general knowledge / high-base dimension.
[0072] By setting primary and secondary tags, the annotation scenarios can be refined to meet refined data requirements.
[0073] The natural language processing-based data labeling method provided in this embodiment labels first data from different information sources using a labeling model. Specifically, data from different information sources can be labeled according to the same preset label classification system, resulting in uniformly labeled information. Furthermore, the preset label classification system includes a classification subsystem, each with a hierarchical description corresponding to the classification category. Under this labeling rule system, accurate and clear labeling information can be obtained.
[0074] In some optional embodiments, the above-mentioned data labeling method based on natural language processing also includes: filtering the labeling information based on the hierarchical priority of the classification category under the classification subsystem to obtain the target labeling information of the first data, and the hierarchical priority is used to represent the priority of the hierarchical content under the same level.
[0075] As described above, in the preset label classification system, the first-level label of the same classification category includes at least one label content, and the first-level label corresponds to at least one second-level label. Therefore, when labeling the first data, the labeling model may output multiple label contents for the same labeling level under the same classification subsystem. Therefore, in order to make the label content more able to focus on the key content in the data, the labeling information output by the labeling model is filtered in combination with the hierarchical priority of the classification category.
[0076] For example, Figure 4 As shown in the following figure, under the scenario content classification subsystem, the hierarchical priorities of the first-level tag content are: business theme analysis > enhanced analysis > advanced analysis > basic interpretation > ...; the hierarchical priorities of the second-level tag content corresponding to the first-level tag content business theme analysis are: path analysis > strategy analysis > ...
[0077] When filtering annotation information using the hierarchical priorities of the classification categories within the classification subsystem, the hierarchical priorities of the first-level tags can be compared first, followed by the hierarchical priorities of the second-level tags. This high-level to low-level filtering method can be used to characterize the accuracy of the resulting target annotation information.
[0078] There may be multiple situations for the annotation information given by the annotation model. The hierarchical priority of the classification categories under the classification system is used to filter the annotation information, so that the final target annotation information pays more attention to the important content in the data.
[0079] In some optional implementations, filtering the annotation information based on the hierarchical priorities of the classification categories under the classification subsystem to obtain target annotation information of the first data includes:
[0080] Step a1: determine whether there are multiple annotation contents at the same level under the same classification category of the same classification subsystem.
[0081] Step a2: If there are multiple annotation contents, compare the priorities of the annotation contents, screen the multiple annotation contents, and obtain target annotation information of the first data.
[0082] For example, if the annotation information obtained after the first data passes through the annotation model is: business topic analysis-path analysis, business topic analysis-strategy analysis, enhanced analysis-attribution analysis.
[0083] When performing hierarchical priority comparison, first compare the label content of the first-level labels, and retain the annotation information corresponding to the label content with the highest priority. In the above example, the label content of the first-level labels involved is business topic analysis and enhanced analysis. Since the hierarchical priority of business topic analysis is greater than the hierarchical priority of enhanced analysis, the retained annotation information is: business topic analysis-path analysis, business topic analysis-strategy analysis. When the label content of the first-level labels is the same, compare the hierarchical priority of the label content of the second-level labels. Since the hierarchical priority of path analysis is greater than the hierarchical priority of strategy analysis, the target annotation information of the first data is business topic analysis-path analysis.
[0084] For example, the annotation information includes one or more annotation paths, and the annotation paths include primary annotation content and secondary annotation content. That is, the annotation content at each annotation level is represented by the annotation path, making the annotation information clear and easy to understand.
[0085] Specifically, the annotation path can be represented as the tag content of the first-level tag - the tag content of the second-level tag. If there are only two tag levels, and the last tag level has multiple tag contents, multiple annotation paths can be used to represent it: the tag content of the first-level tag - the tag content 1 of the second-level tag, and the tag content of the first-level tag - the tag content 2 of the second-level tag; or, a single annotation path can be used to represent it: the tag content of the first-level tag - the tag content 1 / tag content 2 of the second-level tag.
[0086] In some optional implementations, the above-mentioned data annotation method based on natural language processing further includes:
[0087] Step b1: Acquire second data from a preset information source.
[0088] Step b2: Screen the second data using the rejection model, and determine the second data that is business data as the first data.
[0089] Since the second data from the preset information source is diverse, some of which involves business content while others does not. Therefore, in order to reduce the data processing workload of the subsequent annotation model, after obtaining the second data from the preset information source, the rejection model is used to screen the second data and determine whether the second data is business data. If the second data is business data, the second data is determined as the first data subsequently input into the annotation model.
[0090] The rejection model can be built based on the language model. Training the rejection model can be done by fine-tuning the pre-trained language model by adding fine-tuning data. The fine-tuning data includes both positive and negative samples from business scenarios.
[0091] Of course, the rejection model can also be implemented based on other model architectures, and there is no limitation on it here. It only needs to ensure that the trained rejection model can determine whether the second data belongs to business data.
[0092] The rejection model is used to filter the secondary data from the pre-set data source, including marking the secondary data as business data. This means that by identifying and filtering out data that does not meet the analysis requirements, this screening effectively eliminates irrelevant or invalid data, ensuring the accuracy and effectiveness of subsequent processing.
[0093] In some optional implementations, the above-mentioned data annotation method based on natural language processing further includes:
[0094] Step c1: Obtain label requirements of a training data set. The training data set is used to train a target business model. The target business model corresponds to the label requirements.
[0095] Step c2: screening the data pool based on label requirements to obtain sample data and labels of the sample data to form a training data set. The data pool includes the first data and the labeling information of the first data.
[0096] As described above, the labeled first data can be stored in a data pool to facilitate downstream business model training. Specifically, before training the target business model, a training dataset must be obtained. The requirements for the training data in the training dataset are related to the target business model. In other words, the label requirements for the training dataset must be determined.
[0097] The first data stored in the data pool is labeled. By filtering the data pool according to the label requirements, sample data that meets the label requirements can be obtained. Accordingly, these sample data have corresponding labels. These sample data and their labels form the training data set. Of course, the source of the training data in the training data set is not limited to the data pool, but can also include other sources.
[0098] The data pool includes the first data and the label information of the first data, that is, the labeled data is stored in the same data pool, which is convenient for subsequent screening of sample data when the label requirements of the training data set are obtained to meet the corresponding business needs.
[0099] As a specific application embodiment of the present application, a data annotation platform is deployed in the server, and the working principle of the data annotation platform is as follows: Figure 8 As shown. Specifically, the online log 801 corresponds to a preset information source, for example, a resource recommendation platform, a video playback platform. The text data generated by these preset information sources is stored in the online log (unlabeled) 801. The rejection model 802 of the data annotation platform filters 803 whether the text in the online log (unlabeled) 801 is business data. If it is non-business data, no subsequent analysis is performed. If it is business data, the annotation model 804 is used to label the data to obtain annotation information. Among them, the annotation model 804 labels the data based on a preset tag classification system. Combined with the rule screening 805, the annotation information is filtered to obtain the target annotation information of the data, and the target annotation information can be written into another online log (with labels) 806.
[0100] The rejection model 802 is obtained by fine-tuning the first language model using the rejection fine-tuning data, and the labeling model 804 is obtained by fine-tuning the second language model using the marking fine-tuning data.
[0101] The data labeling method based on natural language processing provided in the embodiment of the present application uses a unified preset label classification system to label the data, so that the consistency in the data processing process is improved. The labeling method based on the preset label classification system ensures that the classification results between different teams and systems are more consistent, thereby improving the reliability and accuracy of data analysis and reducing conflicts during data integration and sharing. This labeling method is an automated labeling method that can improve the efficiency of data labeling and reduce dependence on manual labeling. That is, it not only saves time and human resources, but also speeds up data processing, making the management and analysis of large-scale data more efficient and able to meet business needs more promptly.
[0102] Because the rejection and annotation models are fine-tuned using the corresponding fine-tuning data, they can better adapt to the diversity and complexity of natural language, enhancing their ability to respond to diverse language expressions and user needs. This flexibility ensures the model's applicability in different contexts and user groups, improving its practicality in real-world applications and ultimately enhancing the overall user experience.
[0103] In this embodiment, a data annotation device based on natural language processing is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0104] This embodiment provides a data annotation device based on natural language processing, such as Figure 9 Shown, including:
[0105] The first data acquisition module 901 is used to acquire first data, where the first data comes from one or more preset information sources.
[0106] The labeling module 902 is used to label the first data using the labeling model to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes the labeling level and the labeling content under the labeling level.
[0107] In some optional implementations, the classification subsystem includes one or more of scene content, data content, and a knowledge base; and the hierarchical description includes the hierarchies under the classification categories and the hierarchical content under the hierarchies.
[0108] In some optional implementations, the classification categories under the classification subsystem include primary tags and secondary tags; the classification categories under the scenario content include business scenarios; and the classification categories under the data content include time descriptions and visualization descriptions.
[0109] In some optional implementations, the data annotation apparatus based on natural language processing further includes:
[0110] The annotation screening module is used to filter the annotation information based on the hierarchical priority of the classification category under the classification subsystem to obtain the target annotation information of the first data. The hierarchical priority is used to represent the priority of the hierarchical content under the same level.
[0111] In some optional implementations, the annotation screening module includes:
[0112] The annotation content determination unit is used to determine whether the annotation information has multiple annotation contents of the same level under the same classification category of the same classification subsystem.
[0113] The screening unit is configured to compare priorities of the marked contents if there are multiple marked contents, screen the multiple marked contents, and obtain target marked information of the first data.
[0114] In some optional implementations, the annotation information includes one or more annotation paths, and the annotation paths include primary annotation content and secondary annotation content.
[0115] In some optional implementations, the data annotation apparatus based on natural language processing further includes:
[0116] A second data acquisition module, configured to acquire second data from the preset information source;
[0117] The data screening module is used to screen the second data using a rejection model and determine the second data belonging to business data as the first data.
[0118] In some optional implementations, the data annotation apparatus based on natural language processing further includes:
[0119] The label requirement acquisition module is used to obtain the label requirements of the training data set. The training data set is used to train the target business model, and the target business model corresponds to the label requirements.
[0120] The sample data screening module is used to screen the data pool based on label requirements to obtain sample data and labels of the sample data to form a training data set. The data pool includes first data and labeling information of the first data.
[0121] The data labeling device based on natural language processing provided by the embodiment of the present disclosure can execute the data labeling method based on natural language processing provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. The device can label data from different information sources according to the same preset label classification system to obtain labeling information with unified rules. At the same time, the preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. Under this label rule system, accurate and clear labeling information can be obtained. The further functional description of each of the above modules and units is the same as that of the corresponding embodiment above, and will not be repeated here.
[0122] Figure 10 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.
[0123] The following specific reference Figure 10 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a memory 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device are also stored in the RAM 1003. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0124] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.
[0125] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the memory 1008, or installed from the ROM 1002. When the computer program is executed by the processor 1001, the above-mentioned functions defined in the data annotation method based on natural language processing of the embodiment of the present disclosure are performed.
[0126] Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0127] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the data annotation method based on natural language processing shown in the above embodiment is implemented.
[0128] Part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0129] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A data annotation method based on natural language processing, characterized in that: include: Acquire first data, where the first data comes from one or more preset information sources; The first data is labeled using a labeling model to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes a labeling level and the labeling content under the labeling level.
2. The method according to claim 1, characterized in that The classification subsystem includes one or more of scene content, data content and knowledge base; The hierarchical description includes the hierarchical levels under the classification category and the hierarchical contents under the hierarchical levels.
3. The method according to claim 2, characterized in that The classification categories under the classification subsystem include primary labels and secondary labels; The classification categories under the scenario content include business scenarios; The classification categories under the data content include time description and visualization description.
4. The method according to claim 1, wherein Also includes: The annotation information is filtered based on the hierarchical priority of the classification category under the classification subsystem to obtain target annotation information of the first data, wherein the hierarchical priority is used to represent the priority of hierarchical content under the same hierarchy.
5. The method according to claim 4, characterized in that The filtering of the annotation information based on the hierarchical priority of the classification category under the classification subsystem to obtain target annotation information of the first data includes: Determine whether the annotation information has multiple annotation contents at the same level under the same classification category in the same classification subsystem; If the plurality of annotation contents exist, the priorities of the annotation contents are compared, and the plurality of annotation contents are screened to obtain target annotation information of the first data.
6. The method according to claim 5, characterized in that The annotation information includes one or more annotation paths, and the annotation path includes primary annotation content and secondary annotation content.
7. The method according to claim 1, characterized in that Also includes: Acquiring second data from the preset information source; The second data is screened using a rejection model, and the second data belonging to business data is determined as the first data.
8. The method according to any one of claims 1 to 7, characterized in that Also includes: Obtaining label requirements for a training data set, where the training data set is used to train a target business model, and the target business model corresponds to the label requirements; Based on the label requirement, the data pool is screened to obtain sample data and labels of the sample data to form the training data set. The data pool includes the first data and the annotation information of the first data.
9. A data annotation device based on natural language processing, characterized in that: include: A first data acquisition module is used to acquire first data, where the first data comes from one or more preset information sources; A labeling module is used to label the first data using a labeling model to obtain labeling information. The labeling model is configured to label the input data based on a preset label classification system. The preset label classification system includes a classification subsystem, and the classification subsystem has a hierarchical description corresponding to the classification category. The labeling information includes a labeling level and the labeling content under the labeling level.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data labeling method based on natural language processing as described in one of claims 1 to 8 by executing the computer instructions.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the data labeling method based on natural language processing as described in one of claims 1 to 8.
12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data annotation method based on natural language processing according to one of claims 1 to 8.
Citation Information
Patent Citations
Business record classification method for hierarchical system and mixed data type
CN114428855A
Rhythm labeling method and device
CN118471184A
Data labeling method, electronic equipment, storage medium and program product
CN119646523A