An event text intelligent labeling method and device based on a pre-trained model
By combining a pre-trained model and a deterministic policy gradient module in intelligent annotation of event text, the annotation decision network is optimized, solving the problems of insufficient sample data and poor annotation effect, and achieving higher annotation accuracy and improved accuracy for a few categories.
Patent Information
- Application Number
- CN202411881511.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing intelligent annotation technologies for event texts suffer from high training difficulty and poor annotation results due to insufficient or low-quality sample data, especially for a few categories where the annotation accuracy is insufficient.
We employ an intelligent event text annotation method based on a pre-trained model, combining a natural language processing sub-model and a deterministic policy gradient module. We use an annotation decision network to make sequence annotation decisions and optimize the annotation strategy to improve accuracy.
By continuously optimizing the annotation strategy of the deterministic policy gradient module, the accuracy of event text annotation and model performance have been improved, and the annotation accuracy for a few categories has been enhanced.
Smart Images

Figure CN119807394B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent labeling, and in particular to an event text intelligent labeling method and device based on a pre-trained model. BACKGROUND
[0002] Event text intelligent labeling is a method of using artificial intelligence technology to automatically and efficiently label and classify events. Through training models and algorithms, the system can automatically identify and label key information in events such as time, location, participants, and behavior, helping users quickly understand the important content and background information of events.
[0003] Currently, due to the multiple event types and lack of sample data in the field of intelligent labeling, event labeling is difficult, and existing event text intelligent labeling technology has high model training difficulty and poor labeling effect when facing complex event information. The performance of existing intelligent labeling technology largely depends on the quality and quantity of training data. If the training data is insufficient or of poor quality, the model may not be able to learn enough features to accurately perform the labeling task. In some cases, some classes in the dataset may have more samples than other classes, which causes the model to be biased towards the majority class, thereby affecting the labeling accuracy of the minority class. SUMMARY
[0004] The present application provides an event text intelligent labeling method and device based on a pre-trained model to improve the accuracy of event text labeling.
[0005] According to an aspect of the present application, an event text intelligent labeling method based on a pre-trained model is provided, comprising:
[0006] loading a pre-trained event text labeling model, the event text labeling model comprising a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module comprising a labeling decision network;
[0007] obtaining a to-be-labeled event text and state information of the to-be-labeled event text, inputting the to-be-labeled event text into the natural language processing sub-model for recognition to obtain text features of the to-be-labeled event text;
[0008] inputting the text features and the state information into the labeling decision network for sequence labeling decision to determine a labeling decision of the to-be-labeled event text, and labeling the to-be-labeled event text based on the labeling decision.
[0009] According to another aspect of the present application, an event text intelligent labeling device based on a pre-trained model is provided, comprising:
[0010] The event text labeling model loading module is configured to load a pre-trained event text labeling model, wherein the event text labeling model comprises a natural language processing sub-model and a deterministic policy gradient module.
[0011] The text feature recognition module is configured to obtain a to-be-labeled event text and state information of the to-be-labeled event text, input the to-be-labeled event text into the natural language processing sub-model for recognition, and obtain text features of the to-be-labeled event text.
[0012] The event text labeling module is configured to input the text features and the state information into the labeling decision network for sequence labeling decision, determine a labeling decision of the to-be-labeled event text, and label the to-be-labeled event text based on the labeling decision.
[0013] According to another aspect of the present application, an electronic device is provided, which comprises:
[0014] at least one processor; and
[0015] a memory connected to the at least one processor in communication; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the event text intelligent labeling method based on a pre-trained model according to any one of the embodiments of the present application.
[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the event text intelligent labeling method based on a pre-trained model according to any one of the embodiments of the present application when executed by the processor.
[0018] The technical scheme of the embodiments of the present application loads a pre-trained event text labeling model, wherein the event text labeling model comprises a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module comprises a labeling decision network; obtains a to-be-labeled event text and state information of the to-be-labeled event text, inputs the to-be-labeled event text into the natural language processing sub-model for recognition, and obtains text features of the to-be-labeled event text; inputs the text features and the state information into the labeling decision network for sequence labeling decision, determines a labeling decision of the to-be-labeled event text, and labels the to-be-labeled event text based on the labeling decision. The natural language processing sub-model and the deterministic policy gradient module are combined, the labeling strategy of the deterministic policy gradient module is continuously optimized through pre-training, and thus the accuracy of event text labeling is improved.
[0019] It is to be understood that the details set forth in the description contained herein do not limit the scope of the embodiments of the application. Other embodiments of the application will be readily apparent to one of ordinary skill in the art from the description herein, including the working examples. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those of ordinary skill in the art without any creative effort based on these drawings.
[0021] Figure 1 is a flowchart of an event text intelligent labeling method based on a pre-trained model provided by an embodiment of the present application;
[0022] Figure 2 is a framework diagram of an event text labeling model provided by the embodiment one of the present application;
[0023] Figure 3 is a labeling schematic diagram of a BIO sequence labeling mechanism provided by the embodiment one of the present application;
[0024] Figure 4 is a structural schematic diagram of an event text intelligent labeling device based on a pre-trained model provided by the embodiment two of the present application;
[0025] Figure 5 is a structural schematic diagram of an electronic device provided by the embodiment three of the present application. DETAILED DESCRIPTION
[0026] In order to make the technical personnel in the art better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort should be within the scope of protection of the present application.
[0027] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] Embodiment one
[0029] Figure 1 is a flowchart of an event text intelligent labeling method based on a pre-trained model provided by an embodiment of the application. The embodiment can be applicable to the labeling of event text in an event text intelligent labeling task. The method can be executed by an event text intelligent labeling device based on a pre-trained model. The event text intelligent labeling device based on a pre-trained model can be realized in the form of hardware and / or software. The event text intelligent labeling device based on a pre-trained model can be configured in a computer or a server or other electronic device. As shown in the figure, the method comprises: Figure 1
[0030] S110, load a pre-trained event text labeling model, the event text labeling model comprising a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module comprising a labeling decision network.
[0031] In an embodiment of the application, the event text labeling model comprises a natural language processing sub-model and a deterministic policy gradient module. The natural language processing sub-model is used to provide semantic understanding, and the deterministic policy gradient module is used to provide a dynamically optimized labeling strategy. Specifically, the natural language processing sub-model includes but is not limited to BERT, ERNIE, RoBERTa and other models. The deterministic policy gradient module includes a labeling decision network and a value network, which are usually neural networks. The labeling decision network is used to learn a deterministic policy and output a labeling strategy for the event text. The value network is used to evaluate the value of the labeling strategy output by the labeling decision network.
[0032] On the basis of the above-mentioned embodiments, optionally, the event text labeling model training method comprises: obtaining sample data of an event text labeling task, the sample data comprising labeled sample data of a first round and unlabeled sample data of the first round; iteratively performing the following steps, and obtaining a trained event text labeling model in the case of meeting a first iteration stop condition: training the event text labeling model based on the labeled sample data of the current round to obtain an updated event text labeling model; predicting the unlabeled sample data of the current round based on the updated event text labeling model to obtain a label probability distribution of each unlabeled sample data of the current round; sampling and labeling the unlabeled sample data of the current round based on the label probability distribution of each unlabeled sample data of the current round to obtain labeled sample data of the next round.
[0033] In the embodiments of the present application, the training process of the event text labeling model is as follows:
[0034] 1) Obtain sample data of an event text labeling task, the sample data comprising labeled sample data of a first round and unlabeled sample data of the first round. The sample data is event text data, the sample data comprises labeled sample data and unlabeled sample data, the number of the labeled sample data is less than the number of the unlabeled sample data, and the labeled sample data is used as the training data of the first round.
[0035] 2) Iteratively perform the following steps 2.1-2.3, and obtain a trained event text labeling model in the case of meeting a first iteration stop condition:
[0036] 2.1) Train the event text labeling model based on the labeled sample data of the current round to obtain an updated event text labeling model.
[0037] 2.2) Predict the unlabeled sample data of the current round based on the updated event text labeling model to obtain a label probability distribution of each unlabeled sample data of the current round; wherein the label probability distribution comprises the probability that each character or word in the unlabeled sample data belongs to each label category, and the sum of the probabilities of each label category for each character or word in the unlabeled sample data is 1. Specifically, the label probability distribution is related to the labeling mechanism used, for example, in the case of a BIO labeling mechanism, the label probability distribution comprises the probability that each character or word in the unlabeled sample data belongs to the three label categories of "B", "I" and "O".
[0038] 2.3) sample the unlabeled sample data of the current round based on the label probability distribution of each unlabeled sample data of the current round, and label the sampled unlabeled sample data to obtain labeled sample data of the next round. It should be noted that the unlabeled sample data obtained by sampling is the sample most helpful for improving the model.
[0039] It should be noted that the first iteration stopping condition can be that all sample data has been labeled. Specifically, in the case where all sample data has been labeled, the iteration is stopped, and the trained event text labeling model is obtained.
[0040] On the basis of the above embodiment, optionally, the event text labeling model is trained based on the labeled sample data to obtain an updated event text labeling model, including: fine-tuning the natural language processing sub-model based on the labeled sample data, identifying the labeled sample data based on the fine-tuned natural language processing sub-model, and extracting sample features of the labeled sample data; at each time step, input the sample features and the current state information of the labeled sample data to the labeling decision network for sequence labeling decision to obtain the labeling decision of the current state of the labeled sample data, and interact with the labeled sample data based on the labeling decision to determine experience data samples, and store the experience data samples in an experience buffer; wherein the experience data sample includes current state information, the labeling decision, the reward value and the next state information; iteratively execute the following steps: obtain a batch of experience data samples from the experience buffer, determine the loss function of the value network based on the batch of experience data samples; update the network parameters of the value network based on the loss function of the value network; evaluate the labeling decision based on the updated value network to obtain the value of the labeling decision, determine the gradient of the labeling decision network based on the value of the labeling decision, and update the labeling decision network based on the gradient.
[0041] Figure 2 is a framework diagram of the event text labeling model provided by the first embodiment of the present application, which follows the model framework as shown in Figure 2 In the embodiment of the present application, the training process of each round of the event text labeling model is as follows:
[0042] 1) Fine-tune the pre-trained natural language processing sub-model based on the labeled sample data to adapt the natural language processing sub-model to the characteristics of the event text labeling task;
[0043] 2) based on the fine-tuned natural language processing sub-model, identifying the labeled sample data and extracting sample features of the labeled sample data; wherein the sample features contain key information of the labeled text data in terms of semantics and event structure, specifically, the sample features include event triggers, arguments and predicted labels of each word or character in the labeled text data.
[0044] 3) at each time step, inputting the sample features and the current state information of the labeled sample data into the labeling decision network for sequence labeling decision, obtaining the labeling decision of the current state of the labeled sample data, and executing the labeling decision in the labeled sample data, determining the reward value and the next state information based on the executed labeled sample data, forming an experience data sample based on the current state information, the labeling decision, the reward value and the next state information, and storing the experience data sample in the experience buffer.
[0045] Wherein, the current state information is used to provide reference for the current labeling decision, specifically, the current state information includes the length of the text of the labeled sample data, the types of the contained words, the syntax structure, etc., and the accuracy of the labeled part, the labeling style, etc. in the labeled sample data, which is not limited here. The labeling decision is composed of continuous labeling actions, and each labeling action corresponds to the labeling type of a word or character in the labeled text data. The next state information is the state information of the labeled sample data after executing the labeling decision. The reward value as a reward signal is used to guide the learning direction of the labeling decision network, the closer the real label of the labeled sample data and the label corresponding to the labeling decision given by the labeling decision network, the greater the reward value. Specifically, the reward function of the reward value includes:
[0046]
[0047] Wherein, R represents the reward value, n represents the number of words and / or characters in the labeled sample data, y i represents the real label of the i-th word or character in the labeled sample data, x i represents the label corresponding to the labeling decision of the i-th word or character in the labeled sample data, I(·) is an indicator function, when y i = x i , the value of the indicator function is 1, otherwise 0.
[0048] Based on the above embodiments, optionally, the sequence labeling mechanism used in the sequence labeling decision process is a word-level sequence labeling mechanism. Specifically, word-level sequence labeling mechanisms include, but are not limited to, BIO labeling mechanisms, BIOES labeling mechanisms, BMES labeling mechanisms, etc., and are not limited here. Taking the BIO labeling mechanism as an example, in the intelligent annotation task of event text, we use the BIO sequence labeling mechanism to identify entities related to events in the text, including trigger words and arguments. A trigger word is a word in a sentence that causes an event to occur; it is key to understanding the event in the sentence and also represents the event type. The arguments surrounding it provide detailed information about this action or event. An argument refers to a word associated with a trigger word; its edges provide necessary information for the trigger word, such as participants, time, location, etc. The BIO sequence labeling mechanism provides a label of "B", "I", or "O" for each word or phrase in the sentence.
[0049] The event text annotation model identifies event information such as trigger words and arguments in event text according to the following rules in order to complete the task of intelligent annotation:
[0050] B-tag: Used to mark the first word of a trigger word or argument.
[0051] I-tags: Used to tag subsequent words of trigger words or arguments.
[0052] O-tags: Used to label words that do not belong to any trigger words or arguments.
[0053] like Figure 3 As shown, Figure 3 This is the BIO annotation corresponding to the event text "John attended a meeting on the 29th". The BIO sequence annotation mechanism can help the model understand and extract event structure information from the text through intelligent annotation of event text, laying the foundation for further event analysis and information extraction.
[0054] 4) Iteratively execute steps 4.1-4.2. If the second iteration stopping condition is met, the updated event text annotation model is obtained:
[0055] 4.1) Obtain a batch of empirical data samples from the empirical buffer, determine the loss function of the value network based on the batch of empirical data samples, and update the network parameters of the value network based on the loss function of the value network.
[0056] 4.2) Evaluate the labeling decision based on the updated value network to obtain the value of the labeling decision, determine the gradient of the labeling decision network based on the value of the labeling decision, and update the labeling decision network based on the gradient.
[0057] It should be noted that the second iteration stopping condition can be whether the loss function of the value network converges, and specifically, in the case where the loss function of the value network converges, the iteration is stopped to obtain an updated event text labeling model. The updated event text labeling model includes a fine-tuned natural language processing sub-model and an iterative training deterministic policy gradient module.
[0058] On the basis of the above-mentioned embodiments, optionally, the loss function of the value network is determined based on the batch of experience data samples, including: for each experience data sample, evaluating the labeling decision based on the value network to obtain a predicted value of the labeling decision; and determining a target value based on the reward value, the discount factor and the maximum value of the next state; and determining the loss function of the value network based on the predicted value of the labeling decision and the target value.
[0059] In the embodiments of the present application, for each experience data sample, the labeling decision is evaluated based on the value network to obtain a predicted value of the labeling decision; and a target value is determined based on the reward value, the discount factor and the maximum value of the next state. Specifically, the calculation formula of the target value is as follows:
[0060] y = r + γa'maxQ(s', a');
[0061] Wherein, y represents the target value, r represents the reward value, γ is the discount factor, the value range is usually between 0 and 1, and Q(s', a') is the maximum value of the next state s'-labeling decision a'.
[0062] Further, the loss function of the value network is determined based on the predicted value of the labeling decision and the target value. The loss function of the value network is used to measure the error between the predicted value of the value network and the target value. For example, assuming that the loss function of the value network is the mean square error loss function, the loss function of the value network is as follows:
[0063]
[0064] Wherein, N represents the number of experience data samples, y i represents the target value of the i-th experience data sample, Q(s i , a i , θ) represents the predicted value of the i-th experience data sample, s i represents the current state in the i-th experience data sample, a i represents the current state labeling decision in the i-th experience data sample, and θ represents the parameters of the value network.
[0065] On the basis of the above-mentioned embodiments, optionally, the label probability distribution of each unlabeled sample data of the current round is used for sampling and labeling the unlabeled sample data of the current round to obtain labeled sample data of the next round, comprising: determining the uncertainty index value of each unlabeled sample data of the updated event text labeling model based on the label probability distribution of each unlabeled sample data of the current round, wherein the uncertainty index value is used to measure the uncertainty of the label probability distribution of the unlabeled sample data; sampling and labeling the unlabeled sample data of the current round based on the uncertainty index corresponding to each unlabeled sample data to obtain the labeled sample data of the next round.
[0066] Wherein, the uncertainty index value is used to measure the uncertainty of the label probability distribution of the unlabeled sample data, and the uncertainty index can be the entropy of the label probability distribution. Specifically, the calculation formula of the entropy of the label probability distribution is as follows:
[0067] H(P(y|x))=-∑ c∈Y P(y=c|x)logP(y=c|x);
[0068] Wherein, H represents the entropy of the label probability distribution, P(y|x) represents the label probability distribution, y represents the label category, x represents the unlabeled sample data, and Y represents the set of all possible label categories. It should be noted that the higher the entropy, the greater the uncertainty of the updated event text labeling model to the unlabeled sample data.
[0069] In the embodiment of the application, the entropy of the label probability distribution is calculated based on the label probability distribution of each unlabeled sample data of the current round, and the entropy of the label probability distribution is used as the uncertainty index value of each unlabeled sample data of the updated event text labeling model; the unlabeled sample data of the current round is sampled based on the uncertainty index corresponding to each unlabeled sample data, and the unlabeled sample data obtained by sampling is labeled to obtain the labeled sample data of the next round.
[0070] The embodiment of the application realizes active learning by designing a query strategy, selects sample data that is difficult for the current model to distinguish based on uncertainty sampling for labeling, and then inputs the labeled sample data to the model for training, so that the performance of the model and the labeling accuracy of the data are gradually improved through iteration.
[0071] S120, obtaining the event text to be labeled and the state information of the event text to be labeled, inputting the event text to be labeled into the natural language processing sub-model for recognition to obtain the text features of the event text to be labeled.
[0072] Wherein, the to-be-labeled event text refers to an event text to be labeled, and it should be noted that the type of the to-be-labeled event text is the same as the type of sample data used by the event text labeling model in the training process. The state information refers to state information of the to-be-labeled event text when the to-be-labeled event text is not labeled. Specifically, the state information is semantic content and context information describing the to-be-labeled event text. In the embodiment of the present application, the to-be-labeled event text and the state information of the to-be-labeled event text are obtained, the to-be-labeled event text is input into the natural language processing sub-model for recognition, and the text features of the to-be-labeled event text are obtained. The text features of the to-be-labeled event text include event trigger words, arguments of the to-be-labeled event text, and predicted labels of each word or single word.
[0073] In S130, the text features and the state information are input into the labeling decision network for sequence labeling decision, the labeling decision of the to-be-labeled event text is determined, and the to-be-labeled event text is labeled based on the labeling decision.
[0074] In the embodiment of the present application, the text features and the state information are input into the labeling decision network of the deterministic policy gradient module for sequence labeling decision, the labeling decision of the to-be-labeled event text is determined, and the to-be-labeled event text is labeled based on the labeling decision, thereby realizing the labeling of the to-be-labeled event text.
[0075] It should be noted that in some online learning or model updating scenarios, the labeling strategy of the labeling decision network in the deterministic policy gradient module can be updated to adapt to the to-be-labeled event text or task requirements. The updating process of the labeling strategy is the same as the training process of each round of the event text labeling model described above, and will not be described here.
[0076] The technical scheme of the embodiment loads a pre-trained event text labeling model, the event text labeling model includes a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module includes a labeling decision network, obtains a to-be-labeled event text and state information of the to-be-labeled event text, inputs the to-be-labeled event text into the natural language processing sub-model for recognition, obtains text features of the to-be-labeled event text, inputs the text features and the state information into the labeling decision network for sequence labeling decision, determines a labeling decision of the to-be-labeled event text, and labels the to-be-labeled event text based on the labeling decision. The natural language processing sub-model and the deterministic policy gradient module are combined, the labeling strategy of the deterministic policy gradient module is continuously optimized through training, and therefore the accuracy of event text labeling is improved.
[0077] Embodiment Two
[0078] Figure 4 is a structural schematic diagram of an event text intelligent labeling device based on a pre-trained model provided by Embodiment Two of the present application. As shown inFigure 4 The device comprises:
[0079] The event text labeling model loading module 210 is configured to load a pre-trained event text labeling model, wherein the event text labeling model comprises a natural language processing submodel and a deterministic policy gradient module.
[0080] The text feature recognition module 220 is configured to obtain a to-be-labeled event text and state information of the to-be-labeled event text, input the to-be-labeled event text into the natural language processing submodel for recognition, and obtain text features of the to-be-labeled event text.
[0081] The event text labeling module 230 is configured to input the text features and the state information into the labeling decision network for sequence labeling decision, determine a labeling decision of the to-be-labeled event text, and label the to-be-labeled event text based on the labeling decision.
[0082] The technical scheme of the embodiment comprises the following steps: loading a pre-trained event text labeling model, wherein the event text labeling model comprises a natural language processing submodel and a deterministic policy gradient module, the deterministic policy gradient module comprising a labeling decision network; obtaining a to-be-labeled event text and state information of the to-be-labeled event text, inputting the to-be-labeled event text into the natural language processing submodel for recognition, and obtaining text features of the to-be-labeled event text; inputting the text features and the state information into the labeling decision network for sequence labeling decision, determining a labeling decision of the to-be-labeled event text, and labeling the to-be-labeled event text based on the labeling decision. The natural language processing submodel and the deterministic policy gradient module are combined, the labeling strategy of the deterministic policy gradient module is continuously optimized through training, and thus the accuracy of event text labeling is improved.
[0083] On the basis of the above embodiment, the device can further comprise an event text labeling model training module, configured to:
[0084] Obtain sample data of an event text labeling task, wherein the sample data comprises labeled sample data of a first round and unlabeled sample data of the first round.
[0085] Iteratively perform the following steps: in the case where a first iteration stop condition is met, a trained event text labeling model is obtained.
[0086] Train the event text labeling model based on labeled sample data of a current round, and obtain an updated event text labeling model.
[0087] Predict unlabeled sample data of the current round based on the updated event text labeling model, and obtain a label probability distribution of each unlabeled sample data of the current round.
[0088] sampling and labeling the unlabeled sample data of the current round based on the label probability distribution of each unlabeled sample data of the current round to obtain labeled sample data of a next round.
[0089] In the above embodiment, optionally, the deterministic policy gradient module further comprises a value network configured to evaluate the value of the labeling decision output by the labeling decision network.
[0090] The training module of the event text labeling model is configured to:
[0091] Based on the labeled sample data, fine-tune the natural language processing sub-model, and based on the fine-tuned natural language processing sub-model, identify the labeled sample data and extract sample features of the labeled sample data.
[0092] At each time step, input the sample features and the current state information of the labeled sample data into the labeling decision network for sequence labeling decision to obtain a labeling decision of the current state of the labeled sample data, and interact with the labeled sample data based on the labeling decision to determine an experience data sample, and store the experience data sample in an experience buffer; wherein the experience data sample comprises current state information, the labeling decision, the reward value and the next state information.
[0093] Iteratively perform the following steps to obtain the updated event text labeling model when a second iteration stopping condition is met:
[0094] Obtain a batch of experience data samples from the experience buffer, determine a loss function of the value network based on the batch of experience data samples, and update the network parameters of the value network based on the loss function of the value network.
[0095] Based on the updated value network, evaluate the labeling decision to obtain the value of the labeling decision, determine the gradient of the labeling decision network based on the value of the labeling decision, and update the labeling decision network based on the gradient.
[0096] In the above embodiment, optionally, the training module of the event text labeling model is configured to:
[0097] For each experience data sample, evaluate the labeling decision based on the value network to obtain a predicted value of the labeling decision, and determine a target value based on the reward value, a discount factor and a maximum value of the next state.
[0098] Determine the loss function of the value network based on the predicted value of the labeling decision and the target value.
[0099] On the basis of the above-mentioned embodiments, optionally, the reward function of the reward value comprises:
[0100]
[0101] wherein R represents the reward value, n represents the number of characters and / or words in the labeled sample data, y i represents the true label of the i-th character or word in the labeled sample data, represents the label corresponding to the labeling decision of the i-th character or word in the labeled sample data, and I(·) is an indicator function, which has a value of 1 when y i =x i , and 0 otherwise.
[0102] On the basis of the above-mentioned embodiments, optionally, the training module of the event text labeling model is configured to:
[0103] determine, based on the label probability distribution of each unlabeled sample data in the current round, an uncertainty index value of each unlabeled sample data for the updated event text labeling model, the uncertainty index value being used to measure the uncertainty of the label probability distribution of the unlabeled sample data;
[0104] sample and label the unlabeled sample data in the current round based on the uncertainty index corresponding to each unlabeled sample data, to obtain labeled sample data in the next round.
[0105] On the basis of the above-mentioned embodiments, optionally, the sequence labeling mechanism used in the process of the sequence labeling decision is a character-granularity sequence labeling mechanism.
[0106] The event text intelligent labeling device based on the pre-training model provided in the embodiments of the present application can execute the event text intelligent labeling method based on the pre-training model provided in any of the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method.
[0107] Embodiment three
[0108] Figure 5 is a structural schematic diagram of an electronic device provided in Embodiment Three of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0109] As shown in Figure 5 electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0110] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0111] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the event text intelligent labeling method based on a pre-trained model.
[0112] In some embodiments, the event text intelligent labeling method based on a pre-trained model can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the event text intelligent labeling method based on a pre-trained model described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the event text intelligent labeling method based on a pre-trained model by any other appropriate means, such as by means of firmware.
[0113] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0114] Computer programs used to implement the pre-trained model-based event text intelligent labeling method of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / operations specified in the flow diagrams and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, and partially on a remote machine or server.
[0115] Embodiment four
[0116] Embodiment four of the present application also provides a computer readable storage medium, which stores computer instructions for causing a processor to execute a pre-trained model-based event text intelligent labeling method, the method comprising:
[0117] loading a pre-trained event text labeling model, the event text labeling model comprising a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module comprising a labeling decision network;
[0118] obtaining a to-be-labeled event text and state information of the to-be-labeled event text, inputting the to-be-labeled event text into the natural language processing sub-model for recognition to obtain text features of the to-be-labeled event text;
[0119] inputting the text features and the state information into the labeling decision network for sequence labeling decision to determine a labeling decision of the to-be-labeled event text, and labeling the to-be-labeled event text based on the labeling decision.
[0120] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0121] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0122] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0123] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0124] It should be understood that the various forms of flow shown above can be reordered, added to, or have steps deleted. For example, the steps described in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.
[0125] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for intelligent labeling of event text based on a pre-trained model, characterized in that, The method comprises the following steps: loading a pre-trained event text labeling model, the event text labeling model comprising a natural language processing sub-model and a deterministic policy gradient module, the deterministic policy gradient module comprising a labeling decision network; obtaining a to-be-labeled event text and state information of the to-be-labeled event text, inputting the to-be-labeled event text into the natural language processing sub-model for recognition to obtain text features of the to-be-labeled event text; inputting the text features and the state information into the labeling decision network for sequence labeling decision to determine a labeling decision of the to-be-labeled event text, and labeling the to-be-labeled event text based on the labeling decision; wherein the training method of the event text labeling model comprises: obtaining sample data of an event text labeling task, the sample data comprising labeled sample data of a first round and unlabeled sample data of the first round; iteratively performing the following steps to obtain a trained event text labeling model when a first iteration stop condition is met: training the event text labeling model based on the labeled sample data of the current round to obtain an updated event text labeling model; predicting the unlabeled sample data of the current round based on the updated event text labeling model to obtain a label probability distribution of each unlabeled sample data of the current round; sampling and labeling the unlabeled sample data of the current round based on the label probability distribution of each unlabeled sample data of the current round to obtain labeled sample data of the next round; wherein the deterministic policy gradient module further comprises a value network for evaluating the value of the labeling decision output by the labeling decision network; the training of the event text labeling model based on the labeled sample data to obtain an updated event text labeling model comprises: fine-tuning the natural language processing sub-model based on the labeled sample data, identifying the labeled sample data based on the fine-tuned natural language processing sub-model, and extracting sample features of the labeled sample data; at each time step, inputting the sample features and current state information of the labeled sample data into the labeling decision network for sequence labeling decision to obtain a labeling decision of the current state of the labeled sample data, and interacting with the labeled sample data based on the labeling decision to determine an experience data sample, and storing the experience data sample in an experience buffer; wherein the experience data sample comprises current state information, the labeling decision, a reward value and next state information; iteratively performing the following steps to obtain the updated event text labeling model when a second iteration stop condition is met: obtaining a batch of experience data samples from the experience buffer, determining a loss function of the value network based on the batch of experience data samples, and updating network parameters of the value network based on the loss function of the value network; The value network is updated based on the updated value network to evaluate the labeling decision, to obtain a value of the labeling decision, to determine a gradient of the labeling decision network based on the value of the labeling decision, and to update the labeling decision network based on the gradient.
2. The method of claim 1, wherein, The loss function of the value network is determined based on the batch of experience data samples, including: For each experience data sample, the value network is used to evaluate the labeling decision to obtain a predicted value of the labeling decision, and a target value is determined based on the reward value, the discount factor, and the maximum value of the next state; The loss function of the value network is determined based on the predicted value of the labeling decision and the target value.
3. The method of claim 1, wherein, The reward function of the reward value includes: ; wherein R represents a reward value, n represents the number of characters and / or words in the labeled sample data, represents a true label of the i-th character or word in the labeled sample data, represents a label corresponding to the labeling decision of the i-th character or word in the labeled sample data, is an indicator function, and when the value of the indicator function is 1, otherwise 0.
4. The method of claim 1, wherein, The unlabeled sample data of the current round is sampled and labeled based on the label probability distribution of each unlabeled sample data of the current round to obtain labeled sample data of the next round, including: An uncertainty index value of each unlabeled sample data is determined based on the label probability distribution of each unlabeled sample data of the current round, and the uncertainty index value is used to measure the uncertainty of the label probability distribution of the unlabeled sample data; The unlabeled sample data of the current round is sampled and labeled based on the uncertainty index corresponding to each unlabeled sample data to obtain labeled sample data of the next round.
5. The method of claim 1, wherein, The sequence labeling mechanism used in the sequence labeling decision process is a word granularity sequence labeling mechanism.
6. An event text intelligent labeling device based on a pre-trained model, characterized in that, It includes: An event text labeling model loading module is configured to load a pre-trained event text labeling model, wherein the event text labeling model includes a natural language processing sub-model and a deterministic policy gradient module; A text feature recognition module is configured to obtain a to-be-labeled event text and state information of the to-be-labeled event text, input the to-be-labeled event text into the natural language processing sub-model for recognition, and obtain text features of the to-be-labeled event text; An event text labeling module is configured to input the text features and the state information into a labeling decision network for sequence labeling decision, determine a labeling decision of the to-be-labeled event text, and label the to-be-labeled event text based on the labeling decision. The device further includes an event text labeling model training module configured to: Obtain sample data of an event text labeling task, wherein the sample data includes labeled sample data of a first round and unlabeled sample data of the first round; Iteratively execute the following steps to obtain a trained event text labeling model when a first iteration stop condition is met: Train the event text labeling model based on the labeled sample data of the current round to obtain an updated event text labeling model; Predict the unlabeled sample data of the current round based on the updated event text labeling model to obtain a label probability distribution of each unlabeled sample data of the current round; Sample and label the unlabeled sample data of the current round based on the label probability distribution of each unlabeled sample data of the current round to obtain labeled sample data of the next round. The deterministic policy gradient module further comprises a value network configured to evaluate a value of the labeling decision output by the labeling decision network. The training module of the event text labeling model is configured to: fine-tune a natural language processing sub-model based on the labeling sample data, identify the labeling sample data based on the fine-tuned natural language processing sub-model, and extract sample features of the labeling sample data; At each time step, the sample features and current state information of the labeling sample data are input into the labeling decision network for sequence labeling decision to obtain a labeling decision of a current state of the labeling sample data, and the labeling decision network and the labeling sample data are interacted based on the labeling decision to determine an experience data sample, and the experience data sample is stored in an experience buffer; wherein the experience data sample comprises current state information, the labeling decision, a reward value and next state information; The following steps are iteratively performed, and the updated event text labeling model is obtained when a second iteration stop condition is met: a batch of experience data samples are obtained from the experience buffer, a loss function of the value network is determined based on the batch of experience data samples, and network parameters of the value network are updated based on the loss function of the value network; the labeling decision is evaluated based on the updated value network to obtain a value of the labeling decision, a gradient of the labeling decision network is determined based on the value of the labeling decision, and the labeling decision network is updated based on the gradient.
7. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the event text intelligent labeling method based on the pre-trained model according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the event text intelligent labeling method based on the pre-trained model according to any one of claims 1-5 when executed.
Citation Information
Patent Citations
Fusion labeling method and device, equipment and storage medium
CN116452934A
Text marking method and device based on active learning, equipment and storage medium
CN117828088A