Multi-task learning optimized corpus labeling automation system
The automated corpus annotation system optimized through multi-task learning solves the problem of insufficient adaptive optimization in existing technologies and achieves efficient and accurate corpus annotation. It is suitable for fields such as news reporting, social media analysis, and legal document processing.
Patent Information
- Application Number
- CN202510656831.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
AI Technical Summary
Existing multi-task learning methods lack an effective adaptive optimization mechanism in corpus annotation tasks, resulting in the model's inflexible trade-offs between different tasks, affecting the efficiency and accuracy of annotation, especially when dealing with long texts and complex corpora, making it difficult to capture long-distance dependencies.
The corpus annotation automation system adopts multi-task learning optimization, including a shared encoder, a multi-task decoder, a dynamic adaptive module and an online incremental learning module. Through the sliding window mechanism, Transformer encoder, self-attention mechanism, graph attention network and dynamic loss weight adjustment, it achieves efficient and accurate annotation of multiple tasks.
It improves the efficiency and accuracy of corpus annotation, reduces the cost of manual annotation, provides more reliable input data, and lays the foundation for subsequent natural language processing tasks.
Smart Images

Figure CN120597847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a multi-task learning optimized corpus annotation automatic system. Background Art
[0002] With the continuous development of natural language processing technology, corpus annotation, as a fundamental task in natural language processing, has become increasingly important. Corpus annotation refers to the labeling of text data to provide structured information that can be used for training and reasoning machine learning models. The demand for corpus annotation is particularly urgent in areas such as news reporting, social media analysis, and legal document processing. Existing multi-task learning methods have the following problems when handling corpus annotation tasks: First, the lack of an effective adaptive optimization mechanism leads to inflexible trade-offs between different tasks, affecting overall performance; second, traditional methods have difficulty effectively capturing long-range dependencies in text when processing long texts and complex corpora, further affecting the accuracy of annotation. Summary of the Invention
[0003] The present invention provides a multi-task learning optimized corpus annotation automated system to solve the problems of low efficiency, poor accuracy, lack of effective adaptive optimization mechanism and insufficient processing capability of long texts and complex corpora in the prior art.
[0004] The present invention provides a multi-task learning optimized corpus annotation automatic system, the system comprising:
[0005] Shared encoder: used to segment the input text and generate a word sequence T = {t1, t2, ..., t n The word segmentation process adopts the statistic-based Jieba word segmentation algorithm; and the word sequence is divided into text segments of fixed length L based on the sliding window mechanism. The step size of the sliding window is δ, and each text segment is embedded with timestamp information τ i , and generate a global semantic representation E = {e1, e2, ..., e m}, where each d = 768;
[0006] The multi-task decoder includes: an event boundary detection module, a time sequence reconstruction module, and a causal relationship extraction module. The event boundary detection module calculates each word t through the self-attention mechanism based on the output of the shared encoder. j The probability of the event starting p sj and the ending probability p ej , and mark the event boundary B={(s k , e k )}; Time series reconstruction module: extract the time embedding {τ corresponding to the event boundary B sk, τ ek}, generate the event time chain C={c1→c2→…→c k Causal relationship extraction module: construct event node graph G = (V, E), where node V is event c i , edge E is a candidate causal relationship; the edge weight w is calculated by the 3-layer graph attention network GAT ij , filter w ij Edges with a value of ≥0.7 are output as causal chains;
[0007] Dynamic Adaptive Module: This module dynamically adjusts the loss weight of each task based on the consistency scores of the outputs of the event boundary detection task, the time series reconstruction task, and the causal relationship extraction task.
[0008] Online incremental learning module: When the consistency score S is less than 0.6, the expert review process is triggered and the annotation results are manually corrected; the corrected data is mixed with the original samples in a ratio of 7:3, and the model parameters are updated through the Adam optimizer; the model weights are smoothed using exponential moving average.
[0009] Preferably, the specific parameters of the sliding window mechanism in the shared encoder are: window length L = 512, sliding step δ = 128, and timestamp embedding method: event occurrence time τ i Convert to Unix timestamp and normalize to The specific normalization formula is:
[0010] where τ min and τ max They are the earliest and latest timestamps of the event in the text, which are embedded and concatenated with the word vector to input the encoder.
[0011] Preferably, the implementation of the maximum matching sorting algorithm in the time series reconstruction module includes: defining the event time series similarity sim(c i , c j )=cos(τ si -τ ei , τ sj -τ ej ) where τ si and τ ej Event c i The start and end timestamps are in Unix timestamp format; the optimal event arrangement is solved by dynamic programming The state transition equation of dynamic programming is: in Indicates that the first i events are arranged to the last event c j The maximum sum of similarities.
[0012] Preferably, the structure of the graph attention network GAT in the causal relationship extraction module is: input layer dimension 256, hidden layer dimension 128, output layer dimension 64; multi-head attention mechanism: 4 heads, each attention head dimension is 32, edge weight calculation uses LeakyReLU: negative slope α = 0.2, and the specific edge weight calculation formula is: where a h ead is the learnable parameter vector of each attention head, W is the projection matrix, h i and h j are the eigenvectors of nodes i and j.
[0013] Preferably, in the step 1 semantic vector generation of the dynamic adaptive module, the event boundary description template is "the event starts at position s k , ending at position e k , the confidence level is p sk ×p ek where p sk is the probability of the event starting, p ek is the probability of the event ending”; the time series reconstruction description template is “event ID occurs at time τ sk To time τ ek , the confidence level is in
[0014] is the time series reconstruction confidence calculated by the maximum matching sorting algorithm"; the causal relationship description template is "event ID1 causes event ID2, and the confidence is the edge weight wij"; the template text is encoded by the DeBERTa model.
[0015] Preferably, the expert review process in the online incremental learning module includes: marking low consistency samples as D low ={(x i ,y i )|S(x i )<0.6}; The annotation results, time chain and causal diagram are displayed through a web interface built based on the Flask framework. The web interface includes an event annotation overview module, a temporal logic analysis module and a causal relationship visualization module. It supports manual correction of event boundaries, temporal sequence and causal chain structure, and stores the correction results in JSON format.
[0016] Preferably, the parameter update strategy of the incremental learning in the online incremental learning module is: each update selects a batch size B = 32, a training round E = 3; the loss function is weighted cross entropy
[0017] Among them, L1 is the cross entropy loss of the event boundary labeling task, L2 is the RankNet loss of the time series reconstruction task, and L3 is the binary cross entropy loss of the causal relationship extraction task; the learning rate decay strategy is used during training, and the initial learning rate is 1e -5 , decaying to 90% every 10 rounds of training.
[0018] Preferably, the system also includes anomaly detection and fault tolerance mechanism: real-time monitoring of annotation confidence
[0019] When conf is less than 0.4, the current sample annotation is automatically terminated and an alarm is triggered. The alarm information is sent to the preset operation and maintenance mailbox in the form of an email, and the abnormal sample log is recorded at the same time.
[0020] Preferably, the event boundary detection module and the timing reconstruction module in the multi-task decoder share the weight matrix parameters of the first 6 layers of Transformer decoder layers to enhance the cross-modal fusion of timing information and event boundary features.
[0021] Preferably, the dynamically adjusting the loss weight of each task includes:
[0022] Step 1: Convert the output of each task into a structured description and generate a semantic vector using the pre-trained DeBERTa-v3 model.
[0023] Step 2: Calculate the cosine similarity between tasks Take the mean as the global consistency score S;
[0024] Step 3: Dynamically adjust loss weight according to S where λ i represents the initial weight of each task, α = 0.5.
[0025] The multi-task learning-optimized corpus annotation automation system of the present invention achieves efficient and accurate annotation of corpora through multi-task learning and adaptive optimization. First, it can process multiple tasks simultaneously, fully utilize the information in the data, and improve the efficiency of annotation. Second, based on the adaptive optimization mechanism, it can automatically adjust the annotation strategy according to the characteristics of the data, significantly improving the accuracy of annotation. In addition, through multi-task learning and adaptive optimization, the system ensures the consistency of the annotation results and provides more reliable input for subsequent natural language processing tasks. At the same time, through automated annotation, the cost and time of manual annotation are significantly reduced, saving a lot of manpower and material resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a structural diagram of a multi-task learning optimized corpus annotation automatic system of the present invention. DETAILED DESCRIPTION
[0027] This invention relates to a multi-task learning-optimized corpus annotation automation system. Through multi-task learning and adaptive optimization, the system aims to automatically annotate data, including event boundaries, temporal relationships, and causal relationships. By integrating a shared encoder, a multi-task decoder, a dynamic adaptive module, and an online incremental learning module, the system improves the accuracy and consistency of annotations. It is suitable for a variety of fields, including news reporting, social media analysis, and legal document processing. The overall approach employed is as follows:
[0028] Shared encoder: used to segment the input text and generate a word sequence T = {t1, t2, ..., t n The word segmentation process uses the Jieba word segmentation algorithm based on statistics; and the word sequence is divided into text segments of fixed length L based on the sliding window mechanism. The step size of the sliding window is δ, and each text segment is embedded with timestamp information τ i , and generate a global semantic representation E = {e1, e2, ..., e m}, where each d = 768;
[0029] Multi-task decoder, including: event boundary detection module, time series reconstruction module, causal relationship extraction module, event boundary detection module: based on the output of the shared encoder, calculates each word t through the self-attention mechanism j The probability of the event starting p sj and the ending probability p ej , and mark the event boundary B={(s k , e k )}; Time series reconstruction module: extract the time embedding {τ corresponding to the event boundary B sk , τ ek}, generate the event time chain C={c1→c2→…→c k Causal relationship extraction module: construct event node graph G = (V, E), where node V is event c i , edge E is a candidate causal relationship; the edge weight w is calculated by the 3-layer graph attention network GAT ij , filter w ij Edges with a value of ≥0.7 are output as causal chains;
[0030] Dynamic Adaptive Module: This module dynamically adjusts the loss weight of each task based on the consistency scores of the outputs of the event boundary detection task, the time series reconstruction task, and the causal relationship extraction task.
[0031] Online incremental learning module: When the consistency score S is less than 0.6, the expert review process is triggered and the annotation results are manually corrected; the corrected data is mixed with the original samples in a ratio of 7:3, and the model parameters are updated through the Adam optimizer; the model weights are smoothed using exponential moving average.
[0032] The above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods of the specification to better understand the above technical solution. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments used only to explain the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, it should be noted that, for the convenience of description, only the parts related to the present invention, rather than all, are shown in the drawings.
[0033] Example 1
[0034] like Figure 1 As shown in the figure, the specific parameters of the sliding window mechanism in the shared encoder are: window length L = 512, sliding step δ = 128, and the timestamp embedding method is: the event occurrence time τ i Convert to Unix timestamp and normalize to [-1, 1]. The specific normalization formula is:
[0035] where τ min and τ max They are the earliest and latest timestamps of the event in the text, which are embedded and concatenated with the word vector to input the encoder.
[0036] Specifically, the shared encoder uses a Transformer architecture, specifically a BERT-base variant with a 12-layer encoder, 768 hidden layer dimensions, a 3072-dimensional feedforward neural network, and 12 attention heads. The input news text is first tokenized to generate a word sequence using the statistically-based Jieba word segmentation algorithm. For example, for the news report "Tomorrow at 10:00 AM, the company will hold a new product launch at its headquarters," the tokenization result is "Tomorrow at 10:00 AM, the company will hold a new product launch at its headquarters." The word sequence is then segmented into fixed-length 512-bit text segments using a sliding window mechanism with a step size of 128. Timestamp information is embedded in each segment by converting the event time to a Unix timestamp and normalizing it to a value between -1 and 1. The normalization formula is: subtract the earliest event timestamp from the Unix timestamp, divide it by the latest event timestamp minus the earliest timestamp, multiply by 2, and then subtract 1. After embedding, it is concatenated with the word vector and input into the encoder to generate a global semantic representation, where the dimension of each semantic vector is 768.
[0037] Furthermore, the multi-task decoder includes an event boundary detection module, a timing reconstruction module, and a causal relationship extraction module:
[0038] The implementation of the maximum matching sorting algorithm in the time series reconstruction module includes: defining the event time series similarity sim(c i , c j )=cos(τ si -τ ei , τ sj -τ ej ) where τ si and τ ej Event c i The start and end timestamps are in Unix timestamp format; the optimal event arrangement is solved by dynamic programming The state transition equation of dynamic programming is:
[0039] Where dp[i][j] represents the order of the first i events to the last event c j The maximum sum of similarities.
[0040] The structure of the graph attention network GAT in the causal relationship extraction module is: input layer dimension 256, hidden layer dimension 128, output layer dimension 64; multi-head attention mechanism: 4 heads, each attention head dimension 32, edge weight calculation uses LeakyReLU: negative slope α = 0.2, the specific edge weight calculation formula is:
[0041] where a h eadis the learnable parameter vector of each attention head, W is the projection matrix, h i and h j are the eigenvectors of nodes i and j.
[0042] Specifically, the event boundary detection module uses a self-attention mechanism to calculate the event start and end probabilities for each word and mark event boundaries. Specifically, based on the global semantic representation output by the shared encoder, a multi-head self-attention mechanism (12 heads) calculates the relevance of each word to other words. Words with high weights are more likely to be event boundaries. For example, in the news report mentioned above, "new product launch" may be marked as an event boundary because "new product" and "launch" are closely related to the core content of the event and have high self-attention weights. For each detected event boundary, its position in the word sequence and the timestamp of the event are recorded. For example, "new product launch" is located between the 10th and 12th words in the word sequence, and the event timestamp is "2024-05-09 10:00."
[0043] Specifically, the time series reconstruction module extracts the time embeddings corresponding to event boundaries and generates an event time chain using a maximum matching sorting algorithm. Event time series similarity measures the degree of similarity between the temporal information of two events. It is calculated by subtracting the cosine value from the event's start timestamp and the end timestamp. For example, if two events have the timestamps "2024-05-09 10:00" and "2024-05-09 15:00," the cosine value of the start timestamp minus the end timestamp can be used to measure the temporal similarity between the two events. Dynamic programming is used to find the optimal solution under given constraints. In this process, a state transition equation is used to describe how to transition from one known state to another. For example, when finding the optimal order of events, the sum of the maximum similarities from the first i events to the last event in the order is equal to the sum of the maximum similarities from the first i-1 events to the i-1th event in the order plus the similarity between the i-1th event and the i-th event. In this way, a time chain of events is gradually constructed, ensuring temporal consistency of the event chain.
[0044] Specifically, the causal relationship extraction module constructs an event node graph, where nodes represent events and edges represent candidate causal relationships. The Graph Attention Network (GAT) is a neural network model for processing graph data. It determines the importance of nodes by calculating attention weights between them. In this process, the GAT has an input layer dimension of 256, a hidden layer dimension of 128, and an output layer dimension of 64. The multi-head attention mechanism is a key technology in GAT. It calculates attention weights by dividing the input data into multiple "heads," each with a dimension of 32. Edge weights are calculated by performing a dot product of the concatenation of the learnable parameter vector and the projection matrix multiplied by the node feature vector, and then applying the result to the LeakyReLU activation function. LeakyReLU is an activation function that introduces a small negative slope when the input is negative to avoid the vanishing gradient problem. In this way, edges with edge weights greater than or equal to 0.7 are selected as causal chain outputs, thereby constructing causal relationships between events.
[0045] Furthermore, the dynamic adaptive module dynamically adjusts the loss weight of each task including:
[0046] Step 1: Convert the output of each task into a structured description and generate a semantic vector using the pre-trained DeBERTa-v3 model.
[0047] Step 2: Calculate the cosine similarity between tasks Take the mean as the global consistency score S;
[0048] Step 3: Dynamically adjust loss weight according to S where λ i represents the initial weight of each task, α = 0.5.
[0049] Specifically, the dynamic adaptive module converts the output of each task into a structured description. For example, the event boundary description template is “the event starts at position s k , ending at position e k , the confidence level is p sk ×p ek where p sk is the probability of the event starting, p ek is the probability of the event ending”; the time series reconstruction description template is “event ID occurs at time τ sk To time τ ek , the confidence level is in is the time series reconstruction confidence calculated by the maximum matching sorting algorithm; the causal relationship description template is "event ID1 causes event ID2, and the confidence is the edge weight wij". These template texts are encoded using the pre-trained DeBERTa-v3 model to generate semantic vectors. The cosine similarity between tasks is calculated and the average is taken as the global consistency score. Specifically, the cosine similarity is obtained by calculating the dot product between two vectors divided by the product of their moduli, which measures the similarity between the two vectors. The global consistency score is the average of the cosine similarities between the three tasks, thus The consistency between each task was comprehensively evaluated. The loss weight of each task was dynamically adjusted based on the global consistency score. The adjustment formula is: the adjusted loss weight is equal to the initial loss weight multiplied by (1 plus the adjustment factor multiplied by (1 minus the global consistency score) divided by (1 plus the global consistency score)). For example, when the global consistency score is 0.8, the initial weight of the event boundary detection task is 0.3, and the adjustment factor is 0.5. The adjusted weight is 0.3 multiplied by (1 plus 0.5 multiplied by (1 minus 0.8) divided by (1 plus 0.8)), which is approximately equal to 0.3067.
[0050] Furthermore, the online incremental learning module: The expert review process in the online incremental learning module includes: low consistency samples are marked as D low ={(x i ,y i )|S(x i )<0.6}; The annotation results, time chain and causal diagram are displayed through a web interface built based on the Flask framework. The web interface includes an event annotation overview module, a temporal logic analysis module and a causal relationship visualization module. It supports manual correction of event boundaries, temporal sequence and causal chain structure, and stores the correction results in JSON format.
[0051] The parameter update strategy of incremental learning is: the batch size B = 32 and the training round E = 3 are selected for each update; the loss function is weighted cross entropy. Among them, L1 is the cross entropy loss of the event boundary labeling task, L2 is the RankNet loss of the time series reconstruction task, and L3 is the binary cross entropy loss of the causal relationship extraction task; the learning rate decay strategy is used during training, and the initial learning rate is 1e -5 , decays to 90% every 10 rounds of training
[0052] Specifically, when the global consistency score is less than 0.6, the expert review process is triggered and the annotation results are manually corrected. The corrected data is mixed with the original sample in a ratio of 7:3, and the model parameters are updated through the Adam optimizer (learning rate 1e-5). The exponential moving average (EMA, decay rate 0.999) is used to smooth the model weights. Specifically, the batch size is 32 and the training rounds are 3 for each update. The loss function is weighted cross entropy, where the loss of the event boundary labeling task is the cross entropy loss, the loss of the time series reconstruction task is the ranking loss (such as RankNet loss), and the loss of the causal relationship extraction task is the binary cross entropy loss. A learning rate decay strategy is adopted during training, with an initial learning rate of 1e-5, which decays to 90% every 10 rounds of training.
[0053] Furthermore, anomaly detection and fault tolerance mechanism: real-time monitoring of annotation confidence
[0054] When conf is less than 0.4, the current sample annotation is automatically terminated and an alarm is triggered. The alarm information is sent to the preset operation and maintenance mailbox in the form of an email, and the abnormal sample log is recorded at the same time.
[0055] It should be understood that the embodiments disclosed in the present invention and the above description can enable those skilled in the art to use the present invention to implement the present invention. At the same time, the present invention is not limited to the embodiments mentioned above. It should be understood that those skilled in the art can still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention and are all included in the scope of protection of the present invention.
Claims
1. A multi-task learning optimized corpus annotation automation system, characterized by: include: Shared encoder: used to segment the input text and generate a word sequence T = {t1, t2, ..., t n The word segmentation process adopts the statistic-based Jieba word segmentation algorithm; and the word sequence is divided into text segments of fixed length L based on the sliding window mechanism. The step size of the sliding window is δ, and each text segment is embedded with timestamp information τ i , and generate a global semantic representation E = {e1, e2, ..., e m }, where each d = 768; The multi-task decoder includes: an event boundary detection module, a time sequence reconstruction module, and a causal relationship extraction module. The event boundary detection module calculates each word t through the self-attention mechanism based on the output of the shared encoder. j The probability of the event starting p sj and the ending probability p ej , and mark the event boundary B={(s k , e k )}; Time series reconstruction module: extract the time embedding {τ corresponding to the event boundary B sk , τ ek }, generate the event time chain C={c1→c2→…→c k Causal relationship extraction module: construct event node graph G = (V, E), where node V is event c i , edge E is a candidate causal relationship; the edge weight w is calculated by the 3-layer graph attention network GAT ij , filter w ij Edges with a value of ≥0.7 are output as causal chains. When the time sequence reconstruction module determines the temporal relationship of events, the causal relationship extraction module further analyzes the temporal relationship, verifies the causal relationship between events, and generates a complete event chain to ensure the accuracy and consistency of events in terms of timing and causal relationships. Dynamic Adaptive Module: This module dynamically adjusts the loss weight of each task based on the consistency scores of the outputs of the event boundary detection task, the time series reconstruction task, and the causal relationship extraction task. Online incremental learning module: When the consistency score S is less than 0.6, the expert review process is triggered and the annotation results are manually corrected; the corrected data is mixed with the original samples in a ratio of 7:3, and the model parameters are updated through the Adam optimizer; the model weights are smoothed using exponential moving average.
2. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The specific parameters of the sliding window mechanism in the shared encoder are: window length L = 512, sliding step δ = 128, and timestamp embedding method: event occurrence time τ i Convert to Unix timestamp and normalize to The specific normalization formula is: where τ min and τ max They are the earliest and latest timestamps of the event in the text, which are embedded and concatenated with the word vector to input the encoder.
3. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The implementation of the maximum matching sorting algorithm in the time sequence reconstruction module includes: defining the event time sequence similarity sim(c i , c j )=cos(τ si -τ ei , τ sj -τ ej ) where τ si and τ ei Event c i The start and end timestamps are in Unix timestamp format; the optimal event arrangement is solved by dynamic programming The state transition equation of dynamic programming is: in Indicates that the first i events are arranged to the last event c j The maximum sum of similarities.
4. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The structure of the graph attention network GAT in the causal relationship extraction module is as follows: input layer dimension 256, hidden layer dimension 128, output layer dimension 64; multi-head attention mechanism: 4 heads, each attention head dimension 32, edge weight calculation uses LeakyReLU: negative slope α = 0.2, the specific edge weight calculation formula is: in is the learnable parameter vector of each attention head, W is the projection matrix, h i and h j are the feature vectors of nodes i and j.
5. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: In the dynamic adaptive module, step 1 semantic vector generation: the event boundary description template is "the event starts at position s k , ending at position e k , the confidence level is p sk ×p ek where p sk is the probability of the event starting, p ek is the probability of the event ending"; the time series reconstruction description template is "event ID occurs at time τ sk To time τ ek , the confidence level is in The confidence level of time series reconstruction calculated by the maximum matching sorting algorithm is "confidence level of time series reconstruction calculated by the maximum matching sorting algorithm"; the causal relationship description template is "event ID1 causes event ID2, and the confidence level is edge weight wij"; the template text is encoded by the DeBERTa model.
6. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The expert review process in the online incremental learning module includes: low consistency samples are marked as D low ={(x i ,y i )|S(x i )<0.6}; The annotation results, time chain and causal diagram are displayed through a web interface built based on the Flask framework. The web interface includes an event annotation overview module, a temporal logic analysis module and a causal relationship visualization module. It supports manual correction of event boundaries, temporal sequence and causal chain structure, and stores the correction results in JSON format.
7. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The parameter update strategy of the incremental learning in the online incremental learning module is: each update selects a batch size B = 32, a training round E = 3; the loss function is weighted cross entropy Among them, L1 is the cross entropy loss of the event boundary labeling task, L2 is the RankNet loss of the time series reconstruction task, and L3 is the binary cross entropy loss of the causal relationship extraction task; the learning rate decay strategy is used during training, and the initial learning rate is 1e -5 , decaying to 90% every 10 rounds of training.
8. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The system also includes anomaly detection and fault tolerance mechanisms: real-time monitoring of annotation confidence When conf is less than 0.4, the current sample annotation is automatically terminated and an alarm is triggered. The alarm information is sent to the preset operation and maintenance mailbox in the form of an email, and the abnormal sample log is recorded at the same time.
9. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The event boundary detection module and the timing reconstruction module in the multi-task decoder share the weight matrix parameters of the first 6 layers of Transformer decoder layers to enhance the cross-modal fusion of timing information and event boundary features.
10. The multi-task learning optimized corpus annotation automation system according to claim 1, characterized in that: The dynamic adjustment of the loss weight of each task includes: Step 1: Convert the output of each task into a structured description and generate a semantic vector using the pre-trained DeBERTa-v3 model. Step 2: Calculate the cosine similarity between tasks Take the mean as the global consistency score S; Step 3: Dynamically adjust loss weight according to S where λ i represents the initial weight of each task, α = 0.5.
Citation Information
Cited By
Business data processing method and system based on visual lane modeling
CN121052142A
Abnormality detection method and related device, electronic equipment and storage medium
CN121644243A