Time expression extraction method based on deep learning auxiliary rule
By combining deep learning and rule methods in time expression extraction, the complexity and diversity of time expressions are solved, the extraction effect is improved, and more accurate and comprehensive recognition of time expressions is achieved.
Patent Information
- Application Number
- CN202411922683.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is limited by the complexity, diversity and context dependence of time expression in temporal expression extraction, resulting in poor application on text.
Using a method based on deep learning assisted rules, a rule set of time words, time modifiers and connection words is constructed, and a pre-trained model is combined with Bi-LSTM to fuse context information, and time information extraction and decoding is used using attention mechanism and CRF layer.
It effectively solves the complexity and diversity of time expression, and supplements the time expression that cannot be recognized by the deep learning module through the rule module, which significantly improves the extraction ability of time expressions.
Smart Images

Figure CN119990123A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information extraction, and specifically relates to a time expression extraction method based on deep learning auxiliary rules. Background Art
[0002] Temporal information extraction is an important task in the field of information extraction, which aims to identify and extract time-related information from natural language texts, including specific dates, time points, time periods, time frequencies, etc. Temporal information extraction technology has important application value in many fields.
[0003] Extracting time expressions from time information is the key to information processing. Time expressions provide a standardized way to represent time information, eliminating ambiguity and ambiguity in natural language text. Converting time information to a unified format ensures data accuracy and consistency. When processing multi-source data, standardization of time expressions is the key to integration, making time information from different data sets comparable and connectable. In fields such as news and history, extracting time expressions can improve efficiency and reduce workload. For applications such as finance and meteorology that require analysis of event time series, its extraction is even more fundamental, enabling events to be sorted and analyzed in time, and its role is extremely critical and extensive.
[0004] The effective extraction and normalization of temporal expressions is crucial and involves many steps.
[0005] At the initial stage, the technology must have strong recognition capabilities to accurately find various time expressions from the text, whether it is a clear time point such as "November 12, 2024" or a vague relative expression such as "next week" or "the end of last month". After the recognition is completed, it enters the normalization process, in which standardization is a key link, aiming to eliminate ambiguity, such as converting "next Wednesday" to "2024-11-16" based on the specific context. At the same time, in order to achieve higher accuracy, it is indispensable to analyze the context of the time expression in the text, and it is necessary to clarify its relationship with other content and its function in the sentence. For example, in the water conservancy field, standardizing "July rainfall" to the time range of "2024-07" is conducive to data analysis. In short, this process requires technology to accurately identify, cleverly convert, and deeply understand the time information of the text, so as to provide standardized and unified time data for many fields and promote the efficient development of related work in various fields.
[0006] At present, most time expression extraction tasks use the related technologies of named entity recognition because the task objectives of the two are very similar. The latter is to extract named entities of a specific type in the text, while the former can regard time as an entity type, thus transforming it into a special named entity recognition task. The current time expression extraction methods are mainly divided into the following three types: (1) rule-based methods; (2) machine learning-based methods; (3) deep learning-based methods. Rule-based methods use written rules to identify time expressions.
[0007] The rules are mainly formulated for "time words", "time modifiers" and "time conjunctions". As nouns that directly reflect the concept of time, time words can refer to various time elements, such as time points, time periods, etc. Time modifiers are used to modify time words or expressions, increase the accuracy of time description in the form of ordinal words, frequency words, etc., and provide detailed information such as order and frequency. Time conjunctions connect multiple time expressions to clarify the relationship between different time points or segments, such as simultaneous, sequential, etc., which are of great significance in expressing the logical order of time. These rules can be based on vocabulary or grammatical structure. Although rule-based methods can handle special time expressions, they are difficult to cover all forms. In contrast, machine learning-based methods use annotated corpora to train extraction models, while deep learning-based methods can learn the characteristics of time expressions from a large amount of unlabeled data. The three have their own focuses and jointly contribute to the development and improvement of time information processing.
[0008] However, these three extraction methods are limited by the complexity, diversity and context dependency of time expressions in time expression tasks, resulting in poor application effects on text. The specific problems are: 1) Diversity of time expressions: Time can be expressed in many ways, including specific dates, relative time, time periods, periodic time, etc. These different expressions increase the difficulty of recognition. 2) Grammatical and semantic complexity: Time expressions may contain complex grammatical structures and semantic meanings, and time relations involve calculations of multiple events and times. Understanding such sentences requires deep grammatical and semantic analysis. 3) Context dependency: The understanding of time often needs to rely on specific contexts. Some time expressions may need to refer to time points in other sentences to be understood. Summary of the invention
[0009] Purpose of the invention: In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a time expression extraction method based on deep learning auxiliary rules, which can solve the complexity and diversity problems in text time expression extraction.
[0010] Technical solution: The method for extracting time expressions based on deep learning auxiliary rules described in the present invention comprises the following steps:
[0011] (1) Construct a rule set of time words, time modifiers, and conjunctions, use regular expressions to match specific formats, and identify time expressions in text;
[0012] (2) The pre-trained model is used to fuse the context information with Bi-LSTM, and the relative position encoding is added to the intermediate representation learned for each word in the attention mechanism;
[0013] (3) Assign weights to the intermediate representations learned for each word and calculate the correlation between each word and other words;
[0014] (4) Calculate the weighted context vector and send it to the CRF layer to extract time information, determine whether each token is a time marker, and the specific components of time;
[0015] (5) Based on the current weighted context vector c t and the label y of the previous word t-1 Calculate the feature function of each word;
[0016] (6) The weights are learned by maximum likelihood estimation of the CRF layer and decoded using the Viterbi algorithm to achieve optimal and result extraction of the time expression;
[0017] (7) Integrate the extracted time expressions and adopt different processing strategies according to whether the extracted results are consistent.
[0018] Furthermore, the implementation process of step (1) is as follows:
[0019] Construct a rule set to extract the co-occurrence pattern of time expressions from the water conservancy description text to form a preliminary rule set; specifically including time words, time modifiers and time connectors; find uncovered words in the training corpus and manually add them to the rule set;
[0020] Using the time words in the rules to identify the time marker words in the text;
[0021] Use the time modifiers in the rules to determine whether there is any time modification information before and after the directly recognized time expression;
[0022] Combine time markers, modifiers, and time connectors to build a complete time expression.
[0023] Furthermore, the intermediate representation h learned for each word in the attention mechanism in step (2) is t The implementation process of adding relative position encoding is as follows:
[0024] Generate position encoding vector p through sine and cosine functions t , and the intermediate representation h learned for each word tAdding helps the model distinguish words in different positions and understand the relative position relationship between time words, such as the order of precedence. The specific calculation formula is as follows:
[0025] h' t =h t +p t
[0026] Among them, h t is an intermediate representation learned for each word, p t represents the position encoding vector at time t.
[0027] Furthermore, the step (3) is implemented by the following formula:
[0028]
[0029] Among them, α t is the attention weight, e t It is calculated by the scoring function, and the output h of BERT-Bi-LSTM is t and a learnable weight matrix W as input.
[0030] Furthermore, step (4) is implemented by the following formula:
[0031] c=∑ t α t h' t
[0032] Where c is the weighted context vector, α t is the attention weight, h' t is the hidden state of each word at time t.
[0033] Furthermore, step (5) is implemented by the following formula:
[0034] f(c t ,y t-1 )=σ(W c c t +W vocab F vocab +W pos F pos +b)
[0035] Among them, F vocab is the vocabulary feature vector, F pos is the part-of-speech feature vector, W c , W vocab , W pos is the learnable weight matrix and b is the bias vector.
[0036] Furthermore, the learning weights by maximum likelihood estimation in step (6) are implemented as follows:
[0037] The gradient descent optimization algorithm is used to optimize the gradient information of the model parameters, and the parameter values are continuously adjusted to gradually approach the optimal solution, thereby finding the optimal label sequence:
[0038]
[0039] Among them, L(θ) is the log-likelihood function, p(y|x) is the conditional probability of the label sequence y given the input sequence x, and w k is the weight of each feature function, and Z(x) is a normalization factor that ensures that the sum of the probabilities of all possible label sequences is 1.
[0040] Furthermore, the decoding process using the Viterbi algorithm in step (6) is as follows:
[0041] The feature function weight w k and transition probabilities as input, and regard the time expression as a sequence of different labels; calculate the probabilities of all possible time expression paths; and screen out the path with the largest probability value among these paths, which is the most likely time expression parsing result.
[0042] Furthermore, the implementation process of step (7) is as follows:
[0043] The extracted time expressions are converted into word vectors using the Word2Vec model; the average value of the word vectors is calculated to obtain the representation vector of the phrase; the vectors of the two phrases are compared using the cosine similarity formula; based on the calculated similarity value, a threshold is set to determine whether the two time expressions are considered similar; the threshold is 0.8;
[0044] The similarity of the extracted results is high: it is considered that both deep learning and rule matching have identified the same time expression, which is directly used as the final result;
[0045] The similarity of the extracted results is low: the results are analyzed through the context window to determine which result is more appropriate in a specific context; the text in the context window is converted into feature vectors and these vectors are averaged or concatenated; then the SVM classifier is trained using the manually annotated time expressions and labels as data sets, and the feature vectors are analyzed to determine the appropriateness of the extracted results based on the accuracy of label classification and logical association matching;
[0046] Only time expressions identified by deep learning: If the confidence of the time expression is high, then keep the result directly and add rules to better improve the time extraction rules;
[0047] Only temporal expressions identified by the rule set: Use cross-validation to evaluate the performance of rule matching; if the results of rule matching on multiple validation sets are consistent, then the results are considered more likely to be correct.
[0048] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention aims to solve the limitations of complexity, diversity and context dependency of time expressions in the extraction of time expressions from watershed feature texts; the present invention not only solves the problem of time expression dependency in context, but also can identify time expressions that cannot be learned by the deep learning module through rules, thereby greatly improving the ability to extract time expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of the present invention;
[0050] Figure 2 It is a schematic diagram of the deep learning module structure proposed in the present invention. DETAILED DESCRIPTION
[0051] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0052] like Figure 1 As shown, the present invention proposes a method for extracting time expressions based on deep learning auxiliary rules. By combining the rule module and the deep learning module, the complexity and diversity problems in the extraction of time expressions of watershed feature texts are solved. For the text: "In the first quarter of 2024, multiple weather stations reported abnormally high rainfall, especially during the rainstorms from late February to early March. Starting from around 9 am on February 20, 2024, the study will be conducted within a period of one month. The relevant research will start next Wednesday night, aiming to collect data and combine model analysis to deeply explore the causes of rainstorms. "It is explained, specifically including the following steps:
[0053] Step 1: In the rule module, define time words, time modifiers, and time connectors to form a rule set, and use the rule set to identify time expressions in text data.
[0054] Build a rule set: Extract the co-occurrence pattern of time expressions from a large number of water conservancy description texts to form a preliminary rule set (including time words, time modifiers, and time connectors). Find the words that are not covered in the training corpus and add them to the rule set. The added rules must be unambiguous and have strong universality. For example: [YEAR_REGEX]
[0055] [1-2][0-9]{3}, which means matching any year between 1000 and 2999.
[0056] Identify time marker words: Use the time words in the rules to identify the time marker words in the text. For example, in the sentence "In the first quarter of 2024, multiple weather stations reported unusually high rainfall, especially during the heavy rain period from the end of February to the beginning of March.", identify time marker words such as "year" and "month".
[0057] Judge time modifiers and conjunctions: Use the time modifiers in the rules to judge whether there is information modifying the time before and after the directly identified time expression. Numbers modify the year, month, day, hour, and minute. "To" is used to connect "the end of February" and "the beginning of March", indicating the time range.
[0058] Combine time expressions: When there is only a time word, it is directly used as a time expression. If there is a time modifier collocating with a time word, combine them with the modifier in front and the time word behind. For example, in "the first quarter of 2024", "the first quarter" is the time word, and "2024" can be regarded as the time modifier of "the first quarter". The combined time expression is "the first quarter of 2024". When formats like "[time word 1] and [time word 2]", "[time word 1] to [time word 2]", "[time word 1] until [time word 2]" appear, use time conjunctions such as "to", "until", "and" as clues to combine the relevant time words and modifiers before and after together. For example, for "from the end of February to the beginning of March", first determine "the end of February" and "the beginning of March" respectively, and then connect them through the time conjunction "to", finally forming a complete time expression "from the end of February to the beginning of March" that represents the time range.
[0059] Through step 1, the information extracted from the example text is: "the first quarter of 2024" (combined by the time word "quarter" and the modifier "2024"), "from the end of February to the beginning of March" (combined by the time conjunction "to"), "around 9 am on February 20, 2024" (combination of multiple time marker words), "a cycle of 1 month" (combination of a time modifier and a time word).
[0060] Step 2: As Figure 2 shown, use the pre-trained model to fuse context information with Bi-LSTM, calculate the feature function by means of position encoding, introducing lexical features and part-of-speech features, learn the weights through the maximum likelihood estimation of the CRF layer, and finally use the Viterbi algorithm for decoding to achieve the optimal extraction of time expressions and results.
[0061] Word vector conversion: Convert the word sequence obtained by word segmentation into a word vector representation through the BERT model. BERT uses pre-trained context information to obtain the semantic vector representation of each word. For example, convert ("the end of February", "to", "the beginning of March") into a three-dimensional array.
[0062] Capturing contextual information: The word vector obtained by BERT is input into the bidirectional long short-term memory network, and the contextual information of the text is captured through the forward and backward LSTM layers. It is captured that "around 9 am on February 20, 2024" is the starting time point of the research, and "within 1 month" is the time span of the research.
[0063] Add relative position encoding: The intermediate representation h learned for each word in the attention mechanism t Add relative position code p t , relative position encoding allows the model to understand the order of these time information. For example, "around 9 am on February 20, 2024" is used as the starting time. Through the combination of position encoding and its hidden state, the model can clearly see that it is earlier than "next Wednesday night" and is within the large time range of "the first quarter of 2024" and before "late February to early March", thereby constructing the relative position relationship of the timeline, so that the model has a preliminary structured understanding of the time frame of the entire text. The specific calculation formula is as follows:
[0064] h' t =h t +p t
[0065] Among them, h t is the hidden state of each word at time t, p t represents the position encoding vector at time t.
[0066] Calculate the attention weight, assign weights to the intermediate representations learned for each word, calculate the correlation between each word and other words, and help the model focus on time-related words and avoid interference from irrelevant information. The specific calculation formula is as follows:
[0067]
[0068] Among them, e t is calculated through a scoring function, which converts the output h of BERT-Bi-LSTM t and a learnable weight matrix W as input.
[0069] Calculate the weighted context vector: Send c to the CRF layer to extract time information, determine whether each token is a time marker, and the specific components of time. The specific calculation formula is as follows:
[0070] c=∑ t α t h' t
[0071] Where c is the weighted context vector, α t is the attention weight, h't is the hidden state of each word at time t.
[0072] Calculate feature function: The feature function of the CRF layer is based on the current weighted context vector c t and the label y of the previous word t-1 To calculate. The lexical features are filtered and classified by statistical analysis of the corpus. The part-of-speech features are based on part-of-speech tagging and take into account different parts of speech and combination patterns to accurately describe the word label relationship. The specific calculation formula is as follows:
[0073] f(c t ,y t-1 )=σ(W c c t +W vocab F vocab +W pos F pos +b)
[0074] Among them, F vocab is the vocabulary feature vector, F pos is the part-of-speech feature vector, W c , W vocab , W pos is the learnable weight matrix and b is the bias vector.
[0075] Maximum Likelihood Estimation: The CRF layer learns the weight w of each feature function through maximum likelihood estimation k The process of weight learning is to minimize the negative log-likelihood function through optimization algorithms such as gradient descent to maximize the probability of the label sequence on the training data. The specific calculation formula for weight learning is as follows:
[0076]
[0077] Where L(θ) is the log-likelihood function, p(y|x) is the conditional probability of the label sequence y given the input sequence x, and Z(x) is a normalization factor that ensures that the sum of the probabilities of all possible label sequences is 1.
[0078] Decoding: The Viterbi algorithm is used during decoding, starting from the first label, and gradually calculating the maximum probability path to each label, and finally backtracking the most likely label path of the entire sequence. Relying on the learned feature functions and transition probabilities, the logical labels of time expressions in the text time series are accurately labeled, such as "research start time", "research time range definition", "related time points in the research process", etc. In this way, the shortcomings of the rule module in understanding the logical association of time information are made up.
[0079] Through step 2, the information extracted from the sample text is: 2024 (B-Time), first quarter (B-Time), end of February (B-Time), beginning of March (B-Time), February 20, 2024 (B-Time), morning (I-Time), next Wednesday night (B-Time). Among them, (B-Time) indicates the beginning of the time expression, and (I-Time) indicates the middle part of the time expression.
[0080] Step 3: Integrate the time expressions of the two modules and adopt different processing strategies according to whether the extraction results are consistent.
[0081] Calculate similarity: Use the Word2Vec model to convert the time expressions extracted by the two modules into word vectors. Calculate the average value of their word vectors to obtain the representation vector of the phrase. Use the cosine similarity formula to compare the vectors of the two phrases. Based on the calculated similarity value, set a threshold to decide whether the two time expressions are considered similar. The specific calculation formula is as follows:
[0082]
[0083] Among them, A and B are the vector representations of two phrases.
[0084] The similarity of the extraction results is high: For example, the time expressions extracted by the two modules are "the morning of February 20, 2024" and "around 9 am on February 20, 2024". The word vectors are A = [a1, a2, a3, ...] and B = [b1, b2, b3, ...]. The similarity value is calculated to be 0.9 (the threshold is set to 0.8), which means that the two are highly similar. It is believed that both deep learning and rule matching have identified the same time information, which can be directly used as the final result.
[0085] The similarity of the extracted results is low: If the similarity result is less than 0.8, the similarity of the extracted results is considered to be low. For example, the time expression extracted by the rule module is "late February to early March", while the time expression extracted by the deep learning module is "late February" and "early March". Extract the text of its context window (5 words before and after), convert these texts into feature vectors, and assume that the vectors C = [c1, c2, c3, ...], D = [d1, d2, d3, ...]. After averaging or concatenating C and D, use the trained classifier for analysis. The classifier determines that "late February to early March" is more appropriate in this context, and takes it as the final result.
[0086] Only time expressions recognized by the deep learning module: For example, if the recognized time expression is "next Wednesday night" and its confidence is 0.85 (assuming the upper confidence threshold is 0.8), then keep the result "next Wednesday night" and analyze the characteristics of this result, such as its vocabulary composition, position in the text, etc., to update the rules, for example, add matching patterns related to "next [week X] [time period]" to the rules to improve the time extraction rules.
[0087] Only time expressions recognized by the rule module: For example, if there is a rule of "period [number] months", the rule matching module extracts "period 1 month" according to this rule. Use multiple validation sets for verification. If the time expression "period 1 month" can be consistently extracted according to the rule in these validation sets, then this result is considered to be more likely to be correct.
[0088] Through step 3, the information extracted from the sample text is:
[0089] <TIMEX3 tid="t1"type="QUARTER"> First quarter of 2024
[0090] <TIMEX3 tid="t2"type="DURATION"> Late February to early March
[0091] <TIMEX3 tid="t3"type="DATE_TIME"> Around 9:00 am on February 20, 2024
[0092] <TIMEX3 tid="t4"type="DURATION"> Cycle 1 month
[0093] <TIMEX3 tid="t5"type="DATE_TIME"> Next Wednesday night.
[0094] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.
Claims
1. A method for extracting time expressions based on deep learning auxiliary rules, characterized in that: The following steps are involved: (1) Construct a rule set of time words, time modifiers, and conjunctions, use regular expressions to match specific formats, and identify time expressions in text; (2) The pre-trained model is used to fuse the context information with Bi-LSTM, and the relative position encoding is added to the intermediate representation learned for each word in the attention mechanism; (3) Assign weights to the intermediate representations learned for each word and calculate the correlation between each word and other words; (4) Calculate the weighted context vector and send it to the CRF layer to extract time information, determine whether each token is a time marker, and the specific components of time; (5) Based on the current weighted context vector c t and the label y of the previous word t-1 Calculate the feature function of each word; (6) The weights are learned by maximum likelihood estimation of the CRF layer and decoded using the Viterbi algorithm to achieve optimal and result extraction of the time expression; (7) Integrate the extracted time expressions and adopt different processing strategies according to whether the extracted results are consistent.
2. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The implementation process of step (1) is as follows: Construct a rule set to extract the co-occurrence pattern of time expressions from the water conservancy description text to form a preliminary rule set; specifically including time words, time modifiers and time connectors; find uncovered words in the training corpus and manually add them to the rule set; Using the time words in the rules to identify the time marker words in the text; Use the time modifiers in the rules to determine whether there is any time modification information before and after the directly recognized time expression; Combine time markers, modifiers, and time connectors to build a complete time expression.
3. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The intermediate representation h learned for each word in the attention mechanism in step (2) is t The implementation process of adding relative position encoding is as follows: Generate position encoding vector p through sine and cosine functions t , and the intermediate representation h learned for each word t Adding helps the model distinguish words in different positions and understand the relative position relationship between time words, such as the order of precedence. The specific calculation formula is as follows: h' t =h t +p t Among them, h t is an intermediate representation learned for each word, p t represents the position encoding vector at time t.
4. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The step (3) is implemented by the following formula: Among them, α t is the attention weight, e t It is calculated by the scoring function, and the output h of BERT-Bi-LSTM is t and a learnable weight matrix W as input.
5. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The step (4) is implemented by the following formula: c=∑ t a t h' t Where c is the weighted context vector, α t is the attention weight, h' t is the hidden state of each word at time t.
6. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The step (5) is implemented by the following formula: f(c t ,y t-1 )=σ(W c c t +W vocab F vocab +W pos F pos +b) Among them, F vocab is the vocabulary feature vector, F pos is the part-of-speech feature vector, W c , W vocab , W pos is the learnable weight matrix and b is the bias vector.
7. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The process of learning weights by maximum likelihood estimation through the CRF layer in step (6) is as follows: The gradient descent optimization algorithm is used to optimize the gradient information of the model parameters, and the parameter values are continuously adjusted to gradually approach the optimal solution, thereby finding the optimal label sequence: Among them, L(θ) is the log-likelihood function, p(y|x) is the conditional probability of the label sequence y given the input sequence x, and w k is the weight of each feature function, and Z(x) is a normalization factor that ensures that the sum of the probabilities of all possible label sequences is 1.
8. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The decoding process using the Viterbi algorithm in step (6) is as follows: The feature function weight w k and transition probabilities as input, treating the temporal expression as a sequence of different labels; calculating the probabilities of all possible temporal expression paths; The path with the largest probability value is selected from these paths, which is the most likely time expression parsing result.
9. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The implementation process of step (7) is as follows: Use the Word2Vec model to convert the extracted time expressions into word vectors; calculate the average value of their word vectors to obtain the representation vector of the phrase; use the cosine similarity formula to compare the vectors of the two phrases; and set a threshold based on the calculated similarity value to decide whether the two time expressions are considered similar. The similarity of the extracted results is high: it is considered that both deep learning and rule matching have identified the same time expression, which is directly used as the final result; The similarity of the extracted results is low: the results are analyzed through the context window to determine which result is more appropriate in a specific context; the text in the context window is converted into feature vectors and these vectors are averaged or concatenated; then the SVM classifier is trained using the manually annotated time expressions and labels as data sets, and the feature vectors are analyzed to determine the appropriateness of the extracted results based on the accuracy of label classification and logical association matching; Only time expressions identified by deep learning: If the confidence of the time expression is high, then keep the result directly and add rules to better improve the time extraction rules; Only temporal expressions identified by the rule set: Use cross-validation to evaluate the performance of rule matching; if the results of rule matching on multiple validation sets are consistent, then the results are considered more likely to be correct.
10. The method for extracting time expressions based on deep learning auxiliary rules according to claim 1, characterized in that: The threshold is 0.8.