Entity extraction and labeling method and device and storage medium
Through the multi-model collaborative decision-making method of dynamically allocating weights, the problem of low entity recognition and labeling efficiency in the prior art is solved, efficient entity extraction and labeling of non-standardized text is achieved, and the recognition accuracy and processing efficiency of key entities are improved.
Patent Information
- Application Number
- CN202510885888.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The prior art cannot effectively deal with entities associated with non-standardized expressions and contextual semantics, resulting in low efficiency in entity recognition and labeling, and a single model architecture cannot take into account the identification needs of different entity types, and requires manual secondary verification and integration.
By obtaining dynamic indicators in text data, dynamically assigning the weights of the rule engine, statistical model and deep learning model, configuring a hybrid recognition model, performing entity recognition and boundary verification, and outputting structured entity information.
It significantly improves the extraction accuracy and labeling efficiency of key entities, realizes the integrated processing of entity recognition and structured annotation, and improves the processing efficiency of task text.
Smart Images

Figure CN120387455A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text recognition, and in particular, to an entity extraction and annotation method, device, and storage medium. Background Art
[0002] In some document analysis processes, accurately extracting key entity information such as time, location, and relevant individuals is a fundamental link in document analysis and processing. To achieve entity information extraction, a common method is to perform entity recognition through predefined regular expressions and keyword dictionaries. Although such a method has a good extraction effect on highly structured texts (such as fixed-format tabular data), it is difficult to adapt to non-standardized expressions and cannot effectively process entities with context semantic associations. For example, in the text to be analyzed, when there is a non-standard description such as "around 20 o'clock", or when there is a description that requires determining entity relationships based on context semantics, this method requires deducing the relationships between the characters in the technical problem. In specific cases, a large number of rules need to be manually maintained, resulting in high maintenance costs.
[0003] More prominent is that existing solutions generally adopt a single model architecture and cannot meet the recognition requirements of different entity types. This technical limitation directly leads to the need for manual secondary verification and integration of the output results of different models in actual task processing, seriously affecting the task processing efficiency.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide an entity extraction and annotation method, device, and storage medium, aiming to solve the technical problem that the prior art cannot achieve integrated processing of entity recognition and annotation for specified canonical processed texts.
[0006] To achieve the above purpose, this application proposes an entity extraction and annotation method, and the method includes: Obtain text data to be processed; Extract dynamic indicators reflecting text characteristics from the text data; Determine the combined weights of a rule engine, a statistical model, and a deep learning model based on the dynamic indicators; Configure a hybrid recognition model according to the combined weights, and identify entities in the text data through the hybrid recognition model; Perform boundary verification and type annotation on the recognition results, and output structured entity information.
[0007] In one embodiment, the step of extracting dynamic indicators reflecting text characteristics from the text data includes: Analyze the distribution law of the standard paragraphs of relevant clauses in the text data, and determine the target area in the set of standard professional terms; Evaluate the relevance between the text expression and the field to which the text information belongs according to the distribution law and the target area, and extract the dynamic indicators according to the evaluation results.
[0008] In one embodiment, the step of determining the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators includes: Select a corresponding weight allocation strategy according to the application scenario type of the text data; Assign decision weights to the rule engine, statistical model, and deep learning module through the weight allocation strategy and the dynamic indicators; Among them, the step of assigning decision weights to the rule engine, statistical model, and deep learning module through the weight allocation strategy and the dynamic indicators includes: Determine the model features corresponding to the dynamic indicators in the pre-established data relationship table; Increase and / or decrease the decision weights in the data model corresponding to the model features according to the weight allocation strategy.
[0009] In one embodiment, the entity extraction and annotation method further includes: Monitor the recognition efficiency of the rule engine, statistical model, and deep learning model in the text data in real time; Dynamically correct the initial weight allocation of the rule engine, statistical model, and deep learning model according to the determination result of the monitoring result and the preset weight adjustment condition.
[0010] In one embodiment, the step of configuring a hybrid recognition model according to the combined weights and identifying entities in the text data through the hybrid recognition model includes: Analyze the output result of the hybrid recognition model, and the output result is a set of entity candidates output in parallel by each data model; Rank the entities in the set of entity candidates by priority, and filter out entities below the confidence threshold; Use the filtering result as the recognition result of the entity.
[0011] In one embodiment, the step of ranking the entities in the set of entity candidates by priority and filtering out entities below the confidence threshold includes: Rank the entities with high priority by confidence; Adjust the ranking result by using the combined weights, and filter out entities below the confidence threshold in the adjusted ranking result by means of a weighted voting mechanism.
[0012] In one embodiment, the step of performing boundary verification and type annotation on the recognition result and outputting structured entity information includes: Correct the entity boundary position in the recognition result through the punctuation distribution feature in the recognition result; Use a preset entity length threshold to filter out the recognition results that do not meet the requirements to obtain target entities; Perform type annotation on the target entity and output structured entity information.
[0013] In one embodiment, the step of performing type annotation on the target entity and outputting structured entity information includes: Identify the text data through a bidirectional LSTM-CRF model, and predict the entity boundary region of the target entity in the text data; Based on the syntactic analysis tree, correct the position of the entity boundary region to determine the annotation range of the target entity, and perform entity annotation within the annotation range; Execute the step of outputting structured entity information.
[0014] In addition, to achieve the above object, the present application also proposes an entity extraction and annotation device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the entity extraction and annotation method as described above.
[0015] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the entity extraction and annotation method as described above.
[0016] One or more technical solutions proposed by the present application have at least the following technical effects: The technical solution of the present application obtains the text data to be processed; extracts the dynamic indicators reflecting the text characteristics from the text data; determines the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators; configures a hybrid recognition model according to the combined weights, and identifies the entities in the text data through the hybrid recognition model; performs boundary verification and type annotation on the recognition result, and outputs structured entity information. Through the dynamic weight allocation mechanism, the collaborative decision-making of multiple models is optimized, significantly improving the extraction accuracy and annotation efficiency of key entities such as time and location in the task text, and achieving the technical effect of integrating the intelligent recognition and structured annotation of task elements. Description of the Drawings
[0017] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart provided for the first embodiment of the entity extraction and annotation method of this application; Figure 2 It is a schematic diagram of the device structure of the hardware operating environment involved in the entity extraction and annotation method in the embodiments of this application.
[0020] The implementation, functional features, and advantages of the objectives of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0022] To better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and the specific implementation manners.
[0023] The main solution of the embodiments of this application is: obtaining text data to be processed; extracting dynamic indicators reflecting the text characteristics from the text data; determining the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators; configuring a hybrid recognition model according to the combined weights, and identifying entities in the text data through the hybrid recognition model; performing boundary verification and type annotation on the recognition results, and outputting structured entity information.
[0024] Since existing solutions generally adopt a single model architecture, they cannot take into account the recognition requirements of different entity types. This technical limitation directly leads to the need for manual secondary verification and integration of the output results of different models in the actual task text processing, seriously affecting the processing efficiency of the task text.
[0025] This application provides a solution, which optimizes multi-model collaborative decision-making through a dynamic weight allocation mechanism, significantly improves the extraction accuracy and annotation efficiency of key entities such as time and location in task text, and achieves the technical effect of integrating intelligent recognition and structured annotation of task elements.
[0026] Based on this, the embodiments of this application provide an entity extraction and annotation method, referring to Figure 1, Figure 1 This is a flowchart of the first embodiment of the entity extraction and annotation method of the present application. In this embodiment, the entity extraction and annotation method includes steps S10 to S50: Step S10, obtain the text data to be processed; Step S20, extract dynamic indicators reflecting the text characteristics from the text data; Step S30, determine the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators; Step S40, configure a hybrid recognition model according to the combined weights, and identify the entities in the text data through the hybrid recognition model; Step S50, perform boundary verification and type annotation on the recognition result, and output structured entity information.
[0027] This embodiment provides an integrated processing method for entity extraction and annotation in the text data to be processed, which is essentially achieved through a multi-model collaboration mechanism with dynamic weight allocation. When the text data to be processed is obtained, dynamic indicators that can reflect the text features are extracted from the text data to be processed. The text feature dynamic indicators are used to quantitatively reflect a key parameter set of the relevant specified semantic characteristics, structural complexity, and context relevance of the text, and specifically include language complexity, domain term density, and context relevance. Among them, the language complexity can evaluate the text structuring difficulty by calculating indicators such as the proportion of long and difficult sentences and the depth of nested clauses. For example, when the Flesch-Kincaid readability index of the colloquial fragments in the text is lower than 30, it is marked as a text with high language complexity; the domain term density can be used to count the ratio of the frequency of standard entity words to general vocabulary, and when the density > 15%, it is determined as highly professional content; the context relevance is based on the co-occurrence line to analyze the logical relationship strength between entities, such as the co-occurrence frequency of "abnormal behavior individual - specific instrument".
[0028] Furthermore, a sliding window algorithm is used to detect the local feature fluctuations of the text to identify hybrid paragraphs, such as cross-modal texts that simultaneously contain account transaction data and inquiry ratios. Specifically, the step of extracting dynamic indicators reflecting the text characteristics from the text data includes: Analyze the distribution law of the relevant clause standard paragraphs in the text data, and determine the target area in the set of standard professional terms; Evaluate the relevance between the text expression and the field to which the text information belongs according to the distribution law and the target area, and extract the dynamic indicators according to the evaluation results.
[0029] In this embodiment, the standard text features are quantified and analyzed to provide a decision-making basis for subsequent model weight allocation. Specifically, a hierarchical analysis method is adopted to locate the relevant clause standard paragraphs of the text data to be processed, and operations such as structural feature analysis and distribution rule modeling are performed on the positioning results.
[0030] In the structural feature analysis, a paragraph structure fingerprint is established by identifying the starting markers of standard paragraphs, including the citation of standard clauses, such as the declarative type of conclusion found after inspection, such as the conclusion type of "this institution believes", etc. The paragraph structure fingerprint uses regular expression matching of marker words + context window verification (the first 5 words and the last 10 words) to form a three-dimensional fingerprint vector containing position, type, and level.
[0031] In the distribution rule modeling, spatial distribution analysis is used to calculate the standard paragraph spacing and detect the paragraph aggregation degree. For example, a key rule clause appears at an average of every 2 pages; the relevant rule clause paragraphs of specific abnormal events are concentrated in the front part of the text, etc.
[0032] Then, a time series curve of "standard expression intensity" is constructed by time series analysis to identify its periodic pattern. For example, a standard clause confirmation appears once every 3 rounds of Q&A in the question and answer record text.
[0033] According to the structural feature analysis and distribution rule modeling of the above positioning results, the target area where the professional terms are located is located in the processing results. During the positioning process, the professional terms can be matched based on the basic term library. Specifically, an enhanced Trie tree is constructed by loading the standard term knowledge base to achieve efficient matching, so as to locate the professional target area according to the matching results.
[0034] Among them, a multi-level Trie tree structure dedicated to the standard field is constructed, and its core processing includes knowledge base preprocessing, hybrid index construction, and dynamic loading mechanism. Specifically, the knowledge base preprocessing can stratify the existing standard terms into multiple major categories, establish parent-child node associations for specified nested terms, and attach weight factors to each node; the hybrid index construction combines character-level Trie and word-level jump tables, and uses the first character trigger + hash-assisted positioning for long terms to achieve an average query efficiency of O(1) level; the dynamic loading mechanism needs to preload the core term branches according to the application scenario type, and control the memory occupancy within a predetermined memory amount.
[0035] Moreover, the term influence radius of the target area where the located professional terms are located is determined through context expansion detection. Specifically, the specific value of the influence radius can be obtained through empirical research. For example, the influence radius R = 5 sentences, and it can be correspondingly set according to the specific application scenario type represented by the text data to be processed.
[0036] As shown above, the target regions are merged according to the determined influence radius of the terms to generate continuous target regions, and this merging is implemented by the DBSCAN clustering algorithm. Specifically, based on the positioning process of the target regions where the above-mentioned professional terms are located, it can be characterized by the following examples.
[0037] In addition, for the positioning of the target regions where the professional terms are located obtained from the above analysis, it is also necessary to evaluate its relevance to the field to which the text information belongs to improve the accuracy. This evaluation can be achieved through a preset relevance evaluation mechanism for specification clauses. The relevance evaluation mechanism for specification clauses is provided with a multi-dimensional evaluation system, including the expression professionalism score and the field consistency analysis. The expression professionalism score extracts paragraph embeddings using Legal-BERT and calculates the cosine similarity between the positioning of the professional term target regions and the text benchmark set of the specification clauses. Calculate the text perplexity in the language model of the field to which the text information belongs according to the calculation result of the cosine similarity, and obtain the consistency index according to the calculation result.
[0038] After that, dynamic indicators are generated in the consistency index, and the dynamic indicators are output in the form of an index matrix. In the index matrix, the index names include specification density, expression specification degree, field consistency, and structural complexity, and the parameters represented by each index name have corresponding calculation methods and normalization ranges. The specific calculation methods include TermDensity / MAX_Density, Sigmoid(Ps*2 - 1), Consistency, and 1 - (standard paragraph spacing / MAX_spacing). And the normalization range of the calculation method is related to the application scenario type represented by the text data to be processed.
[0039] In this embodiment, through a hybrid detection method of structural analysis and semantic understanding, relevant analysis is performed on the text data to be processed, and the region where the dynamic indicators are located is accurately selected with the dynamic parameters that stimulate the corpus characteristics, improving the recognition accuracy of standard paragraphs and the positioning accuracy of term regions. Moreover, the evaluation of the processing target through the multi-granularity evaluation system also improves the consistency between the relevance evaluation and manual judgment.
[0040] According to the dynamic indicators extracted from the text data to be processed as shown above, determine the dynamic allocation of the weights of multiple models. Specifically, the multiple modules include a rule engine, a statistical model, and a deep learning model, that is, weights are allocated to the rule engine, the statistical model, and the deep learning model based on the dynamic indicators to form a combined weight.
[0041] Among them, a weight allocation policy library is pre-constructed, and a data table based on the mapping relationship between application scenario types and weight policies is set in the weight allocation policy library, which contains multiple weight allocation policies, so as to allocate decision weights to each data model based on the weight allocation policies. That is, the step of determining the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators includes: Select the corresponding weight allocation policy according to the application scenario type of the text data; Allocate decision weights to the rule engine, statistical model, and deep learning module through the weight allocation policy and the dynamic indicators; Among them, the step of allocating decision weights to the rule engine, statistical model, and deep learning module through the weight allocation policy and the dynamic indicators includes: Determine the model features corresponding to the dynamic indicators in the pre-established data relationship table; Increase and / or decrease the decision weights in the data model corresponding to the model features according to the weight allocation policy.
[0042] In the mapping relationship data table of application scenario types and weight policies set in the weight allocation policy library, the multiple weight allocation policies stored can be divided into three types: term-dominated policies, statement-sensitive policies, and association-intensive policies, which are respectively applicable to different application scenario types.
[0043] Specifically, each weight allocation policy is set with an initial weight allocation, and the initial weight allocation includes the rule engine benchmark weight, the statistical model benchmark weight, and the deep learning model benchmark weight. The basic weight values of each benchmark weight are different, and the basic weight values are dynamically adjusted according to the complexity of the usage characteristics and the term density.
[0044] During the dynamic adjustment process, it can be defined as a dynamic weight calculation process. During the dynamic weight calculation process, the basic policy can be determined according to the application scenario type classifier. Among them, the application scenario type classifier can obtain the initial weights of each data model by inputting the dynamic indicators of the text data to be processed, that is, the initial weights of the rule engine, statistical model, and deep learning model. The dynamic indicators to be input into the application scenario type classifier are the event reasons marked in the document header of the text data to be processed, the high-frequency entity type distribution, and the data flow pattern features, and the initial weights of its data models are set to 0.3 - 0.5 - 0.2 respectively.
[0045] Further, the initial weights are adaptively adjusted according to the dynamic metrics. Since each data model has different focuses in data processing, the adaptive adjustment is essentially a corresponding dynamic adjustment of the initial weights based on the data calculation tendencies of each data model (rule engine, statistical model, and deep learning model). Specifically as follows: The dynamic weight adjustment based on the rule engine is as follows: for term density, regulatory clause citation, and complexity, the initial weight of the rule engine is increased based on the benchmark weight value. For example, for each fixed percentage increase in term density, such as 5%, the weight is increased by a preset value, such as 0.1 (ceiling + 0.3); when a regulatory clause citation is detected, the weight is instantaneously increased by a preset value, such as 0.15; when the complexity is greater than the preset complexity threshold comparison, the weight is decreased by a preset unit value. For example, when the complexity > 0.8, the weight is decreased by 0.05 / 0.1 unit. Among them, the benchmark weight value based on the rule engine can be set according to the information represented by the initial weight. For example, the 5% increase of 0.1 as described above. An example of the weight adjustment based on the rule engine is as follows: the initial weight of a certain task is 0.6. When the term density reaches the preset upper limit value, such as 25%, the final weight rises to the upper limit weight value, such as 0.75.
[0046] The dynamic weight adjustment based on the statistical model is to increase or decrease the initial weight value based on complexity, infrequent word combinations, and context stability. Among them, the numerical range where the complexity is determined, and the initial weight of the statistical model is adjusted according to the data relationship between the benchmark weight value and the complexity. The benchmark weight value and the complexity are linearly correlated. For example, when the complexity is within the preset numerical range, such as the 0.5 - 0.7 range, the weight is linearly positively correlated with the complexity; when an infrequent word combination is found, the initial weight is decreased by the benchmark weight value, such as the weight is decreased by 0.1; and according to the comparison result between the context stability and the preset stability threshold, the initial weight is increased, such as the weight compensation is increased by 0.05; in practical applications, the application scenario of the dynamic weight adjustment based on the statistical model can be: when a specific word appears in a certain type of task text, the weight of the statistical model drops from 0.5 to 0.4.
[0047] The dynamic weight adjustment based on the deep learning model is to adjust the initial weight of the deep learning model based on the text relevance, cross - paragraph reference, and rule coverage of the text data to be processed. For example, for each 0.1 increase in relevance, the weight is increased by 0.08 (ceiling + 0.25); when a cross - paragraph reference is detected, the weight is instantaneously increased by 0.1; when the rule coverage > 70%, the weight is decreased by 0.05.
[0048] Finally, constraint processing is performed on the adjusted weights of each data model (rule engine, statistical model, and deep learning model). The constraint conditions include, but are not limited to, setting the upper and lower limit values of the weights of each data model and the total weight of each data model. For example, ensure that the weight of the rule engine ∈ [0.15, 0.8]; the weight of the statistical model ∈ [0.1, 0.6]; the weight of the deep learning model ∈ [0.2, 0.75]; the sum of the three is strictly equal to 1. Specifically, it can be obtained according to the application scenario type of the text data to be processed or detailed analysis.
[0049] In this embodiment, in the weight allocation for each data model, by providing prior knowledge of the application scenario type through the basic allocation strategy and adjusting the initial weights of each data model with dynamic indicators, it can more accurately reflect the real-time characteristics of the text. Furthermore, through the weighted decision fusion of the two, it not only maintains the type commonality but also captures the individual characteristics, further improving the entity recognition accuracy of the text data to be processed.
[0050] Furthermore, the entity extraction and annotation method further includes: Real-time monitoring of the recognition efficiency of the rule engine, statistical model, and deep learning model in the text data; Dynamically correcting the initial weight allocation of the rule engine, statistical model, and deep learning model according to the determination result of the monitoring result and the preset weight adjustment condition.
[0051] In this embodiment, during the text data recognition process, the fixed weight allocation scheme cannot adapt to the dynamic changes of entities in the text features. Therefore, when the above-mentioned dynamic weight allocation is applied, the initial weights of each data model are further adjusted by real-time monitoring the efficiency detection of each data model, so that the entity recognition effect of each data model conforms to the diversity of the text data to be processed.
[0052] In the specific implementation process, an efficiency detection index system and a weight adjustment trigger condition are set, and it is determined whether the detection result of the efficiency detection index system meets the weight adjustment trigger condition, and the initial weights of each data model (rule engine, statistical model, and deep learning model) are dynamically adjusted according to the determination result. Among them, multi-dimensional detection indexes are set in the efficiency detection index system to respectively perform efficiency detection on the rule engine, statistical model, and deep learning model, as follows: Rule engine: Real-time statistics of the rule matching rate (number of successful matches / number of attempted matches) and the number of rule conflicts; Statistical model: Monitoring the entity boundary accuracy rate (evaluated based on a sliding window) and the type confusion matrix; Deep learning model: Tracking the attention focus degree (key entity attention weight) and the associated inference accuracy rate.
[0053] Specifically, in the weight adjustment trigger conditions set for the multi-dimensional detection indicators, the multi-dimensional detection indicators are judged through the set response mechanism. The trigger conditions of each level of the response mechanism are different. Specifically, for example, when a certain core indicator of a model is lower than the benchmark value by 10% (such as the rule matching rate < 65%), a first-level response is triggered; when the indicator is lower than the benchmark value by 20% for 3 consecutive paragraphs (such as the statistical model boundary accuracy rate < 58%), a second-level response is triggered; when a systematic failure occurs (such as the associated inference of the deep learning model being completely wrong), a third-level response is triggered. Specifically, it can be set according to the entity recognition results of each data model for the text data.
[0054] In the process of dynamically adjusting the initial weights of each data model according to the above-mentioned judgment results, specifically, it can be adjusted based on the response mechanism set in the weight adjustment trigger conditions, that is, the initial weights of each data model are differentially adjusted according to the warning level, as shown below: The initial weight adjustment mechanism of the rule engine can be limited to: in the case of a first-level warning, the weight is reduced and the similar rule extension search is activated; in the case of a second-level warning, the weight is lowered and the alternative rule set is enabled; in the case of a third-level warning, the engine is suspended and the weight is transferred to other models; The initial weight adjustment mechanism of the statistical model adjustment can be limited to: in the case of a first-level warning, the weight can be compensated and increased and the context window is narrowed; in the case of a second-level warning, the feature engineering reconstruction is triggered to dynamically float the weight value; in the case of a third-level warning, the feature template is switched and the weight is reset to the benchmark value; The initial weight adjustment mechanism of the deep learning model adjustment can be limited to: in the case of a first-level warning, the weight is increased to enhance the attention bias; in the case of a second-level warning, the adversarial sample filtering is activated to lock the weight upper limit; in the case of a third-level warning, the domain adaptation copy is loaded and the initial weight is doubled stage by stage.
[0055] Specifically, in this embodiment, by constructing a closed-loop feedback system, the initial weights of each data model can dynamically evolve with the practice of the specification clauses, retaining both the specification and rigor of the rule engine and giving full play to the adaptive advantages of the data-driven model, further improving the technical effect of entity recognition accuracy. The dynamic weight adjustment values shown above can be obtained according to the real-time detection results of each data model. And according to the initial weight adjustment results of each data model, the final weight values are integrated to obtain the combined weights of the rule engine, the statistical model, and the deep learning model to perform entity recognition on the text data to be processed.
[0056] Further, when the rule engine, statistical model, and deep learning model perform entity recognition, they perform parallel processing, that is, these three data models simultaneously perform entity recognition on the text data and output the recognition results. Considering the differences in the data processing methods of each data model, a hybrid recognition model can be configured based on the combined weights of the rule engine, statistical model, and deep learning model, and entity recognition is performed on the text data in the form of the hybrid recognition model.
[0057] Specifically, in the hybrid recognition model, the data structures of the rule engine, statistical model, and deep learning model are respectively configured according to the combined weights. Alternatively, the combined weights of the rule engine, statistical model, and deep learning model are merged to form the data processing layer of the hybrid recognition model, and then corresponding processing channels are set for the configured rule engine, statistical model, and deep learning model. The settings of the processing channels are all adapted to the characteristics of the recognized entities of each data model. Specifically, it can be as follows: The working mode of the rule engine processing channel is hierarchical matching based on a specification knowledge base. For example, the first layer is forced rule matching (such as the numbers of specification clauses and specification terms); the second layer is the execution of inference rules; the third layer is semantic rule verification; the output characteristics of this data model focus on high accuracy and low recall rate, and it has built-in specification logic constraints.
[0058] The feature processing setting parameters of the statistical model processing channel include character-level features, context window features, and application scenario type features; its output characteristics focus on balance determination, stability of regular entities, and whether it can handle out-of-vocabulary words, etc.
[0059] The network structure of the deep learning model channel includes a Legal-BERT pre-training layer, a BiLSTM sequence encoding layer, and a multi-head attention mechanism; the output characteristics focus on high recall rate (91%-93%) and the ability to discover potential associations. This data structure requires a large amount of computing power support.
[0060] As shown above, the entities of the text data are recognized according to the set hybrid recognition model, and the entity recognition results are output. The entity recognition results are output in parallel and can be output in the form of an entity dataset. Then, based on the output results of the entity recognition results, an analysis is performed, and specific entity content is obtained according to the analysis results. That is, the step of configuring the hybrid recognition model according to the combined weights and recognizing the entities in the text data through the hybrid recognition model includes: Analyze the output results of the hybrid recognition model, and the output results are the entity candidate sets output in parallel by each data model; Perform priority sorting on the entities in the entity candidate set, and filter out entities below the confidence threshold; Use the filtering result as the recognition result of the entity.
[0061] As shown above, entity recognition is performed on the text data based on the configured hybrid recognition model, and the entity recognition result is output. In a specific implementation, each data model independently outputs an entity set, and the output entity sets are formed into an original candidate pool for specific entity analysis. Among them, due to the recognition features of each data model, the output entity sets have different text features. For example, the output result of the rule engine is an entity with a mandatory mark; the output result of the statistical model is an entity segment with a probability value; the output result of the deep learning is an entity with an attention weight and its relationship.
[0062] In the original candidate pool, the output results of each data model are used as entity candidate sets for priority ranking. In this priority ranking, a three-level priority system is set up, and the ranking is carried out respectively through the canonical key entity, the task feature entity, and the basic description entity. Among them, the canonical key entity is an entity related to the constitutive elements of non-standard behavior; the task feature entity is an entity that reflects the particularity of the task; the basic description entity is a general task element entity, such as time and location.
[0063] According to the ranking result of the priority ranking shown, the confidence of each entity in the priority ranking is calculated respectively. In this embodiment, each entity is calculated using a preset entity final confidence calculation formula. The entity final confidence Score calculation formula is: Score = αS_rule + βS_stat + γ*S_dl, where α / β / γ are the current model combination weights (such as 0.5 / 0.3 / 0.2), S_rule is the output score of the rule engine (1 for successful matching, otherwise 0), S_stat is the statistical model probability value (0-1), and S_dl is the deep learning normalized attention value (0-1).
[0064] The operation of calculating the confidence of the entities in the priority ranking queue described above can be limited to calculating the confidence only for the entities with a high priority, that is, the entities with a low priority are directly excluded. Specifically, it can be determined according to the specific settings of the priority ranking. After that, after calculating the confidence of the high-priority entities in the priority ranking queue, further ranking is performed according to the calculation result of the confidence calculation, that is, the step of ranking the entities in the entity candidate set by priority and filtering out the entities below the confidence threshold includes: Rank the entities with a high priority by confidence; Use the combination weight to adjust the ranking result, and filter out the entities below the confidence threshold in the adjusted ranking result by a weighted voting mechanism.
[0065] Entities with high priority are sorted according to the confidence calculation results, and the combined weights used are used to adjust the sorting results. Essentially, it is to set different filtering thresholds according to the application scenario type. Since the application scenario type can also be expressed by the combined weights, it can also be defined as setting different filtering thresholds according to the combined weights. The entities sorted by confidence are adjusted according to the different filtering thresholds. Specifically, the different filtering thresholds are set as the thresholds corresponding to the application scenario type, and the thresholds corresponding to each application scenario type can be set according to the application scenario. Therefore, according to the numerical limit of the different filtering thresholds, the entities sorted by confidence are adjusted.
[0066] Furthermore, overlapping entity adjudication and type conflict resolution processing are performed according to the sorting results of the confidence sorting to exclude duplicate entities and conflicting entities. Among them, the overlapping entity adjudication can adjudicate the same entity by comparing the contribution weights of each data model, logical rationality verification, context consistency, and the reserved manual review marks to exclude duplicate entities; and conflicting entities are excluded through a pre-set type priority matrix. In the rules of the type priority matrix, it can be set that the rule engine type annotation has the highest authority. When there is a conflict between the statistical model and deep learning, the annotation matching the application scenario type is preferred, and the expert knowledge base query is initiated for the types that cannot be adjudicated. After excluding the entities that do not meet the requirements in the sorting results of the confidence sorting, the final entity recognition result is obtained.
[0067] Boundary verification and type annotation are performed on the text data to be processed according to the entity recognition results, and structured entity information is output. Specifically, based on the entities represented in the recognition results, the text information that conforms to the entities is marked in the text data to be processed. For example, if the entity is the event occurrence time, the date is marked in the text data, and it is indicated that the entity parameter is night. Specifically, it can be defined that the entity is the task item name represented by the text data to be processed, and the mark as the entity is the filling parameter reflecting the task item name.
[0068] Among them, the steps of performing boundary verification and type annotation on the recognition results and outputting structured entity information include: Correct the entity boundary positions in the recognition results through the punctuation distribution characteristics in the recognition results; Use a preset entity length threshold to filter out the recognition results that do not meet the requirements to obtain the target entities; Perform type annotation on the target entities and output structured entity information.
[0069] Since the text data of the text data to be processed is the task content in literal form, a boundary verification processing flow is set up based on the text data. In the boundary verification processing flow, there are punctuation-guided boundary correction and entity length threshold filtering. Among them, the punctuation-guided boundary correction analyzes the punctuation features of the entity context and establishes three types of correction rules, including special punctuation for standard documents: forcibly retaining the content within "《》" intact, enumeration punctuation: identifying parallel entities separated by "、" and ";", and statement termination punctuation: using "。" and "?" as hard constraints for entity boundaries. In addition, there is a punctuation position weight matrix to adjust the probabilities of the first and last characters of the entity. When the entity boundary falls between quotation marks or parentheses, it is automatically extended to a complete symbol pair.
[0070] In the entity length threshold filtering, there is a dynamic length threshold system, including basic threshold setting, such as for personal names (2-4 characters), etc., application scenario type adjustment, which can relax the numerical description length for financial abnormal behaviors, and text position compensation, which can relax the main text content by a certain proportion compared to the title. In addition, a sliding window is used to evaluate the rationality of the entity, and split, semantic compression, and whole sentence annotation processing are performed on over-long entities. That is, according to the processing of the entity length threshold filtering, the entities in the recognition result are filtered to obtain target entities. Furthermore, type annotation is performed on the target entities to output structured entity information.
[0071] Furthermore, based on the results of the boundary verification and type annotation for the above entities, when outputting the structured entity information of the target entities, it is achieved through a preset entity recognition mechanism. That is, the steps of type annotation for the target entities and outputting structured entity information include: Identifying the text data through a bidirectional LSTM-CRF model and predicting the entity boundary region of the target entity in the text data; Based on the syntactic analysis tree, correcting the position of the entity boundary region to determine the annotation range of the target entity, and performing entity annotation within the annotation range; Performing the step of outputting structured entity information.
[0072] In this embodiment, a domain optimization annotation system is pre-constructed, which can perform corresponding annotation processing for specified entity tags and special markers. Specifically, the standard entity tags include multiple main types + multiple sub-types (such as "related item - item name"), and the special markers are used to mark controversial entities and cross-sentence entities. Among them, the marked objects of the controversial entities need to be manually verified, and the cross-sentence entities can be processed with special associations to clarify the marked scope of the entities. Further, to improve the annotation effect of the domain optimization annotation system, the domain optimization annotation system can be trained based on attention-guided training, and its training scheme includes applying multiple loss weights to standard terms, establishing attention biases for the constituent elements of risk behaviors, and enhancing the robustness of the model using adversarial samples. Specifically, it can be implemented based on specific training schemes.
[0073] In this way, according to the constructed domain optimization annotation system, the entity boundary region of the target entity is marked in the text data, and the entity boundary region is the marked scope selected for the target entity by the domain optimization annotation system. Further, the specific position of the entity boundary region is assisted and corrected by a syntactic analysis tree to further determine the annotation scope of the target entity.
[0074] During the process of assisting and correcting by the syntactic analysis tree, it is implemented based on the analysis mechanism of dependency syntax. The analysis mechanism includes but is not limited to core word alignment, modifier adsorption, and ambiguity structure processing. Among them, core word alignment is used to ensure that the entity contains a complete dependency subtree, modifier adsorption can automatically expand the boundary to include necessary modifiers, and ambiguity structure processing preferentially selects the parsing path that conforms to the standard semantics. In practical applications, the analysis mechanism can be set by a syntactic analyzer enhanced with the domain of the text information to further correct the position of the entity boundary region of the target entity. When setting the syntactic analyzer enhanced with the domain of the text information, specific sentence pattern rules in the domain of the text information need to be added for the analysis mechanism to optimize the long-distance dependency parsing ability and support cross-paragraph anaphora analysis to achieve the position correction of the entity boundary region of the target entity. For example, the original boundary is "electronic account" (lacking key modifiers), and after correction, it becomes "a foreign domain network account used by a certain user" (a complete semantic unit).
[0075] As shown above, after correcting the position of the entity boundary region of the target entity according to the syntactic analysis tree, the final annotation scope of the target entity is obtained, and entity annotation is performed based on the annotation scope. In practical applications, after annotating the annotation scope, the target entity needs to be output in a structured form in combination with the annotated annotation scope to clearly indicate the entity object.
[0076] Specifically, the target entity and its annotation range are output through a structured output specification. The structured output specification needs to clarify the core entity attributes, including the core entity attributes, including text content (original surface form), standardized expression (standardized standard expression), and location information (document ID + start and end offsets); and, for the core entity attributes, indicate their type system, represented as main types (persons / objects / actions, etc.), sub-types (related participating persons / related objects, etc.), and extended attributes (time range, etc.); finally, highlight the association relationship between the target entity and its marked range and other target entities and their marked ranges, mainly including entity relationships and evidence chain pointers, etc.
[0077] Further, before the structured entity information of the target entity and its annotation range is output, screening for output quality control is also required. The output quality control specifically includes a verification mechanism and error tolerance processing, which respectively perform specification compliance checks, logical consistency verification, and evidence chain integrity assessment on the target entity and its annotation range. Among them, the specification compliance check needs to conform to the composition of specification risk behaviors, the logical consistency verification is used to verify the timeline and causal relationship of the target entity and its annotation range, and the evidence chain integrity assessment is used to determine the coverage of key elements of the target entity and its annotation range.
[0078] According to the verification result of the target entity and its marked range by the verification mechanism, corresponding error tolerance processing is performed on the verification result. For example, when it is determined that its confidence level is lower than the threshold, a dispute mark is added; when there are polysemous entities, multiple marked ranges are retained and marked as alternative entity marked ranges; when the target entity cannot be determined, it enters the manual review channel for forced review of the target entity. Finally, the structured entity information of the target entity and its marked range after successful verification and error tolerance processing is output.
[0079] In this embodiment, the multi-model collaborative decision-making is optimized through a dynamic weight distribution mechanism, significantly improving the extraction accuracy and annotation efficiency of key entities such as time and location in the task text data, and realizing the integration of intelligent recognition and structured annotation of task elements.
[0080] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the entity extraction and annotation method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0081] This application provides an entity extraction and annotation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the entity extraction and annotation method in the first embodiment above.
[0082] Refer to the following Figure 2 , which shows a schematic structural diagram of an entity extraction and annotation device suitable for implementing the embodiments of the present application. The entity extraction and annotation device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 2 The entity extraction and annotation device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0083] As Figure 2 shown, the entity extraction and annotation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the entity extraction and annotation device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the entity extraction and annotation device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an entity extraction and annotation device with various systems, it should be understood that it is not required to implement or have all the systems shown. Instead, more or fewer systems may be implemented or had.
[0084] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0085] The entity extraction and annotation device provided by the present application adopts the entity extraction and annotation method in the above-mentioned embodiment, and can solve the technical problem that the prior art cannot achieve the integrated processing of entity recognition and annotation for the specified canonical text data to be processed. Compared with the prior art, the beneficial effects of the entity extraction and annotation device provided by the present application are the same as those of the entity extraction and annotation method provided by the above-mentioned embodiment, and other technical features in the entity extraction and annotation device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0086] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0087] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0088] The present application provides a storage medium, which is a computer-readable storage medium, and has computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the entity extraction and annotation method in the above-mentioned embodiment.
[0089] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0090] The above computer-readable storage medium can be included in the entity extraction and annotation device; it can also exist separately without being assembled into the entity extraction and annotation device.
[0091] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the entity extraction and annotation device, the entity extraction and annotation device implements the technical content of the entity extraction and annotation method embodiment as shown above.
[0092] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0095] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned entity extraction and annotation method, which can solve the technical problem that the prior art cannot implement integrated processing of entity recognition and annotation for the specified text data to be processed in accordance with the specification. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the entity extraction and annotation method provided in the above embodiments, and will not be elaborated here.
Claims
1. An entity extraction and annotation method, characterized in that, The entity extraction and annotation method includes the following steps: Obtain the text data to be processed; Extract dynamic indicators reflecting the text characteristics from the text data; Determine the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators; Configure a hybrid recognition model according to the combined weights, and identify entities in the text data through the hybrid recognition model; Perform boundary verification and type annotation on the recognition results, and output structured entity information.
2. The entity extraction and annotation method according to claim 1, characterized in that The step of extracting dynamic indicators reflecting the text characteristics from the text data includes: Analyze the distribution law of relevant clause standard paragraphs in the text data, and determine the target area in the standardized professional term set; Evaluate the relevance between the text expression and the field to which the text information belongs according to the distribution law and the target area, and extract the dynamic indicators according to the evaluation results.
3. The entity extraction and annotation method according to claim 1, characterized in that, The step of determining the combined weights of the rule engine, statistical model, and deep learning model based on the dynamic indicators includes: Select the corresponding weight allocation strategy according to the application scenario type of the text data; Allocate decision weights to the rule engine, statistical model, and deep learning module through the weight allocation strategy and the dynamic indicators; Among them, the step of allocating decision weights to the rule engine, statistical model, and deep learning module through the weight allocation strategy and the dynamic indicators includes: Determine the model features corresponding to the dynamic indicators in the pre-established data relationship table; Increase and / or decrease the decision weights in the data model corresponding to the model features according to the weight allocation strategy.
4. The entity extraction and annotation method according to claim 3, wherein The entity extraction and annotation method further includes: Monitor the recognition efficiency of the rule engine, statistical model, and deep learning model in the text data in real time; Dynamically correct the initial weight allocation of the rule engine, statistical model, and deep learning model according to the determination result of the monitoring result and the preset weight adjustment condition.
5. The entity extraction and annotation method according to claim 1, wherein The step of configuring a hybrid recognition model according to the combined weights and identifying entities in the text data through the hybrid recognition model includes: Analyze the output result of the hybrid recognition model, and the output result is a set of entity candidates output in parallel by each data model; Sort the entities in the entity candidate set by priority, and filter out entities below the confidence threshold; Use the filtering result as the recognition result of the entity.
6. The entity extraction and annotation method according to claim 5, wherein The step of sorting the entities in the entity candidate set by priority and filtering out entities below the confidence threshold includes: Sort the entities with high priority by confidence; Adjust the sorting result by using the combined weights, and filter out entities below the confidence threshold in the adjusted sorting result by means of a weighted voting mechanism.
7. The entity extraction and annotation method according to claim 1, characterized in that, The step of performing boundary verification and type annotation on the recognition results and outputting structured entity information includes: Correct the entity boundary position in the recognition result through the punctuation distribution feature in the recognition result; Filter out the recognition results that do not meet the requirements by using a preset entity length threshold to obtain the target entity; Perform type annotation on the target entity and output structured entity information.
8. The entity extraction and annotation method according to claim 7, characterized in that The step of performing type annotation on the target entity and outputting structured entity information includes: Identify the text data through a bidirectional LSTM-CRF model, and predict the entity boundary region of the target entity in the text data; Based on the syntactic analysis tree, correct the position of the entity boundary region to determine the annotation range of the target entity, and perform entity annotation within the annotation range; Execute the step of outputting structured entity information.
9. An entity extraction and annotation device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the entity extraction and annotation method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the entity extraction and annotation method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Named entity identification method and device
CN109710925A
Medical record analysis method and device and medical record analysis system
CN115688787A
Aviation text data labeling method and labeling system thereof
CN116244445A
Online consultation method and system based on gynecological nursing knowledge base
CN118113854A
Text element recognition method and system of adaptive integration technology based on multilevel feature fusion
CN119358545A
Cited By
Knowledge point labeling method and system of natural language processing technology, and electronic equipment
CN121388193A