A Rule-Enhanced Chinese Word Segmentation and Semantic Unit Parsing Method and System

CN122311199BActive Publication Date: 2026-09-01TIANJIN FEIPENG SHENGYUAN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610762741.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-01
Estimated Expiration
2046-05-29

AI Technical Summary

Technical Problem

[0004]本申请提供了一种基于规则增强的中文分词与语义单元解析方法及系统,解决了现有分词方法在推理阶段无法利用下游槽位验证失败信号反向修正分词边界、以及规则库无法从运行错误中自动演化的问题,提高了专业垂直场景下专业术语识别的完整性与跨领域分词结果的自校正能力

Benefits of technology

[0004]本申请提供了一种基于规则增强的中文分词与语义单元解析方法及系统,解决了现有分词方法在推理阶段无法利用下游槽位验证失败信号反向修正分词边界、以及规则库无法从运行错误中自动演化的问题,提高了专业垂直场景下专业术语识别的完整性与跨领域分词结果的自校正能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122311199B_ABST
    Figure CN122311199B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology and discloses a method and system for Chinese word segmentation and semantic unit parsing based on rule enhancement. The method includes: constructing a multi-level domain rule base from forced matching rules, pattern template rules, and context constraint rules to generate a set of rule triples; loading the set of rule triples into an AC automaton to linearly scan the input text to obtain a set of candidate segments and trigger word position indices; constructing a candidate directed acyclic graph, and obtaining a structured word segmentation sequence and a pruned candidate pool through dynamic programming after dynamic modulation and fusion scoring; verifying the integrity of downstream slots, and locating backup candidate edges in the pruned candidate pool using gap signals when failure occurs, and obtaining a corrected word segmentation sequence through renegotiation and scoring. This application improves the completeness of professional terminology recognition in professional vertical scenarios and the self-correction capability of cross-domain word segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and system for Chinese word segmentation and semantic unit parsing based on rule enhancement. Background Technology

[0002] Chinese word segmentation is a fundamental step in the natural language processing pipeline. Existing Chinese word segmentation techniques mainly fall into three categories: dictionary-based segmentation methods using mechanical segmentation rules such as forward maximum matching and backward maximum matching; segmentation methods based on statistical models such as conditional random fields and bidirectional long short-term memory networks; and general-purpose segmentation tools such as Jieba and HanLP that integrate the above two types of methods. These methods have achieved relatively mature engineering implementations in general text scenarios and have obtained high segmentation accuracy on open-domain evaluation sets.

[0003] However, in specialized vertical scenarios such as healthcare, law, and finance, existing word segmentation technologies have the following shortcomings: dictionary-based word segmentation methods lack a confidence quantification mechanism, making it impossible to perform effective disambiguation when terminology boundaries conflict, resulting in fragmented segmentation of specialized terms; while statistical models have a certain context-awareness capability, they suffer from insufficient training data in low-resource specialized domains, leading to a significant drop in word segmentation accuracy when transferring to cross-domain applications; and the user dictionaries of general-purpose tools only support precise string forced matching, with matching results not carrying semantic tags, making it impossible to form data linkage with subsequent semantic parsing tasks, and the rule base relies on manual maintenance, failing to automatically evolve from erroneous cases during operation. Summary of the Invention

[0004] This application provides a rule-enhanced Chinese word segmentation and semantic unit parsing method and system, which solves the problems of existing word segmentation methods being unable to use downstream slot verification failure signals to reverse correct word segmentation boundaries during the inference stage, and the rule base being unable to automatically evolve from runtime errors. It improves the completeness of professional terminology recognition in professional vertical scenarios and the self-correction capability of cross-domain word segmentation results.

[0005] Firstly, this application provides a rule-enhanced Chinese word segmentation and semantic unit parsing method, which includes: Step S1: Construct a multi-level domain rule base from forced matching rules, pattern template rules and context constraint rules, and generate a set of rule triples carrying semantic tags and rule confidence from the multi-level domain rule base; Step S2: Load the set of rule triples into the AC automaton, perform a linear scan on the input text, and obtain the candidate fragment set and the trigger word position index; Step S3: Construct a candidate directed acyclic graph based on the candidate fragment set, dynamically modulate the rule confidence using the trigger word position index, calculate the edge weight for each candidate edge according to the fusion scoring formula, and save the low-scoring candidate edges that are abandoned by the dynamic planning path as a pruning candidate pool to obtain the structured word segmentation sequence. Step S4: The downstream slot integrity constraint is used to perform slot verification on the structured word segmentation sequence. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. The gap signal is used to locate the backup candidate edge of the corresponding range in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence.

[0006] Secondly, this application provides a rule-enhanced Chinese word segmentation and semantic unit parsing system, which includes: The generation module is used to construct a multi-level domain rule base from forced matching rules, pattern template rules and context constraint rules, and generate a set of rule triples carrying semantic tags and rule confidence from the multi-level domain rule base; The loading module is used to load the set of rule triples into the AC automaton, perform linear scanning on the input text, and obtain a set of candidate segments and trigger word position indices. The construction module is used to construct a candidate directed acyclic graph based on the candidate fragment set, dynamically modulate the rule confidence based on the trigger word position index, calculate the edge weight for each candidate edge according to the fusion scoring formula, save the low-scoring candidate edges that are abandoned by the dynamic planning path as a pruning candidate pool, and obtain the structured word segmentation sequence. The verification module is used to perform slot verification on the structured word segmentation sequence under the downstream slot integrity constraint. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. The gap signal is used to locate the backup candidate edge of the corresponding range in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence.

[0007] The technical solution provided in this application constructs a multi-level domain rule base consisting of forced matching rules, pattern template rules, and context constraint rules. Each rule is stored as a triple carrying semantic labels and rule confidence, thereby transforming rule knowledge into structured data that can participate in quantitative calculations. Based on this, the set of rule triples is loaded into the AC automaton to complete the scanning of candidate segments of the input text in linear time complexity, and trigger word position indexes are constructed simultaneously. This allows the confidence increment of the context constraint rules to be accurately superimposed on the corresponding candidate edges in the subsequent scoring stage, realizing the contextual dynamic modulation of rule confidence. After calculating the edge weight of each candidate edge in the candidate directed acyclic graph using the fusion scoring formula, the dynamic programming algorithm selects the word segmentation path with the best cumulative score in the global path space. At the same time, the low-scoring candidate edges that are abandoned are completely retained in the pruning candidate pool, so that the backup paths pruned by dynamic programming can still be retrieved and used in subsequent stages, rather than being irreversibly discarded.

[0008] Based on the above word segmentation results, this application further introduces downstream slot integrity constraints to perform slot verification on the structured word segmentation sequence. When verification fails, a gap signal carrying the gap character interval and gap slot type is generated. The gap signal is used to accurately locate the backup candidate edge of the corresponding interval in the pruning candidate pool. By introducing slot feasibility scores, the backup candidate edge is re-negotiated and scored, so that the correction criterion of word segmentation boundary is extended from simple language model scoring to downstream task structure constraint scoring. This mechanism makes the cross-task failure signal in the inference stage an effective input for triggering local correction of word segmentation boundary for the first time. The correction range is strictly limited to the gap character interval and does not affect the other determined word segmentation boundaries, thus balancing the accuracy of correction and computational efficiency. At the same time, successful amendment examples are automatically written into the rule base after being triggered by a threshold, so that the rule base can be continuously expanded without human intervention, fundamentally breaking the dual dilemma of silent propagation of word segmentation errors and static solidification of the rule base in the existing technology. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of an embodiment of the rule-enhanced Chinese word segmentation and semantic unit parsing method in this application. Figure 2 This is a simulation diagram illustrating the relationship between the fusion scoring parameter space and the slot filling accuracy in an embodiment of this application. Detailed Implementation

[0011] This application provides a rule-enhanced Chinese word segmentation and semantic unit parsing method and system. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0012] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the rule-enhanced Chinese word segmentation and semantic unit parsing method in this application includes: Step S1: Construct a multi-level domain rule base from forced matching rules, pattern template rules, and context constraint rules; generate a set of rule triples carrying semantic tags and rule confidence from the multi-level domain rule base. Specifically, the confidence values ​​of the three types of rules in the multi-level domain rule base are stratified according to the ambiguity risk of each type of rule: the confidence value of the forced matching rule ranges from 0.95 to 1.0 because the term string is a precise character sequence with minimal ambiguity; the confidence value of the pattern template rule ranges from 0.80 to 0.90 because regular expression templates have a probability of cross-context mismatch; and the confidence increment of the context constraint rule ranges from 0.10 to 0.30. It exists as an additive item rather than an independent confidence value, and its upper limit of 0.30 is set to avoid context constraints dominating the disambiguation result alone. The confidence fields of the three types of rules have a unified dimension, all being real numbers in the range of 0 to 1, and directly participate in the weighted calculation in the fusion scoring formula in step S3 without additional normalization processing.

[0013] Step S2: Load the set of rule triples into the AC automaton, perform a linear scan on the input text, and obtain the set of candidate segments and the trigger word position index; Specifically, in the fusion scoring formula, the rule weight coefficient is set to 0.7, and the slot feasibility weight coefficient is set to 0.4. Both are configurable hyperparameters rather than fixed constants. The rule weight coefficient ranges from 0.5 to 0.9, and the slot feasibility weight coefficient ranges from 0.3 to 0.5. The parameters are determined as follows: the rule weight coefficient is incremented by 0.1 within the range of 0.5 to 0.9, and the slot feasibility weight coefficient is incremented by 0.1 within the range of 0.3 to 0.5. The downstream slot filling accuracy is used as the evaluation index. The evaluation index is calculated group by group on labeled text of the same type as the target domain. The parameter combination corresponding to the highest evaluation index is selected as the final value. This process is exhaustive verification and does not involve model training, which is a conventional experimental method in this field.

[0014] Step S3: Construct a candidate directed acyclic graph based on the candidate fragment set, dynamically modulate the rule confidence using the trigger word position index, calculate the edge weight for each candidate edge according to the fusion scoring formula, and save the low-scoring candidate edges that are abandoned by the dynamic planning path as a pruning candidate pool to obtain the structured word segmentation sequence. Specifically, the data structure of the pruned candidate pool is a dictionary nested list: the key of the dictionary is a tuple of character intervals consisting of the start and end character positions, and the value of the dictionary is a list of candidate segments sorted in descending order of edge weights. Each element in the list contains three fields: candidate segment string, semantic label, and edge weight. During slot verification in step S4, the corresponding candidate segment list in the pruned candidate pool is directly retrieved using the gap character interval tuple in the gap signal as the key, eliminating the need to re-execute the AC automaton scan; the retrieval complexity is constant.

[0015] Step S4: The downstream slot integrity constraint is used to perform slot verification on the structured word segmentation sequence. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. The gap signal is used to locate the backup candidate edge of the corresponding range in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence.

[0016] Specifically, the slot verification check range is the eight consecutive character intervals after the anchor term. Eight characters is an empirical value, determined based on the average character length of structured expressions such as medical orders in the medical field and legal citations in the legal field. It can be adjusted according to the specific field during implementation. When the number of remaining characters at the end of the input text is less than eight, the text termination position is used as the check boundary. When the termination position of the selected backup candidate edge after renegotiation is inconsistent with the starting position of the adjacent term, the termination position of the backup candidate edge is used as the new starting point and the text termination position is used as the ending point. The dynamic programming recursion of step S3 is re-executed within the remaining character interval. The cumulative score of the new starting point node is initialized as the sum of the renegotiation score of the backup candidate edge and the cumulative score of its predecessor path. The local optimal path within this interval is obtained by backtracking and concatenating it with the previously determined word segmentation result of the gap character interval to obtain the complete corrected word segmentation sequence.

[0017] In one specific embodiment, step S1 includes: High-frequency professional terms in the professional field are associated with corresponding semantic tags. The longest priority principle is used as the matching constraint. Each term is assigned a rule confidence value ranging from 0.95 to 1.0 to obtain a set of forced matching rules. The data structure of each rule in the forced matching rule set is a triple consisting of a term string, a semantic tag, and a rule confidence value. The regular expression templates for structured information such as dosage, date, and number are associated and labeled with corresponding semantic tags. Each template is configured with a rule confidence level ranging from 0.80 to 0.90 to obtain a pattern template rule set. The data structure of each rule in the pattern template rule set is a triple consisting of a regular expression template, a semantic tag, and a rule confidence level. The contextual trigger words in the domain are associated with their target semantic label types. A confidence increment ranging from 0.10 to 0.30 is configured for each trigger word to obtain a set of contextual constraint rules. The data structure of each rule in the set of contextual constraint rules is a triple consisting of a trigger word, a target semantic label, and a confidence increment. The forced matching rule set, pattern template rule set, and context constraint rule set are organized according to hierarchical index to obtain a multi-level domain rule base, from which a rule triple set is derived.

[0018] Specifically, the term strings in the forced matching rule set are precise character sequences of high-frequency professional terms, the semantic tags are predefined semantic category identifiers, and the rule confidence values ​​range from 0.95 to 1.0, manually determined by domain experts based on the ambiguity rate of the terms in professional literature; the lower the ambiguity rate, the higher the confidence. The longest-priority principle is implemented as follows: when a character range in the input text matches multiple term strings of different lengths simultaneously, the matching result with the largest number of characters covered is retained, and the shorter matching result completely covered by it is discarded. When two matching results cover the same number of characters, the rule with the higher confidence is retained.

[0019] The regular expression templates in the pattern template rule set are regular expression strings describing structured text patterns. For dosage information, the regular expression template matches character combinations of consecutive numbers followed by mass units; for date information, it matches character combinations of a four-digit year followed by one or two digits for the month, and then one or two digits for the day; and for number information, it matches character combinations of one or three uppercase letters followed by five to seven consecutive numbers. The regular expression templates are scanned independently of the AC automaton index structure of the forced matching rule set, and the candidate fragments generated by both are uniformly subjected to overlap resolution in subsequent steps.

[0020] In the context constraint rule set, trigger words are words with clear contextual references in professional texts, target semantic labels are the semantic categories of subsequent words pointed to by the trigger words, and confidence increments are corrections superimposed on the confidence of the candidate segment's base rules, ranging from 0.10 to 0.30. The upper limit of the confidence increment is set to 0.30 because when the base rule confidence is at its lowest value of 0.80, the result after adding the maximum increment of 0.30 is 1.10. Even after truncating to 1.0, it still does not exceed the upper bound of the confidence, thus avoiding the context constraint alone dominating the disambiguation result. The three rule sets are indexed in a hierarchical order: forced matching rule set, pattern template rule set, and context constraint rule set. The purpose of the hierarchical index is to load the rules into the corresponding processing modules according to type in subsequent steps, resulting in a multi-level domain rule library. All triple entries are then derived from the multi-level domain rule library to obtain a rule triple set.

[0021] In one specific embodiment, step S2 loads the set of rule triples into the AC automaton, performs a linear scan on the input text, and obtains a set of candidate segments and trigger word position indices, including: Insert each term string from the set of rule triples that force matching the set of rule sets into the trie structure of the AC automaton, calculate the mismatch pointer on the trie structure, and obtain the index structure of the AC automaton. The regular expression templates of the pattern template rule set in the rule triple set are separated from the AC automaton index structure. The regular expression scan is performed on each line of the input text to obtain a set of regular expression matching fragments carrying character ranges and semantic labels. Based on the AC automaton index structure, a single linear scan is performed on the input text to obtain a set of term matching segments carrying character ranges, semantic labels, and rule confidence. The term matching segment set and the regular expression matching segment set are overlapped and eliminated according to the longest priority principle to obtain a set of candidate segments. Perform string location scanning on the input text using all trigger words in the context constraint rule set of the rule triple set. Record the start and end character positions of each trigger word, the corresponding target semantic label, and the confidence increment into a hash table to obtain the trigger word position index.

[0022] Specifically, the construction process of the trie structure of the AC automaton is as follows: all term strings in the forced matching rule set are inserted into the trie character by character, with each character corresponding to a node in the trie. Term strings sharing the same prefix share the prefix path in the trie, and the terminating node of a term string carries the corresponding semantic label and rule confidence. The mismatch pointer is calculated as follows: for each non-root node in the trie, the corresponding node in the trie for its longest proper suffix is ​​found, and the mismatch pointer is set to that node. The calculation of the mismatch pointer is completed by breadth-first traversal of the trie. After the mismatch pointer is calculated, when the AC automaton index structure performs a single linear scan of the input text, if the current character fails to match, it jumps along the mismatch pointer instead of backtracking to the beginning of the text, thus ensuring that the scanning time complexity is a linear function of the length of the input text and is independent of the number of term string entries in the rule base.

[0023] The overlap resolution rules for the regular expression matching fragment set and the term matching fragment set are as follows: when the character intervals of two candidate fragments overlap arbitrarily, the candidate fragment with more covered characters is retained; when the two candidate fragments cover the same number of characters, the candidate fragment with higher rule confidence is retained; when both the number of covered characters and the rule confidence are the same, candidate fragments from the forced matching rule set are retained, and candidate fragments from the pattern template rule set have lower priority than those from the forced matching rule set. After overlap resolution, there are no entries in the candidate fragment set with completely overlapping character intervals, but different candidate fragments are allowed to share some character positions. Candidate fragments sharing some character positions form mutually exclusive competing paths when constructing the candidate directed acyclic graph in subsequent step S3.

[0024] The hash table for trigger word position indexes uses the starting character position of the trigger word as the key and the combination of the target semantic label and the confidence increment as the value. When multiple trigger words exist at the same starting character position, multiple combinations of target semantic labels and confidence increments are stored in list form. The string positioning scan of the trigger word performs a complete traversal of the input text, recording all occurrence positions of each trigger word, not just the first occurrence position. When the same trigger word appears multiple times in different positions in the input text, the start and end character positions of each occurrence are independently recorded in the hash table. During the context dynamic modulation in step S3, the hash table is queried according to the starting character position of each candidate edge, and the entries in the query results where the target semantic label and the candidate edge semantic label are consistent are used to perform confidence increment superposition.

[0025] In one specific embodiment, step S3 involves constructing a candidate directed acyclic graph based on the candidate fragment set, and dynamically modulating the rule confidence using the trigger word position index, including: Using all character positions of the input text as the node set, each candidate segment in the candidate segment set is mapped to a directed edge from the start character position to the end character position. Each directed edge carries the semantic label and rule confidence of the corresponding candidate segment, resulting in a candidate directed acyclic graph. Based on the trigger word position index, dynamic context modulation is performed on each directed edge in the candidate directed acyclic graph: query the trigger word records in the five character intervals before the starting character position of each directed edge. When the target semantic label of the trigger word is consistent with the semantic label of the directed edge, the corresponding confidence increment is added to the rule confidence of the directed edge. When the added result exceeds 1.0, it is truncated to 1.0 to obtain the modulated rule confidence.

[0026] Specifically, the node set of the candidate directed acyclic graph covers all character positions of the input text, including the start and end positions, and the number of nodes is equal to the number of characters in the input text plus one. For character intervals not covered by the forced matching rule set or the pattern template rule set, the basic word segmentation tool is called to perform a fallback segmentation on the interval. The fallback segmentation result is also mapped as directed edges and added to the candidate directed acyclic graph. The semantic labels of the directed edges generated by the fallback segmentation are recorded as non-entity labels, and the rule confidence is recorded as zero, thereby ensuring that there is at least one complete path from the start node to the end node of the text in the candidate directed acyclic graph, without path breaks.

[0027] In context-based dynamic modulation, the query range is the five-character interval preceding the starting character position of each directed edge. These five characters are an empirical value determined based on the average number of characters between trigger words and target terms in professional text, and can be adjusted according to the specific domain. When the starting character position of a directed edge is less than five, the query range is bounded to the left by the text's starting position. The condition for confidence increment stacking is that the target semantic label of the trigger word and the semantic label of the directed edge must be strictly consistent. Semantic label consistency is determined by exact string matching, not fuzzy matching. When multiple trigger word records satisfying the conditions exist within the same query range, the confidence increments of all satisfying conditions are accumulated and then stacked onto the rule confidence of the directed edge. If the stacked result exceeds 1.0, it is truncated to 1.0. The truncation operation is performed once after all increments are accumulated, rather than truncating each increment individually.

[0028] In one specific embodiment, step S3 calculates the edge weight for each candidate edge according to the fusion scoring formula, including: The candidate fragments corresponding to each directed edge in the candidate directed acyclic graph are input into the lightweight semantic model. After performing basic word segmentation on the character content of the candidate fragments, the weights of each sub-word are accumulated and normalized using the keyword weight table to obtain the semantic score. Based on the modulated rule confidence and semantic score, the edge weight of each directed edge is calculated according to the fusion scoring formula. The fusion scoring formula is: the edge weight is equal to the product of the rule weight coefficient and the modulated rule confidence, plus the product of the difference between the rule weight coefficient and one and the semantic score, where the rule weight coefficient is 0.7, thus obtaining the candidate edge weight set.

[0029] Specifically, the lightweight semantic model is executed as follows: The candidate segment character content corresponding to each directed edge is input into a basic word segmentation tool to obtain a sub-word list consisting of one or more sub-words. For each sub-word in the sub-word list, a keyword weight table is queried. The keyword weight table is a pre-constructed dictionary structure, where the key is a domain keyword string, and the value is the importance weight of that keyword in the domain text, ranging from 0 to 1. Sub-words not in the keyword weight table have a weight of zero. The query weights of all sub-words in the sub-word list are summed and divided by the product of the sub-word list length and the maximum weight in the keyword weight table to obtain the normalized semantic score. When the normalization result exceeds 1.0, it is truncated to 1.0. The keyword weight table is manually calibrated by domain experts based on the word frequency and discriminative power of keywords in professional literature. It represents domain knowledge input at the same level as the multi-level domain rule base. Those skilled in the art can complete the semantic score calculation by following the above steps; it does not involve neural network training.

[0030] In the fusion scoring formula, the rule weight coefficient is set to 0.7. This means that the modulated rule confidence accounts for 70% of the weight in the edge weight calculation, and the semantic score accounts for the remaining 30%, with the sum of their weights being 1.0. The reason for setting the rule weight coefficient to 0.7 is that in professional vertical scenarios, rule confidence has higher reliability than semantic score, and therefore it is given higher weight. The adjustable range of the rule weight coefficient is 0.5 to 0.9, with a step size of 0.1. It is verified group by group on labeled text of the same type as the target domain, using the downstream slot filling accuracy as the evaluation index. The value corresponding to the highest evaluation index is selected as the final rule weight coefficient. This process is exhaustive verification and does not involve model training. When a directed edge comes from the bottom-line split and the modulated rule confidence is zero, the edge weight is equal to the product of 0.3 and the semantic score. The range of edge weight values ​​is necessarily lower than that of directed edges with rule support, thus putting it at a disadvantage in dynamic programming path selection.

[0031] Figure 2 This is a simulation diagram illustrating the relationship between the fusion scoring parameter space and the slot filling accuracy in an embodiment of this application. Figure 2 This diagram illustrates the three-dimensional response surface formed by the slot filling accuracy when the rule weighting coefficient α and the slot feasibility weighting coefficient β change jointly within their respective ranges. α ranges along the horizontal axis in steps of 0.05 from 0.50 to 0.90, and β ranges along the vertical axis in steps of 0.025 from 0.30 to 0.50. The surface height corresponds to the slot filling accuracy under this parameter combination, and the color depth corresponds one-to-one with the accuracy value; lighter colors indicate higher accuracy. Figure 2 It can be seen that when α is 0.70 and β is 0.40, the surface reaches the peak region at the corresponding position, which is consistent with the parameter selection of 0.7 for the rule weight coefficient and 0.4 for the slot feasibility weight coefficient.

[0032] In one specific embodiment, step S3 involves saving the low-scoring candidate edges that are abandoned by the dynamic programming path selection as a pruning candidate pool to obtain a structured word segmentation sequence, including: Using the candidate edge weight set as input, perform dynamic programming to solve the global path in the candidate directed acyclic graph: starting with the cumulative score of the text start node initialized to zero, traverse the nodes one by one in ascending order of their positions, select the incoming edge that maximizes the cumulative score and its predecessor node for each node, and obtain the optimal path cumulative score array and predecessor node record array. The global optimal word segmentation path is obtained by backtracking from the text termination node to the start node using the predecessor node record array; semantic tags, rule confidence and rule source fields are added to each word in the global optimal word segmentation path to obtain a structured word segmentation sequence; All directed edges in the candidate directed acyclic graph that are not selected by the globally optimal word segmentation path are recorded in an ordered list using the character range as the index key and the candidate fragment and its edge weight as the index value, thus obtaining the pruning candidate pool.

[0033] Specifically, the execution process of dynamic programming for global path finding is as follows: The length of the optimal path cumulative score array is equal to the number of characters in the input text plus one. Each element in the array corresponds to the current optimal cumulative score of a node position. Initially, the element corresponding to the starting node of the text has a value of zero, and the elements corresponding to other nodes have a value of negative infinity. The predecessor node record array has the same length as the optimal path cumulative score array, and initially all elements are recorded as null values. During traversal, for each node, all directed edges in the candidate directed acyclic graph that terminate at that node are searched. For each directed edge, the sum of the cumulative score of its starting node and the edge weight is calculated. The largest sum among all directed edges is selected and written into the element corresponding to the current node in the optimal path cumulative score array. The starting node position of the corresponding directed edge is written into the element corresponding to the current node in the predecessor node record array. When a node has no incoming edges, its cumulative score remains negative infinity, and the node is not visited during backtracking.

[0034] In the structured word segmentation sequence, the rule source field of each term records the rule entry identifier that generates the directed edge corresponding to that term. The directed edge generated by the forced matching rule set records the rule number of the corresponding term string. The directed edge generated by the pattern template rule set records the rule number of the corresponding regular expression template. The rule source field of the directed edge generated by the fallback segmentation records the fallback identifier. The ordered list in the pruning candidate pool uses character range tuples as keys and candidate fragment lists as values. The candidate fragment list is sorted in descending order of edge weight. When there are multiple unselected directed edges in the same character range, all directed edges are recorded in the candidate fragment list of the corresponding character range. No unselected directed edges are discarded, thus ensuring that all backup candidate fragments in the character range can be retrieved when renegotiation is performed in step S4.

[0035] In one specific embodiment, in step S4, the downstream slot integrity constraint is used to perform slot verification on the structured word segmentation sequence. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. The gap signal is used to locate the corresponding interval's backup candidate edge in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence, including: Slot verification is performed on the structured word segmentation sequence based on the downstream slot integrity constraint: the correspondence between semantic tags and anchor words in the structured word segmentation sequence is used as the verification benchmark. The semantic tags of the target slot are checked within the eight consecutive character intervals after the anchor words. When the semantic tags of the target slot are missing, the character intervals of the missing position and the missing slot type are encapsulated into a gap signal. The candidate edges are retrieved from the pruning candidate pool based on the gap character interval in the gap signal. The candidate edges whose semantic labels are consistent with the gap slot type are selected and sorted in descending order of edge weight to obtain the candidate edge list. The slot feasibility score is introduced into the renegotiation scoring formula, and the renegotiation score is recalculated for each alternative candidate edge in the alternative candidate edge list: the renegotiation score is equal to the sum of the edge weight obtained from the original fusion scoring formula, the product of the slot feasibility weight coefficient and the slot feasibility score, where the slot feasibility weight coefficient is 0.4, and the slot feasibility score is 1 when the semantic label of the alternative candidate edge is consistent with the gap slot type, and 0 otherwise. The alternative candidate edge with the highest renegotiation score is selected to replace the original words in the gap character interval in the structured word segmentation sequence, and the corrected word segmentation sequence is obtained.

[0036] Specifically, the definition of downstream slot integrity constraints is as follows: Slot combination rules are pre-defined for structured information units in the target domain. Each slot combination rule includes an anchor semantic tag and one or more target slot semantic tags. The anchor semantic tag serves as the starting point identifier for triggering slot verification, and the target slot semantic tag is the semantic category that must appear within a consecutive eight-character interval after the anchor term. Taking the medical field as an example, the anchor semantic tag is "drug," and the target slot semantic tag is "dosage," which together constitute a slot combination rule. The slot verification execution process is as follows: Traverse the structured word segmentation sequence. When the semantic tag of a term matches the anchor semantic tag, record the position of the term's ending character as the starting point for inspection. Traverse subsequent terms in the structured word segmentation sequence within the character interval from the starting point to the starting point plus eight characters, checking for terms whose semantic tags match the target slot semantic tags. When the number of remaining characters at the end of the inspection interval is less than eight, use the text termination position as the inspection boundary. When the inspection result is negative, encapsulate the character interval from the starting point to the inspection boundary and the target slot semantic tag into a gap signal.

[0037] In the gap signal, the gap character interval is the start and end character position tuple of the check range in step S4 slot verification, and the gap slot type is the semantic label string of the target slot where the corresponding term could not be found within the check range. Using the gap character interval tuple as the key, an exact key matching search is performed in the pruning candidate pool. When a key that completely matches the gap character interval exists in the pruning candidate pool, the corresponding candidate fragment list is retrieved; when no completely matching key exists, all keys whose start and end character positions are completely contained within the gap character interval are retrieved, and the corresponding candidate fragment lists are merged and rearranged in descending order of edge weight. The filtering condition is an exact match between the semantic label of the candidate fragment and the gap slot type string, resulting in a list of backup candidate edges.

[0038] In the renegotiation scoring formula, the slot feasibility score is a binary variable, taking the value of one or zero. Since all entries in the alternative candidate edge list have been filtered for consistency between semantic labels and gap slot types, the slot feasibility score of all alternative candidate edges in the list is one. Therefore, the renegotiation score is equal to the edge weight obtained from the original fusion scoring formula plus the slot feasibility weight coefficient of 0.4. After the alternative candidate edges in the list are sorted in descending order of renegotiation score, the first entry is taken as the replacement term. When the terminating character position of a candidate edge is inconsistent with the starting character position of an adjacent word in the structured word segmentation sequence, the dynamic programming recursion of step S3 is re-executed within the remaining character interval, with the terminating character position of the candidate edge as the new starting point and the text termination position as the ending point. The cumulative score of the new starting point node is initialized as the sum of the renegotiation score of the candidate edge and the value of the corresponding element at the starting character position of the candidate edge in the optimal path cumulative score array. The optimal path cumulative score array is an array generated and retained during the dynamic programming solution in step S3, and the corresponding element can be directly read without recalculation. After backtracking to obtain the local optimal path, it is concatenated with the word segmentation result previously determined in the gap character interval to obtain the complete corrected word segmentation sequence. When there are no candidate edges in the pruning candidate pool that meet the screening conditions, the words in the original structured word segmentation sequence are retained in the character interval corresponding to the gap signal without replacement, and the corrected word segmentation sequence is consistent with the structured word segmentation sequence in this interval.

[0039] The rule-enhanced Chinese word segmentation and semantic unit parsing method in the embodiments of this application has been described above. The rule-enhanced Chinese word segmentation and semantic unit parsing system in the embodiments of this application is described below. One embodiment of the rule-enhanced Chinese word segmentation and semantic unit parsing system in the embodiments of this application includes: The generation module is used to construct a multi-level domain rule base from forced matching rules, pattern template rules and context constraint rules, and generate a set of rule triples carrying semantic tags and rule confidence from the multi-level domain rule base; The loading module is used to load the set of rule triples into the AC automaton, perform linear scanning on the input text, and obtain a set of candidate segments and trigger word position indices. The construction module is used to construct a candidate directed acyclic graph based on the candidate fragment set, dynamically modulate the rule confidence based on the trigger word position index, calculate the edge weight for each candidate edge according to the fusion scoring formula, save the low-scoring candidate edges that are abandoned by the dynamic planning path as a pruning candidate pool, and obtain the structured word segmentation sequence. The verification module is used to perform slot verification on the structured word segmentation sequence under the downstream slot integrity constraint. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. The gap signal is used to locate the backup candidate edge of the corresponding range in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence.

[0040] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A rule-enhanced Chinese word segmentation and semantic unit parsing method, characterized in that, The method includes: Step S1: Construct a multi-level domain rule base using forced matching rules, pattern template rules, and context constraint rules. Generate a set of rule triples carrying semantic tags and rule confidence levels from this multi-level domain rule base. This includes: associating high-frequency professional terms within the professional domain with their corresponding semantic tags, using the longest priority principle as the matching constraint, and configuring a rule confidence level ranging from 0.95 to 1.0 for each term, thus obtaining a forced matching rule set. The data structure of each rule in the forced matching rule set is a triple consisting of a term string, a semantic tag, and a rule confidence level; associating regular expression templates for structured information such as dosage, date, and number with their corresponding semantic tags, and configuring a value ranging from 0.80 to 0.90 for each template. The rule confidence is used to obtain a pattern template rule set. Each rule in the pattern template rule set has a data structure consisting of a triplet composed of a regular expression template, a semantic label, and a rule confidence. Contextual trigger words within the domain are associated with their target semantic label types. A confidence increment ranging from 0.10 to 0.30 is configured for each trigger word to obtain a context constraint rule set. Each rule in the context constraint rule set has a data structure consisting of a triplet composed of a trigger word, a target semantic label, and a confidence increment. The forced matching rule set, the pattern template rule set, and the context constraint rule set are organized according to a hierarchical index to obtain the multi-level domain rule library. The rule triplet set is derived from the multi-level domain rule library. Step S2: Load the set of rule triples into the AC automaton, perform a linear scan on the input text, and obtain the candidate fragment set and the trigger word position index; Step S3: Construct a candidate directed acyclic graph based on the candidate fragment set, and dynamically modulate the rule confidence using the trigger word position index. Here, the set of all character positions of the input text is used as the node set, and each candidate fragment in the candidate fragment set is mapped to a directed edge from the start character position to the end character position. Each directed edge carries the semantic label and rule confidence of the corresponding candidate fragment, thus obtaining the candidate directed acyclic graph. Based on the trigger word position index, context dynamic modulation is performed on each directed edge in the candidate directed acyclic graph: query the trigger word records in the five character intervals before the starting character position of each directed edge; when the target semantic label of the trigger word is consistent with the semantic label of the directed edge, the corresponding confidence increment is added to the rule confidence of the directed edge; when the addition result exceeds 1.0, it is truncated to 1.0 to obtain the modulated rule confidence. The edge weight is calculated for each candidate edge according to the fusion scoring formula. The low-scoring candidate edges that are abandoned by the dynamic programming path are saved as a pruning candidate pool to obtain the structured word segmentation sequence. The fusion scoring formula is: the edge weight is equal to the product of the rule weight coefficient and the modulated rule confidence, plus the product of the difference between the rule weight coefficient and one and the semantic score, where the rule weight coefficient is 0.7, and the candidate edge weight set is obtained. Step S4: Perform slot verification on the structured word segmentation sequence based on downstream slot integrity constraints. When verification fails, generate a gap signal carrying the gap character range and gap slot type. Specifically, the slot verification on the structured word segmentation sequence is performed based on the downstream slot integrity constraints: using the correspondence between semantic tags and anchor words in the structured word segmentation sequence as the verification benchmark, check whether the target slot semantic tag exists in the eight consecutive character ranges after the anchor word. When the target slot semantic tag is missing, encapsulate the character range at the missing position and the missing slot type into a gap signal. The gap signal is used to locate the backup candidate edge of the corresponding interval in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence. The slot feasibility score is one when the semantic label of the backup candidate edge is consistent with the gap slot type, and zero otherwise.

2. The rule-enhanced Chinese word segmentation and semantic unit parsing method according to claim 1, characterized in that, In step S2, the set of rule triples is loaded into the AC automaton, and the input text is linearly scanned to obtain a set of candidate segments and trigger word position indices, including: Insert each term string from the set of rule triples that force the matching rule set into the trie structure of the AC automaton, calculate the mismatch pointer on the trie structure, and obtain the index structure of the AC automaton. The regular expression templates of the pattern template rule set in the rule triple set are separated from the AC automaton index structure, and regular expression scanning is performed on each line of the input text to obtain a set of regular expression matching fragments carrying character ranges and semantic tags. Based on the AC automaton index structure, a single linear scan is performed on the input text to obtain a set of term matching segments carrying character ranges, semantic labels, and rule confidence; the set of term matching segments and the set of regular expression matching segments are overlapped and eliminated according to the longest priority principle to obtain the set of candidate segments. Perform string location scanning on the input text using all trigger words in the context constraint rule set of the rule triple set, and record the start and end character positions of each trigger word, the corresponding target semantic label, and the confidence increment in a hash table to obtain the trigger word position index.

3. The rule-enhanced Chinese word segmentation and semantic unit parsing method according to claim 1, characterized in that, Step S3 involves calculating the edge weight for each candidate edge according to the fusion scoring formula, including: The candidate fragments corresponding to each directed edge in the candidate directed acyclic graph are input into a lightweight semantic model. After performing basic word segmentation on the character content of the candidate fragments, the weights of each sub-word are accumulated and normalized using a keyword weight table to obtain a semantic score. The specific execution process of the lightweight semantic model is as follows: The character content of the candidate fragments corresponding to each directed edge is input into a basic word segmentation tool to obtain a sub-word list consisting of one or more sub-words; For each sub-word in the sub-word list, the keyword weight table is queried. The keyword weight table is a pre-constructed dictionary structure, where the key is a domain keyword string and the value is the importance weight of the keyword in the domain text, ranging from 0 to 1. The weight of sub-words not in the keyword weight table is recorded as zero; The query weights of all sub-words in the sub-word list are accumulated and divided by the product of the length of the sub-word list and the maximum weight in the keyword weight table to obtain the normalized semantic score. When the normalization result exceeds 1.0, it is truncated to 1.

0. Based on the modulated rule confidence and the semantic score, the edge weight of each directed edge is calculated according to the fusion scoring formula.

4. The rule-enhanced Chinese word segmentation and semantic unit parsing method according to claim 3, characterized in that, In step S3, the low-scoring candidate edges that are abandoned by the dynamic programming path selection are saved as a pruning candidate pool to obtain a structured word segmentation sequence, including: Using the candidate edge weight set as input, perform dynamic programming global path solving on the candidate directed acyclic graph: starting with the cumulative score of the text start node initialized to zero, traverse the nodes one by one from smallest to largest position, select the incoming edge that maximizes the cumulative score and its predecessor node for each node, and obtain the optimal path cumulative score array and predecessor node record array. The global optimal word segmentation path is obtained by backtracking from the text termination node to the start node using the predecessor node record array; semantic tags, rule confidence, and rule source fields are added to each word in the global optimal word segmentation path to obtain the structured word segmentation sequence; All directed edges in the candidate directed acyclic graph that are not selected by the globally optimal word segmentation path are recorded in an ordered list using character ranges as index keys and candidate segments and their edge weights as index values ​​to obtain the pruning candidate pool.

5. The Chinese word segmentation and semantic unit parsing method based on rule enhancement according to claim 1, characterized in that, In step S4, the gap signal is used to locate the backup candidate edge of the corresponding interval in the pruning candidate pool. A slot feasibility score is introduced to re-negotiate and score the backup candidate edge, resulting in a corrected word segmentation sequence, including: The gap character range in the gap signal is used to retrieve all alternative candidate edges in the corresponding range in the pruning candidate pool. Alternative candidate edges whose semantic labels are consistent with the gap slot type are filtered out and sorted in descending order of edge weight to obtain a list of alternative candidate edges. The slot feasibility score is incorporated into the renegotiation scoring formula, and the renegotiation score is recalculated for each backup candidate edge in the backup candidate edge list: the renegotiation score is equal to the sum of the edge weight obtained from the original fusion scoring formula, the product of the slot feasibility weight coefficient and the slot feasibility score, where the slot feasibility weight coefficient is 0.4; the backup candidate edge with the highest renegotiation score is selected to replace the original word in the gap character interval of the structured word segmentation sequence, thus obtaining the corrected word segmentation sequence.

6. A rule-enhanced Chinese word segmentation and semantic unit parsing system, characterized in that, For implementing the rule-enhanced Chinese word segmentation and semantic unit parsing method as described in any one of claims 1-5, the rule-enhanced Chinese word segmentation and semantic unit parsing system comprises: The generation module is used to construct a multi-level domain rule base from forced matching rules, pattern template rules, and context constraint rules. It generates a set of rule triples carrying semantic tags and rule confidence levels from this multi-level domain rule base. This includes: associating high-frequency professional terms within the professional domain with their corresponding semantic tags, using the longest-first-come-last-serves principle as the matching constraint, and configuring a rule confidence level ranging from 0.95 to 1.0 for each term, resulting in a forced matching rule set. The data structure of each rule in the forced matching rule set is a triple consisting of a term string, a semantic tag, and a rule confidence level; and associating regular expression templates for structured information such as dosage, date, and number with their corresponding semantic tags, configuring a value ranging from 0.80 to 0.9 for each template. A rule confidence score of 0 is used to obtain a pattern template rule set. Each rule in the pattern template rule set has a data structure consisting of a triplet composed of a regular expression template, a semantic label, and a rule confidence score. Contextual trigger words within the domain are associated with their target semantic label types. A confidence increment ranging from 0.10 to 0.30 is configured for each trigger word to obtain a context constraint rule set. Each rule in the context constraint rule set has a data structure consisting of a triplet composed of a trigger word, a target semantic label, and a confidence increment. The forced matching rule set, the pattern template rule set, and the context constraint rule set are organized according to a hierarchical index to obtain the multi-level domain rule library. The rule triplet set is derived from the multi-level domain rule library. The loading module is used to load the set of rule triples into the AC automaton, perform linear scanning on the input text, and obtain a set of candidate segments and trigger word position indices. A construction module is used to construct a candidate directed acyclic graph (DAG) based on the candidate fragment set, and to dynamically modulate the rule confidence based on the trigger word position index. Specifically, the entire set of characters in the input text is used as the node set, and each candidate fragment in the candidate fragment set is mapped to a directed edge pointing from the starting character position to the ending character position. Each directed edge carries the semantic label and rule confidence of the corresponding candidate fragment, resulting in the candidate DAG. Based on the trigger word position index, dynamic context modulation is performed on each directed edge in the candidate DAG: the trigger word records within the five character intervals before the starting character position of each directed edge are queried. When the target semantic label of the trigger word matches the semantic label of the directed edge, the corresponding confidence increment is added to the rule confidence of the directed edge. If the added result exceeds 1.0, it is truncated to 1.0 to obtain the modulated rule confidence. The edge weight is calculated for each candidate edge according to the fusion scoring formula. The low-scoring candidate edges that are abandoned by the dynamic programming path are saved as a pruning candidate pool to obtain the structured word segmentation sequence. The fusion scoring formula is: the edge weight is equal to the product of the rule weight coefficient and the modulated rule confidence, plus the product of the difference between the rule weight coefficient and one and the semantic score, where the rule weight coefficient is 0.7, and the candidate edge weight set is obtained. The verification module is used to perform slot verification on the structured word segmentation sequence based on downstream slot integrity constraints. When the verification fails, a gap signal carrying the gap character range and gap slot type is generated. Specifically, the slot verification on the structured word segmentation sequence is performed based on the downstream slot integrity constraints: the correspondence between semantic tags and anchor words in the structured word segmentation sequence is used as the verification benchmark. The existence of the target slot semantic tag is checked within eight consecutive character ranges after the anchor word. When the target slot semantic tag is missing, the character range at the missing position and the missing slot type are encapsulated into a gap signal. The gap signal is used to locate the backup candidate edge of the corresponding interval in the pruning candidate pool. The slot feasibility score is introduced to re-negotiate and score the backup candidate edge to obtain the corrected word segmentation sequence. The slot feasibility score is one when the semantic label of the backup candidate edge is consistent with the gap slot type, and zero otherwise.

7. The system according to claim 6, characterized in that, The set of rule triples is loaded into the AC automaton, and the input text is linearly scanned to obtain a set of candidate segments and trigger word position indices, including: Insert each term string from the set of rule triples that force the matching rule set into the trie structure of the AC automaton, calculate the mismatch pointer on the trie structure, and obtain the index structure of the AC automaton. The regular expression templates of the pattern template rule set in the rule triple set are separated from the AC automaton index structure, and regular expression scanning is performed on each line of the input text to obtain a set of regular expression matching fragments carrying character ranges and semantic tags. Based on the AC automaton index structure, a single linear scan is performed on the input text to obtain a set of term matching segments carrying character ranges, semantic labels, and rule confidence; the set of term matching segments and the set of regular expression matching segments are overlapped and eliminated according to the longest priority principle to obtain the set of candidate segments. Perform string location scanning on the input text using all trigger words in the context constraint rule set of the rule triple set, and record the start and end character positions of each trigger word, the corresponding target semantic label, and the confidence increment in a hash table to obtain the trigger word position index.

Citation Information

Patent Citations

  • Heterogeneous data preprocessing method, system and equipment and storage medium

    CN121327057A

  • Intelligent agent task construction method driven by output type

    CN122088550A