A method for constructing an efficacy matrix based on topic clustering and two-round question answering
By constructing an efficacy matrix based on topic clustering and two-round question answering, the problem of inconsistency between topic granularity and efficacy dimension in patent texts was solved, achieving stable information extraction and matrix construction, and ensuring stability and consistency across chapters.
Patent Information
- Application Number
- CN202511803137.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-24
- Estimated Expiration
- 2045-12-03
AI Technical Summary
In existing intelligent analysis of patent texts, topic clustering methods lack linkage constraints with efficacy dimensions, resulting in inconsistencies between topic granularity and downstream extraction boundaries. Information extraction results are not stable enough when migrating across documents or chapters. Furthermore, traditional information extraction schemes lack evidence binding and backtracking links, making it difficult to form a stable and reliable technology-efficacy correspondence.
We employ a efficacy matrix construction method based on topic clustering and two-round question answering. Through corpus standardization, semantic vectorization, and evidence triple construction, we generate hierarchical topics and perform granular gating. We combine hierarchical topics and granular gating to conduct the first round of question answering, extract efficacy phrases and bind them to evidence. We complete the second round of question answering and synonym merging through dimensional knowledge injection, construct a technology efficacy matrix and perform coverage consistency verification.
Stable question-and-answer and counting are achieved within the same text location and semantic expression framework, reducing mismatches and omissions. The scope of questions and answers is consistent with the extraction boundary, and cross-chapter mismerging is constrained. A unified source coordinate system and evidence backtracking chain are formed to ensure the contextual authenticity and item stability of the matrix structure.
Smart Images

Figure CN121256033B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of patented text intelligent analysis and information extraction technology, and in particular to a method for constructing an efficacy matrix based on topic clustering and two-round question answering. Background Technology
[0002] Patent texts consist of heterogeneous parts from multiple sources, including titles, abstracts, claims, and specifications. Their structure is complex, with redundant and synonymous terminology. Technical and efficacy elements are scattered across different chapters and paragraphs, lacking a direct semantic and positional correspondence. Manual merging or simple keyword-based statistics are prone to mismatches and omissions, making it difficult to establish a stable and reliable technical-efficacy correspondence. Existing topic clustering methods often focus on planar topics or fixed-granularity clustering, lacking constraints linked to the efficacy dimension. The topic granularity often differs from downstream extraction boundaries, leading to imbalances in subsequent phrase extraction and matrix statistics. Traditional information extraction schemes rely heavily on rules, dictionaries, or single-round question-and-answer, lacking evidence binding and backtracking links with source paragraphs. This makes it difficult to solidify the "statement-position-semantics" triadic correspondence at the sentence / segment level, resulting in insufficient stability of extraction results when migrating across documents or chapters. Statistical modeling for technology-efficacy relationships often involves co-occurrence counting at the document or paragraph level without considering hierarchical themes, granular gating, and evidence-level verification. The counting results are easily affected by high-frequency terms and differences in document length. At the same time, there is a lack of systematic verification and feedback mechanisms for coverage and consistency, making it impossible to conduct closed-loop revisions within the process for weak thematic nodes, easily confused dimensions at the boundaries, and evidence gaps. Summary of the Invention
[0003] This invention provides a method for constructing a performance matrix based on topic clustering and two-round question answering, in order to solve the problem of how to construct a technical performance matrix and complete the coverage consistency test and feedback flow based on second-round question answering, dimension alignment, synonym merging and hierarchical topics, while following granular gating and relying on corpus standardization, semantic vectorization and evidence triples.
[0004] To address the aforementioned technical problems, this invention provides a method for constructing an efficacy matrix based on topic clustering and two-round question answering, comprising:
[0005] The patent text is obtained, and the title, abstract, claims and description are hierarchically divided according to fixed fields. Corpus standardization, semantic vectorization and evidence triple construction are performed to obtain evidence triples that retain paragraph numbers and claim numbers.
[0006] Hierarchical topic generation is performed based on semantic vectorization and corpus standardization, establishing a hierarchical structure of parent-child and sibling relationships. Granular gating calculation is performed by combining layer depth threshold and entry threshold. Consistency integration is completed through node coverage range and boundary stability verification to obtain hierarchical topics and granular gating.
[0007] In terms of hierarchical topics and granular gating, the first round of question and answer is executed based on the entry threshold in the granular gating, extracting efficacy phrases and completing evidence binding through text boundaries, paragraph positions and semantic expressions;
[0008] Based on the first round of question and answer and evidence binding, adjacent merging of efficacy phrases within the nodes of the hierarchical topic is performed to generate temporary dimension units, and dimension knowledge entries containing positive and negative examples are injected. Boundary constraints are updated through boundary prompts and layer depth thresholds.
[0009] Based on efficacy dimension aggregation and dimension knowledge injection, the second round of question and answer is completed through representative expressions and boundary prompts in dimension knowledge items. Dimension alignment is performed by combining semantic vectorized item alignment fragments. Synonym merging is completed through text inclusion relationship and semantic high consistency rules, and consistent mapping processing of primary key and subordinate expression is performed.
[0010] Based on the second round of question answering, dimension alignment and synonym merging, a technical effectiveness matrix is constructed by combining the row space nodes of hierarchical topics. Coverage consistency is checked by counting each certificate and verifying the consistency of sources. Feedback is completed by triggering additional questions or boundary fine-tuning based on gap markers and conflict markers.
[0011] Furthermore, hierarchical topic generation based on semantic vectorization and corpus standardization includes:
[0012] Based on semantic vectorization, adjacent processing units from the same patent text and with similar semantics are clustered. During the clustering process, the paragraph position and chapter level recorded in the corpus standardization are used as boundary constraints to avoid disordered merging across chapters. The candidate set formed by each clustering and its corresponding source identifier and position identifier are fixed at the same time.
[0013] Repeated or synonymous segments are uniformly merged, while maintaining the original order and subordinate relationship, and the original position is recorded in the merged entries;
[0014] The representative fragment and its context fragment are included together as the core content of the set. Based on this core content, the inclusion relationship, parallel relationship and inheritance relationship between different candidate sets are determined, thus forming a hierarchical structure of parent and child and siblings.
[0015] A bidirectional mapping to corpus standardization is established at each node of the hierarchy, so that any topic node can be located to the precise position of the original text, and can be located from the original text back to the corresponding topic node.
[0016] Record text fragments that co-occur with evidence triples within the hierarchical nodes;
[0017] The density, span, and coverage of nodes at each level are checked for consistency. If any node is found to be too wide or too narrow in coverage, it is fine-tuned without breaking the original chapter boundaries.
[0018] Furthermore, the granular gating calculation, which combines the layer depth threshold and the entry threshold, includes:
[0019] Based on the representative fragments, context fragments and the mapping to the corpus standardization at each node of the hierarchical topic, the coverage and boundary stability of that node are determined.
[0020] For nodes with uneven coverage or weak boundary stability, small-scale merging or splitting within the hierarchy can be performed to ensure that the description of the node is consistent with the text fragment it maps to.
[0021] Using the evidence triples that have a reference relationship with the node as a reference, the source text fragments and related text fragments in the evidence triples are labeled in the context of the node. The labeling results are used to express the questionable and answerable range of the node in subsequent question and answer.
[0022] In the vertical direction of the hierarchy, the depth distribution of each layer is calculated based on the parent-child relationship, and in the horizontal direction, the relative distance between nodes in the same layer is calculated based on the sibling relationship. Based on the vertical and horizontal relationships, several gating boundaries are determined without introducing new semantics.
[0023] By combining the chapter tags in the corpus standardization, gating constraints are set for nodes in different chapters, so that nodes in the abstract scope, nodes in the claim scope, and nodes in the specification scope have different entry strategies when they are called in the future.
[0024] Nodes that co-occur frequently with evidence triples are marked as priority entry targets;
[0025] The gating layer depth threshold, entry threshold, and boundary rules are uniformly encapsulated into a readable structural description.
[0026] Furthermore, consistency integration is achieved through node coverage and boundary stability checks, including:
[0027] Read the parent-child and sibling relationships of hierarchical topics and the bidirectional mapping to corpus standardization. At the same time, read the layer depth threshold, entry threshold and boundary rules in the granular gating and compare each node item by item in the same structural view.
[0028] If it is found that the coverage of a node is partially excluded by the gating rules, the coverage description of the node is adjusted to match the entry threshold of the granular gating without changing the parent-child and sibling relationships of the hierarchical topics.
[0029] If a slight misalignment is found between the node boundary and the gating boundary, the boundary is fine-tuned based on the paragraph position in the corpus standardization.
[0030] Establish a consistency label at the node level. The consistency label is used to record the synchronization status of the node in terms of hierarchical structure, gating boundary and text mapping.
[0031] For nodes that need to be corrected, the correction will be completed within this integration, and the consistency label will be updated after the correction is completed;
[0032] A list of evidence citations is created for each node, which is derived from the annotation results of the evidence triples in the context of that node.
[0033] Generate a set of constraint rules, which describes the entry order, jump conditions, and termination conditions that should be followed when a node is invoked to enter the question-and-answer and extraction process.
[0034] The integration results record the mapping integrity identifier between each node and corpus standardization, the expression integrity identifier between each node and semantic vectorization, and the citation integrity identifier between each node and the evidence triple.
[0035] Furthermore, the first round of question-and-answer is performed based on the entry threshold in the granular gating, including:
[0036] In the hierarchical topic, select a set of target nodes with clear parent-child and sibling relationships. The determination of the target node set is constrained by the layer depth threshold and entry threshold in the granularity gating. Any node that does not meet the entry threshold is not included.
[0037] For nodes that meet the entry threshold, the paragraph boundaries and sentence numbers mapped to the node are read from the corpus standardization to form a questionable interval that corresponds one-to-one with the node.
[0038] For each target node, based on the entry strategy and termination conditions recorded in the granular gating, question and answer prompt text is generated within the questionable range. The question and answer prompt text is composed of the node's representative fragment, context fragment, and limiting boundary.
[0039] Set response format requirements and source guidance in the prompt text so that subsequent responses can include the source paragraph number, the start and end positions of the context, and the corresponding phrase fragments;
[0040] For text fragments marked as co-occurring with technical and efficacy elements in the evidence triples obtained in S130, a reference reference is inserted in the question and answer prompt text;
[0041] Sort all target nodes by chapter level and relative node position, generate question and answer batches, and initiate questions and answers in sequence;
[0042] Each question and answer session outputs a question and answer record entry containing the question text, the answer text, the source paragraph position, the corresponding topic node, and the gating strategy used.
[0043] After each question and answer session, a consistency check is immediately performed to verify whether the source identifiers and paragraph position identifiers in the question and answer record entries are consistent with the corpus standardization, and whether the question and answer content falls within the questionable range.
[0044] Furthermore, efficacy phrases are extracted and evidence binding is completed by linking them to text boundaries, paragraph positions, and semantic expressions, including:
[0045] For each question and answer record entry, the response text is segmented into sentences and words are standardized. The boundaries of sentence segmentation are based on the sentence segment number in the corpus standardization, and word standardization is based on the terminology standardization table in the corpus standardization.
[0046] In the completed and standardized response text, possible efficacy phrase candidates are identified. The identification of candidates is based on the response format requirements set in the question and answer prompt text, and cross-checks are performed in combination with the source paragraph position recorded in the question and answer record entries. Candidates that cannot be located in the original source paragraph are not adopted.
[0047] For candidates that can be located, read their corresponding semantic vectorized entries, calculate the semantic similarity with other segments in the same entry, eliminate candidates with too broad or too narrow coverage, and mark the remaining candidates as efficacy phrase items.
[0048] For each functional phrase item, text boundaries, paragraph positions, and semantic expressions are registered simultaneously. Text boundaries are derived from the start and end positions of the original paragraphs, paragraph positions are derived from the paragraph and sentence numbers standardized by the corpus, and semantic expressions are derived from the corresponding entries of semantic vectorization.
[0049] Deduplication and merging of efficacy phrases from the same topic node. The deduplication rule is that the text is consistent or the text contains the same content and the semantic expression is highly consistent. The merging rule is that the text is similar and the semantic expression is similar.
[0050] For each topic node, a list of effective phrase items is generated. Each item in the list has three types of information: text boundary, paragraph position, and semantic expression, and corresponds to a unique question and answer record item.
[0051] Reserve evidence location markers in the list of efficacy phrase items, which will be filled in one by one when S330 calls the evidence triplet;
[0052] For each efficacy phrase item, a search key pointing to the semantically vectorized entry is provided.
[0053] Furthermore, evidence binding is achieved through binding text boundaries, paragraph positions, and semantic expressions, including:
[0054] Read the list of efficacy phrase items by node. For each item in the list, retrieve the item in the evidence triple that is consistent with the position of its source paragraph and has overlapping context.
[0055] For each entry in the candidate set, read the source text fragment, associated text fragment, and supporting text fragment in the entry, and compare them character by character with the text boundary of the function phrase item. Any entry that can locate the complete phrase boundary within the source text fragment or supporting text fragment is marked as a matching entry.
[0056] For each function phrase and its set of matching entries, a binding entry is created. The binding entry records the text boundaries, paragraph positions and semantic expressions of the function phrase, as well as the source text fragments, related text fragments and supporting text fragments in the matching entry.
[0057] Register the question and answer record entry number and topic node number on each bound entry;
[0058] When there are multiple matching entries for the same function phrase, they are sorted according to the degree of coverage of their text boundaries and the degree of consistency with their semantic expression. The top one or several entries are marked as primary bindings, and the remaining entries are marked as backup bindings.
[0059] Perform a consistency check on all bound entries within the node. The consistency check includes whether the source identifier is consistent, whether the paragraph position is continuous, whether the supporting text fragment is complete, and whether the function phrase item is fully presented within the bound range.
[0060] All the binding entries of all nodes are collected to form a binding set that matches the first round of question and answer and the extraction of efficacy phrases.
[0061] Furthermore, within the nodes of the hierarchical topic, adjacency merging of efficacy phrases is performed to generate temporary dimensional units, including:
[0062] For each target node in the hierarchical topic, read the list of efficacy phrase items and the evidence items paired with each efficacy phrase item in that node. Define the processable range according to the entry threshold and termination condition in the granular gating. Based on this, merge adjacent efficacy phrase items within the range. During the merging process, the semantic similarity relationship indicated by the semantic vectorized items is used as a reference, while keeping the paragraph order and sentence boundaries in the corpus standardization from being crossed.
[0063] After completing the adjacent merging within the node, the representative expression of the merged efficacy phrase group is determined. The representative expression is the one with the most stable coverage and the most sufficient evidence items in the group as the core expression. The core expression, along with its corresponding evidence source paragraph, context position, and question and answer source, is fixed and saved to form a temporary dimension unit of the group.
[0064] While maintaining the node hierarchy, temporary dimension units that are highly similar between different nodes are compared at the same level. Temporary dimension units that are highly similar in semantic vectorization and are in adjacent chapters in corpus standardization are merged or paralleled according to the boundary rules for cross-node merging in granular gating.
[0065] Generate a mapping relationship for each temporary dimension unit. The mapping relationship records the correspondence between the efficacy phrase item and the temporary dimension unit, the correspondence between the temporary dimension unit and its node, and the correspondence between the temporary dimension unit and the evidence item.
[0066] Perform a consistency check on all temporary dimension units within a node. The consistency check includes whether the representative expression corresponds completely to its source of evidence, whether the representative expression is located within the processable range, and whether the mapping relationship is self-consistent.
[0067] Two types of entry information are generated on the node-level dimensional unit list. One type of entry points to the representative expression and its main evidence item of each dimensional unit, and the other type of entry points to all efficacy phrase items and all evidence items within each dimensional unit.
[0068] Furthermore, inject dimensional knowledge entries containing both positive and negative examples, including:
[0069] Read the representative expression and main evidence entry of each dimension unit, and form a context summary that can explain the boundary features based on the position interval of the representative expression in its main evidence entry and the preceding and following text adjacent to that interval.
[0070] Select several fragments that can serve as positive and negative examples from all the evidence entries associated with this dimension unit. Positive examples are used to indicate typical segments that are consistent with the representative expression, while negative examples are used to indicate easily confused segments that are similar to the representative expression but do not belong to this dimension unit.
[0071] By combining the distribution of efficacy phrases within the same dimensional unit with the representative expressions of adjacent dimensional units, a set of boundary prompts is compiled;
[0072] By combining the parent-child and sibling relationships in the hierarchical topics, the subordinate relationship with the upper-level dimension and the parallel relationship with the same-level dimension are recorded for this dimension unit;
[0073] The above information, along with the representative expression, source location, question and answer source, list of efficacy phrases, and mapping relationships, are encapsulated into dimensional knowledge entries;
[0074] For each dimension of knowledge entry, establish a retrieval key pointing to the semantically vectorized entry and a position key pointing to the corpus-standardized entry;
[0075] Register the corresponding layer depth and entry threshold for each dimension of knowledge entry with granular gating;
[0076] A consistency check is performed on the set of entries. The consistency check includes whether the representative expression matches the context summary, whether the positive and negative examples come from evidence entries related to the dimension unit, whether the boundary cue is consistent with the boundary rules of granular gating, and whether the key value corresponds completely to the source.
[0077] Furthermore, boundary constraints are updated using boundary cue statements and layer depth thresholds, including:
[0078] Read the context summary and boundary prompts in each dimension knowledge item one by one, and compare them with the representative expression and evidence coverage of the corresponding dimension unit in the efficacy dimension aggregation. If it is found that the paragraph covered by the context summary exceeds the processable range limited by the granular gating, then based on the entry threshold and termination condition in the granular gating, the range of efficacy phrases covered by the representative expression of the dimension unit is narrowed.
[0079] For cases where different dimensional units are similar in semantic vectorization and adjacent but not completely overlapping in corpus standardization, cross-checking is performed according to the positive and negative examples recorded in the dimensional knowledge entries. If the positive examples of the two dimensional units can be placed in the context summary of the other and do not violate the boundary rules of granular gating, then it is determined that there is a risk of overlap between the two.
[0080] Guided by boundary prompts, make a fine-tuning within the original node range. Fine-tuning methods include shrinking the phrases representing the expression, adjusting the allocation of efficacy phrases, and switching the primary and backup evidence items.
[0081] For each dimension of knowledge entry, placeholder information for entry strategy, jump conditions and termination conditions is added. The entry strategy is set according to the depth of granular gating and the entry threshold. The jump conditions are set according to the parent-child and sibling relationships of the hierarchical topics. The termination conditions are set according to the completeness of the coverage of the evidence entries.
[0082] A global consistency review is conducted at the dimensional level. The review includes whether the representative expressions of all dimensional units fall within the processable range, whether all mapping relationships are self-consistent, whether the boundary descriptions of all dimensional knowledge items and their corresponding dimensional units are consistent, and whether all entry strategies and jump paths can be executed under hierarchical topics and granular gating.
[0083] The key innovations of this invention include:
[0084] (1) The collaborative generation and consistency integration mechanism of dynamic hierarchical topics and granular gating supports the integrated organization of source coordinate system and evidence-level backtracking.
[0085] (2) The evidence triple-driven bidirectional mapping framework establishes a stable alignment between topic nodes, paragraph positions and semantic expressions, ensuring that question answering and counting are performed within the same structure.
[0086] (3) The linkage process of two rounds of question and answer and dimensional knowledge injection first fixes the basis for efficacy expression, and then forms a standardized column space through dimensional alignment and synonym merging.
[0087] (4) Counting is carried out based on the admission rules of consistent source and overlapping context, and the item management method of coexisting primary and backup items is adopted in conjunction with the correction strategy of length and source distribution.
[0088] The following are its main beneficial effects:
[0089] (1) Forming a unified source coordinate system and evidence backtracking chain: Corpus standardization, semantic vectorization and evidence triple construction serve as the foundation. The entire process of first-round question and answer, efficacy phrase extraction, evidence binding, efficacy dimension aggregation, dimension knowledge injection, second-round question and answer and dimension alignment, synonym merging, and finally the construction of the technical efficacy matrix is located in the same text position and semantic expression framework. Question and answer records, phrase items, dimension units and counting units maintain integrated mapping and traceable source, reduce mismatch and omission, and facilitate item-level verification and backtracking revision.
[0090] (2) Dynamic hierarchical topics and granular gating are integrated: hierarchical topics give parent-child and sibling relationships, and granular gating sets entry and termination boundaries; after the two are integrated in a consistent manner, the question and answer range, extraction boundary and subsequent counting row space are consistent, cross-chapter boundary and cross-topic erroneous phenomena are constrained, matrix row index and topic node correspond one-to-one, and the statistical range is synchronized with the text structure.
[0091] (3) Two-round question-and-answer driven column space standardization: The first round of question-and-answer combines evidence binding to fix the original text fragments and source locations of efficacy expressions; efficacy dimension aggregation and dimension knowledge injection provide contextual summaries, positive and negative examples and boundary prompts; the second round of question-and-answer completes dimension alignment and synonym merging under the guidance of the above knowledge, generates a list of standard expressions and subordinate expressions, and the column space is standardized. The same set of expressions is used for subsequent counting and testing stages.
[0092] (4) Evidence-level counting and item correction in parallel: The technical effectiveness matrix construction only accepts aligned segments with consistent sources and overlapping contexts. The counting items record source identifiers, paragraph positions, context start and end and expression types. Correction strategies are implemented for length and source distribution. Main items and backup items are stored in layers. The matrix structure takes into account both contextual authenticity and item stability. Attached Figure Description
[0093] Figure 1 This is a flowchart illustrating a method for constructing an efficacy matrix based on topic clustering and two-round question answering, provided in an embodiment of this application. Detailed Implementation
[0094] Example 1: Refer to Figure 1 This is a flowchart illustrating a method for constructing an efficacy matrix based on topic clustering and two-round question answering, provided by an embodiment of the present invention. The process may include at least steps S100-S600:
[0095] S100. Obtain the patent text, divide the title, abstract, claims and description into hierarchical categories according to fixed fields, perform corpus standardization, semantic vectorization and evidence triple construction to obtain evidence triples that retain paragraph numbers and claim numbers;
[0096] S200: Based on semantic vectorization and corpus standardization, hierarchical topic generation is performed to establish a hierarchical structure of parent-child and sibling relationships. Granular gating calculation is performed by combining layer depth threshold and entry threshold. Consistency integration is completed through node coverage range and boundary stability verification to obtain hierarchical topics and granular gating.
[0097] S300. In terms of hierarchical topics and granular gating, the first round of question and answer is executed based on the entry threshold in the granular gating, extracting efficacy phrases and completing evidence binding through text boundaries, paragraph positions and semantic expressions;
[0098] S400. Based on the first round of question and answer and evidence binding, adjacent merging of efficacy phrases within the nodes of the hierarchical topic is performed to generate temporary dimension units. Dimensional knowledge entries containing positive and negative examples are injected, and boundary constraints are updated through boundary prompts and layer depth thresholds.
[0099] S500, based on efficacy dimension aggregation and dimension knowledge injection, completes the second round of question answering through representative expressions and boundary prompts in dimension knowledge items, performs dimension alignment by combining semantic vectorized item alignment fragments, completes synonym merging through text inclusion relationship and semantic high consistency rules, and performs consistent mapping processing of primary key hook subordinate expression;
[0100] S600 constructs a technical effectiveness matrix based on the second round of question answering, dimension alignment and synonym merging, combined with the row space nodes of the hierarchical topics. It performs coverage consistency verification by counting each certificate and checking the consistency of sources, and completes feedback flow by triggering additional questions or boundary fine-tuning based on gap markers and conflict markers.
[0101] Step S100 includes at least steps S110-S130:
[0102] S110. Obtain the patent text, standardize the corpus, and obtain standardized corpus;
[0103] The input for this step is patent text, and the output is corpus standardization. Specifically, firstly, patent texts from public databases, internal enterprise databases, and project materials are uniformly accessed. The title, abstract, claims, and specification are hierarchically divided according to fixed fields, establishing record entries corresponding to the source. The order, hierarchy, and positional relationship of each entry are maintained to ensure retrieval and citation based on positional relationships in subsequent steps. Furthermore, punctuation, spaces, line breaks, header / footer markers, and meaningless separators in the patent text are uniformly processed. Character encoding differences caused by different sources are unified, breaks caused by line breaks and page breaks are eliminated, and continuous paragraphs are connected. The original paragraph numbers, claim numbers, and hierarchy markers are retained, creating a parsable hierarchical sequence in the patent text. Further, the terminology involved in the patent text is standardized for consistency. Different spellings of the same term in different paragraphs are unified, maintaining a consistent correspondence between the original and standardized expressions to avoid segmentation errors caused by inconsistent terminology in subsequent processing. After the above processing is completed, segments with cross-references between the claims and the specification are paired to establish a traceable correspondence between the limiting statements in the claims and the descriptions of embodiments in the specification, ensuring that the same technical element has a stable annotation method in different parts. To ensure the reproducibility of subsequent processing, a searchable source identifier and paragraph position identifier are generated for each record, recording the start and end positions, chapters, and subordinate relationships of each text in the entire patent text, and storing them in the same data structure as the corresponding entries. Understandably, corpus standardization includes not only the consistency of characters and format, but also the segmentation standardization at the paragraph and sentence levels. After completing the identification of sentence and paragraph boundaries, a sequential number and a hierarchical number are established for each segment, so that the corpus standardization can be directly called in subsequent processing. Furthermore, for repeated expressions in the abstract and claims, deduplication rules are used to merge highly repeated segments, retaining the first occurrence of the expression, and using citation tags to record its repetition in other positions, thereby forming a non-redundant corpus standardization. Finally, the corpus standardization, source identifier, and paragraph position identifier are output together as direct input for subsequent semantic vectorization from corpus standardization. They are also directly referenced when generating hierarchical topics by obtaining semantic vectorization and corpus standardization in S210, and provide a traceable textual basis for subsequent construction of evidence triples.
[0104] S120. Perform semantic vectorization on the standardized corpus to obtain semantic vectorization;
[0105] The input to this step is the standardized corpus, and the output is semantic vectorization. Specifically, based on corpus standardization, processing units are established at the paragraph and sentence levels. Each processing unit maintains a one-to-one correspondence with its source identifier and paragraph position identifier, and this correspondence remains unchanged during subsequent transmission. Furthermore, the processing units in the corpus standardization are linguistically standardized, including the annotation of word boundaries, phrase boundaries, and syntactic components, ensuring that the same processing unit has a stable semantic carrying structure in subsequent expressions. Based on this, semantic representations are generated for each processing unit, giving each unit a numerical expression consistent with its semantic meaning. This numerical expression, along with the source identifier and paragraph position identifier, is packaged and stored to form a semantic vectorized entry that can be retrieved and compared. Furthermore, to ensure comparability between different processing units, the scales involved in the semantic representation generation process are unified, providing a comparable metric space for semantic vectorization within the same patent text. The scale information of this metric space is saved along with the semantic vectorization to avoid inconsistencies in subsequent multi-document analysis. In the above processing, the semantic representations of the abstract, claims, and specification are organized hierarchically. This ensures that the semantic vectorization of the abstract layer, claims layer, and specification layer are independent yet mutually referential, allowing for hierarchical aggregation during subsequent generation of hierarchical topics based on semantic vectorization and corpus standardization. After generation, the semantic vectorization undergoes integrity and consistency checks. These checks include whether all processing units are covered, whether source identifiers and paragraph position identifiers are consistent with corpus standardization, and whether there are any missing or misaligned entries. If inconsistencies are found, the process reverts to corpus standardization for correction and regenerates the corresponding entries to ensure a strict correspondence between semantic vectorization and corpus standardization. Furthermore, to facilitate range control in subsequent hierarchical topic generation and granular gating calculations, corpus level and chapter tags are attached to the semantic vectorization entries. This accurately limits the source range of any entry, preventing aggregation from crossing segments that should not be aggregated. Ultimately, semantic vectorization forms a list of directly searchable entries, maintaining a one-to-one or one-to-many mapping relationship with corpus standardization. This serves as the input carrier for S210 to obtain semantic vectorization and corpus standardization for hierarchical topic generation, as well as the retrieval basis for S610 when constructing the technical efficacy matrix, and provides a callable scope description for S310 to limit the scope of questions in the first round of question answering.
[0106] S130. Based on the semantic vectorization and the corpus standardization, construct the evidence triples to obtain the evidence triples;
[0107] The inputs to this step are the semantic vectorization and the corpus standardization, and the output is the evidence triple. Specifically, using semantic vectorization as the retrieval entry point and corpus standardization as the text foundation, segments involving technical elements and efficacy elements are targeted for localization. First, processing units related to technical elements and processing units related to efficacy expression are retrieved from the semantic vectorization entries. The retrieval scope is limited based on corpus level tags and chapter tags, ensuring that the technical limitations and efficacy descriptions of the same entry remain spatially adjacent or traceable. After localization, the processing units related to technical elements are used as the starting point, and the processing units related to efficacy expression are used as the ending point. A directional relationship from the starting point to the ending point is established using paragraph position identifiers as intermediaries. Within this directional relationship, key sentences that can support the association are retrieved, forming a reproducible contextual link between the key sentences and the starting and ending points. Furthermore, the text in the aforementioned contextual links is standardized and sliced to determine the boundaries of text segments used for subsequent binding, ensuring that the text segments neither cross the original chapter boundaries nor truncate key information. Simultaneously, each text segment, along with its source identifier and paragraph position identifier, is permanently stored, forming the basic evidence information that can be cited line by line in subsequent steps. Further, based on the above, an entry is created for each contextual link. Each entry includes three parts: a starting text segment, an ending text segment, and a key sentence segment, along with the corresponding source identifier and paragraph position identifier, arranged in a fixed order to ensure consistent identification and parsing in subsequent processing. Based on these entries, evidence triples are constructed, comprising a source text segment, a related text segment, and supporting text segments. Each part retains the source identifier and paragraph position identifier consistent with the corpus standardization, allowing each item in the evidence triple to be verified back to its original position in the corpus standardization. To avoid overlap and confusion among evidence triplets across different patent texts, the evidence triplets are grouped and managed according to the patent text during the construction process. This ensures that evidence triplets under the same patent text form an ordered sequence within the same group, and that different patent texts are distinguished by groups, thus maintaining clear boundaries for the evidence set. The evidence triplets retain their original paragraph numbers, claim numbers, and hierarchical markers. Furthermore, for multiple evidence triplets arising from the same technical element but different expressions of efficacy, they are arranged according to the order of paragraph position identifiers, preserving the original order so that they can be traversed and cited in the natural narrative order during the initial round of questioning. Items that are obviously duplicated or completely identical in expression are deduplicated, retaining the first item to appear, and recording the position of duplicate items after the first item in a citation format to ensure no omissions during subsequent citations.Furthermore, to facilitate the subsequent first-round question-and-answer session and efficacy phrase extraction, the supporting text fragments in the evidence triples are defined as materials that can be directly used for questioning and answering. This allows subsequent processing to directly extract information and bind evidence within this material range, avoiding the uncertainty caused by out-of-bounds retrieval.
[0108] Step S200 includes at least steps S210-S230:
[0109] S210. Obtain the semantic vectorization and the corpus standardization, perform hierarchical topic generation, and obtain hierarchical topics;
[0110] This step uses semantic vectorization as the input for aggregation and comparison, and corpus standardization as the foundation for text localization and backtracking. It combines the evidence triples formed in S130 as a reference to establish a searchable candidate set of topics at the paragraph and sentence levels of the patent text. Specifically, firstly, based on semantic vectorization, processing units originating from the same patent text and possessing similar semantics are clustered adjacently. During the clustering process, the paragraph positions and chapter levels recorded in the corpus standardization serve as boundary constraints to avoid disordered merging across chapters. The candidate set formed by each cluster, along with its corresponding source and position identifiers, are simultaneously fixed to ensure that subsequent backtracking along the corpus standardization can proceed. Further, within the candidate set, repetitive or synonymous segments are uniformly merged, maintaining the original order and subordinate relationships. The original position is recorded in the merged entries for source tracing during subsequent citations. Subsequently, representative segments that represent the semantic center of each candidate set are identified. These representative segments and their context segments are included as the core content of the set. Based on this core content, the inclusion, parallel, and inheritance relationships between different candidate sets are determined, thus forming a hierarchical structure of parent-child and sibling relationships. Furthermore, to enable the hierarchical structure to interface with subsequent question-and-answer and efficacy extraction, a bidirectional mapping to the standardized corpus is established at each node of the hierarchy. This allows any topic node to be located at its precise position in the original text, and simultaneously allows for reverse location from the original text to the corresponding topic node. At the same time, text segments that co-occur with the evidence triples are recorded within the hierarchical nodes, so that when the hierarchical topic is subsequently invoked, it can directly obtain the reference context related to the technical and efficacy elements. To ensure the stability of the hierarchical division, this step checks the consistency of the density, span, and coverage of nodes at each level after generating the initial hierarchy. If nodes with excessively wide or narrow coverage are found, fine-tuning is performed without disrupting the original chapter boundaries to obtain a computable structural carrier for subsequent granular gating calculations from the hierarchical topics. After the aforementioned processing, the hierarchical topics are output. The hierarchical topics maintain a one-to-one or one-to-many mapping relationship with the semantic vectorization and corpus standardization, and establish optional reference relationships with the evidence triples. This output serves as one of the sole inputs in subsequent granular gating calculations from the hierarchical topics. It also serves as the direct source of scope limitation when obtaining the hierarchical topics and granular gating for the first round of question answering in S310, and as the structural basis for the row space when constructing the technical efficacy matrix in S610.
[0111] S220. Perform granular gating calculations from the hierarchical topics to obtain granular gating;
[0112] This step uses the hierarchical topic as input, refers to the chapter level and paragraph position in the corpus standardization, and optionally references text fragments and source identifiers involving technical and efficacy elements in the evidence triples to form a granular gating system that can be directly invoked in subsequent steps. Specifically, firstly, for each node of the hierarchical topic, based on the representative fragments, context fragments, and mappings to the corpus standardization within the node, the coverage and boundary stability of the node are determined. Coverage describes the actual span of the node in the original text, and boundary stability describes whether the boundary between the node and adjacent nodes is clear. For nodes with uneven coverage or weak boundary stability, small-scale merging or splitting within the hierarchy is performed to ensure that the description of the node is consistent with the text fragments it maps to. Further, using the evidence triples that have a reference relationship with the node as a reference, the source text fragments and related text fragments in the evidence triples are labeled within the context of the node. The labeling results are used to express the questionable and answerable range of the node in subsequent question-and-answer sessions, so as to constrain the scope of question-and-answer during the first round of question-and-answer in S310. Subsequently, in the vertical direction of the hierarchy, the depth distribution of each layer is calculated based on parent-child relationships, and in the horizontal direction, the relative distance between nodes in the same layer is calculated based on sibling relationships. Based on the vertical and horizontal relationships, several gating boundaries are determined without introducing new semantics. The gating boundaries are used to express the allowed range of question answering and extraction at what layer depth and relative distance. Furthermore, this step combines the chapter tags in the corpus standardization to set gating constraints for nodes in different chapters, so that nodes in the abstract scope, claim scope, and specification scope have different entry strategies when subsequently invoked, avoiding contextual shifts caused by cross-scope invocation; at the same time, nodes that co-occur more frequently with the evidence triples are marked as priority entry objects, so that nodes more closely related to technical and efficacy elements are covered first in the initial round of question answering. To ensure the granular gating's executableness when invoked in subsequent steps, this step encapsulates the gating's depth threshold, entry threshold, and boundary rules into a readable structural description. This description is stored alongside the hierarchical topic, allowing it to be directly read and compared item by item during subsequent consistency integration of the hierarchical topic and the granular gating. After the aforementioned processing, the granular gating is output, establishing a one-to-one association with the hierarchical topic and participating in consistency integration together with the hierarchical topic in S230. Simultaneously, the granular gating is invoked as a range constraint during the first round of question-and-answer processing of the hierarchical topic and the granular gating in S310, referenced as an extraction boundary during efficacy phrase extraction from the first round of question-and-answer in S320, and used to limit the statistical range of certificate-by-certificate counting during the construction of the technical efficacy matrix in S610. This ensures that the output of this step forms a stable input in subsequent processes.
[0113] S230. Perform consistent integration of the hierarchical topic and the granular gating to obtain the hierarchical topic and granular gating;
[0114] This step uses the hierarchical topics and granular gating as the sole inputs, and verifies and unifies the structural correspondence, scope boundary relationships, and reference relationships between them one by one to output hierarchical topics and granular gating that can be directly used for question answering and extraction. Specifically, it first reads the parent-child and sibling relationships of the hierarchical topics and the bidirectional mapping to the corpus standardization, and simultaneously reads the layer depth threshold, entry threshold, and boundary rules in the granular gating. Then, it compares each node item by item within the same structural view. The comparison includes whether the node's coverage is fully encompassed by the gating rules, whether the node's boundary is consistent with the gating boundary, and whether the node's questionable and answerable ranges match the gating entry strategy. Specifically, consistency integration is achieved through node coverage and boundary stability checks. If a node's coverage is found to be partially excluded by gating rules, its coverage description is adjusted to match the entry threshold of the granular gating, without changing the parent-child and sibling relationships of the hierarchical topics. If a slight misalignment is found between the node's boundary and the gating boundary, the boundary is fine-tuned based on the paragraph position in the corpus standardization to ensure consistency at the text level. Furthermore, to ensure that subsequent question answering and extraction follow the same path, this step establishes a consistency label at the node level. The consistency label records the node's synchronization status in terms of hierarchical structure, gating boundary, and text mapping. A node is marked as consistent when all three aspects are aligned, and as requiring correction when any aspect is not aligned. For nodes requiring correction, this step completes the correction within this integration process, and the consistency label is updated after correction to ensure that the integrated structure does not require re-alignment when invoked. Subsequently, to facilitate subsequent evidence-by-evidence citation, this step establishes an evidence citation list for each node. This list originates from the annotation results of the evidence triple in the context of that node, used to directly locate the reference fragment during the first round of question-and-answer in S310, and to identify each fragment one by one during evidence binding based on the efficacy phrase extraction and the evidence triple in S330. To ensure the controllability of the invocation, this step also generates a set of constraint rules. The set of constraint rules describes the entry order, jump conditions, and termination conditions that should be followed when a node is invoked into question-and-answer and extraction, and is bound to the node in a readable form within the structure. Furthermore, to ensure the connection with subsequent statistics and verification, this step records the mapping integrity identifier between each node and the corpus standardization, the expression integrity identifier between each node and the semantic vectorization, and the citation integrity identifier between each node and the evidence triple in the integration results. This allows subsequent steps to directly determine whether a node meets the preconditions for being statistically analyzed and verified based on these three types of identifiers.
[0115] Step S300 includes at least steps S310-S330:
[0116] S310. Obtain the hierarchical topic and the granularity gating, perform the first round of question and answer, and obtain the first round of question and answer;
[0117] This step uses the hierarchical topics and granular gating obtained in S230 as the sole input, and the corpus standardization and semantic vectorization formed in step 110 as the foundation for localization and retrieval. Simultaneously, it refers to the evidence triples obtained in S130 as optional contextual evidence to determine the topic nodes, question-and-answer scope, and question-and-answer order for the first round of question-and-answer. Specifically, firstly, a set of target nodes with clear parent-child and sibling relationships is selected from the hierarchical topics. The determination of this target node set is constrained by the layer depth threshold and entry threshold in the granular gating; nodes that do not meet the entry threshold are not included. For nodes that meet the entry threshold, the paragraph boundary and sentence / segment number mapped to the node are read from the corpus standardization to form a questionable interval corresponding to that node. This interval maintains the same source identifier and paragraph position identifier mapping relationship with the semantic vectorization entries in step 1120, ensuring bidirectional backtracking at both the textual and semantic levels during the question-and-answer process. Subsequently, for each target node, based on the entry strategy and termination conditions recorded in the granular gating, question-and-answer prompt text is generated within the questionable range. This prompt text consists of a representative segment of the node, a contextual segment, and a limiting boundary. To ensure the usability of the question-and-answer output, this step sets response format requirements and source guidance in the prompt text, enabling subsequent responses to include the source paragraph number, the start and end positions of the context, and the corresponding phrase segment. This facilitates boundary identification during the next step of extracting efficacy phrases. Furthermore, for text segments marked as co-occurring with technical and efficacy elements in the evidence triples obtained in S130, this step inserts referential citations into the question-and-answer prompt text. This allows the response content to be limited within the scope of the referential citations, thereby ensuring that the question-and-answer process closely follows the traceable evidence context without altering the hierarchical theme and the granular gating. Next, all target nodes are sorted by chapter level and relative node position to generate question-and-answer batches, which are then initiated sequentially. Each question-and-answer session outputs a question-and-answer record entry containing the question text, answer text, source paragraph position, corresponding topic node, and the gating strategy used. To ensure the integrity of the question-and-answer record entries, this step immediately performs a consistency check after each question-and-answer session. The check verifies whether the source identifier and paragraph position identifier in the question-and-answer record entry are consistent with the standardized corpus, and whether the question-and-answer content falls within the acceptable question range. If the check fails, a supplementary question-and-answer session is initiated within the original node range, replacing the original entry.
[0118] S320. Extract efficacy phrases from the first round of questions and answers to obtain efficacy phrase extraction results;
[0119] Following the initial question-and-answer session, this step uses the question-and-answer record entries from the first round as the sole input, and the corpus standardization and semantic vectorization as auxiliary inputs for alignment and retrieval. It performs phrase-level identification, boundary delineation, and source fixation on efficacy-related expressions in the response text, forming a structured efficacy phrase extraction result. Specifically, firstly, the response text in each question-and-answer record entry is segmented into sentences and words are standardized. The boundaries of sentence segmentation are based on the sentence segment numbers in the corpus standardization, ensuring that no segmentation crosses the original paragraph boundaries. Word standardization is based on the terminology unification table in the corpus standardization, unifying the writing forms of synonymous expressions and common variants, so that subsequent phrase boundary identification is not disturbed by synonym differences. Subsequently, potential efficacy phrase candidates are identified from the standardized response text. Candidate identification is based on the response format requirements set in the question-and-answer prompt text, and is cross-checked using the source paragraph positions recorded in the question-and-answer record entries. Candidates that cannot be located within the original source paragraph are not adopted. For candidates that can be located, their corresponding semantic vectorized entries are read, and their semantic similarity with other segments within the same entry is calculated to eliminate candidates with overly broad or narrow coverage. The remaining candidates are marked as efficacy phrase items. Furthermore, to ensure consistency between the text and semantics of efficacy phrase items, this step simultaneously registers the text boundary, paragraph position, and semantic expression for each efficacy phrase item. The text boundary comes from the start and end positions of the original paragraph, the paragraph position comes from the standardized paragraph and sentence numbers of the corpus, and the semantic expression comes from the corresponding entry of the semantic vectorization. All three are stored with the same source identifier for synchronous referencing in subsequent steps. Based on this, duplicate and merged efficacy phrases from the same topic node are performed. The deduplication rule is textual consistency or textual inclusion with highly consistent semantic expression, while the merging rule is textual similarity and semantic similarity. Deduplication and merging only occur within the same node and do not cross nodes to maintain the hierarchical correspondence between nodes and efficacy phrases. After the above processing, an efficacy phrase list is formed for each topic node. Each item in the list has three types of information: text boundary, paragraph position, and semantic expression, and corresponds to a unique question and answer record entry. This step compiles the efficacy phrase list of all nodes into efficacy phrase extraction. The efficacy phrase extraction is stored on a node-by-node basis and maintains bidirectional referencing with the first round of question and answer. To facilitate evidence binding in the next step, this step reserves evidence location identifiers in the efficacy phrase item list for filling in the evidence triples one by one when S330 calls the evidence triples, realizing direct placement from efficacy phrase items to evidence items; at the same time, to ensure a smooth connection with subsequent efficacy dimension aggregation, this step provides a search key for each efficacy phrase item pointing to the semantic vectorized item, so that it can be directly referenced when the efficacy phrase items are aggregated in S320.
[0120] S330. Based on the efficacy phrase extraction and the evidence triple, perform evidence binding to obtain the first round of question and answer, efficacy phrase extraction and evidence binding;
[0121] After extracting the efficacy phrases, this step uses the efficacy phrase extraction and evidence triples as parallel inputs, and uses the corpus standardization and semantic vectorization as the foundation for localization and consistency verification to construct a mapping relationship from efficacy phrase items to evidence items, forming a binding result that can be directly used in subsequent steps. Specifically, firstly, the list of efficacy phrase items is read on a node-by-node basis. For each item in the list, entries with the same position as its source paragraph and overlapping context are retrieved from the evidence triples. Entries that satisfy both source consistency and context overlap are included in the candidate set. For each entry in the candidate set, the source text fragment, related text fragment, and supporting text fragment are read, and compared character by character with the text boundary of the efficacy phrase item. Entries whose complete phrase boundary can be located within the source text fragment or supporting text fragment are marked as suitable entries, and those that cannot be located are discarded. Subsequently, for each functional phrase and its matching set of entries, a binding entry is established. This binding entry records the text boundaries, paragraph positions, and semantic expressions of the functional phrase, as well as the source text fragments, related text fragments, and supporting text fragments in the matching entries, maintaining consistency with the standardized source identifiers of the corpus. To ensure the traceability of the binding entries, this step registers the question-and-answer record entry number and the topic node number on each binding entry, enabling the binding entry to trace back to the question-and-answer source and topic position that generated the functional phrase. Furthermore, for cases where multiple matching entries exist for the same functional phrase, this step sorts them according to the degree of coverage of its text boundaries and the degree of consistency with its semantic expression. The sorting is completed within the same node, and the top one or more entries are marked as primary bindings, while the remaining entries are marked as backup bindings. Both primary and backup bindings are stored in a structured format to ensure quick switching during retrospective calls in subsequent steps. Next, a consistency check is performed on all bound entries within a node. This check includes verifying the consistency of source identifiers, the continuity of paragraph positions, the completeness of supporting text fragments, and the complete presentation of the efficacy phrase within the bound scope. Bound entries that fail consistency checks are returned to the candidate set from the previous step for re-filtering until consistency is achieved or no usable entries are available for that efficacy phrase. After completing the node-level check, all bound entries are aggregated to form a binding set that matches the first round of question answering and the efficacy phrase extraction; that is, the first round of question answering and efficacy phrase extraction are bound to evidence. This set is stored in nodes and maintains a reference relationship with the hierarchical topics and granular gating obtained in S230, enabling subsequent steps to be read within a unified hierarchy and boundary.To ensure smooth transition to subsequent steps, this step generates an aggregation entry point in the binding set that can be read by S320. This entry point points to the efficacy phrase item and its main binding entry within each node, allowing direct referencing of representative binding content during efficacy dimension aggregation. Simultaneously, a counting entry point is generated in the binding set that can be read by S610. This entry point points to all binding entries corresponding to each node and each efficacy phrase item, enabling evidence-by-evidence counting and scope limitation during technical efficacy matrix construction. Furthermore, a coverage verification entry point is recorded in the set that can be read by S620. This entry point points to all efficacy phrase items that failed to complete binding and their respective nodes, used to locate gaps during coverage consistency checks and to feed back gap information to S310 and S320 in S630 for supplementation. At this point, this step outputs the first round of question-and-answer, efficacy phrase extraction, and evidence binding. This output simultaneously possesses a one-to-one or one-to-many mapping relationship between efficacy phrase items and evidence items, and structurally aligns with the source identifier and paragraph position identifier mappings of the hierarchical topics, granular gating, corpus standardization, and semantic vectorization.
[0122] Step S400 includes at least steps S410-S430:
[0123] S410. Obtain the first round of questions and answers, extract efficacy phrases and bind evidence, perform efficacy dimension aggregation, and obtain efficacy dimension aggregation.
[0124] This step uses the first round of question-and-answer sessions, efficacy phrase extraction, and evidence binding obtained in S330 as the sole input, and the hierarchical topics and granular gating obtained in S230 as range constraints. It also uses the corpus standardization formed in step 110 and the semantic vectorization formed in step 120 as the foundation for positioning and alignment. Without altering the existing hierarchical structure and boundary rules, it performs aggregation processing around efficacy phrase items and their corresponding evidence items. Specifically, for each target node in the hierarchical topics, it first reads the list of efficacy phrase items within that node and the evidence items paired with each efficacy phrase item. It then defines the processable range according to the entry threshold and termination conditions in the granular gating, and merges adjacent efficacy phrase items within this range. During the merging process, the semantic similarity relationship indicated by the semantic vectorization items is used as a reference, while maintaining the paragraph order and sentence boundaries in the corpus standardization. Candidates that do not meet the boundary consistency requirements are not included in the merging. Furthermore, after merging adjacent nodes, representative expressions are determined for the resulting groups of efficacy phrases. The representative expression is the one with the most stable coverage and the most sufficient evidence items within the group, which is then used as the core expression. This core expression, along with its corresponding evidence source paragraph, contextual position, and question-and-answer source, is fixed and saved to form a temporary dimensional unit for that group. Further, while maintaining the node hierarchy, highly similar temporary dimensional units between different nodes are compared at the same level. Temporary dimensional units that are highly similar in semantic vectorization and located in adjacent chapters in corpus standardization are merged or paralleled according to the boundary rules for cross-node merging in granularity gating. During merging, the representative expression is unified and the evidence list is expanded; during paralleling, each unit's representative expression is maintained, but mutual references are established to ensure that subsequent calls can expand the scope without disrupting the hierarchical theme. Subsequently, a mapping relationship is generated for each temporary dimension unit. This mapping relationship records the correspondence between efficacy phrases and temporary dimension units, between temporary dimension units and their respective nodes, and between temporary dimension units and evidence entries. It maintains the same source identifier as the paragraph positions in corpus standardization and the expression entries in semantic vectorization, ensuring that subsequent backtracking and cross-referencing can be performed along this mapping. Next, a consistency check is performed on all temporary dimension units within a node. The consistency check includes whether the representative expression completely corresponds to its evidence source, whether the representative expression falls within the processable range, and whether the mapping relationship is self-consistent. If inconsistencies are found, the process reverts to the adjacent merging stage to perform a supplementary merging within the original range, replacing the original entry with the supplementary result. After the above checks pass, a list of node-level dimension units is formed.To ensure continuity with subsequent steps, this step generates two types of entry information on the node-level dimensional unit list. One type of entry points to the representative expression and its main evidence entry for each dimensional unit, used to directly extract definitions and positive / negative examples when performing dimensional knowledge injection on the efficacy dimension aggregation in S420. The other type of entry points to all efficacy phrases and their complete evidence entries within each dimensional unit, used for evidence-by-evidence counting and scope limitation when constructing the technical efficacy matrix in S610. The output of this step is the efficacy dimension aggregation, which consists of the node-level dimensional unit list, its mapping relationships, and the entry information. It maintains boundary constraints consistent with hierarchical topics and granular gating, and source correspondence consistent with corpus standardization and semantic vectorization. The result will serve as one of the sole inputs for dimensional knowledge injection in the next step, and will be continuously referenced in subsequent rounds of question answering, dimension alignment and synonym merging, and technical efficacy matrix construction.
[0125] S420. Perform dimensional knowledge injection on the aforementioned efficacy dimension aggregation to obtain dimensional knowledge injection;
[0126] Following the aforementioned aggregation of efficacy dimensions, this step uses the list of node-level dimensional units in the efficacy dimension aggregation as the sole input, and the first round of question-and-answer sessions, efficacy phrase extraction, and evidence binding obtained in S330 as the source of instances. The hierarchical topics and granular gating obtained in S230 serve as the scope constraints, and steps 110 and 120 respectively act as the backtracking foundation for text and semantics. Dimensional knowledge entries are constructed around each dimensional unit, which can be directly invoked by subsequent question-and-answer sessions and alignment. Specifically, the representative expression and main evidence entry of each dimensional unit are first read. Based on the location interval of the representative expression in its main evidence entry and the preceding and following text adjacent to that interval, a context summary that illustrates the boundary characteristics is formed. This context summary is only extracted within the processable interval of the node it belongs to, without crossing the boundary defined by the granular gating. Simultaneously, the position of the source paragraph related to this context summary is recorded for subsequent source verification in different steps. Furthermore, without altering the representative text, several fragments are selected from all evidence entries associated with this dimension unit as positive and negative examples. Positive examples indicate typical phrases consistent with the representative expression, while negative examples indicate easily confused phrases similar to the representative expression but not belonging to this dimension unit. Each example fragment is appended with its source location and question-and-answer source marker to ensure no confusion occurs during subsequent citations. Further still, combining the distribution of efficacy phrases within the same dimension unit with the representative expressions of adjacent dimension units, a set of boundary prompts is compiled. These boundary prompts are used to limit the scope of questions and responses in subsequent rounds of question-and-answer sessions, preventing cross-dimensional diffusion. Simultaneously, combining the parent-child and sibling relationships within the hierarchical topics, the subordinate relationship with the upper-level dimension and the parallel relationship with the same-level dimension are recorded for this dimension unit, providing a calling path when cross-level retrieval is required later. Subsequently, after compiling the context summary, positive and negative examples, and boundary prompts, this information, along with representative expressions, source locations, question and answer sources, a list of efficacy phrases, and mapping relationships, is encapsulated into dimensional knowledge entries. Each dimensional knowledge entry corresponds one-to-one with a dimensional unit, and placeholder information for entry strategies, jump conditions, and termination conditions is reserved in the structure. This placeholder information will be filled in the next step based on the boundary constraint update results. Next, to ensure efficient retrieval of dimensional knowledge entries in subsequent steps, this step establishes a retrieval key pointing to the semantic vectorized entry and a position key pointing to the corpus-standardized entry for each dimensional knowledge entry. Both types of keys are consistent with the source, ensuring that alignment can be completed on both the textual and semantic sides in any call. Simultaneously, the corresponding layer depth and entry threshold for granular gating are registered on each dimensional knowledge entry, allowing subsequent calls to directly determine whether to allow entry for questioning or alignment based on the gating.After the dimensional knowledge entries are generated, this step performs a consistency check on the set of entries. The consistency check includes whether the representative expression matches the context summary, whether the positive and negative examples come from evidence entries related to the dimensional unit, whether the boundary prompts are consistent with the boundary rules of the granularity gating, and whether the key values completely correspond to the source. If the check fails, the entry information of the dimensional unit is backtracked to re-extract the required fragments and replace the corresponding content in the entries. After the check is completed, the dimensional knowledge injection is output. The dimensional knowledge injection consists of a set of dimensional knowledge entries oriented towards the dimensional unit, and maintains a one-to-one correspondence with the efficacy dimension aggregation, and maintains the same hierarchical and boundary constraints as the hierarchical topic and granularity gating. This output will be directly called in the boundary constraint update in the next step, and will also be read as a prompt basis during the second round of question answering in S510. It will be continuously used as a constraint and reference during the dimensional alignment, synonym merging, and consistency mapping processing in S520 and S530, and finally used as the domain anchor point for column space and evidence-by-evidence counting when constructing the technical efficacy matrix in S610.
[0127] S430. Based on the dimensional knowledge injection and the efficacy dimension aggregation, perform boundary constraint update to obtain the efficacy dimension aggregation and dimensional knowledge injection.
[0128] After constructing the dimensional knowledge entries, this step uses dimensional knowledge injection and efficacy dimension aggregation as parallel inputs. It continues to use hierarchical topics and granular gating as the scope framework, and corpus standardization and semantic vectorization as the backtracking foundation. It performs unified boundary constraint updates around the boundary consistency of dimensional units and cross-dimensional relationships, ensuring consistency between the aggregation structure and knowledge entries in terms of boundaries, scope, and invocation paths. Specifically, it first reads the context summary and boundary prompts in each dimensional knowledge entry, comparing them with the representative expression and evidence coverage of the corresponding dimensional unit in the efficacy dimension aggregation. If the context summary covers a segment exceeding the processable range defined by the granular gating, it uses the entry threshold and termination conditions in the granular gating to shrink the scope of the efficacy phrases covered by the representative expression of that dimensional unit. After shrinking, it synchronously adjusts the corresponding entries in the mapping relationship to ensure that the representative expression, efficacy phrases, and evidence entries are presented within the same scope. Furthermore, for cases where different dimensional units are similar in semantic vectorization and adjacent but not completely overlapping in corpus standardization, this step cross-checks the positive and negative examples recorded in the dimensional knowledge entries. If the positive examples of both dimensional units can be placed in the context summary of the other and do not violate the boundary rules of granular gating, then it is determined that there is a risk of overlap between the two. At this time, a fine-tuning is carried out within the original node range guided by the boundary prompt. The fine-tuning methods include shrinking the phrases representing the expression, adjusting the assignment of efficacy phrase items, and switching the primary and backup evidence items. After the fine-tuning is completed, the records of the parallel or subordinate relationships between the two are updated. Furthermore, to ensure that dimensional knowledge entries provide clear entry paths when directly invoked in subsequent question-and-answer and alignment processes, this step, after completing boundary consistency and overlap resolution, completes placeholder information for entry strategies, jump conditions, and termination conditions for each dimensional knowledge entry. Entry strategies are set according to the depth of granular gating and entry thresholds; jump conditions are set based on the parent-child and sibling relationships of hierarchical topics; and termination conditions are set based on the completeness of evidence entries' coverage. After these three types of information are completed, they are written back to the dimensional knowledge entry, and a one-to-one update marker is established between the dimensional unit and the knowledge entry to ensure that only the updated content is used in subsequent invocations. Subsequently, a global consistency review is performed at the dimensional level. The review includes whether the representative expressions of all dimensional units fall within the processable range, whether all mapping relationships are self-consistent, whether the boundary descriptions of all dimensional knowledge entries and their corresponding dimensional units are consistent, and whether all entry strategies and jump paths can be executed under hierarchical topics and granular gating. If the review fails, the corresponding dimensional unit is rolled back to the previous processing stage for local correction and rewriting until the review passes.After the review is completed, the output is the aggregation of the effectiveness dimension and the injection of dimensional knowledge. The output is an aggregated structure and a set of knowledge items updated with boundary constraints. The two are consistent in structure and boundaries, and the source and positional relationships with hierarchical topics, granular gating, corpus standardization, and semantic vectorization remain unchanged. The calling path of this output in subsequent steps is as follows: In the second round of question answering in S510, the representative expressions, context summaries, and boundary prompts in the dimensional knowledge items are directly used as the basis for asking questions and the scope of the response; In the dimension alignment, synonym merging, and consistent mapping processing in S520 and S530, synonym merging and consistent mapping are performed with the updated mapping relationship and entry strategy as constraints; In the construction of the technical effectiveness matrix in S610, the column space and the range of certificate-by-certificate counting are determined by the effectiveness dimension aggregation, and the positive and negative examples recorded in the dimensional knowledge items are used as the reference for counting verification; In the coverage consistency check and feedback feedback in S620 and S630, the records of global consistency review are used as the verification entry and feedback guide.
[0129] Step 500 includes at least steps S510-S530:
[0130] S510. Obtain the efficacy dimension aggregation and dimension knowledge injection, perform a second round of question and answer, and obtain the second round of question and answer and dimension alignment.
[0131] This step uses the list of dimensional units and mapping relationships in the efficacy dimension aggregation as the sole semantic entry point, and the dimensional knowledge items in the dimensional knowledge injection as the basis for question answering and alignment prompts. At the same time, the hierarchical topics and the granularity gating are used as range constraints, and the corpus standardization and semantic vectorization are used as the foundation for text backtracking and expression alignment. When necessary, the source information recorded in the first round of question answering, efficacy phrase extraction and evidence binding is referred to to ensure that the question answering and alignment process is carried out within the same source coordinate system. Specifically, firstly, for each dimension unit in the efficacy dimension aggregation, the corresponding dimension knowledge entry is read. The dimension knowledge entry includes structured content such as representative expressions, context summaries, positive and negative examples, boundary prompts, entry strategies, jump conditions, and termination conditions. Then, within the processable range defined by the granularity gating, the target node and question-and-answer order for this question and answer are determined according to the entry strategy, based on the parent-child and sibling relationships of the hierarchical topics. Question texts corresponding to each dimension unit are generated. The question texts carry representative expressions and boundary prompts, and mark the source paragraph position and context range required for the response, so that the response can be located in both content and position. Next, a response is initiated for each question text, resulting in a question-and-answer record entry containing the response text, source identifier, paragraph position, context start and end, and representative expression. To ensure the usability of the record entries, this step immediately checks whether the paragraph position and context start and end are within the processable range based on the corpus standardization, and checks the semantic consistency between the response text and the representative expression based on the semantic vectorization. If it is found that the response text is outside the processable range or the semantic consistency is insufficient, a supplementary question-and-answer session is performed within the same target node according to the jump conditions, and the original record entry is replaced. Furthermore, after the question-and-answer record entries are formed, this step performs alignment processing according to the dimension unit: using the representative expression as the anchor point, key segments with semantic consistency with the representative expression are marked within the context range of the response text, and these key segments are bidirectionally bound to the paragraph position in the corpus standardization and the expression entry in the semantic vectorization, forming an aligned segment that matches the dimension unit. The generation of the aligned segment follows the termination conditions, does not cross boundaries or paragraphs, and maintains a one-to-one or one-to-many correspondence with the function phrase item of the dimension unit. After processing all target nodes, the question-and-answer record entries and aligned fragments are aggregated to form a second-round question-and-answer and dimension alignment. Internally, data is organized by dimension unit, and the source identifiers and positional mapping relationships between the efficacy dimension aggregation, the dimension knowledge injection, the hierarchical topics, the granularity gating, the corpus standardization, and the semantic vectorization are preserved. This output serves as the direct input for the next step of synonym merging of the second-round question-and-answer and dimension alignment, and is also reserved and directly searchable as a column space candidate and counting entry point for subsequent technical efficacy matrix construction.
[0132] S520. Perform synonym merging on the second round of question answering and dimension alignment to obtain the synonym merging result;
[0133] Building upon the previous round of question-and-answer and dimension alignment, this step uses the question-and-answer records and alignment fragments as parallel inputs. It uses representative expressions, positive and negative examples, and boundary prompts recorded in the dimension knowledge injection as the judgment criteria, continuously adhering to the boundary constraints of the hierarchical topics and granular gating. Simultaneously, it uses corpus standardization and semantic vectorization as the verification basis for textual and semantic consistency, conducting synonym merging around the diverse expressions within each dimension unit. Specifically, firstly, all efficacy-related expressions from the previous round of question-and-answer and dimension alignment are collected within each dimension unit. These expressions are all bound to source identifiers, paragraph positions, and alignment fragments. Then, based on representative expressions and positive and negative examples, expressions consistent with or highly similar to the representative expressions are marked as priority candidates, while expressions similar to negative examples or crossing boundary prompts are marked as exclusion candidates. A position check is then performed within the paragraph level of the corpus standardization; expressions outside the processable range of that dimension unit are not included in the merging scope. Next, the priority candidate sequences are merged: when two expressions are completely identical in text, they are directly merged into a single standard expression; when two expressions have an inclusion relationship in text and are highly consistent in semantic vectorization, they are merged into a standard expression with broader coverage, and the merged entry is recorded as a subordinate expression; when multiple expressions are similar but have slightly different boundaries, they are fine-tuned within their original paragraph positions based on boundary prompts, unified into a standard expression with consistent boundaries, and the original aligned fragment is retained as supplementary evidence. Furthermore, to prevent cross-node merging from disrupting the boundaries between the hierarchical topic and the granularity gating, this step only performs synonym merging within the same dimensional unit and its belonging node, without crossing sibling nodes or parent-child nodes; when cross-node referencing is necessary, only the referencing relationship is established without changing the attribution. After completing the merging within the unit, a list of standard and subordinate expressions is formed for each dimensional unit. Each entry in the list is bound to the source identifier, paragraph position, aligned fragment, and semantic expression, and the merging path is recorded in a structured manner for complete invocation in subsequent consistent mapping processing. Subsequently, a consistency check is performed at the dimensional unit level. This check includes verifying whether the standard expression falls within the processable range of the dimensional unit, whether all subordinate expressions are correctly positioned in the corpus standardization, and whether all expressions correspond one-to-one with the expression entries in the semantic vectorization. Lists failing the check are reverted to the aforementioned merging process for partial correction and re-checking. After the check is complete, a synonym merge is output. This synonym merge is organized by dimensional units, containing lists of standard and subordinate expressions, their source mappings, and merging paths, and establishing bidirectional references with the second-round question-and-answer and dimensional alignment. This output will be directly invoked in the next step based on the synonym merge and the second-round question-and-answer and dimensional alignment consistency mapping process, while also being reserved as a unified entry point for the column space standard expressions and evidence counting during subsequent technical efficacy matrix construction.
[0134] S530. Based on the synonym merging and the second round of question answering and dimension alignment, perform consistent mapping processing to obtain the second round of question answering, dimension alignment and synonym merging;
[0135] After merging the efficacy-related expressions, this step uses the list of standard and subordinate expressions from the synonym merging as the main input, and the question-and-answer record entries and alignment fragments from the second round of question-and-answer and dimension alignment as secondary input. It continuously references the list of dimension units and mapping relationships from the efficacy dimension aggregation, and is constrained by the boundaries of the hierarchical topics and granular gating. The corpus standardization and semantic vectorization serve as the verification basis for position and expression, constructing a unified mapping from standard expressions to dimension units, efficacy phrase items, evidence items, and topic nodes. Specifically, first, the list of standard and subordinate expressions is read within each dimension unit, and the response text and its alignment fragment corresponding to each expression are located in the question-and-answer record entries. Using the standard expression as the primary key, all subordinate expressions and their alignment fragments within that dimension unit are linked one by one, forming a mapping entry centered on the standard expression. Subsequently, based on the mapping relationships in the efficacy dimension aggregation, each mapping entry is linked to the corresponding efficacy phrase item and evidence item, ensuring that any standard expression can be traced back along the mapping path to the original efficacy phrase item and its evidence source, and is positioned in the corpus standardization and aligned in the semantic vectorization. Furthermore, to handle situations where standard representations cross-reference within sibling or parent-child nodes, this step distinguishes between merge mappings and reference mappings based on the structural relationships of the hierarchical topics and the entry threshold of the granularity gating: merge mappings are established for parallel elements at the same level with consistent boundaries, using the same standard representation and merging evidence entries; reference mappings are established for elements with only vertical relationships or inconsistent boundaries, retaining their respective standard representations and allowing mutual references within the mapping entries without changing their affiliation. Next, a global consistency review is performed at the dimensional unit set level. This review includes checking whether each standard representation corresponds to only one or a set of explicit dimensional units, whether each subordinate representation has been attached to the primary key, whether each evidence entry is counted only within the allowed column space, and whether each mapping is compatible with the boundary rules of the hierarchical topics and the granularity gating. Entries that fail the review are reverted to the aforementioned attachment and differentiation steps for revision, and the revised entries replace the original records. Subsequently, to ensure seamless integration with subsequent processing, this step generates three types of callable entry points in the consistent mapping results: the first is a counting entry point for S610, pointing to all evidence entries and their source locations under each standard representation, for evidence-by-evidence counting and column space limitation during the construction of the technical efficacy matrix; the second is a verification entry point for S620, pointing to entries for which no merge mapping has been established or only a reference mapping has been established, for locating potential gaps and unclear boundary locations during coverage consistency checks; and the third is a feedback entry point for S630, pointing to dimension units and representation pairs that still have ambiguities after review, for guiding supplementary questions and answers or boundary fine-tuning during feedback feedback.After completing the above processing, the output includes second-round question answering, dimension alignment, and synonym merging. The output structurally unifies synonym merging, second-round question answering, and dimension alignment, and maintains consistent source identification and position mapping with the efficacy dimension aggregation, dimension knowledge injection, hierarchical topics, granular gating, corpus standardization, and semantic vectorization. This allows it to be directly read and executed in subsequent statistical construction and verification backflow.
[0136] Step S600 includes at least steps S610-S630:
[0137] S610. Obtain the second round of question answering, dimension alignment, synonym merging, and hierarchical topics, and construct a technical efficacy matrix to obtain the technical efficacy matrix construction.
[0138] This step uses the second round of question-and-answer, dimension alignment, and synonym merging as the sole source of the column space, and the hierarchical topic as the sole source of the row space. Throughout the process, it continuously adheres to the boundary constraints of the granular gating, using corpus standardization and semantic vectorization as the foundation for text positioning and expression alignment. Simultaneously, it references the source identifiers recorded in the evidence triples and the first round of question-and-answer, efficacy phrase extraction, and evidence binding to ensure that counting units can be traced back item by item. Specifically, firstly, the target node set in the hierarchical topic is read. This target node set consists of processable nodes after consistent integration, maintaining a one-to-one or one-to-many correspondence with the paragraph positions of the corpus standardization and the expression items of the semantic vectorization. Simultaneously, the standard and subordinate expression lists, question-and-answer record items, and aligned fragments from the second round of question-and-answer, dimension alignment, and synonym merging are read to establish a correspondence with the dimension units of the efficacy dimension aggregation, enabling any expression to be traced back along the mapping path to the representative expression, efficacy phrase item, and evidence item. Subsequently, within each target node, based on the entry threshold and termination conditions defined by the granular gating, the standard and subordinate expressions associated with that node are traversed one by one. Aligned segments that can be placed within the original paragraphs of the corpus standardization and are consistent with the semantic vectorized expression are included in the counting range. To ensure the verifiability of the evidence-by-evidence counting, this step only accepts entries that have the same source and overlapping context with the evidence triple, and simultaneously registers the source identifier, paragraph position, context start and end, and expression type in the counted entries. Furthermore, after completing the collection of entries within a node, multiple aligned segments of the same expression within the same node are deduplicated and merged. Deduplication is based on completely consistent text boundaries, and merging is based on text inclusion and consistent source. After merging, the segment that best represents the context is retained as the main entry, and the remaining segments are registered as reserve entries for subsequent verification. Repeated citations across paragraphs are not included in the statistics to avoid boundary conflicts with the granular gating. Next, with both row and column spaces fixed, a set of counting units is formed, using nodes and descriptions as coordinates. Each counting unit corresponds to a set of main entries and several backup entries. Row and column indexes are generated on the set, allowing any unit to be directly retrieved. To reduce offsets caused by differences in length and source concentration, this step corrects the counting units based on the length of the paragraph and the distribution of sources without changing the content of the entries. The correction results are stored together with the original count values. Simultaneously, all counting units are arranged according to node order and description order to generate a unit list for constructing the technical efficacy matrix. The list retains back references to the efficacy dimension aggregation and the dimension knowledge injection entries for cross-checking in subsequent verification and backflow.After completing the above processing, a technical efficacy matrix is output. The technical efficacy matrix consists of a list of counting units, row indexes, column indexes, and source mappings. It maintains a consistent source identification and position mapping relationship with the hierarchical topics, the second round of question answering and dimension alignment and synonym merging, the efficacy dimension aggregation, the dimension knowledge injection, the corpus standardization, and the semantic vectorization. This output will serve as the sole input for the next step of performing a coverage consistency check on the technical efficacy matrix construction, and will also provide a structured carrier for the subsequent output of a visualization report.
[0139] S620. Perform a coverage consistency check on the constructed technical efficacy matrix to obtain the coverage consistency check result;
[0140] Following the aforementioned construction of the technical efficacy matrix, this step uses the list of counting units, row index, and column index as the sole inputs. It continuously references the hierarchical topics and granular gating as scope boundaries, the efficacy dimension aggregation and dimensional knowledge injection as semantic and instance references, and the corpus standardization and semantic vectorization as the verification base for position and expression. The coverage and consistency relationships are checked layer by layer. Specifically, first, the coverage is checked in the row space according to the node order. The processable range and parent-child and sibling relationships of each target node are read. It is checked whether there are counting units under the node. If not, a gap marker is generated on the node, and the source position, parent node, and sibling node of the node are recorded in the gap marker to clarify the subsequent backflow direction. For nodes with counting units, the distribution of their main entries and backup entries is statistically analyzed. It is checked whether the distribution is concentrated in a small number of sources. If it is excessively concentrated, it is marked as a source concentration marker to indicate the key scope for subsequent verification. Subsequently, the coverage is checked by dimension unit within the column space. The count distribution of each standard representation and subordinate representation on different nodes is read to check for any missing representations or representations that are only located on a single node. Missing representations are marked with column gaps, and representations that are only located on a single node are marked with weak coverage. For representations with column gaps and weak coverage, the context summary and positive and negative examples in the dimensional knowledge injection are further checked. If the context summary indicates a processable interval that is not located in the matrix, the indication is written into the gap mark so that it can be used as a basis for questioning during subsequent reflow. Furthermore, to verify consistency, this step compares the mapping paths of the second round of question-and-answer and dimension alignment and synonym merging within the row-column intersection counting unit. It checks whether the main entry of the counting unit is consistent with the standard expression, whether the backup entry is consistent with the subordinate expression, and whether the source of the entry is consistent with the evidence triple. If there is any inconsistency between the main entry and the standard expression, or between the source and the evidence entry, or if the entry goes out of bounds, an alignment conflict mark, a source conflict mark, and a boundary conflict mark are generated respectively. The backtracking path is written in the mark, including the source position of the corresponding node, the corresponding dimension unit, the corresponding question-and-answer record, and the corresponding evidence entry, so as to ensure that it can be directly located later. Next, the aforementioned gap markers, source concentration markers, weak coverage markers, and three types of conflict markers are summarized to form a problem list for coverage consistency verification. The list is then grouped and sorted according to problem type and severity. To facilitate subsequent processing, this step also generates three types of entry points on the problem list: one type points to the target nodes and expressions that need to be supplemented in the first and second rounds of question and answer, triggering additional question and answer during reflow; another type points to the dimension units and boundary prompts that need to be adjusted in the effectiveness dimension aggregation and the dimension knowledge injection, triggering boundary fine-tuning during reflow; and the third type points to the counting units that need to be corrected at the counting level, triggering recounting and replacement during reflow.After completing the above verification and organization, a coverage consistency check is output. The coverage consistency check consists of a problem list and various tags, various entry points and global verification records. It maintains an item-level one-to-one mapping relationship with the construction of the technical effectiveness matrix, and maintains a node-level and dimension-level consistency relationship with the hierarchical topics, the granular gating, the effectiveness dimension aggregation and the dimension knowledge injection. This output will serve as the direct basis for the next step of feedback flow based on the coverage consistency check and the construction of the technical effectiveness matrix.
[0141] S630. Based on the coverage consistency test and the technical efficacy matrix, a feedback loop is performed to obtain the feedback loop.
[0142] After completing the coverage and consistency verification, this step uses the issue list and various markers as trigger objects, and the count unit list and row and column indexes as revision objects. Following the boundary constraints of the hierarchical topics and granular gating, it calls the entry information of the first round of question-and-answer, the second round of question-and-answer, and the dimensional knowledge injection, forming a set of return instructions that can be directly executed by the preceding steps. If necessary, it also performs local updates to the technical effectiveness matrix construction. Specifically, firstly, for gap markers and weak coverage markers, based on the recorded node positions and dimensional units, it generates additional question-and-answer instructions for the first round and the second round of question-and-answer. These additional question-and-answer instructions include the target node, the questionable range, representative expressions, and boundary prompts. When the instructions are passed to S310 and S510, they can directly initiate question-and-answer sessions and return new question-and-answer record entries and aligned fragments. After returning, this step, without changing the original node boundaries, attaches the new entries to the corresponding dimensional units and performs supplementary counting and placement replacement on the relevant units at the counting level. Subsequently, for source set markers and source conflict markers, based on the source location and evidence entries in the markers, boundary fine-tuning instructions are generated for the aggregation of the efficacy dimension and the injection of the dimension knowledge. The boundary fine-tuning instructions include representative expressions that need to be shrunk or expanded, positive and negative examples that need to be replaced or supplemented, and entry strategies, jump conditions and termination conditions that need to be updated. When the instructions are passed to S320, S420 and S430, fine-tuning and write-back can be performed directly. After receiving the write-back, this step updates the original dimension units and knowledge entries and performs synchronous replacement on the row and column index accordingly. Furthermore, for alignment conflict markers and boundary conflict markers, merge and dispatch instructions are generated for the second round of question answering, dimension alignment, and synonym merging. These instructions include subordinate expressions to be merged, standard expressions to be split, primary and backup entries to be switched, and reference relationships to be established. When the instructions are passed to S520 and S530, merging, splitting, and mapping updates can be directly executed. After receiving the update, this step recounts the relevant counting units, removes out-of-bounds entries, and replaces them with new entries, ensuring that the count is self-consistent in the new column space. Next, to maintain the traceability of the reflow process, this step registers the source identifier and time sequence on each reflow instruction and each local update, and writes update records into the counting unit list constructed by the technical effectiveness matrix, so that any reflow can be fully reviewed. Simultaneously, to avoid repeated triggering of reflow at the same node, this step adds processing markers to the issue list, marking processed and unprocessed entries, and triggers a new round only for unprocessed entries in the next round of verification.After completing the above instruction generation, execution, and synchronous update, a feedback loop is output. The feedback loop consists of append question instructions, boundary fine-tuning instructions, and merging and assignment instructions and their execution records. It maintains a one-to-one mapping with the construction of the technical efficacy matrix and the coverage consistency check, and is structurally connected with the hierarchical topics, the granular gating, the efficacy dimension aggregation, the dimension knowledge injection, the second round of question answering, and the dimension alignment and synonym merging.
Claims
1. A method for constructing an efficacy matrix based on topic clustering and two-round question answering, characterized in that, include: The patent text is obtained, and the title, abstract, claims and description are hierarchically divided according to fixed fields. Corpus standardization, semantic vectorization and evidence triple construction are performed to obtain evidence triples that retain paragraph numbers and claim numbers. Hierarchical topic generation is performed based on semantic vectorization and corpus standardization, establishing a hierarchical structure of parent-child and sibling relationships. Granular gating calculation is performed by combining layer depth threshold and entry threshold. Consistency integration is completed through node coverage range and boundary stability verification to obtain hierarchical topics and granular gating. In terms of hierarchical topics and granular gating, the first round of question and answer is executed based on the entry threshold in the granular gating, extracting efficacy phrases and completing evidence binding through text boundaries, paragraph positions and semantic expressions; Based on the first round of question and answer and evidence binding, adjacent merging of efficacy phrases within the nodes of the hierarchical topic is performed to generate temporary dimension units, and dimension knowledge entries containing positive and negative examples are injected. Boundary constraints are updated through boundary prompts and layer depth thresholds. Based on efficacy dimension aggregation and dimension knowledge injection, the second round of question and answer is completed through representative expressions and boundary prompts in dimension knowledge items. Dimension alignment is performed by combining semantic vectorized item alignment fragments. Synonym merging is completed through text inclusion relationship and semantic high consistency rules, and consistent mapping processing of primary key and subordinate expression is performed. Based on the second round of question answering, dimension alignment and synonym merging, a technical effectiveness matrix is constructed by combining the row space nodes of hierarchical topics. Coverage consistency is checked by counting each certificate and verifying the consistency of sources. Feedback is completed by triggering additional questions or boundary fine-tuning based on gap markers and conflict markers.
2. The method according to claim 1, characterized in that, Hierarchical topic generation based on semantic vectorization and corpus standardization includes: Based on semantic vectorization, adjacent processing units from the same patent text and with similar semantics are clustered. During the clustering process, the paragraph position and chapter level recorded in the corpus standardization are used as boundary constraints to avoid disordered merging across chapters. The candidate set formed by each clustering and its corresponding source identifier and position identifier are fixed at the same time. Repeated or synonymous segments are uniformly merged, while maintaining the original order and subordinate relationship, and the original position is recorded in the merged entries; The representative fragment and its context fragment are included together as the core content of the set. Based on this core content, the inclusion relationship, parallel relationship and inheritance relationship between different candidate sets are determined, thus forming a hierarchical structure of parent and child and siblings. A bidirectional mapping to corpus standardization is established at each node of the hierarchy, so that any topic node can be located to the precise position of the original text, and can be located from the original text back to the corresponding topic node. Record text fragments that co-occur with evidence triples within the hierarchical nodes; The density, span, and coverage of nodes at each level are checked for consistency. If nodes with excessively wide or narrow coverage are found, they are fine-tuned without disrupting the original chapter boundaries.
3. The method according to claim 1, characterized in that, Granular gating calculations combining layer depth thresholds and entry thresholds include: Based on the representative fragments, context fragments and the mapping to the corpus standardization at each node of the hierarchical topic, the coverage and boundary stability of that node are determined. For nodes with uneven coverage or weak boundary stability, small-scale merging or splitting within the hierarchy can be performed to ensure that the description of the node is consistent with the text fragment it maps to. Using the evidence triples that have a reference relationship with the node as a reference, the source text fragments and related text fragments in the evidence triples are labeled in the context of the node. The labeling results are used to express the questionable and answerable range of the node in subsequent question and answer. In the vertical direction of the hierarchy, the depth distribution of each layer is calculated based on the parent-child relationship, and in the horizontal direction, the relative distance between nodes in the same layer is calculated based on the sibling relationship. Based on the vertical and horizontal relationships, several gating boundaries are determined without introducing new semantics. By combining the chapter tags in the corpus standardization, gating constraints are set for nodes in different chapters, so that nodes in the abstract scope, nodes in the claim scope, and nodes in the specification scope have different entry strategies when they are called in the future. Nodes that co-occur frequently with evidence triples are marked as priority entry targets; The gating layer depth threshold, entry threshold, and boundary rules are uniformly encapsulated into a readable structural description.
4. The method according to claim 1, characterized in that, Consistent integration is achieved through node coverage and boundary stability checks, including: Read the parent-child and sibling relationships of hierarchical topics and the bidirectional mapping to corpus standardization. At the same time, read the layer depth threshold, entry threshold and boundary rules in the granular gating and compare each node item by item in the same structural view. If it is found that the coverage of a node is partially excluded by the gating rules, the coverage description of the node is adjusted to match the entry threshold of the granular gating without changing the parent-child and sibling relationships of the hierarchical topics. If a slight misalignment is found between the node boundary and the gating boundary, the boundary is fine-tuned based on the paragraph position in the corpus standardization. Establish a consistency label at the node level. The consistency label is used to record the synchronization status of the node in terms of hierarchical structure, gating boundary and text mapping. For nodes that need to be corrected, the correction will be completed within this integration, and the consistency label will be updated after the correction is completed; A list of evidence citations is created for each node, which is derived from the annotation results of the evidence triples in the context of that node. Generate a set of constraint rules, which describes the entry order, jump conditions, and termination conditions that should be followed when a node is invoked to enter the question-and-answer and extraction process. The integration results record the mapping integrity identifier between each node and corpus standardization, the expression integrity identifier between each node and semantic vectorization, and the citation integrity identifier between each node and the evidence triple.
5. The method according to claim 1, characterized in that, Evidence binding is accomplished by linking text boundaries, paragraph positions, and semantic expressions, including: Read the list of efficacy phrase items by node. For each item in the list, retrieve the item in the evidence triple that is consistent with the position of its source paragraph and has overlapping context. For each entry in the candidate set, read the source text fragment, associated text fragment, and supporting text fragment in the entry, and compare them character by character with the text boundary of the function phrase item. Any entry that can locate the complete phrase boundary within the source text fragment or supporting text fragment is marked as a matching entry. For each function phrase and its set of matching entries, a binding entry is created. The binding entry records the text boundaries, paragraph positions and semantic expressions of the function phrase, as well as the source text fragments, related text fragments and supporting text fragments in the matching entry. Register the question and answer record entry number and topic node number on each bound entry; When there are multiple matching entries for the same function phrase, they are sorted according to the degree of coverage of their text boundaries and the degree of consistency with their semantic expression. The top one or several entries are marked as primary bindings, and the remaining entries are marked as backup bindings. Perform a consistency check on all bound entries within the node. The consistency check includes whether the source identifier is consistent, whether the paragraph position is continuous, whether the supporting text fragment is complete, and whether the function phrase item is fully presented within the bound range. All the binding entries of all nodes are collected to form a binding set that matches the first round of question and answer and the extraction of efficacy phrases.
6. The method according to claim 1, characterized in that, Within the nodes of a hierarchical topic, adjacent merging of efficacy phrases is performed to generate temporary dimension units, including: For each target node in the hierarchical topic, read the list of efficacy phrase items and the evidence items paired with each efficacy phrase item in that node. Define the processable range according to the entry threshold and termination condition in the granular gating. Based on this, merge adjacent efficacy phrase items within the range. During the merging process, the semantic similarity relationship indicated by the semantic vectorized items is used as a reference, while keeping the paragraph order and sentence boundaries in the corpus standardization from being crossed. After completing the adjacent merging within the node, the representative expression of the merged efficacy phrase group is determined. The representative expression is the one with the most stable coverage and the most sufficient evidence items in the group as the core expression. The core expression, along with its corresponding evidence source paragraph, context position, and question and answer source, is fixed and saved to form a temporary dimension unit of the group. While maintaining the node hierarchy, temporary dimension units that are highly similar between different nodes are compared at the same level. Temporary dimension units that are highly similar in semantic vectorization and are in adjacent chapters in corpus standardization are merged or paralleled according to the boundary rules for cross-node merging in granular gating. Generate a mapping relationship for each temporary dimension unit. The mapping relationship records the correspondence between the efficacy phrase item and the temporary dimension unit, the correspondence between the temporary dimension unit and its node, and the correspondence between the temporary dimension unit and the evidence item. Perform a consistency check on all temporary dimension units within a node. The consistency check includes whether the representative expression corresponds completely to its source of evidence, whether the representative expression is located within the processable range, and whether the mapping relationship is self-consistent. Two types of entry information are generated on the node-level dimensional unit list. One type of entry points to the representative expression and its main evidence item of each dimensional unit, and the other type of entry points to all efficacy phrase items and all evidence items within each dimensional unit.
7. The method according to claim 1, characterized in that, Updating boundary constraints using boundary prompts and layer depth thresholds includes: Read the context summary and boundary prompts in each dimension knowledge item one by one, and compare them with the representative expression and evidence coverage of the corresponding dimension unit in the efficacy dimension aggregation. If it is found that the paragraph covered by the context summary exceeds the processable range limited by the granular gating, then based on the entry threshold and termination condition in the granular gating, the range of efficacy phrases covered by the representative expression of the dimension unit is narrowed. For cases where different dimensional units are similar in semantic vectorization and adjacent but not completely overlapping in corpus standardization, cross-checking is performed according to the positive and negative examples recorded in the dimensional knowledge entries. If the positive examples of the two dimensional units can be placed in the context summary of the other and do not violate the boundary rules of granular gating, then it is determined that there is a risk of overlap between the two. Guided by boundary prompts, make a fine-tuning within the original node range. Fine-tuning methods include shrinking the phrases representing the expression, adjusting the allocation of efficacy phrases, and switching the primary and backup evidence items. For each dimension of knowledge entry, placeholder information for entry strategy, jump conditions and termination conditions is added. The entry strategy is set according to the depth of granular gating and the entry threshold. The jump conditions are set according to the parent-child and sibling relationships of the hierarchical topics. The termination conditions are set according to the completeness of the coverage of the evidence entries. A global consistency review is conducted at the dimensional level. The review includes whether the representative expressions of all dimensional units fall within the processable range, whether all mapping relationships are self-consistent, whether the boundary descriptions of all dimensional knowledge items and their corresponding dimensional units are consistent, and whether all entry strategies and jump paths can be executed under hierarchical topics and granular gating.
Citation Information
Patent Citations
Hydropower station knowledge graph construction method and system
CN119990274A
Machine learning intelligent question and answer analysis method and question and answer system based on government affair service
CN120198086A