Policy tracing dynamic retrieval method based on multi-level intention recognition and contrastive learning
Through the method of multi-level intent recognition and comparative learning, the problems of insufficient relevance and accuracy in existing policy tracing methods are solved, and efficient and accurate retrieval and generation optimization of complex policy texts are achieved, which is suitable for scenarios such as policy formulation and compliance assessment.
Patent Information
- Application Number
- CN202411953517.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing policy traceability and information retrieval methods rely on keyword matching or rule-based retrieval models, resulting in insufficient relevance and accuracy of retrieval results, making it difficult to process complex policy texts and provide deep semantic understanding and dynamic optimization.
A multi-level intent recognition and contrastive learning method is adopted, and text data processing and high-dimensional semantic vector representation are performed through a pre-trained semantic embedding model. Combined with multiple rounds of dynamic retrieval and fuzzy semantic extension matching, the core content is extracted using multi-layer verification and Attention mechanism, ultimately generating high-quality policy tracing results.
It achieves accurate understanding and efficient retrieval of complex policy texts, improves the relevance and accuracy of retrieval results, reduces information redundancy, and provides policy traceability results with added value. It is suitable for high-precision application scenarios such as policy formulation and compliance assessment.
Smart Images

Figure CN119884313B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, information retrieval and generation technology, and more specifically to a policy tracing dynamic retrieval method based on multi-level intent recognition and comparative learning. Background Art
[0002] With the rapid growth in the number of policy documents and regulatory texts and the significant increase in the complexity and diversity of their content, how to efficiently and accurately trace policies and retrieve information has become an important technical challenge.
[0003] Existing retrieval methods primarily rely on keyword matching or rule-based retrieval models. These methods often struggle to accurately understand user intent when handling complex queries, resulting in insufficient relevance and accuracy in search results. Furthermore, existing methods lack the ability to deeply understand semantics, perform dynamic optimization, or perform contextual analysis when processing large-scale, complex policy texts. The resulting provenance results are relatively simplistic and lack substantial added value.
[0004] Therefore, there is an urgent need for a method that integrates multi-level intent recognition, comparative learning, and generative optimization to achieve accurate policy traceability and dynamic retrieval optimization to meet the needs of high-precision application scenarios such as complex policy review and regulatory tracing. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem that existing retrieval methods mainly rely on keyword matching or rule-based retrieval models, resulting in insufficient relevance and accuracy of retrieval results.
[0006] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0007] The policy traceability dynamic retrieval method based on multi-level intent recognition and contrastive learning includes the following steps:
[0008] S1. Receive and standardize the query content input by the user to generate preprocessed text data;
[0009] S2. Use the pre-trained semantic embedding model to embed the pre-processed text data, convert the text data into a high-dimensional semantic vector representation, generate an embedding vector and store it in the knowledge base;
[0010] S3. Through multi-level intent recognition, we deeply analyze the query content entered by the user, extract the core needs, and dynamically select the knowledge base based on the domain classification, adjusting the search scope of the knowledge base according to different core needs;
[0011] S4. In the knowledge base, the user's query content is retrieved using multiple rounds of dynamic retrieval driven by embedded vectors and fuzzy semantic expansion matching to obtain the retrieval content;
[0012] S5. Optimize the search content through multi-layer verification and comparative learning mechanisms to obtain optimized search content;
[0013] S6: Using a multi-level attention mechanism combined with paragraph focusing technology, we conduct in-depth analysis of the optimized search content and extract the core content that is most relevant to the user's query.
[0014] S7. Generate and optimize the core content of the paragraph focus and extraction, and generate policy tracing results through multiple rounds of dynamic adjustment strategies;
[0015] S8. Display the generated policy tracing results to the user.
[0016] Furthermore, the query content input by the user in step S1 includes but is not limited to text data of policy documents, legal provisions, and management points.
[0017] Furthermore, the standardization preprocessing in step S1 includes: word segmentation, stop word removal, syntactic analysis, synonym replacement, capitalization standardization, and punctuation cleaning.
[0018] Furthermore, the semantic embedding model in step S2 includes but is not limited to BERT and RoBERTa.
[0019] Furthermore, the step S3 mainly includes the following steps:
[0020] S31, multi-level intent recognition, uses a pre-trained semantic embedding model to semantically vectorize user queries, converting query content into a high-dimensional semantic representation. These semantic vectors are then clustered using hierarchical clustering, grouping vectors with similar semantic features into the same level. Semantic grouping strategies are then applied to refine and optimize the hierarchical groups. Finally, an adaptive weight adjustment mechanism assigns appropriate priorities to intent features at each level, thereby extracting the core intent of the user query in a hierarchical manner.
[0021] S32. Combining the results of multi-level intent recognition, a multi-label classification model is used to accurately match the user's query content to the most relevant field. Based on the classification results, the corresponding field knowledge base is automatically selected for subsequent retrieval and content matching.
[0022] Furthermore, the step S4 mainly includes the following steps:
[0023] S41: Multiple rounds of dynamic retrieval, performing preliminary retrieval based on the embedding vector generated in step S2, and extracting the search content that is most similar to the query content entered by the user from the knowledge base obtained in step S3 by calculating the similarity of the embedding vectors;
[0024] S42, fuzzy semantic expansion matching, captures content that is semantically similar to the query entered by the user but expressed differently through fuzzy matching, and combines semantic expansion with synonym expansion and semantic networks to optimize search conditions;
[0025] S43. After the initial search, it enters the multi-round dynamic optimization stage, gradually adjusting and optimizing the search conditions. Each round of search will dynamically adjust the search parameters based on the feedback of the previous round of results, expand the search coverage, and obtain the final search content.
[0026] Furthermore, the step S5 mainly includes the following steps:
[0027] S51: After the initial search is completed, the search content obtained in step S4 is subjected to multi-layer verification, and the search results are evaluated and screened by calculating the similarity between the search content and the query content input by the user;
[0028] S52. Construct a contrastive learning model to distinguish relevant and irrelevant parts of the search content. For results that are semantically similar but actually irrelevant, use the contrastive learning mechanism to optimize them to improve screening accuracy.
[0029] S53. For irrelevant parts, a feedback loop mechanism is started, and the verification results are directly fed back to step S4 to adjust the search content and perform new search content iteration.
[0030] Furthermore, step S6 mainly includes the following steps:
[0031] S61, performing paragraph focusing on the optimized search content obtained in step S5, calculating the semantic similarity between the paragraphs and the query, analyzing the semantics and structure of each paragraph, prioritizing paragraphs that are highly relevant to the user query, and identifying the most valuable text fragments;
[0032] S62. Introducing a multi-level attention mechanism. Based on paragraph focus, the attention mechanism allocates attention to text segments at different levels to identify high-attention information.
[0033] S63. Extract high-interest information to obtain core content that is most relevant to the query content input by the user.
[0034] Furthermore, step S7 mainly includes the following steps:
[0035] S71, focusing on the paragraphs and extracting the core content most relevant to the query input by the user, and continuously optimizing the generation process through multiple rounds of dynamic adjustment strategies;
[0036] S72. Use content fusion to organically combine search content with generated content, and filter and integrate policy tracing results through semantic analysis and contextual relevance assessment;
[0037] S73. Through multiple rounds of optimization and content fusion, high-quality policy tracing results are ultimately generated;
[0038] S74. Conduct a quality assessment on the generated high-quality policy traceability results and optimize the generation strategy based on user feedback. The assessment process includes semantic similarity detection, content redundancy check, and logical consistency analysis to ensure the high standards of the final output.
[0039] S75. During the generation process, the generated policy tracing result is re-verified for relevance to the query content initially input by the user, and if necessary, it is fed back to step S71 to re-optimize the generation conditions.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1. By integrating multi-level intent recognition, comparative learning and generative optimization methods, the present invention solves the problems of insufficient understanding of complex policy texts and poor relevance of retrieval results in existing policy tracing and retrieval methods.
[0042] 2. The present invention uses multi-level intent recognition technology to accurately analyze the core requirements of the query. Combined with the dynamic field classification mechanism, it can automatically select the most relevant knowledge base, reduce the interference of irrelevant information, and improve the pertinence and accuracy of the retrieval.
[0043] 3. By employing an embedded multi-round search strategy and fuzzy semantic expansion technology, this invention dynamically adjusts search criteria and expands search coverage to ensure comprehensive coverage of content relevant to the query. Combined with a comparative learning mechanism, this method performs multi-level validation of search results, eliminating semantically similar but irrelevant content and optimizing result quality.
[0044] 4. By combining a multi-level attention mechanism with paragraph focusing technology, the present invention can accurately identify and extract core content that is highly relevant to the query in complex texts, avoid information redundancy and omission of key information, and ensure that the generated results are highly relevant and valuable.
[0045] 5. The present invention uses a dynamically adjusted generation optimization mechanism to perform multiple rounds of dynamic adjustments on the extracted core content, and organically integrates the search content with the newly generated content, avoiding simple information piling, providing more coherent and value-added tracing results, and improving output quality.
[0046] 6. This invention continuously adjusts and optimizes the processing flow by displaying traceability results and collecting user feedback. User feedback is directly used to improve the method, forming a closed-loop optimization loop that can dynamically adapt to different user query requirements and continuously improve performance and output accuracy.
[0047] 3. This invention can be effectively applied in various complex scenarios such as policy tracing and compliance checking. It is particularly suitable for retrieval needs that require high accuracy and relevance, providing strong support for policy formulation, compliance assessment and other work. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is the overall flow chart of the present invention;
[0049] Figure 2 This is a flow chart of step S5 in the present invention;
[0050] Figure 3 This is a flow chart of step S6 in the present invention;
[0051] Figure 4 This is a flow chart of step S7 in the present invention;
[0052] Figure 5 Schematic diagram of the policy traceability dynamic retrieval method with multi-level intent recognition and contrastive learning in the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] See also Figure 1-4 , a policy traceability dynamic retrieval method based on multi-level intent recognition and contrastive learning, specifically including the following steps:
[0055] S1. Receive and standardize the query content input by the user to generate preprocessed text data;
[0056] The front-end interface receives user input, which may contain complex text data such as policy documents, regulations, and management points. To ensure the effectiveness and accuracy of subsequent processing, the input text data is pre-processed in a standardized manner, including the following steps:
[0057] S11, data reception:
[0058] The front end receives the query content input by the user and collects the input text data. This data may be unstructured policy documents, regulatory content, or management key points, with complex and diverse data.
[0059] S12. Word Segmentation:
[0060] Use natural language processing (NLP) technology to perform word segmentation on the received text, splitting the text into independent lexical units for subsequent analysis and processing.
[0061] S13. Stop Word Removal:
[0062] Remove common stop words that have no actual semantic contribution (such as "de", "le", "zai", etc.), reduce text noise, and retain keyword vocabulary that contributes to semantic understanding.
[0063] S14. Syntactic Analysis:
[0064] Perform syntactic analysis to understand the structure of the text, identify the relationships between words, enhance the depth of understanding of the text content, and provide structured syntactic information for subsequent semantic embedding generation.
[0065] S15. Data Normalization:
[0066] Perform normalization processing on the input text, such as unifying case, cleaning punctuation marks, and replacing synonyms, etc., to improve the consistency and standardization of text processing.
[0067] S16. Preprocessing Output:
[0068] The preprocessed text data is converted into a more normalized and structured input form, preparing for subsequent semantic embedding generation and efficient retrieval.
[0069] Through the above steps, the query input by the user can be effectively preprocessed in a standardized manner, reducing noise and redundant information in the text, and providing a solid foundation for the in-depth semantic understanding and subsequent retrieval of policy texts.
[0070] S2. Use a pre-trained semantic embedding model to generate embeddings for the preprocessed text data, convert the text data into a high-dimensional semantic vector representation, generate embedding vectors and store them in the knowledge base. The semantic embedding model includes but is not limited to BERT, RoBERTa;
[0071] Use the pre-trained semantic embedding model BERT to generate embeddings for the preprocessed text data, converting the text content into a high-dimensional semantic vector representation. Through the embedding model, capture the deep semantic features of the text, including context relationships, dependencies between words, and implicit semantic associations, ensuring that each policy document, regulatory clause, or user query is accurately mapped to the high-dimensional vector space.
[0072] The resulting high-dimensional semantic vectors retain the core semantic information of the text, placing similar textual content closer together in the vector space, facilitating subsequent rapid retrieval and comparison. Embedding vectors enable a deeper understanding of textual content and provide a unified semantic representation for different types of text, enhancing the method's ability to handle diverse inputs.
[0073] The generated embedding vectors are efficiently stored in the knowledge base and, combined with corresponding domain classifications and metadata, provide a data foundation for subsequent dynamic retrieval, comparative learning, and content generation optimization. Vectorized storage enables rapid similarity calculation and content matching, improving retrieval efficiency and response speed, ensuring accurate and relevant traceability results when processing complex policy queries.
[0074] S3. Through multi-level intent recognition, we deeply analyze the query content entered by the user, extract the core needs, and dynamically select the knowledge base based on the domain classification, adjusting the search scope of the knowledge base according to different core needs;
[0075] Through multi-layered intent recognition technology, we deeply analyze user queries, accurately extract core requirements, and combine domain classification to enable dynamic knowledge base selection. This process accurately understands user query intent and enables targeted subsequent searches.
[0076] S31, Multi-level Intent Recognition:
[0077] A pre-trained semantic embedding model is used to semantically vectorize user queries and convert the query content into a high-dimensional semantic representation. Hierarchical clustering is used to perform hierarchical clustering analysis on these semantic vectors, and vectors with similar semantic features are grouped into the same level. On this basis, a semantic grouping strategy is applied to refine and optimize the hierarchical groups. Finally, an adaptive weight adjustment mechanism is used to assign reasonable priorities to the intent features of each level, thereby hierarchically extracting the core intent of the user query.
[0078] S32. Field classification and dynamic knowledge base selection:
[0079] Combining the results of multi-level intent recognition, a multi-label classification model is used to accurately match user queries to the most relevant fields (such as finance, agriculture, justice, science and education, etc.). This multi-label classification effectively distinguishes the dominant intent in complex multi-intent queries, setting an accurate domain scope for subsequent searches. Based on the classification results, the corresponding domain knowledge base is automatically selected for subsequent searches and content matching. The dynamic domain knowledge base selection mechanism automatically adjusts the search scope based on different intents, reducing interference from irrelevant fields, thereby improving search efficiency and accuracy.
[0080] S4. In the knowledge base, the user's query content is retrieved using multiple rounds of dynamic retrieval driven by embedded vectors and fuzzy semantic expansion matching to obtain the retrieval content;
[0081] An embedded vector-driven multi-round dynamic retrieval strategy is adopted, combined with fuzzy semantic expansion technology to ensure comprehensive retrieval of relevant policy content in user queries, thereby improving retrieval coverage and accuracy.
[0082] S41. Embedded multi-round search strategy:
[0083] First, a preliminary search is performed using the Dense Passage Retrieval (DPR) technique based on the embedding vector generated in step S2. In this stage, the search content that is most similar to the user query is extracted from the knowledge base by calculating the similarity of the embedding vectors.
[0084] S42, fuzzy matching and semantic expansion:
[0085] During the multi-round search process, a fuzzy matching strategy was introduced to address the shortcomings of the initial search. This fuzzy matching strategy captures content that is semantically similar to the user's query but expressed differently, avoiding the omission of relevant policies due to the limitations of exact matching. Furthermore, semantic expansion technology was combined with synonym expansion and semantic networks (such as WordNet) to enrich search criteria. Semantic expansion can identify synonyms, antonyms, and related terms, expanding the search scope and ensuring comprehensive coverage of all potentially relevant policy texts.
[0086] S43. After the initial search, enter the multi-round dynamic optimization stage to gradually adjust and optimize the search conditions:
[0087] After the initial search, the search process continues with multiple rounds of dynamic optimization, gradually adjusting and optimizing the query conditions. Each round of search dynamically adjusts search parameters (such as similarity thresholds and query expansion ranges) based on the results of the previous round, expanding the search coverage and ensuring that all potentially relevant content is retrieved, ultimately resulting in the final search results.
[0088] S44. Dynamic adjustment and iterative optimization:
[0089] The combination of an embedded multi-round search strategy and fuzzy semantic expansion enables dynamic, iterative retrieval within the embedded space, gradually optimizing search content. Each round relies on feedback from the previous round, and through continuous adjustment of search criteria and semantic expansion, it gradually approaches the most relevant policy content. This approach not only improves search coverage but also optimizes content matching in complex query scenarios, enabling more intelligent responses to diverse user queries.
[0090] By combining multiple rounds of dynamic retrieval and fuzzy semantic expansion, this step achieves comprehensive coverage and accurate retrieval of policy content, providing users with more relevant and complete traceability results.
[0091] S5. Optimize the search content through multi-layer verification and comparative learning mechanisms to obtain optimized search content;
[0092] The retrieved content is optimized through multi-layered validation and comparative learning mechanisms to ensure the relevance and accuracy of the final output. This process continuously improves the quality of the results through meticulous screening and feedback loops of the search results.
[0093] S51. Content verification and optimization:
[0094] After the initial search, the search results from step S4 undergo multiple layers of validation to ensure their relevance to the user's query. This validation goes beyond surface similarity and includes in-depth analysis at the semantic level. By calculating the semantic similarity between the content and the query, the search results are evaluated and filtered to eliminate irrelevant or falsely detected content.
[0095] S52. Contrastive learning driven optimization process:
[0096] For semantically similar but irrelevant results, a contrastive learning mechanism is used to optimize and improve screening accuracy. This contrastive learning mechanism plays a key role in this optimization process. By building a contrastive learning model, it is possible to effectively distinguish between relevant and irrelevant parts of the search content, particularly when faced with semantically similar interference items, accurately eliminating irrelevant content. Through continuous iterative training and adjustment of search content, contrastive learning continuously improves recognition capabilities in complex semantic scenarios, ensuring that the final output results are both highly relevant and accurately reflect the user's query intent.
[0097] S53. Feedback loop and search strategy adjustment:
[0098] If the content verification and comparative learning process reveals any discrepancies between the results and the user's query intent, a feedback loop is initiated. The verification results are fed directly into the embedded multi-round search and fuzzy matching in step S4 to adjust search conditions and initiate new search iterations. This closed-loop feedback mechanism dynamically adjusts search strategies, continuously optimizing the accuracy and relevance of results, enabling self-learning and adaptation, and ultimately enhancing overall intelligence.
[0099] S54. Content Verification and Retrieval Feedback Integration:
[0100] Verification results are used not only to filter and optimize current search content but also to adjust subsequent search strategies. By directly integrating content verification feedback into the search strategy adjustment process, the search and optimization processes complement each other, forming an adaptive closed-loop optimization. This mechanism maintains efficiency and accuracy in complex policy text environments, ensuring that users receive the most relevant and optimized search content for multi-purpose, multi-domain queries.
[0101] Through the combination of multi-layer verification and comparative learning, the search results can be deeply screened and optimized, the relevance and accuracy of the content can be continuously improved, and ultimately users can be provided with high-quality policy traceability results that meet the query intent.
[0102] S6: Using a multi-level attention mechanism combined with paragraph focusing technology, we conduct in-depth analysis of the optimized search content and extract the core content that is most relevant to the user's query.
[0103] Using a multi-level attention mechanism combined with paragraph-focusing technology, we conduct in-depth analysis of search results to accurately extract the core content most relevant to the user's query. This process involves multi-level processing of the paragraph's semantics and structure to ensure that the output accurately reflects the user's query intent.
[0104] S61, paragraph focus technology:
[0105] First, we perform paragraph focusing on the retrieved text results, analyzing the semantics and structure of each paragraph and prioritizing those that are highly relevant to the user's query. Paragraph focusing calculates the semantic similarity between the paragraph and the query to identify the most valuable text fragments.
[0106] During the focusing process, not only the matching of keywords is considered, but also the contextual association and logical structure of the paragraph are deeply analyzed to ensure that the extracted paragraph can fully reflect the core theme of the user's query.
[0107] S62, Multi-level Attention Mechanism:
[0108] To accurately identify the core content of a paragraph, a multi-level attention mechanism is introduced. Based on paragraph focus, the attention mechanism allocates different levels of attention to sentences and words within a paragraph, identifying information that is of high interest.
[0109] Multi-level attention allows for deep focus at the sentence, word, and semantic levels. By focusing on key sentences and words, it improves the ability to extract relevant information. This mechanism effectively enhances understanding of complex texts and avoids missing key details.
[0110] S63, core content extraction:
[0111] After paragraph focusing and multi-level attention processing, core content extraction is performed on high-interest information. Core content extraction aims to mine the text fragments most closely related to the user's query, ensuring that the generated results accurately reflect the user's query intent.
[0112] Combining contextual semantics and attention allocation results, content with high relevance and information value in the paragraph is extracted, avoiding information redundancy and interference from irrelevant content.
[0113] S64. Intent-based precise positioning:
[0114] The combination of paragraph focusing and a multi-level attention mechanism can accurately locate relevant text segments within complex policy documents based on user intent. This intent-based localization approach significantly improves text analysis capabilities and content extraction accuracy, ensuring highly relevant tracing results.
[0115] Through the dual alignment of semantics and intent, paragraph focus is no longer limited to superficial similarities, but rather a deep understanding of the inherent connections of the text, enabling accurate responses to user queries.
[0116] Through paragraph focusing and core content extraction using a multi-level attention mechanism, this step effectively improves the ability to deeply analyze text and extract relevant information, providing users with more accurate policy tracing results and enhancing the response effect in complex query scenarios.
[0117] S7. Generate and optimize the core content of the paragraph focus and extraction, and generate policy tracing results through multiple rounds of dynamic adjustment strategies;
[0118] We optimize the results of paragraph focus and core content extraction, generating high-quality policy tracing results through multiple rounds of dynamic adjustment strategies. We also use content fusion methods to organically combine retrieved content with generated new content to improve the accuracy and logical coherence of the output results.
[0119] S71. Generation optimization and dynamic adjustment:
[0120] Initially, the text generated from the paragraph focusing and core content extraction phases is generated. The generation process is then continuously optimized through multiple rounds of dynamic adjustment strategies. This generation optimization mechanism combines contextual understanding, semantic consistency checks, and logical structure adjustments to ensure that the generated content is more aligned with the user's query intent. Dynamic adjustment involves repeated iterations of the generated content, gradually revising the generation strategy by evaluating its relevance, semantic integrity, and logical coherence. Each round of adjustment refines the output of the previous round, ensuring a more consistent and accurate final result.
[0121] S72. Content integration strategy:
[0122] An innovative content fusion strategy is employed to organically integrate search results with generated content. This fusion strategy goes beyond simply splicing the two pieces of content together. Instead, it uses semantic analysis and contextual relevance assessment to filter and integrate the most valuable information. During the content fusion process, search results are reorganized through semantic matching and logical associations, allowing the generated content to flow naturally with existing text and avoiding duplication and clutter. This process not only improves the overall quality of the content but also enhances the added value of the output.
[0123] S73, high-quality output generation:
[0124] Through multiple rounds of optimization and content fusion, high-quality policy tracing results are ultimately generated. The output not only accurately reflects the core intent of the user's query but also provides in-depth analysis and comprehensive information, providing users with value beyond the search results themselves. The generation process comprehensively considers policy background, relevant regulations, and contextual information to ensure clear logic and semantic consistency in the output, allowing users to more intuitively understand the key points of policy tracing.
[0125] S74. Output quality assessment and feedback optimization:
[0126] The final generated output is evaluated for quality, and the generation strategy is optimized based on user feedback. This evaluation process includes semantic similarity detection, content redundancy checking, and logical consistency analysis to ensure the high quality of the final output. The feedback mechanism allows for continuous improvement of the generation model and fusion strategy based on user feedback on the generated content, achieving adaptive content generation optimization. This ensures continuous improvement in output quality in response to diverse queries.
[0127] S75, Feedback Optimization Mechanism:
[0128] During the generation process, the relevance of the generated content to the query content initially input by the user will be verified again. If necessary, it will be fed back to step S71 to re-optimize the generation conditions to ensure that the final output content is logically coherent and semantically clear.
[0129] S8, display feedback of traceability results and overall improvement based on user feedback;
[0130] The generated policy tracing results are presented to users, and user feedback is collected to continuously improve performance.
[0131] like Figure 5 As shown, the present invention also proposes a policy traceability dynamic retrieval process summary based on multi-level intent recognition and contrastive learning, including:
[0132] The user query processing process 410 is used to receive the query content input by the user and perform standardized preprocessing on the input query content, including operations such as word segmentation, removal of stop words, and syntactic analysis, to generate standardized text data suitable for subsequent semantic processing.
[0133] The semantic embedding generation process 420 is used to use the pre-trained model to perform high-dimensional semantic embedding on the standardized text data, generate high-dimensional vectors that reflect the deep semantics of the text, and store these vectors in the knowledge base to provide basic data support for subsequent retrieval and comparative learning.
[0134] The multi-level intent recognition and classification process 430 is used to perform multi-level intent recognition on the query content input by the user, combine semantic embedding, hierarchical clustering and adaptive weight adjustment technologies to accurately extract the core intent of the user query, and map the recognition results to the knowledge base of the corresponding field for classification.
[0135] The embedding vector retrieval and optimization process 440 is used to perform multiple rounds of dynamic retrieval in the embedding space, optimize the preliminary retrieval results through the DensePassage Retrieval technology, and dynamically adjust the retrieval conditions in combination with the fuzzy matching strategy and semantic expansion technology to expand the search scope and improve the accuracy of the retrieval results.
[0136] The content verification and comparative learning process 450 is used to perform multi-level verification on the initially retrieved content, optimize the search results through the comparative learning mechanism, exclude semantically similar but irrelevant content, and dynamically adjust the search strategy based on the verification feedback to continuously optimize the relevance of the output results.
[0137] The paragraph focusing and core extraction process 460 is used to perform paragraph focusing analysis on the retrieved text through a multi-level attention mechanism, identify the most relevant paragraphs, and extract the core content to ensure that the generated tracing results can accurately reflect the user's query intention.
[0138] The generation optimization and content fusion process 470 is used to perform multiple rounds of dynamic adjustments on the results of paragraph focusing and core content extraction, combine the retrieved content with the newly generated content through the generation optimization mechanism, implement the content fusion strategy, and generate coherent and value-added policy tracing results.
[0139] The result display and feedback process 480 is used to display the generated policy traceability results to users and collect user feedback information. Based on user feedback, each processing process is adjusted and optimized to continuously improve the overall performance and accuracy of the output results.
[0140] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. The scope of patent protection of the present invention shall be based on the claims. Any equivalent structural changes made using the description and drawings of the present invention should also be included in the scope of protection of the present invention.
Claims
1. A policy traceability dynamic retrieval method based on multi-level intent recognition and contrastive learning, characterized by: The following steps are involved: S1. Receive and standardize the query content input by the user to generate preprocessed text data; S2. Use the pre-trained semantic embedding model to embed the pre-processed text data, convert the text data into a high-dimensional semantic vector representation, generate an embedding vector and store it in the knowledge base; S3. Through multi-level intent recognition, we deeply analyze the query content entered by the user, extract the core needs, and dynamically select the knowledge base based on the domain classification, adjusting the search scope of the knowledge base according to different core needs; S31, multi-level intent recognition, uses a pre-trained semantic embedding model to semantically vectorize user queries, converting query content into a high-dimensional semantic representation. These semantic vectors are then clustered and analyzed using hierarchical clustering, grouping vectors with similar semantic features into the same level. Semantic grouping strategies are then applied to refine and optimize the hierarchical groups. Finally, an adaptive weight adjustment mechanism is used to assign appropriate priorities to intent features at each level, thereby extracting the core intent of the user query in a hierarchical manner. S32. Combine the results of multi-level intent recognition and use a multi-label classification model to accurately match the user's query content to the most relevant domain. Based on the classification results, the corresponding domain knowledge base is automatically selected for subsequent retrieval and content matching. S4. In the knowledge base, the user's query content is retrieved using multiple rounds of dynamic retrieval driven by embedded vectors and fuzzy semantic expansion matching to obtain the retrieval content; S5. Optimize the search content through multi-layer verification and comparative learning mechanisms to obtain optimized search content; S51: After the initial search is completed, the search content obtained in step S4 is subjected to multi-layer verification, and the search results are evaluated and screened by calculating the similarity between the search content and the query content input by the user; S52. Construct a contrastive learning model to distinguish relevant and irrelevant parts of the search content. For results that are semantically similar but actually irrelevant, use the contrastive learning mechanism to optimize them to improve screening accuracy. S53: For irrelevant parts, a feedback loop mechanism is started, and the verification results are directly fed back to step S4 to adjust the search content and perform new search content iteration; S6: Using a multi-level attention mechanism combined with paragraph focusing technology, we conduct in-depth analysis of the optimized search content and extract the core content that is most relevant to the user's query. S7. Generate and optimize the core content of the paragraph focus and extraction, and generate policy tracing results through multiple rounds of dynamic adjustment strategies; S71, focusing on the paragraphs and extracting the core content most relevant to the query input by the user, and continuously optimizing the generation process through multiple rounds of dynamic adjustment strategies; S72. Use content fusion to organically combine search content with generated content, and filter and integrate policy tracing results through semantic analysis and contextual relevance assessment; S73. Through multiple rounds of optimization and content fusion, high-quality policy tracing results are ultimately generated; S8. Display the generated policy tracing results to the user.
2. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The query content input by the user in step S1 includes text data of policy documents, legal provisions, and management points.
3. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The standardization preprocessing in step S1 includes: word segmentation, stop word removal, syntactic analysis, synonym replacement, capitalization standardization, and punctuation cleaning.
4. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The semantic embedding models in step S2 include BERT and RoBERTa.
5. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The step S4 mainly includes the following steps: S41: Multiple rounds of dynamic retrieval, performing preliminary retrieval based on the embedding vector generated in step S2, and extracting the search content that is most similar to the query content entered by the user from the knowledge base obtained in step S3 by calculating the similarity of the embedding vectors; S42, fuzzy semantic expansion matching, captures content that is semantically similar to the query entered by the user but expressed differently through fuzzy matching, and combines semantic expansion with synonym expansion and semantic networks to optimize search conditions; S43. After the initial search, it enters the multi-round dynamic optimization stage, gradually adjusting and optimizing the search conditions. Each round of search will dynamically adjust the search parameters based on the feedback of the previous round of results, expand the search coverage, and obtain the final search content.
6. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The step S6 mainly includes the following steps: S61, performing paragraph focusing on the optimized search content obtained in step S5, calculating the semantic similarity between the paragraphs and the query, analyzing the semantics and structure of each paragraph, prioritizing paragraphs that are highly relevant to the user query, and identifying the most valuable text fragments; S62. Introducing a multi-level attention mechanism. Based on paragraph focus, the attention mechanism allocates attention to text segments at different levels to identify high-attention information. S63. Extract high-interest information to obtain core content that is most relevant to the query content input by the user.
7. The policy tracing dynamic retrieval method based on multi-level intent recognition and contrastive learning according to claim 1 is characterized by: The step S7 further comprises the following steps: S74. Conduct a quality assessment on the generated high-quality policy traceability results and optimize the generation strategy based on user feedback. The assessment process includes semantic similarity detection, content redundancy check, and logical consistency analysis to ensure the high standards of the final output. S75. During the generation process, the generated policy tracing result is used to verify its relevance to the query content initially input by the user, and this can be fed back to step S71 to re-optimize the generation conditions.
Citation Information
Patent Citations
Intelligent question and answer method, device and equipment and storage medium
CN114416927A
Government affair service field multi-strategy fusion dialogue method based on knowledge graph
CN116628172A
Cited By
Complex teaching task-oriented mixed intention recognition and dynamic knowledge retrieval method
CN121743441A