Resource recommendation method and system based on hybrid retrieval RAG

By building a hybrid retrieval RAG system that combines a semantic vector library and a tag library, the problems of recall redundancy and matching bias in existing RAG systems in fields with high semantic complexity are solved, and resource recommendations with high accuracy and high coverage are achieved.

CN120780916AActive Publication Date: 2025-10-14ZHEJIANG DAGU TECH CO LTD

Patent Information

Application Number
CN202511202623.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-14
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing RAG systems are unable to effectively utilize the structured label information of resource data for precise conditional filtering in areas with high semantic complexity, resulting in redundant or mismatched recall content. They are also unable to automatically identify user intent types or perform diversified expression rewriting, resulting in insufficient recall coverage and the mixing of low-quality text.

Method used

A hybrid retrieval RAG method is adopted, combined with a semantic vector library and a tag library. By constructing a vector library and a tag library, dual modeling of resource text is performed, and a language model is used for semantic expansion and intent recognition. Structured query statements are generated, and semantic encoding and similarity calculation are combined for precise screening and splicing.

Benefits of technology

It improves the accuracy and context matching of resource recommendations, ensures the relevance and quality of recalled content, and enhances the efficiency and controllability of the resource recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780916A_ABST
    Figure CN120780916A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent recommendation, and discloses a hybrid retrieval RAG-based resource recommendation method and system, and the method comprises the steps: collecting resource text data, and constructing a vector library and a tag library; expanding the user question based on the language model to obtain a plurality of semantic extension questions; performing intention recognition, judging whether the user question is a resource recommendation question, and if yes, determining a target classification type; screening the data according to the field definition in the tag library to obtain a candidate knowledge fragment set; obtaining candidate vectors, mapping the user question and the semantic extension question into query vectors, calculating the similarity between the query vectors and each candidate vector, and selecting knowledge supplement content; and performing resource splicing on all the knowledge supplement contents to generate resource recommendation answers. According to the method, a structured label screening mechanism and a semantic vector fine arrangement mechanism are fused, and the problems of recall redundancy, matching deviation and the like caused by the fact that an existing RAG system only depends on semantic similarity retrieval are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent recommendation technology, and in particular to a resource recommendation method and system based on hybrid retrieval RAG. Background Art

[0002] With the development of artificial intelligence, particularly Large Language Model (LLM) technology, Retrieval-Augmented Generation (RAG) has become an important tool for implementing intelligent question-answering and knowledge recommendation systems. Combining information retrieval and natural language generation techniques, RAG can inject external knowledge as context into language models, thereby improving the accuracy, contextual relevance, and content coverage of question-answering in specific domains.

[0003] In areas of high semantic complexity, such as government policies, scientific research services, and scientific and technological resources, users often raise highly targeted resource query requirements. Existing RAGs mostly perform similarity recall based on semantic vectors, and fail to use the structured label information already available in the resource data for precise conditional filtering, resulting in redundant returned content or matching deviations. It is impossible to automatically identify the user's intent type or the resource category to which it belongs, nor is it possible to rewrite user questions in a diversified expression, resulting in insufficient recall coverage. User questions are difficult to automatically convert into operational structured retrieval logic, resulting in the system being unable to perform field-level screening of the resource database. Simply relying on vector space recall and then directly using it to generate answers without setting clear semantic thresholds and refined sorting strategies can easily introduce irrelevant or low-quality text.

[0004] Therefore, it is necessary to design a resource recommendation method and system based on hybrid retrieval RAG to solve the problems existing in current technology. Summary of the Invention

[0005] In view of this, the present invention proposes a resource recommendation method and system based on hybrid retrieval RAG, aiming to solve the current problems of lack of structural constraints, semantic extension and intent judgment capabilities, and low semantic screening accuracy.

[0006] In one aspect, the present invention proposes a resource recommendation method based on hybrid retrieval RAG, comprising: Collect resource text data, and construct a vector library and a label library based on the resource text data; Collect user questions, expand the user questions based on the language model, and obtain several semantically expanded questions; Performing intent recognition on all the semantically extended questions to determine whether the user question is a resource recommendation question, and if so, determining the target classification type to which the resource belongs; Convert the user questions and semantic expansion questions into structured query statements, filter the data according to the field definitions in the tag library, and obtain a set of candidate knowledge fragments that meet the tag conditions; Perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector; map the user question and the semantic expansion question into a query vector based on the semantic encoding, calculate the similarity between the query vector and each candidate vector, and select segments with a similarity greater than a similarity threshold as knowledge supplementary content; All the supplementary knowledge contents are combined to generate resource recommendation answers.

[0007] Furthermore, when collecting resource text data and constructing a vector library and a label library based on the resource text data, the process includes: Preprocessing the resource text data, including text segmentation, formatting, and noise removal; Constructing the vector library: dividing the pre-processed resource text data into blocks according to semantic logic, wherein the block division includes division by title level and division by parent-child semantic structure; Perform vector encoding on each text block based on the semantic embedding model to obtain a semantic vector; Establishing a mapping relationship between the semantic vector and the original resource text data, and storing the mapping relationship in a database to obtain the vector library; Build a tag library: Extract keywords and entities related to application conditions from preprocessed resource text data, and build tag names and tag values; Label resource content based on rule templates to form label expressions; Combine several tag expressions according to business logic, build a tag condition tree, and bind the tag content to the corresponding resource fragment; The bound tag information is organized into a structured table format and stored in a database to obtain the tag library.

[0008] Furthermore, user questions are collected and expanded based on the language model to obtain several semantically expanded questions, including: The user question is rewritten based on the language model to generate multiple query questions with different semantic expressions. Figure 1 the correct form of problem expression; Before executing question rewriting, guiding the language model to generate diversified questions by constructing a prompt word template, wherein the prompt word template includes instruction information, language style requirements, and context memory loading fields; During question rewriting, the current session history information is combined.

[0009] Furthermore, the intent of all the semantically extended questions is identified to determine whether the user question is a resource recommendation question, including: Construct a set of predefined intent categories, including resource recommendation and non-resource consultation, input all the semantic expansion questions, and generate intent classification results. A majority vote is performed on the recognition results of multiple semantic expansion questions. If the result belongs to the resource recommendation intention, the target classification type of the resource is determined based on the semantic keywords; otherwise, the recommendation process is terminated and a rejection is triggered.

[0010] Furthermore, the user question is converted into a structured query statement, and the data is filtered according to the field definition in the tag library to obtain a set of candidate knowledge fragments that meet the tag conditions, including: Extracting table structure information of the tag library, wherein the table structure information includes: an enumeration set of field name, field type, field comment, tag name, and tag value; Semantically matching the field descriptions and annotations in the tag library based on the user question and the semantically extended question, generating a structured SQL statement that meets the semantic intent, wherein the structured SQL statement includes a WHERE statement for specifying a matching condition between the tag field and the tag value; If the structured SQL statement cannot be parsed or the query fails, the fallback strategy rule is enabled to match the rule of keyword and tag name mapping, and the SQL query condition is manually constructed; A query operation is performed in the tag library based on the structured SQL statement to obtain a set of candidate knowledge fragments that match the semantic conditions of the user question.

[0011] Furthermore, semantic vector encoding is performed on each resource text in the candidate knowledge fragment set to obtain a candidate vector, including: Normalizing each resource text in the candidate knowledge fragment set, wherein the normalization includes removing HTML tags, normalizing punctuation, normalizing spaces, and trimming the length of the text; Encode the candidate knowledge fragments using a semantic embedding model consistent with the vector library construction phase; The semantic embedding model maps each selected knowledge fragment into a semantic vector of fixed dimension to obtain the candidate vector, and the candidate vector is obtained by using mean pooling or CLS token pooling strategy during the encoding process.

[0012] Furthermore, mapping the user question and the semantically expanded question into a query vector based on semantic coding, calculating the similarity between the query vector and each candidate vector, and selecting a segment with a similarity greater than a similarity threshold as the knowledge supplement content includes: Encoding the user question and the semantically expanded question based on the semantic embedding model, mapping the user question into a query vector, wherein the query vector and the candidate vector have the same dimension and semantic space; Calculating the cosine similarity between each candidate vector and the query vector to obtain the similarity of each candidate vector; All the similarities are compared with the similarity threshold respectively to obtain the knowledge supplement content.

[0013] Furthermore, all the similarities are compared with the similarity threshold respectively to obtain the supplementary knowledge content, including: When the similarity is greater than or equal to the similarity threshold, the resource text corresponding to the similarity is used as the knowledge supplement content. If there are multiple resource texts that meet the conditions, a semantic relevance analysis is performed on the multiple resource texts, and they are sorted from high to low in terms of similarity and from high to low in terms of relevance, and the top n items are selected as the knowledge supplement content.

[0014] Furthermore, all the supplementary knowledge contents are combined to generate resource recommendation answers, including: The selected supplementary knowledge content is sorted from high to low according to its similarity to the query vector, and the sorted pieces of content are sequentially spliced ​​as context input. During the splicing process, separators are inserted between each piece of content, and the original user question and semantically expanded question are embedded in the spliced ​​context; The generated answer output result is post-processed, and the post-processing includes removing invalid sentences, unifying terminology styles, and de-duplicating classifications to obtain the resource recommendation answer.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: the present application integrates structured label screening and semantic vector precision ranking mechanism, overcoming the problems of recall redundancy and matching deviation caused by relying solely on semantic similarity retrieval in the existing RAG system. The resource text is dual-modeled: on the one hand, a semantic vector library is constructed to capture the deep semantic information of the resource content; on the other hand, a label library is constructed to extract the conditional attribute information of policies or resources and support precise matching based on field rules. The user questions are diversified and semantically expanded through the language model to improve the recall coverage; at the same time, the accuracy of the recommendation logic is guaranteed through intent recognition and resource type classification. The user intention is automatically converted into SQL form to realize conditional screening of the label field and obtain highly relevant candidate knowledge fragments. The candidate content is semantically encoded and similarity calculated, a threshold is set and precision ranking is performed, and a recommended answer with a reasonable structure and strong semantic pertinence is generated through context splicing. The accuracy of resource recommendations, context matching and controllability of generated content are improved.

[0016] On the other hand, the present application also provides a resource recommendation system based on hybrid retrieval RAG, which is used to apply the above-mentioned resource recommendation method based on hybrid retrieval RAG, including: The collection unit is configured to collect resource text data and construct a vector library and a label library according to the resource text data; a processing unit configured to collect user questions, expand the user questions based on a language model, and obtain a plurality of semantically expanded questions; an identification unit configured to perform intent recognition on all the semantically extended questions, determine whether the user question is a resource recommendation question, and if so, determine the target classification type to which the resource belongs; a screening unit configured to convert the user question and the semantically expanded question into a structured query statement, screen the data according to the field definition in the tag library, and obtain a set of candidate knowledge fragments that meet the tag conditions; a matching unit configured to perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector; map the user question and the semantic expansion question into a query vector based on the semantic encoding, calculate the similarity between the query vector and each candidate vector, and select segments with a similarity greater than a similarity threshold as knowledge supplementary content; The output unit is configured to perform resource splicing on all the knowledge supplementary contents and generate resource recommendation answers.

[0017] It is understandable that the above-mentioned resource recommendation method and system based on hybrid retrieval RAG have the same beneficial effects, which will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 A flowchart of a resource recommendation method based on hybrid retrieval RAG provided in an embodiment of the present invention; Figure 2 This is a functional block diagram of a resource recommendation system based on hybrid retrieval RAG provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood, and so that the scope of the present disclosure can be conveyed to those skilled in the art. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0020] In the existing retrieval enhancement generation system, the resource recommendation process has the defects of insufficient multi-modal data fusion and single retrieval logic. The recall mechanism based on pure semantic vectors cannot effectively integrate structured label information, resulting in the absence of key field constraints in resource screening conditions, and a large amount of irrelevant redundant text in the returned results. User query intent recognition only relies on a single question expression form, and fails to cover diversified expression modes through semantic expansion, resulting in potential relevant resource missing detection. The mapping relationship between the query statement and the database field lacks an automatic conversion mechanism, so that the structured label cannot participate in accurate condition filtering, directly affecting the recall accuracy of the candidate knowledge fragments. The semantic similarity threshold setting lacks dynamic adjustment capability, resulting in the mixing of low-quality text into the generated context, affecting the information density and accuracy of the final recommendation results.

[0021] For example, in the government scientific and technological resource intelligent recommendation platform, when the user submits a complex query request containing multiple condition constraints, only the vector space calculation is used to match similar text fragments. When the user query involves "biological and medical field research and development equipment sharing policy", the existing method cannot identify the corresponding resource classification label of "biological and medical", nor can it map "research and development equipment sharing" to the policy type field for structured screening, resulting in a large number of irrelevant policy documents such as medical talent introduction and drug approval process in the returned results. The lack of semantic expansion mechanism makes the system unable to generate related expressions such as "biological and medical experimental instrument open use guide", further reducing the recall rate of potential high-value resources. Vector similarity calculation does not combine label condition filtering, so that device maintenance procedures and other non-policy texts are mistakenly judged as effective results due to similar semantics.

[0022] If the above problems are not solved, the dual bottleneck of retrieval accuracy and recall rate will be faced for a long time. The idle of structured labels results in the inactivation of database field level screening capability, and the resource recommendation process loses the accurate control dimension. The lack of semantic expansion causes incomplete coverage of user intent, and the key resource missing detection rate increases exponentially with the query complexity. The mixing of low-quality text into the generated context will cause the decline of content credibility, and the user needs to spend extra time for manual screening and verification. Finally, it will greatly reduce the practical value of the resource recommendation system, and it is difficult to meet the service needs of professional fields with high precision and high efficiency.

[0023] For this, see Figure 1 As shown, this application proposes a resource recommendation method based on hybrid retrieval RAG, including: S100: Collect resource text data, and construct a vector library and a label library based on the resource text data.

[0024] S200: Collect user questions, expand the user questions based on the language model, and obtain several semantically expanded questions.

[0025] S300: Perform intent recognition on all semantic extension questions to determine whether the user question is a resource recommendation question. If so, determine the target classification type to which the resource belongs.

[0026] S400: Convert user questions and semantically extended questions into structured query statements, filter data according to field definitions in the tag library, and obtain a set of candidate knowledge fragments that meet tag conditions.

[0027] S500: Perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector. Based on the semantic encoding, the user question and the semantically expanded question are mapped into a query vector. The similarity between the query vector and each candidate vector is calculated, and segments with similarity greater than a similarity threshold are selected as supplementary knowledge content.

[0028] S600: All supplementary knowledge contents are spliced ​​together to generate resource recommendation answers.

[0029] Specifically, hybrid retrieval (RAG) refers to a search-enhanced generation method that combines vector retrieval with structured tag retrieval. This method can be implemented using semantic vector matching and structured query language (SQL) to improve the accuracy and contextual relevance of resource recommendations. A vector library refers to a database that stores semantic vectors of resource text. This method can be implemented by encoding text blocks using a semantic embedding model to support semantic similarity calculations. A tag library refers to a database that stores structured tag information for resources. This method can be implemented using keyword extraction and rule template annotation to support structured conditional filtering. Semantic question expansion refers to the generation of diverse question expressions using a language model. This method can be implemented using prompt word templates to guide generation and expand the coverage of user intent. Intent recognition refers to determining whether a user question falls within the scope of a resource recommendation request. This method can be implemented using predefined intent classifications and a majority voting mechanism to determine whether to trigger the recommendation process. Structured query statements refer to converting user questions into database-executable filtering criteria. This method can be implemented by generating SQL statements based on semantically matched tag fields to accurately filter candidate resources. Semantic vector encoding refers to converting text into high-dimensional vector representations. This method can be implemented using an embedding model consistent with the vector library to calculate semantic similarity. The similarity threshold is the minimum matching criterion for screening candidate vectors. This can be achieved through cosine similarity calculation and threshold setting, and is used to filter out low-quality text. Resource splicing is the process of integrating filtered knowledge fragments into coherent answers. This can be achieved through sorting, splicing, and post-processing strategies to generate structured recommendation results.

[0030] This application uses a hybrid retrieval mechanism to integrate semantic vector matching and structured label filtering, and combines intent recognition and question expansion technology to solve the problems of insufficient accuracy and redundancy caused by a single retrieval mode in the existing RAG resource recommendation scenario, and achieve high-accuracy and high-coverage resource recommendation.

[0031] The working process and principles of this application are as follows: resource text data is collected and a vector library and a tag library are constructed based on the resource text data. The vector library is used to store semantic vector representations of resource text, and the tag library is used to store structured tag information for resources. User questions are collected and semantically expanded based on a language model to generate multiple expanded questions with similar semantics but different expressions. This step aims to increase question coverage and improve the recall rate of relevant resources. Intent recognition is then performed on all semantically expanded questions to determine whether the user question is a resource recommendation question. If it is a resource recommendation question, the target classification type of the resource is further determined. This step can filter out non-recommendation queries and improve efficiency. User questions and semantically expanded questions are converted into structured query statements. The data is filtered according to the field definitions in the tag library to obtain a set of candidate knowledge fragments that meet the tag requirements. This step uses the structured tag information for precise screening and narrows the candidate range. Semantic vector encoding is performed on each resource text in the candidate knowledge fragment set to obtain a candidate vector. Simultaneously, the user question and semantically expanded question are mapped into a query vector. The similarity between the query vector and each candidate vector is calculated, and fragments with a similarity greater than a threshold are selected as knowledge supplementary content. This step, combined with vector similarity matching, further improves retrieval accuracy. Finally, all supplementary knowledge content is spliced ​​together to generate the final resource recommendation answer. This entire process combines structured tag screening and semantic vector matching, realizing a hybrid retrieval RAG resource recommendation method.

[0032] As a preferred embodiment, the solution of this application is specifically implemented as follows: Collect text data from resources, including policy documents, scientific research reports, and patent documents. Preprocess the collected text data, including removing HTML tags, standardizing punctuation, and normalizing spaces. Then, construct a vector library and a tag library. To build the vector library, use the BERT model to encode the text in chunks, generating a 768-dimensional vector for each chunk. To build the tag library, extract keywords and entities from the text, such as policy type, applicable field, and release date, to form structured tags.

[0033] After collecting user questions, the T5 model is used to rewrite the questions, generating 5-10 semantically similar expanded questions. For example, the original question "What are the equipment sharing policies in the biopharmaceutical field?" might be expanded to include "What are the regulations for open use of instruments and equipment in the biopharmaceutical industry?"

[0034] We identify the intent of the extended question and use the pre-trained BERT classifier to determine whether it is a resource recommendation question. If so, we further identify the target resource type, such as "policy document."

[0035] Convert the user question into an SQL query, such as "SELECT * FROM policy WHERE field=biomedicine AND type=equipment sharing." Execute the SQL query to obtain candidate knowledge fragments.

[0036] The candidate segments are vector-encoded using the same BERT model as the vector library. The cosine similarity between the user question vector and the candidate vector is calculated, and segments with a similarity greater than 0.8 are selected as supplementary knowledge content.

[0037] Finally, the selected knowledge content is sorted and spliced ​​by similarity to generate recommended answers. The answers include information such as the policy name, main content overview, and scope of application.

[0038] Through the above scheme, this application realizes the hybrid retrieval of structured tags and semantic vectors, which improves the accuracy of resource recommendations. The generation of semantic expansion questions increases the recall rate of potentially relevant resources. The intent recognition mechanism avoids invalid retrieval of non-recommendation questions. The automatic generation of structured query statements realizes accurate label condition screening. The setting of semantic similarity threshold ensures the quality of knowledge supplement content. Overall, this scheme improves the performance of resource recommendation in complex query scenarios.

[0039] In some of the above-mentioned solutions of this application, a solution of constructing a vector library and a label library is proposed to support resource recommendation. However, in this process, the resource text data may have problems such as format confusion, noise interference, and unreasonable semantic segmentation, resulting in insufficient accuracy of subsequent vector encoding and label extraction. At the same time, the label library is not structured enough, making it difficult to support accurate conditional screening.

[0040] The present application further proposes a scheme for collecting resource text data and constructing a vector library and a label library respectively according to the resource text data, including: preprocessing the resource text data, the preprocessing including text segmentation, formatting and noise removal. When constructing the vector library, the preprocessed resource text data is divided into blocks according to semantic logic, and the block segmentation includes segmentation according to title level and parent-child semantic structure segmentation. Each block of text is vectorized and encoded based on the semantic embedding model to obtain a semantic vector. A mapping relationship is established between the semantic vector and the original resource text data, and the mapping relationship is stored in a database to obtain a vector library. When constructing the label library, keywords and entities related to the application conditions are extracted from the preprocessed resource text data, and label names and label values ​​are constructed. The resource content is labeled based on the rule template to form a label expression. According to the business logic, several label expressions are combined to establish a label condition tree, and the label content is bound to the corresponding resource fragment. The bound label information is organized into a structured table form and stored in the database to obtain a label library.

[0041] Among them, the preprocessing step divides long texts into logical paragraphs through text segmentation, unifies the text encoding and typesetting format through regular formatting, and removes noise to filter out non-text symbols and irrelevant content. Block segmentation uses title level segmentation to identify chapter structure, and parent-child semantic structure segmentation to identify the logical nested relationship between paragraphs. Vectorized encoding uses pre-trained semantic embedding models, such as BERT or Sentence-BERT, to map text blocks into 768-dimensional vectors. Label names and label values ​​are extracted through named entity recognition and keyword extraction algorithms, such as the entity recognition model based on BERT-CRF. The label condition tree combines label expressions through logical operators, such as the nested combination of AND, OR, and NOT.

[0042] Specifically, the preprocessing phase first cleans the original resource text, for example, by removing HTML tags, standardizing punctuation, and merging excess spaces, ensuring that subsequent processing is based on normalized text. During the segmentation phase, for policy documents or technical documents, the text blocks are divided into chapter-level blocks, for example, "Chapter 1" through "Chapter 3" are treated as separate blocks. Parent-child relationships between paragraphs are also identified, such as merging a definition paragraph and its explanation paragraph into the same block. During vector library construction, each text block is encoded using a semantic embedding model, indexed and mapped to the original text, and stored in a vector database such as Milvus or FAISS. During tag library construction, rule templates are used to match conditional fields in resource content. For example, "Applicable to: Enterprise" is annotated with the tag name "Applicable to" and the tag value "Enterprise." Multiple tag expressions are combined using a conditional tree, for example, "(Label A = Value 1 AND Label B = Value 2) OR Label C = Value 3." Finally, the tag names, tag values, and associated resource fragment IDs are stored in a structured table. Therefore, the vector library supports semantic similarity retrieval, and the tag library supports structured conditional screening. The two work together to improve the accuracy and efficiency of resource recall.

[0043] As a preferred embodiment, the solution of this application is specifically implemented as follows: When collecting resource text data and building a vector library and a label library based on the resource text data, the following steps are included: Preprocess resource text data. Preprocessing includes text segmentation, formatting, and noise removal. Specifically, text segmentation can be done using natural paragraphs or fixed-length segments. Formatting includes standardizing fonts, font sizes, and line spacing. Noise removal includes removing special characters and HTML tags.

[0044] Build a vector library. Segment the preprocessed resource text data into blocks based on semantic logic. This segmentation includes segmentation by title level and parent-child semantic structure. For example, you can segment based on document title level or parent-child structure based on semantic relationships between paragraphs.

[0045] Each block of text is vectorized and encoded based on a semantic embedding model to obtain a semantic vector. Pre-trained language models such as BERT and RoBERTa can be used as semantic embedding models to map text into a fixed-dimensional vector representation.

[0046] Establish a mapping relationship between the semantic vector and the original resource text data, and store the mapping relationship in the database to obtain a vector library. This can be stored in the form of key-value pairs, with the vector as the key and the original text as the value.

[0047] Build a tag library. Extract keywords and entities related to the application conditions from the preprocessed resource text data and construct tag names and tag values. Use named entity recognition technology to extract key entities such as time, location, and organization name.

[0048] The resource content is labeled based on the rule template to form a label expression. The rule template can include an "if-then" structure to match the corresponding label according to the text content characteristics.

[0049] Combine multiple tag expressions based on business logic to create a tag condition tree, and bind the tag content to the corresponding resource fragment. A decision tree structure can be used to organize multiple tag conditions.

[0050] Organize the bound tag information into a structured table format and store it in a database to obtain a tag library. You can use a relational database to store it, and create fields such as resource ID, tag name, and tag value.

[0051] Through the above technical solutions, this application achieves efficient preprocessing and structured storage of resource text data. The construction of the vector library supports subsequent semantic similarity retrieval, while the tag library provides the basis for precise conditional screening. This hybrid indexing method improves the accuracy and efficiency of resource retrieval, effectively addressing the limitations of a single retrieval method. At the same time, tagging makes resource content easier to manage and classify, providing reliable data support for subsequent personalized recommendations.

[0052] In some of the above-mentioned solutions of this application, the user question expression form is single, resulting in insufficient semantic expansion and failure to cover diverse query intentions, affecting the recall rate and accuracy of subsequent retrieval.

[0053] This application further proposes collecting user questions, expanding user questions based on language models, and obtaining several semantically extended questions, including: rewriting user questions based on language models, generating multiple semantically different but query-meaningful questions. Figure 1Before question rewriting, a prompt word template is constructed to guide the language model to generate diverse questions. The prompt word template includes instruction information, language style requirements, and context memory loading fields. During the question rewriting process, the current session history information is combined.

[0054] Among them, question rewriting processing uses a language model to generate an extended question that is semantically equivalent to the original question but expressed in a different form. For example, "How to apply for scientific research funding" can be rewritten as "What are the steps in the scientific research funding application process" or "What materials are needed to obtain scientific research funding." The prompt word template constrains the generation direction through preset instruction information, such as requiring the generation of interrogative sentences or declarative sentences containing keywords. The language style requirement specifies the degree of formality or colloquialism of the generated question. The context memory loading field is used to inject domain terminology or business rules. The conversation history information obtains the content that has been interacted in the current conversation through a caching mechanism. For example, the previous question "Application conditions for scientific and technological innovation projects" is used as context to associate the relevant content of the "Application materials list" when generating the extended question.

[0055] Specifically, after receiving the user's original question, the language model first loads the instruction information from the prompt word template, such as "Generate three different question expressions, including the keywords 'application conditions' and 'material list'." The model then adjusts the generated sentence structure based on the language style requirements of the template, using, for example, the interrogative sentence "What conditions are required to apply for a science and technology innovation project?" or the declarative sentence "Please explain the qualifications required to apply for a science and technology innovation project." The context memory loading field uses the "national project type" mentioned in the current conversation as a constraint to ensure that the generated expanded questions are limited to a specific category. During the question rewriting process, the model combines the previous conversation content stored in the cache, for example, linking "application process" with "material list" to generate expanded questions such as "the specific steps and required documents from application to approval for a national science and technology innovation project." The multiple semantic expanded questions generated are then subjected to majority voting by the subsequent intent recognition module to select valid queries that meet the resource recommendation intent, thereby improving the coverage of the search criteria and the matching accuracy of the candidate knowledge fragments.

[0056] As a preferred embodiment, the solution of this application is specifically implemented as follows: After collecting user questions, we rewrite them based on the language model. The question rewriting process generates multiple queries with different semantic expressions but different query meanings through the GPT-3.5 model. Figure 1The question expression form is constructed. Before performing question rewriting, a prompt word template is constructed to guide the language model to generate diversified questions. The prompt word template includes instruction information, language style requirements, and context memory loading fields. The instruction information is "please rewrite the following question and generate 5 questions with different expressions but similar semantics". The language style requirement is "use concise and colloquial expressions". The context memory loading field includes the user's historical question and answer records. In the question rewriting process, the current session history information is combined to ensure that the generated question is consistent with the user's query intent.

[0057] For example, for the user question "What are the policy supports for college students to start a business?", after question rewriting processing, the following semantic expansion questions may be generated: 1. What are the preferential policies for college students to start a business? 2. What are the support measures for college students to start a business? 3. How does the government support college students to start a business? 4. What policy bonuses can college students starting a business enjoy? 5. What help can college students starting a business get? Through the above technical solutions, the present application can generate multiple questions with similar semantics but different expressions, expanding the semantic coverage of the question. By introducing the prompt word template and context information, the generated question is closer to the user's actual query intent. This method improves the accuracy and comprehensiveness of subsequent resource recommendation, effectively solving the problem of insufficient resource matching caused by single question expression.

[0058] In some of the above schemes of the present application, the intent recognition of the user question only depends on the single question input, which is easy to misjudge due to user expression ambiguity or diversity, and then cannot accurately filter the target resource category or trigger the refusal, affecting the recommendation accuracy and user experience.

[0059] The present application further proposes to construct a set of pre-defined intent category list, the intent category includes resource recommendation and non-resource class consultation, input all semantic expansion questions, generate intent classification results, multiple voting on the recognition results of multiple semantic expansion questions, if the result belongs to the resource recommendation class intent, then determine the target classification type of the resource according to the semantic keywords. Otherwise, terminate the recommendation process and trigger the refusal.

[0060] Among them, the intent category list distinguishes resource recommendation class from non-resource class consultation in a predefined manner, covering the possible core demand types of users. The semantic expansion problem inputs the intent classification model through diversified expression forms, improving the input diversity of intent recognition. The majority voting mechanism reduces the probability of misjudgment of a single problem by counting the classification results of multiple expansion problems. The semantic keyword matching determines the specific target classification type based on the keyword library in the resource classification system and the intent result output by the classification model.

[0061] Specifically, after the user question is generated into multiple variants through semantic expansion, each expansion question is input into a pre-trained intent classification model, which outputs the intent category probability distribution of each question based on the classification label list. The classification results of multiple questions are counted by majority voting. If the proportion of resource recommendation class intent exceeds the preset threshold, it is determined that the user demand belongs to the resource recommendation class. Further, for the user question determined as the resource recommendation class, the keywords related to resource classification are extracted from the original question and the expansion question, such as “policy”, “scientific research service”, “scientific and technological resource”, etc. The target classification type of the resource is determined by matching the keywords with the preset classification label. If the majority voting result shows a non-resource class intent, the recommendation process is terminated and a refusal prompt is returned. This process reduces the risk of misjudgment of a single question through multiple question input and majority voting, improves the classification accuracy through keyword matching, and ensures that the resource screening conditions are consistent with the real needs of users.

[0062] As a preferred embodiment, the scheme of the present application is implemented as follows: A set of predefined intent category list is constructed, including resource recommendation and non-resource class consultation. The semantic expansion question is input to generate the intent classification result. The recognition results of multiple semantic expansion questions are counted by majority voting. If the result belongs to the resource recommendation class intent, the target classification type of the resource is determined according to the semantic keywords. Otherwise, the recommendation process is terminated and a refusal is triggered.

[0063] Specifically, the predefined intent category list includes “resource recommendation”, “policy consultation”, “business handling” and other categories. Each semantic expansion question is processed by an intent classification model to obtain the corresponding intent category. For example, for the expansion question “What are the suitable technological project funding for start-ups”, the intent classification result is “resource recommendation”. For “How to apply for a business license”, the intent classification result is “business handling”.

[0064] Furthermore, the intent classification results for all extended questions are tallied, and a simple majority vote is used to determine the final intent. If more than half of the extended questions are classified as "resource recommendation," the user's question is considered to be in the resource recommendation category. Keywords such as "startup" and "technological project" are extracted from the extended questions and matched to the pre-set resource classification system, determining the target category as "technological innovation funding."

[0065] If the final intent does not fall under the resource recommendation category, the subsequent resource search and recommendation process will be terminated. Instead, the rejection mechanism will be triggered to explain to the user that the current question is not applicable to resource recommendation, and the user will be guided to ask the question again or be transferred to other service modules.

[0066] Through the above technical solutions, this application achieves accurate identification and classification of user question intent. Through the voting mechanism of multiple semantically extended questions, the robustness of intent judgment is improved, and the misjudgment that may be caused by a single question statement is avoided. At the same time, for non-resource recommendation questions, unnecessary retrieval processes are terminated in a timely manner, thereby improving response efficiency. In addition, the determination of target classification lays the foundation for subsequent precise resource matching, which helps to improve the relevance and pertinence of recommendation results.

[0067] In some of the above-mentioned solutions of this application, it is proposed to convert user questions and semantic expansion questions into structured query statements for filtering tag library data. However, when generating structured query statements, the statement may not be parsed or the query may fail due to complex tag library field definitions or semantic understanding deviations, resulting in inaccurate or incapable of execution of filtering conditions.

[0068] The present application further proposes to extract the table structure information of the tag library, and the table structure information includes an enumeration set of field names, field types, field comments, tag names and tag values. According to the semantic matching of the field descriptions and annotations in the tag library of the user question and the semantic extension question, a structured SQL statement that meets the semantic intent is generated. The structured SQL statement includes a WHERE statement, which is used to specify the matching conditions of the tag field and the tag value. If the structured SQL statement cannot be parsed or the query fails, the fallback strategy rule is enabled to match the rules for mapping keywords to tag names, and the SQL query conditions are manually constructed. Query operations are performed in the tag library based on the structured SQL statement to obtain a set of candidate knowledge fragments that match the semantic conditions of the user question.

[0069] When extracting table structure information, the tag library metadata is parsed to obtain field names, types, and annotations, and then combined with the enumerated set of tag values ​​to form a complete set of field definitions. When generating structured SQL statements, natural language understanding technology is used to map the semantic elements in the user's question to the tag library fields, and the WHERE conditional expressions are automatically combined. The fallback strategy rules are based on a preset tag name and keyword mapping table. When automatic generation fails, the query conditions are manually constructed through keyword matching to ensure that the filtering logic is executable.

[0070] Specifically, the extraction of table structure information provides field constraints for structured queries to avoid syntax errors caused by mismatches in field types or value ranges. When semantic matching generates SQL statements, the query conditions are dynamically combined by associating field descriptions with user question keywords to improve screening accuracy. The fallback strategy is triggered when automatic generation fails, and keyword mapping rules are used to construct query conditions to ensure the robustness of the screening process. For example, when a user question contains "scientific research funding support", semantic matching maps it to the "funding type" field of the tag library and generates an SQL statement containing "WHERE funding type = scientific research funding". If the field mapping fails, the label name is matched according to the "scientific research funding" keyword to construct a conditional expression. Through the above steps, ensure that the screening conditions of the candidate knowledge fragment set are consistent with the user's intentions. Figure 1 This improves the efficiency and accuracy of subsequent vector retrieval.

[0071] As a preferred embodiment, the solution of this application is specifically implemented as follows: Extract the tag library's table structure, including field names, field types, field annotations, and an enumeration of tag names and tag values. For example, for a policy resource tag library, the table structure might include fields such as "Policy Type," "Applicable to," and "Release Date." Each field has a corresponding data type, annotation, and possible value range.

[0072] Based on the semantics of the user question and the semantically extended question, the tag library's field descriptions and annotations are matched to generate a structured SQL statement that meets the semantic intent. The SQL statement includes a WHERE clause, which specifies the matching conditions between tag fields and tag values. For example, for the user question "What are the science and technology innovation policies released in 2023?", the following SQL query might be generated: SELECT * FROM policy_table WHERE policy_type = Technological Innovation AND release_year = 2023; If the structured SQL statement cannot be parsed or the query fails, the fallback strategy rule is enabled to match the rule mapping keywords and tag names, and the SQL query conditions are manually constructed. For example, if the "Technological Innovation" tag cannot be directly matched, the fuzzy matching rule may be used to expand the query conditions: SELECT * FROM policy_table WHERE policy_type LIKE %Technology% OR policy_type LIKE %Innovation% AND release_year = 2023; Execute query operations in the tag library based on structured SQL statements to obtain a set of candidate knowledge fragments that match the semantic conditions of the user's question. These candidate knowledge fragments will serve as input for subsequent semantic vector matching to further improve retrieval accuracy.

[0073] Through the above technical solution, this application realizes the automatic conversion of user questions into structured queries, thereby improving the accuracy of resource retrieval. By combining the field definitions of the tag library for data screening, the interference of irrelevant resources is reduced and the retrieval efficiency is improved. At the same time, by setting a fallback strategy, it is ensured that relevant results can still be returned in the case of complex queries. This method effectively solves the limitations of traditional RAG systems in processing resource queries in fields with high semantic complexity, and provides users with more accurate and relevant resource recommendations.

[0074] In some of the above-mentioned schemes of this application, the resource texts in the candidate knowledge fragment set have format differences, redundant noise and length redundancy problems, which lead to inconsistencies in the semantic vector encoding process, affect the accuracy of vector similarity calculation, and thus reduce the relevance of the recall results.

[0075] The present application further proposes to encode semantic vectors for each resource text in the candidate knowledge fragment set. When obtaining the candidate vectors, the process includes: normalizing each resource text in the candidate knowledge fragment set, wherein the normalization process includes removing HTML tags from the text, normalizing punctuation, normalizing spaces, and trimming the length. The candidate knowledge fragments are encoded using a semantic embedding model that is consistent with the vector library construction phase. The semantic embedding model maps each candidate knowledge fragment into a semantic vector of fixed dimension to obtain a candidate vector. The candidate vector is obtained during the encoding process using a mean pooling or CLStoken pooling strategy.

[0076] Among them, the normalization processing contains four operation layers: the HTML tag removal layer deletes the hyper-text markup language elements in the document through regular expression matching. The punctuation standardization layer converts full-width symbols into half-width symbols and unifies different forms of punctuation expressions. The space normalization layer combines consecutive spaces and eliminates tab characters. The length clipping layer sets a maximum character number threshold and truncates the part exceeding the threshold. The semantic embedding model adopts the same pre-training model parameters as the vector library construction stage to ensure that the candidate vector and the semantic vector in the vector library are in the same spatial coordinate system. The pooling strategy is selected according to the model type, and for the model based on the Transformer, the CLS token pooling is preferred, and for the model based on the BiLSTM, the mean pooling is adopted.

[0077] Specifically, before the candidate knowledge fragment enters the encoding stage, the HTML tag removal layer identifies the hyper-text markup language elements in the text and deletes them. The punctuation standardization layer converts full-width symbols into half-width symbols and unifies different forms of punctuation expressions. The space normalization layer combines consecutive spaces and eliminates tab characters. The length clipping layer sets a maximum character number threshold and truncates the part exceeding the threshold. <tag>Format string and perform deletions, such as replacing " Policy Interpretation ” is converted to “policy interpretation”. The punctuation normalization layer converts full-width Chinese commas "," to half-width commas "", and unifies wavy lines "~" into dashes "—". The space normalization layer compresses consecutive spaces in the text into single spaces, eliminating the blank spaces caused by tabs. The length clipping layer sets the maximum number of characters to 512 and truncates any excess characters at the end. The normalized text is input into the semantic embedding model. After the model outputs the hidden layer vector for each token, the pooling method is selected based on the model architecture: if the model contains a CLS token, the vector corresponding to the CLS position is extracted as the candidate vector. If the model uses a recurrent neural network structure, the mean of all token vectors is calculated as the candidate vector. This process ensures that the candidate vector maintains dimensional and spatial consistency with the original encoding in the vector library, avoiding semantic shifts caused by preprocessing differences.

[0078] As a preferred embodiment, the solution of this application is specifically implemented as follows: Each resource text in the candidate knowledge fragment set is normalized. Normalization includes removing HTML tags, punctuation standardization, space standardization, and length trimming. Specifically, first use regular expressions to match and remove HTML tags, such as 、 Then, we standardized Chinese and English punctuation to a standard format, such as converting full-width commas "," to half-width commas "". We then normalized the spaces in the text, removing excess spaces and adding appropriate spaces after punctuation marks. Finally, we trimmed the text to the first 512 characters.

[0079] Furthermore, the candidate knowledge fragments are encoded using a semantic embedding model consistent with the vector library construction phase. The semantic embedding model uses a pre-trained BERT model, which takes normalized text as input and outputs a 768-dimensional vector representation.

[0080] The semantic embedding model then maps each selected knowledge fragment into a fixed-dimensional semantic vector, obtaining a candidate vector. During the encoding process, these candidate vectors are obtained using either mean pooling or CLS token pooling strategies. Specifically, the mean pooling strategy averages the hidden states of all tokens in the last layer of BERT, while the CLS token strategy directly uses the hidden state corresponding to the [CLS] tag as the representation of the entire sequence.

[0081] Through the above technical solutions, the present application realizes the normalization processing and vectorized representation of candidate knowledge fragments. Normalization processing improves the text quality, eliminates format inconsistencies, and lays the foundation for subsequent vectorization. The use of a semantic embedding model consistent with the vector library ensures the consistency of the vector space, which is conducive to the accuracy of subsequent similarity calculations. The mean pooling or CLS token strategy is used to obtain a fixed-dimensional vector representation, which not only captures the semantic information of the entire text, but also facilitates subsequent vector operations. This series of processing improves the representation quality of candidate knowledge fragments, provides a reliable foundation for subsequent semantic matching and similarity calculations, and thus enhances the accuracy and relevance of resource recommendations.

[0082] In some of the above-mentioned schemes of the present application, when recalling knowledge fragments based on semantic vector similarity, screening only by a single similarity threshold may result in fragments with high similarity but repeated content being selected at the same time, or the presence of multiple fragments with similarity but scattered semantics, affecting the accuracy and information density of the final recommended content.

[0083] This application further proposes encoding user questions based on a semantic embedding model, mapping the user question and semantically expanded question into a query vector. The query vector and candidate vector share the same dimensionality and semantic space. Cosine similarity is calculated between each candidate vector and the query vector to obtain a similarity score. All similarities are then compared against a similarity threshold to obtain supplementary knowledge content.

[0084] The semantic embedding model uses the same model architecture and parameter configuration as the vector library construction phase, ensuring that the query vector and candidate vectors are in the same semantic space. Cosine similarity is calculated using the vector dot product and the module length ratio, and the calculation process is accelerated by matrix operations. The similarity threshold is dynamically adjusted based on the recall accuracy of historical data, and the threshold range is set between 0.7 and 0.9. When multiple resource texts meet the threshold conditions, a secondary sorting strategy based on semantic relevance is adopted. By calculating the semantic similarity matrix between candidate segments, redundant content is removed while retaining diversity.

[0085] Specifically, the user question is encoded by the semantic embedding model to generate a 768-dimensional query vector, and each text in the candidate knowledge fragment set is encoded by the same model to generate a 768-dimensional candidate vector. The cosine similarity between the query vector and each candidate vector is calculated to form a similarity score list. The score list is compared with the preset similarity threshold of 0.8 to filter out candidate fragments with scores higher than the threshold. When there are multiple candidate fragments that meet the conditions, the semantic similarity between the fragments is further calculated. If the similarity between two fragments exceeds 0.95, they are judged as duplicate content, and only the fragment with the highest similarity is retained. Finally, they are sorted in descending order by similarity score, and the top 3 non-duplicate fragments are selected as knowledge supplementary content. This process improves content quality through a double filtering mechanism, in which threshold screening eliminates low-relevance content, and semantic deduplication and sorting optimize information density to ensure that the recommendation results meet both accuracy and diversity requirements.

[0086] As a preferred embodiment, the solution of this application is specifically implemented as follows: The user question is encoded using a semantic embedding model, and the user question and semantically expanded question are mapped into a query vector. The query vector and candidate vectors have the same dimension and semantic space. Cosine similarity is calculated between each candidate vector and the query vector to obtain the similarity of each candidate vector. All similarities are compared with a similarity threshold to obtain supplementary knowledge content.

[0087] Specifically, we first use the pre-trained BERT model as a semantic embedding model. The user question is input into the BERT model, and the hidden state corresponding to the last layer [CLS] tag is extracted as a 768-dimensional query vector. For each candidate vector, its cosine similarity with the query vector is calculated. The similarity threshold is set to 0.7, and the resource text corresponding to the candidate vector with a similarity greater than 0.7 is used as the knowledge supplement content. If multiple resource texts meet the conditions, a semantic relevance analysis is performed, and they are sorted from high to low similarity and from high to low relevance. The top three are selected as the final knowledge supplement content.

[0088] Through the above technical solution, this application achieves precise knowledge matching based on semantic similarity. This improves the accuracy and relevance of resource recommendations and avoids the inclusion of irrelevant or low-quality text. Furthermore, by setting similarity thresholds and semantic relevance ranking, the quality and diversity of recommended content are ensured, enhancing the user experience.

[0089] In some of the above-mentioned schemes of this application, when there are multiple resource texts with similarities greater than or equal to a threshold in the candidate knowledge fragment set, directly selecting all qualified fragments may result in content redundancy or uneven quality, affecting the accuracy of the recommendation results and user satisfaction.

[0090] This application further proposes that when the similarity is greater than or equal to the similarity threshold, the resource text corresponding to the similarity is used as the knowledge supplement content. If there are multiple resource texts that meet the conditions, the multiple resource texts are subjected to semantic relevance analysis, and are sorted from high to low in terms of similarity and from high to low in terms of relevance, and the top n items are selected as the knowledge supplement content.

[0091] Semantic relevance analysis is performed by calculating the semantic distance between candidate segments, using a pre-trained model-based approach to identify inter-sentence relationships. The number of selected top n recommendations is dynamically adjusted based on the preset length of the recommended results, with n ranging from 3 to 10. Similarity and relevance are weighted using a weighting mechanism, with a similarity weight of 0.7 and a relevance weight of 0.3. A sliding window strategy is used during the sorting process to perform a secondary check on candidate segments with adjacent scores.

[0092] Specifically, after obtaining candidate segments that meet the similarity threshold, they are first sorted in descending order according to the cosine similarity value to form an initial queue. Subsequently, the semantic relationship between each segment in the queue and other segments is modeled, and the content duplication is evaluated by calculating the Euclidean distance between sentence vectors, and redundant segments with a duplication higher than 0.9 are eliminated. The remaining segments are re-sorted according to the weighted sum of the similarity score and the relevance score. The weighted total score is calculated as: total score = similarity × 0.7 + (1-duplication) × 0.3. Finally, the top 5 with the highest total score are selected as supplementary knowledge content. This number is automatically adapted according to the length of the user's question and automatically expanded to 8 when the question contains more than 3 entities. This dual sorting mechanism ensures that high-matching segments are presented first, while eliminating semantically overlapping content through relevance filtering, so that the recommendation results are both accurate and diverse.

[0093] As a preferred embodiment, the solution of the present application is specifically implemented as follows: after calculating the similarity between the candidate vector and the query vector, all resource texts in the candidate knowledge fragment set whose similarity reaches a preset threshold of 0.85 are screened as a preliminary candidate set. For the case where the preliminary candidate set contains three resource texts, the semantic relevance score between each text and the user question is calculated separately: the title keywords of the first resource text match the user question by 92%, the contextual logical coherence score of the second resource text reaches 0.78, and the entity coverage rate of the third resource text reaches 85%. After arranging in descending order of similarity, the similarity of the first resource text is 0.91, the second is 0.88, and the third is 0.86. The sorting is further adjusted according to the semantic relevance score, and the second resource text is promoted to the first place. Finally, the first two resource texts are selected as knowledge supplementary content, and a recommended answer containing the two texts is generated.

[0094] Through the above technical solution, this application solves the existing problem of low-quality text recall caused by relying solely on vector similarity. It effectively improves the relevance and accuracy of recommended content through a dual sorting mechanism. Specifically, when multiple resource texts meet the similarity threshold, a secondary sorting is performed in conjunction with semantic relevance analysis, prioritizing content with more complete semantic logic and more comprehensive keyword coverage. This avoids the semantic bias caused by relying solely on vector space distance and ensures that the final recommendation results meet both similarity and semantic coherence requirements.

[0095] In some of the above-mentioned solutions of this application, during the process of generating resource recommendation answers, directly splicing knowledge supplementary content may lead to redundant answers, inconsistent terminology or chaotic structure, affecting the accuracy and readability of the recommendation results.

[0096] This application further proposes to splice all supplementary knowledge content to generate resource recommendation answers, including sorting the selected supplementary knowledge content from high to low according to their similarity to the query vector, and sequentially splicing the sorted content into the context input. The splicing process inserts delimiters between each piece of content and embeds the original user question and semantically expanded question into the spliced ​​context. The generated answer output is post-processed, including removing invalid sentences, unifying terminology, and de-duplicating to obtain resource recommendation answers.

[0097] The sorting operation establishes priorities through similarity metrics, ensuring that highly relevant content is presented first. Inserting separators prevents semantic confusion between different knowledge fragments and maintains a clear text structure. Embedding original and expanded questions maintains contextual coherence and prevents generated content from deviating from user intent. The post-processing stage uses a sentence filtering mechanism to remove duplicate or meaningless text, unifies professional vocabulary expressions through a term mapping table, and merges similar information based on semantic clustering.

[0098] Specifically, when generating resource recommendation answers, the candidate knowledge fragments are first sorted in descending order of similarity to form a priority queue. A vertical bar symbol is used as a separator to insert between adjacent content, and the original user question and its extended question are embedded in the header of the spliced ​​text as a generation prompt. For example, when processing 5 candidate contents, the top 3 with the highest similarity are selected and spliced ​​in order, with a "|" separator added between each one. After generating the preliminary answer, regular expression matching is used to delete sentences containing uncertain words such as "maybe" and "probably", and a domain term dictionary is used to uniformly replace "LLM" with "large language model", and content involving the same policy terms is merged. The final output is a recommended answer with complete paragraph structure, standardized terminology, and no redundant information.

[0099] As a preferred embodiment, the solution of this application is specifically implemented as follows: After the candidate knowledge fragments are sorted from high to low by similarity, the top five sorted pieces of content are concatenated in sequence to form the contextual input, using the horizontal dividing line in Markdown syntax as a delimiter. During the concatenation process, a metadata identifier of "resource number-similarity value" is added to the header of each knowledge fragment, and the text content of the original user question and three semantically expanded questions is inserted at the beginning of the concatenated text. The generated contextual input is fed into the pre-trained language model for text generation. The output is processed by the post-processing module, and the following operations are performed: regular expression matching is used to remove sentences containing "may contain errors" or "for reference only." The terminology of policy document content is uniformly adjusted to the expression form defined in the "State Council Standardization Format Specification." For repeated occurrences of the same resource entry, only the most similar version is retained.

[0100] Through the above technical solutions, this application effectively solves the problems of redundant content, loose structure, and inconsistent terminology in resource recommendation answers. By standardizing sorting and separators, it ensures that highly relevant knowledge fragments are prioritized in the generation process, improving the logical coherence of the answers. Through metadata identification and question embedding mechanisms, the generation model's ability to capture contextual relevance is enhanced. Through deduplication and terminology standardization operations in the post-processing process, the probability of interference from invalid information is reduced, ensuring that the output content conforms to the expression standards of the professional field.

[0101] In some of the above-mentioned solutions of this application, the existing technology cannot effectively utilize the structured tag information in the resource data for precise conditional filtering, resulting in redundant returned content or matching deviation. At the same time, user questions are difficult to automatically convert into operational structured retrieval logic, and relying solely on vector space recall is prone to introducing irrelevant or low-quality text.

[0102] In the above embodiment, by integrating structured label screening and semantic vector precision ranking mechanism, the problems of recall redundancy and matching deviation caused by relying solely on semantic similarity retrieval in the existing RAG system are overcome. The resource text is dual-modeled: on the one hand, a semantic vector library is constructed to capture the deep semantic information of the resource content. On the other hand, a label library is constructed to extract the conditional attribute information of policies or resources and support precise matching based on field rules. The user questions are diversified and semantically expanded through the language model to improve the recall coverage. At the same time, the accuracy of the recommendation logic is guaranteed through intent recognition and resource type classification. The user intention is automatically converted into SQL form to realize conditional screening of the label field and obtain highly relevant candidate knowledge fragments. The candidate content is semantically encoded and similarity calculated, the threshold is set and precision ranking is performed, and the recommended answers with reasonable structure and strong semantic pertinence are generated through context splicing. The accuracy of resource recommendation, context matching and controllability of generated content are improved.

[0103] In another preferred embodiment based on the above embodiment, refer to Figure 2 As shown, this embodiment provides a resource recommendation system based on hybrid retrieval RAG, which is used to apply the above-mentioned resource recommendation method based on hybrid retrieval RAG, including: The collection unit is configured to collect resource text data and construct a vector library and a label library according to the resource text data.

[0104] The processing unit is configured to collect user questions, expand the user questions based on the language model, and obtain a plurality of semantically expanded questions.

[0105] The recognition unit is configured to perform intent recognition on all semantic extension questions, determine whether the user question is a resource recommendation question, and if so, determine the target classification type to which the resource belongs.

[0106] The screening unit is configured to convert user questions and semantically extended questions into structured query statements, screen the data according to the field definitions in the tag library, and obtain a set of candidate knowledge fragments that meet the tag conditions.

[0107] The matching unit is configured to perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector. Based on the semantic encoding, the user question and semantically expanded question are mapped into a query vector. The similarity between the query vector and each candidate vector is calculated, and segments with similarity greater than a similarity threshold are selected as supplementary knowledge content.

[0108] The output unit is configured to perform resource splicing on all knowledge supplementary contents and generate resource recommendation answers.

[0109] Specifically, the collection unit preprocesses the resource text data, divides and encodes the vector library according to semantic logic blocks, extracts keywords and entities to build a label library, and stores the binding relationship between the label condition tree and the resource segment in a structured table. The processing unit generates semantic expansion questions based on language models, combines prompt word templates and conversation history information to guide the generation process. The recognition unit classifies multiple semantic expansion questions using a predefined intent category list, determines the intent type through majority voting, and identifies the target classification type if it belongs to the resource recommendation category. The screening unit extracts structured information from the label library table, converts the user question into a structured SQL statement for label field matching query, and uses rule matching to construct query conditions if the query fails. The matching unit normalizes the candidate knowledge segments, uses the same semantic embedding model as the vector library to encode candidate vectors, calculates the cosine similarity between the user question query vector and the candidate vector, and selects the knowledge supplement content according to the threshold screening and sorting. The output unit sorts and concatenates the knowledge supplement content according to the similarity, inserts a separator and embeds the user question, removes invalid statements through post-processing, and generates the final answer by unifying the term style.

[0110] Specifically, the system implements a hybrid retrieval mechanism by constructing a structured label library and a semantic vector library. The processing unit generates diversified semantic expansion questions to improve the robustness of intent recognition, and the recognition unit reduces the probability of misjudgment through majority voting. The screening unit converts natural language questions into structured query statements and uses field definitions in the label library for accurate condition filtering to reduce redundant data. The matching unit combines semantic vector similarity calculation and threshold screening to ensure the relevance of the recalled content. The output unit sorts and concatenates multiple source content and performs post-processing to improve the coherence and accuracy of the answer. Each unit works together to combine structured label retrieval and semantic vector matching, solving the problems of insufficient condition filtering, intent recognition bias, and low-quality recall in existing technologies, and achieving high-precision resource recommendation.

[0111] As a preferred embodiment, the solution of this application is implemented as follows: The resource recommendation system comprises six collaborative modules. The collection unit uses a distributed crawler program to retrieve policy documents from the government data platform. It then uses a text cleaning tool to segment the raw data and divide the cleaned text into semantically coherent paragraphs. The paragraph segmentation utilizes a three-level segmentation strategy based on the heading hierarchy. The processing unit uses a pre-trained language model interface to convert the user input, "How to apply for high-tech enterprise subsidies," into five expanded questions expressed in different sentence structures. This expansion process incorporates the user's previous R&D investment inquiries in the current session. The recognition unit uses an integrated classification model to analyze the intent of the expanded questions. When three expanded questions are identified as qualification certification inquiries, the resource recommendation process is triggered. The screening unit converts the user question into an SQL query containing "policy type = science and technology support" and "applicable to = enterprise" and selects knowledge fragments from the tag library that contain tax incentive clauses and application procedures. The matching unit uses a BERT-based model to vectorize candidate texts, calculates cosine similarity with the user question vector, and selects text fragments with a similarity exceeding 0.82. The output unit will sort the screened policy clauses according to the effectiveness level, insert page breaks between clauses when splicing, and finally generate a recommended reply including policy basis, application conditions and processing channels.

[0112] Through the above technical solutions, this application effectively solves the defects of traditional RAG systems in structured data utilization and intent recognition. By constructing a dual retrieval mechanism of tag library and vector library, a hybrid recall based on field condition screening and semantic similarity matching is achieved, reducing the false recall rate of irrelevant policy clauses. A combined strategy of multi-question expansion and integrated intent classification is adopted to increase the recognition accuracy of resource recommendation questions from 78% of single-question analysis to 92% of multi-question voting mechanism. The automatic generation mechanism of structured query statements enables the system to accurately match key fields such as applicable objects and validity periods of policy documents, solving the technical bottleneck that traditional semantic retrieval cannot handle precise condition filtering. The dual screening and sorting mechanism of knowledge fragments effectively avoids the mixing of low-quality texts, ensuring that the final recommended content is both in line with policy specifications and semantically relevant.

[0113] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0115] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention. < / tag>

Claims

1. A resource recommendation method based on hybrid retrieval RAG, characterized in that: include: Collect resource text data, and construct a vector library and a label library based on the resource text data; Collect user questions, expand the user questions based on the language model, and obtain several semantically expanded questions; Performing intent recognition on all the semantically extended questions to determine whether the user question is a resource recommendation question, and if so, determining the target classification type to which the resource belongs; Convert the user questions and semantic expansion questions into structured query statements, filter the data according to the field definitions in the tag library, and obtain a set of candidate knowledge fragments that meet the tag conditions; Perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector; map the user question and the semantic expansion question into a query vector based on the semantic encoding, calculate the similarity between the query vector and each candidate vector, and select segments with a similarity greater than a similarity threshold as knowledge supplementary content; All the supplementary knowledge contents are combined to generate resource recommendation answers.

2. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: Collecting resource text data and building a vector library and a label library based on the resource text data includes: Preprocessing the resource text data, including text segmentation, formatting, and noise removal; Constructing the vector library: dividing the pre-processed resource text data into blocks according to semantic logic, wherein the block division includes division by title level and division by parent-child semantic structure; Perform vector encoding on each text block based on the semantic embedding model to obtain a semantic vector; Establishing a mapping relationship between the semantic vector and the original resource text data, and storing the mapping relationship in a database to obtain the vector library; Build a tag library: Extract keywords and entities related to application conditions from preprocessed resource text data, and build tag names and tag values; Label resource content based on rule templates to form label expressions; Combine several tag expressions according to business logic, build a tag condition tree, and bind the tag content to the corresponding resource fragment; The bound tag information is organized into a structured table format and stored in a database to obtain the tag library.

3. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: Collect user questions, expand the user questions based on the language model, and obtain several semantically expanded questions, including: Rewrite the user question based on the language model to generate multiple question expressions with different semantic expressions but consistent query intent; Before executing question rewriting, guiding the language model to generate diversified questions by constructing a prompt word template, wherein the prompt word template includes instruction information, language style requirements, and context memory loading fields; During question rewriting, the current session history information is combined.

4. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: Performing intent recognition on all semantically extended questions to determine whether the user question is a resource recommendation question includes: Construct a set of predefined intent categories, including resource recommendation and non-resource consultation, input all the semantic expansion questions, and generate intent classification results. A majority vote is performed on the recognition results of multiple semantic expansion questions. If the result belongs to the resource recommendation intention, the target classification type of the resource is determined based on the semantic keywords; otherwise, the recommendation process is terminated and a rejection is triggered.

5. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: Converting the user question and semantic expansion question into a structured query statement, filtering the data according to the field definition in the tag library, and obtaining a set of candidate knowledge fragments that meet the tag conditions includes: Extracting table structure information of the tag library, wherein the table structure information includes: an enumeration set of field name, field type, field comment, tag name, and tag value; Semantically matching the field descriptions and annotations in the tag library based on the user question and the semantically extended question, generating a structured SQL statement that meets the semantic intent, wherein the structured SQL statement includes a WHERE statement for specifying a matching condition between the tag field and the tag value; If the structured SQL statement cannot be parsed or the query fails, the fallback strategy rule is enabled to match the rule of keyword and tag name mapping, and the SQL query condition is manually constructed; A query operation is performed in the tag library based on the structured SQL statement to obtain a set of candidate knowledge fragments that match the semantic conditions of the user question.

6. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: Performing semantic vector encoding on each resource text in the candidate knowledge fragment set to obtain a candidate vector includes: Normalizing each resource text in the candidate knowledge fragment set, wherein the normalization includes removing HTML tags, normalizing punctuation, normalizing spaces, and trimming the length of the text; Encode the candidate knowledge fragments using a semantic embedding model consistent with the vector library construction phase; The semantic embedding model maps each selected knowledge fragment into a semantic vector of fixed dimension to obtain the candidate vector, and the candidate vector is obtained by using mean pooling or CLS token pooling strategy during the encoding process.

7. The resource recommendation method based on hybrid retrieval RAG according to claim 6, characterized in that: Mapping the user question and the semantically expanded question into a query vector based on semantic coding, calculating the similarity between the query vector and each candidate vector, and selecting a segment with a similarity greater than a similarity threshold as knowledge supplementary content includes: Encoding the user question and the semantically expanded question based on the semantic embedding model, mapping the user question into a query vector, wherein the query vector and the candidate vector have the same dimension and semantic space; Calculating the cosine similarity between each candidate vector and the query vector to obtain the similarity of each candidate vector; All the similarities are compared with the similarity threshold respectively to obtain the knowledge supplement content.

8. The resource recommendation method based on hybrid retrieval RAG according to claim 7, characterized in that: Comparing all the similarities with the similarity threshold respectively to obtain the supplementary knowledge content includes: When the similarity is greater than or equal to the similarity threshold, the resource text corresponding to the similarity is used as the knowledge supplement content. If there are multiple resource texts that meet the conditions, a semantic relevance analysis is performed on the multiple resource texts, and they are sorted from high to low in terms of similarity and from high to low in terms of relevance, and the top n items are selected as the knowledge supplement content.

9. The resource recommendation method based on hybrid retrieval RAG according to claim 1, characterized in that: When combining all the supplementary knowledge content to generate resource recommendation answers, the following is included: The selected supplementary knowledge content is sorted from high to low according to its similarity to the query vector, and the sorted pieces of content are sequentially spliced ​​as context input. During the splicing process, separators are inserted between each piece of content, and the original user question and semantically expanded question are embedded in the spliced ​​context; The generated answer output result is post-processed, and the post-processing includes removing invalid sentences, unifying terminology styles, and de-duplicating classifications to obtain the resource recommendation answer.

10. A resource recommendation system based on hybrid retrieval RAG, used to apply the resource recommendation method based on hybrid retrieval RAG according to any one of claims 1 to 9, characterized in that: include: The collection unit is configured to collect resource text data and construct a vector library and a label library according to the resource text data; a processing unit configured to collect user questions, expand the user questions based on a language model, and obtain a plurality of semantically expanded questions; an identification unit configured to perform intent recognition on all the semantically extended questions, determine whether the user question is a resource recommendation question, and if so, determine the target classification type to which the resource belongs; a screening unit configured to convert the user question and the semantically expanded question into a structured query statement, screen the data according to the field definition in the tag library, and obtain a set of candidate knowledge fragments that meet the tag conditions; a matching unit configured to perform semantic vector encoding on each resource text in the candidate knowledge segment set to obtain a candidate vector; map the user question and the semantic expansion question into a query vector based on the semantic encoding, calculate the similarity between the query vector and each candidate vector, and select segments with a similarity greater than a similarity threshold as knowledge supplementary content; The output unit is configured to perform resource splicing on all the knowledge supplementary contents and generate resource recommendation answers.

Citation Information

Patent Citations

  • Multi-source and multi-mode fused knowledge reasoning method, system and device and medium

    CN119005340A

  • Data reordering retrieval method and system based on RAG

    CN120086307A

  • Knowledge graph-based traffic engineering large model intelligent question-answering system and method

    CN120407752A

  • Generating a unified metadata graph via a retrieval-augmented generation (RAG) framework systems and methods

    US12135740B1

  • Question answer recommendation method, storage medium, and electronic device

    WO2025025953A1

Cited By

  • AI (artificial intelligence)-based geological GIS (geographic information system) model quick retrieval method and system

    CN121029906A

  • RAG application-oriented context poisoning attack defense method

    CN121234911A

  • Large-scale language model knowledge extraction method and system without retrieval assistance

    CN121255967A

  • RAG intelligent retrieval question-answering system and method based on enhanced metadata

    CN121256007A

  • A RAG intelligent retrieval question and answer system and method based on enhanced metadata

    CN121256007B