Dynamic knowledge retrieval and intelligent generation system driven by RAG model
By designing a dynamic knowledge retrieval and intelligent generation system in the RAG model-driven knowledge retrieval system, using the multi-domain sub-knowledge base and Transformer model, combining the semantic similarity of the content requested by users, dynamically adjusting the generation rules, the problem of inconsistent with user needs in the existing system is solved, and more accurate and relevant content generation is achieved.
Patent Information
- Application Number
- CN202510469149.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing RAG model-driven knowledge retrieval system may experience insufficient recall or matching errors in handling language ambiguity, contextual context understanding, and the generated content may not be consistent with user needs, resulting in information confusion.
A dynamic knowledge retrieval and intelligent generation system driven by RAG model was designed. By establishing a multi-domain sub-knowledge base, using the pre-trained generation model of the Transformer architecture, combining the semantic similarity of the requested content of the user input twice, dynamically adjusting the generation rules, optimizing the answer strategy, and ensuring that the generated content meets user needs.
It improves user knowledge retrieval experience, ensures the accuracy and relevance of generated content, avoids invalid or duplicate feedback, reduces information redundancy, and improves the system's response quality and user satisfaction.
Smart Images

Figure CN119988574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of retrieval intelligent generation technology, and in particular to a RAG model-driven dynamic knowledge retrieval and intelligent generation system. Background Art
[0002] The RAG (Retrieval-Augmented Generation) model is a model architecture that combines information retrieval and text generation technologies. The RAG model uses a large external knowledge base to flexibly generate natural language answers that conform to the context. It aims to provide more accurate, real-time, and context-related answers or information summaries by dynamically calling external knowledge bases, thereby achieving better results in various NLP tasks.
[0003] Although RAG-based knowledge retrieval can use multiple knowledge bases to accurately locate relevant information in massive data, it may still have insufficient recall or matching errors in dealing with language ambiguity and contextual understanding. In addition, different users have different needs for retrieval results. In many scenarios, when the generative model integrates massive data information, information confusion may occur, that is, the generated content contains knowledge that is inconsistent with the retrieval information required by the user or is even completely wrong. How to ensure the correctness of the generated results in meeting the user's retrieval needs is a current research hotspot. Summary of the invention
[0004] In view of the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a RAG model-driven dynamic knowledge retrieval and intelligent generation system, so as to be able to intelligently generate retrieval content that meets user needs based on the user's actual retrieval needs, thereby improving the user's knowledge retrieval experience.
[0005] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: a RAG model-driven dynamic knowledge retrieval and intelligent generation system, comprising:
[0006] The knowledge base establishment module establishes the general knowledge base of the RAG model. The general knowledge base includes sub-knowledge bases in multiple fields, and each sub-knowledge base content has a practical and automated collection mechanism, which can be reflected in the sub-knowledge base in a timely manner after the data source is updated.
[0007] Knowledge retrieval module: after receiving the user's input request content, the system searches for document content related to the input question or context in the general knowledge base, and merges the document content in the most relevant sub-knowledge base with the user's original input request content to form a fused text content;
[0008] The content generation module uses a pre-trained generative model based on the Transformer architecture, combines its own pre-trained knowledge to perform reasoning and text generation, and refers to the context of the fused text content to form coherent, reliable, and uniformly styled answer content for feedback to users;
[0009] The dynamic adjustment module records the next request content entered by the user after receiving the feedback content. When the user enters a new request content, the new request content is compared with the old request content, and the appropriate sub-knowledge base is selected to generate content for feedback to the user based on the correlation between the two request contents.
[0010] In some embodiments, when a user inputs new request content, the system will compare the new request content with the old request content, and first perform text preprocessing on the old request content and the new request content, including word segmentation, stop word removal, word form restoration and other operations to eliminate noise, and use a pre-trained text encoder to convert each request into a vector representation of a fixed dimension, and then calculate the semantic similarity between the two requests.
[0011] In some embodiments, a request similarity threshold is set. After obtaining the semantic similarity between the contents of two requests, the semantic similarity is compared with the request similarity threshold. If the semantic similarity is less than or equal to the request similarity threshold, the system continues to execute the normal content generation method and feedback to the user; if the semantic similarity is greater than the request similarity threshold, the system executes the generation rule adjustment strategy.
[0012] In some embodiments, the generation rule adjustment strategy includes locating the document content used for generating feedback content last time and the sub-knowledge base to which it belongs, and using a pre-trained text encoder to vectorize the used document content, and marking all document contents in the sub-knowledge base to which it belongs as candidate documents, extracting metadata information of all candidate documents, and using the same pre-trained text encoder to vectorize each candidate document, by calculating the cosine similarity between the document content vector used for generating feedback content last time and each candidate document vector, a similarity score for each candidate document is obtained, and a candidate document similarity threshold is set. In the sub-knowledge base, all candidate documents with similarity scores higher than the candidate document similarity threshold are set to a blocked state.
[0013] In some embodiments, three levels of judgment intervals are divided between the upper limit of semantic similarity and the request similarity threshold. A higher level of judgment interval indicates that the request content input by the user this time is more similar to the last time. After the user inputs the request content for the second time, when the semantic similarity is greater than the request similarity threshold, the request content input this time is matched to the judgment interval of the corresponding level according to the semantic similarity of the two request contents, and different content generation methods are selected according to different levels of judgment intervals.
[0014] In some embodiments, when the system matches the request content input this time to the first-level judgment interval, the system does not execute the generated rule adjustment strategy, but analyzes the difference between the two request contents input by the user, and increases the importance of the difference when executing the knowledge retrieval module;
[0015] When the system matches the request content input this time to the secondary judgment interval, the system executes the generated rule adjustment strategy normally;
[0016] When the system matches the request content input this time to the third-level judgment interval, the system does not execute the generation rule adjustment strategy, but will execute the shielding sub-knowledge base strategy.
[0017] In some embodiments, the way to increase the importance of differential content is: use a set comparison algorithm to compare the keyword sets after two request processing, identify differential keywords that exist in the current request content but not in the previous request content, and weight the differential keywords. When the system executes the knowledge retrieval module, it searches for the most relevant content in the general knowledge base based on each keyword in the request content, and selects the corresponding content for use based on the relevance score order. After weighting the differential keywords, the document content containing these weighted differential keywords can obtain a higher relevance score.
[0018] In some embodiments, the strategy of blocking sub-knowledge bases includes: locating the sub-knowledge base used to generate feedback content last time, setting the sub-knowledge base to a blocked state, and excluding the document content in the sub-knowledge base set to a blocked state when executing the knowledge retrieval module to match the document content in the general knowledge base after the request content input by the user this time, and automatically clearing the blocked state after the feedback content is generated this time.
[0019] In some embodiments, after shielding the corresponding sub-knowledge base and retrieving the document content with the highest relevance score, the document content retrieved this time is vectorized using a pre-trained text encoder, and the document content used in the last retrieval is vectorized using a general pre-trained text encoder to calculate the cosine similarity between the document content vectors of the two retrievals, and mark it as content similarity, set a minimum relevance threshold, and compare the content similarity with the minimum relevance threshold: when the content similarity is greater than or equal to the minimum relevance threshold, no additional response is made; when the content similarity is less than the minimum relevance threshold, the system combines the document content retrieved this time with the document content retrieved last time when executing the knowledge retrieval module, and merges the two contents with the user's original input request content to form fused text content, and generates feedback content when executing the subsequent content generation module.
[0020] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned RAG model-driven dynamic knowledge retrieval and intelligent generation system.
[0021] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0022] Firstly, by obtaining the semantic similarity of the request contents inputted twice by the user, the system of the present invention can more accurately judge the changes in user needs. If the two request contents have a high similarity in subject or context, the system will realize that the user's needs are not met, and then optimize the answer strategy to avoid providing invalid or repetitive feedback.
[0023] Secondly, the present invention can eliminate useless content that is duplicated with user needs through the design of the generation rule adjustment strategy. The system can generate new answers more quickly and accurately, avoid the accumulation of meaningless answers, and reduce unnecessary information redundancy.
[0024] Third, by dividing the user request content into different judgment intervals, the system can flexibly adapt to the diversity of user needs. Whether the user slightly adjusts the request or repeatedly inputs a similar request, the system can accurately judge its intention and adjust the direction of content generation.
[0025] Fourthly, the present invention can avoid the problem of new document content not matching the core topic of the user's original request, resulting in a complete shift in the semantic center, by setting a minimum relevance threshold, so as to avoid feedback errors caused by excessive deviation of the information source due to shielding the sub-knowledge base, thereby improving the accuracy of answers and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the principle of the dynamic knowledge retrieval and intelligent generation system driven by the RAG model of the present invention;
[0027] Figure 2 This is a module diagram of the RAG model-driven dynamic knowledge retrieval and intelligent generation system of the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] It is to be understood that the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the element may be multiple, and the term "one" should not be understood as a limitation on the quantity.
[0030] The present invention provides a RAG model driven dynamic knowledge retrieval and intelligent generation system, such as Figure 1 and Figure 2 As shown, including:
[0031] The knowledge base establishment module establishes a general knowledge base of the RAG model. The general knowledge base includes sub-knowledge bases in multiple fields. The field division of each sub-knowledge base can determine the fields that need to be covered according to application requirements (such as technology, medical care, law, finance, entertainment, etc.). Each field knowledge base should have a clear theme and content scope. And for different fields, select highly authoritative and credible data sources. For example, in the field of science and technology, you can refer to professional paper databases, technical blogs and forums; in the field of medicine, you can cite medical journals, clinical guidelines and information published by authoritative institutions; in the field of law, you can collect government regulations, legal cases and industry reviews. And each sub-knowledge base content has a practical and automated collection mechanism, including the establishment of automatic crawling, API docking and other channels to obtain the latest information from public networks, databases, journals and other sources, so that it can be reflected in the sub-knowledge base in a timely manner after the data source is updated.
[0032] Knowledge retrieval module: after receiving the user's input request content (such as questions and dialogue context), the system searches for document content related to the input question or context in the general knowledge base, obtains the relevance of each sub-knowledge base to the request content, and merges the document content in the sub-knowledge base with the user's original input request content to form a fused text content;
[0033] The content generation module uses a pre-trained generative model based on the Transformer architecture. While receiving external knowledge prompts, the generative model combines its own pre-trained knowledge for reasoning and text generation. During the generation process, the model will refer to the context of the fused text content to form a coherent, reliable and uniformly styled answer content for feedback to the user.
[0034] Dynamic adjustment module: after receiving feedback from the user, the next request content entered by the user is recorded. When the user enters a new request content, the new request content is compared with the old request content, the relevance of the two request contents is analyzed, and the appropriate sub-knowledge base is selected according to the relevance to generate content for feedback to the user;
[0035] When the user enters new request content, the system will compare the new request content with the old request content. First, the old request content and the new request content will be preprocessed separately, including word segmentation, stop word removal, word form restoration and other operations to eliminate noise. The pre-trained text encoder is used to convert each request into a vector representation of fixed dimension, and then the semantic similarity between the two requests is calculated. The specific calculation method is to calculate the cosine similarity of the two request vectors. The cosine similarity can reflect the angular relationship between the two requests in the semantic space and is suitable for judging the subject similarity of the request content.
[0036] Set the request similarity threshold. After obtaining the semantic similarity between the two request contents, compare the semantic similarity with the request similarity threshold. If the semantic similarity is less than or equal to the request similarity threshold, it can be considered that the two request contents entered by the user are not highly correlated in terms of subject or context, that is, the user has not repeatedly entered similar request contents. The system then continues to execute the normal content generation method to feedback to the user. If the semantic similarity is greater than the request similarity threshold, it can be considered that the two request contents entered by the user are highly correlated in terms of subject or context, that is, the user has repeatedly entered similar request contents. This means that the knowledge content generated and fed back to the user by the system last time was not recognized by the user and did not meet the user's needs. The system then executes the generation rule adjustment strategy. For example, if the request similarity threshold is set to 0.7, and the semantic similarity between the user's request content and the previous request content is 0.5 after the user enters the request content next time, the system then continues to execute the normal content generation method to feedback to the user.
[0037] The generation rule adjustment strategy includes locating the document content used in the last generation of feedback content and the sub-knowledge base to which it belongs, and using a pre-trained text encoder to vectorize the document content used, and marking all the document content in the sub-knowledge base to which it belongs as candidate documents, extracting metadata information of all candidate documents, and using the same pre-trained text encoder to vectorize each candidate document. Then, by calculating the cosine similarity between the document content vector used in the last generation of feedback content and each candidate document vector, the similarity score of each candidate document is obtained, and the candidate document similarity threshold (for example, 0.8) is set. In the sub-knowledge base, all candidate documents with similarity scores higher than the candidate document similarity threshold are set to a shielded state. When the knowledge retrieval module is executed to match the document content in the general knowledge base after the user inputs the request content this time, the document content set to the shielded state will be excluded, and the shielded state will be automatically cleared after the feedback content is generated this time. In this way, the system can effectively filter out document content with a high similarity to the last feedback content, prevent repeated retrieval, and thus avoid reusing content that failed to meet user needs before when generating new answers.
[0038] The upper limit of the semantic similarity of the two request contents is 1, that is, the request contents entered by the two users are 100% identical. The upper limit of the semantic similarity and the request similarity threshold are divided into three levels of judgment intervals. The higher the level of the judgment interval, the more similar the request content entered by the user this time is to the last time. After the user enters the request content for the second time, when the semantic similarity is greater than the request similarity threshold, the request content entered this time is matched to the judgment interval of the corresponding level according to the semantic similarity of the two request contents. For example, when the request similarity threshold is set to 0.7, the range of the first-level judgment interval is [0.7-0.8), the range of the second-level judgment interval is [0.8-0.9), and the range of the third-level judgment interval is [0.9-1]. When the user enters the request content for the second time, the semantic similarity of the two request contents is 0.95, and the system matches the request content entered this time to the third-level judgment interval. Different content generation methods are selected according to different levels of judgment intervals. Specifically:
[0039] When the system matches the request content input this time to the first-level judgment interval, it indicates that the user is not satisfied with the feedback content generated before, and the request content input last time has been modified, that is, the user adds or corrects keywords to try to make the system generate content that is more in line with the user's own needs. The system does not execute the generation rule adjustment strategy, but analyzes the difference in the content of the user's two requests input before and after, and increases the importance of the difference when executing the knowledge retrieval module; first, the keyword sets after the two request processing are compared using the set comparison algorithm to identify the difference keywords that exist in the current request content but not in the previous request content. Then, the difference keywords are weighted. When the system executes the knowledge retrieval module, it will find the most relevant content in the general knowledge base according to each keyword in the request content, and select the corresponding content for use according to the relevance score order. After weighting the difference keywords, the document content containing these weighted difference keywords can obtain a higher relevance score and thus be used first.
[0040] When the system matches the request content entered this time to the secondary judgment interval, it indicates that the user is not satisfied with the previously generated feedback content, and the modification of the request content entered last time is also limited, and the system executes the generation rule adjustment strategy normally;
[0041] When the system matches the request content entered this time to the third-level judgment interval, it indicates that the user is not satisfied with the feedback content generated previously, and has not made any changes to the request content entered last time. If the request content entered twice is basically the same, the system will not execute the generation rule adjustment strategy, but will execute the sub-knowledge base shielding strategy;
[0042] The reason for the above design is that when the user inputs the request content again, when the semantic similarity is greater than the request similarity threshold, the system determines that the user has input similar and repeated content, indicating that the user is not satisfied with the content of the previous system-generated feedback. At this time, the higher the semantic similarity between the two request contents, the more it indicates that the request content input by the user has not been changed, and the easier it is for the system to retrieve the same content when executing the knowledge retrieval module; and the lower the semantic similarity between the two request contents, the more it indicates that although the main content of the request content input by the user is similar, the user has supplemented, improved and corrected the original content, and the system should focus more on the part of the request content added and modified by the user.
[0043] When the request content matches the third-level judgment interval, it means that the user has input almost the same request content again after not receiving satisfactory feedback from the last request. This usually indicates that the user believes that the request content he input is already quite clear and tries to let the system output search content that meets his needs. At this time, if the content generated by the system deviates from the user's expectations again and again, it means that there is a deviation between the sub-knowledge base currently used and the user's needs, resulting in the system output direction being inconsistent with the user's inner recognition. At this time, the field and direction of the content generated by the system should be changed by executing the sub-knowledge base shielding strategy. The sub-knowledge base shielding strategy specifically locates the sub-knowledge base used to generate feedback content last time, sets the sub-knowledge base to a shielded state, and executes the knowledge retrieval module to match the document content in the general knowledge base after the user inputs the request content this time. The document content in the sub-knowledge base set to a shielded state will be excluded, and the shielding state will be automatically cleared after the feedback content is generated this time.
[0044] In addition, after shielding the corresponding sub-knowledge base and retrieving the document content with the highest relevance score, the document content retrieved this time is vectorized using the pre-trained text encoder, and the document content used in the last retrieval is vectorized using the general pre-trained text encoder, so as to be able to calculate the cosine similarity between the document content vectors of the two retrievals, and mark it as content similarity, set a minimum relevance threshold (for example, 0.4), compare the content similarity with the minimum relevance threshold, and make corresponding responses according to the comparison results: when the content similarity is greater than or equal to the minimum relevance threshold, it means that the document retrieved this time is After the relevant sub-knowledge base is blocked, the content is reasonably different from the document content used last time. This indicates that after the relevant sub-knowledge base is blocked, the content generated by the system is dynamically adjusted to provide to the user without additional response. When the content similarity is less than the minimum relevance threshold, it means that the document content retrieved this time is very different from the document content used last time. This shows that after the relevant sub-knowledge base is blocked, the content obtained may be wrong. This is because after the sub-knowledge base used last time is blocked, the system searches for relevant content in other sub-knowledge bases, but if the new document content does not match the core topic of the user's original request, it will cause the semantic center of gravity to shift completely. Ideally, after the relevant sub-knowledge base is blocked, the system should be able to retrieve information that still maintains a certain degree of continuity in semantics. Even if there are adjustments, the document content should have a reasonable continuation in the core topic. If the similarity of the two retrieved contents is too low, it means that the newly retrieved content is out of touch with the main semantics of the user's request. Therefore, when executing the knowledge retrieval module, the system combines the document content retrieved this time with the document content retrieved last time, and fuses the two contents with the user's original input request content to form a fused text content, and generates feedback content when executing the subsequent content generation module. This approach can minimize feedback errors caused by excessive deviation from a single information source, and also provides the system with more information as a reference, improving the accuracy of answers and user satisfaction.
[0045] In general, the present invention aims to design a dynamic knowledge retrieval and intelligent generation system driven by a RAG model. In view of the problem that the generated content may not meet the user's retrieval requirements in the retrieval process driven by the RAG model, the present invention establishes a sub-knowledge base containing multiple fields to ensure that a wide range of knowledge fields are covered and the application requirements of different fields are met. The latest information from sources such as public networks, databases, and journals is obtained in real time through automatic crawling, API docking, etc., ensuring the timeliness and accuracy of the sub-knowledge base content, and being able to reflect new knowledge and updated data in a timely manner. By obtaining the semantic similarity of the request content input by the user twice, the system can more accurately judge the changes in user needs. If the two request contents have a high similarity in subject or context, the system will realize that the user's needs are not met, and then optimize the answer strategy to avoid providing invalid or repeated feedback. By designing the generation rule adjustment strategy, useless content that repeats the user's needs can be eliminated, and the system can generate new answers more quickly and accurately, avoiding the accumulation of meaningless answers, reducing unnecessary information redundancy, and making the feedback obtained by the user each time more in line with actual needs, thereby improving the response quality of the system and user satisfaction. And by dividing the user's request content into different judgment intervals, the system can flexibly adapt to the diversity of user needs. Whether the user slightly adjusts the request or repeatedly enters a similar request, the system can accurately judge its intention and adjust the content generation strategy to achieve more personalized services. In addition, under the setting of the minimum relevance threshold, feedback errors caused by excessive deviation from a single information source can be avoided as much as possible, while also providing the system with more information as a reference, improving the accuracy of the answer and user satisfaction.
[0046] The embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. The embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable medium mentioned above in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of an electrical, magnetic, optical, electromagnetic, infrared segment, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination of the above.
[0047] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0048] Those skilled in the art should understand that the above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. RAG model-driven dynamic knowledge retrieval and intelligent generation system, characterized by: include: The knowledge base establishment module establishes the general knowledge base of the RAG model. The general knowledge base includes sub-knowledge bases in multiple fields, and each sub-knowledge base has a practical and automated collection mechanism, which can be reflected in the sub-knowledge base in a timely manner after the data source is updated; Knowledge retrieval module: after receiving the user's input request content, the system searches for document content related to the input question or context in the general knowledge base, and merges the document content in the most relevant sub-knowledge base with the user's original input request content to form a fused text content; The content generation module uses a pre-trained generative model based on the Transformer architecture, combines its own pre-trained knowledge to perform reasoning and text generation, and refers to the context of the fused text content to form coherent, reliable, and uniformly styled answer content for feedback to users; The dynamic adjustment module records the next request content entered by the user after receiving the feedback content. When the user enters a new request content, the new request content is compared with the old request content, and the appropriate sub-knowledge base is selected to generate content based on the correlation between the two request contents to feedback to the user.
2. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 1, characterized in that: When the user enters new request content, the system will compare the new request content with the old request content, first perform text preprocessing on the old request content and the new request content respectively, and use the pre-trained text encoder to convert each request into a vector representation of fixed dimension, and then calculate the semantic similarity between the two requests.
3. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 2, characterized in that: Set the request similarity threshold. After obtaining the semantic similarity between the two request contents, compare the semantic similarity with the request similarity threshold. If the semantic similarity is less than or equal to the request similarity threshold, the system continues to execute the normal content generation method and feedback to the user; if the semantic similarity is greater than the request similarity threshold, the system executes the generation rule adjustment strategy.
4. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 3 is characterized in that: The generation rule adjustment strategy includes locating the document content used for generating feedback content last time and the sub-knowledge base to which it belongs, and using a pre-trained text encoder to vectorize the used document content, and marking all document contents in the sub-knowledge base to which it belongs as candidate documents, extracting metadata information of all candidate documents, and using the same pre-trained text encoder to vectorize each candidate document, by calculating the cosine similarity between the document content vector used for generating feedback content last time and each candidate document vector, a similarity score for each candidate document is obtained, and a candidate document similarity threshold is set. In the sub-knowledge base, all candidate documents with similarity scores higher than the candidate document similarity threshold are set to a shielded state.
5. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 4, characterized in that: The upper limit of semantic similarity and the request similarity threshold are divided into three levels of judgment intervals. The higher the level of the judgment interval, the more similar the request content entered by the user this time is to the last time. After the user enters the request content for the second time, when the semantic similarity is greater than the request similarity threshold, the request content entered this time is matched with the judgment interval of the corresponding level according to the semantic similarity of the two request contents, and different content generation methods are selected according to different levels of judgment intervals.
6. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 5, characterized in that: When the system matches the request content input this time to the first-level judgment interval, the system does not execute the generation rule adjustment strategy, but analyzes the difference between the two request contents input by the user and increases the importance of the difference when executing the knowledge retrieval module; When the system matches the request content input this time to the secondary judgment interval, the system executes the generated rule adjustment strategy normally; When the system matches the request content input this time to the third-level judgment interval, the system does not execute the generation rule adjustment strategy, but will execute the shielding sub-knowledge base strategy.
7. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 6, characterized in that: The way to increase the importance of differential content is: use a set comparison algorithm to compare the keyword sets after the two request processing, identify the differential keywords that exist in the current request content but not in the previous request content, and weight the differential keywords. When the system executes the knowledge retrieval module, it searches for the most relevant content in the general knowledge base based on each keyword in the request content, and selects the corresponding content for use based on the order of relevance scores. After weighting the differential keywords, the document content containing these weighted differential keywords can obtain a higher relevance score.
8. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 6, characterized in that: The sub-knowledge base shielding strategy includes: locating the sub-knowledge base used to generate feedback content last time, setting the sub-knowledge base to a shielded state, and excluding the document content in the sub-knowledge base set to a shielded state when executing the knowledge retrieval module to match the document content in the general knowledge base after the user inputs the request content this time, and automatically clearing the shielding state after the feedback content is generated this time.
9. The RAG model-driven dynamic knowledge retrieval and intelligent generation system according to claim 8, characterized in that: After shielding the corresponding sub-knowledge base and retrieving the document content with the highest relevance score, the document content retrieved this time is vectorized using a pre-trained text encoder, and the document content used in the last retrieval is vectorized using a general pre-trained text encoder to calculate the cosine similarity between the document content vectors of the two retrievals, and mark it as content similarity. A minimum relevance threshold is set to compare the content similarity with the minimum relevance threshold: when the content similarity is greater than or equal to the minimum relevance threshold, no additional response is made; when the content similarity is less than the minimum relevance threshold, the system combines the document content retrieved this time with the document content retrieved last time when executing the knowledge retrieval module, and merges the two contents with the user's original input request content to form fused text content, and generates feedback content when executing the subsequent content generation module.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the RAG model-driven dynamic knowledge retrieval and intelligent generation system described in any one of claims 1 to 9.
Citation Information
Patent Citations
Fault-tolerant text query method and equipment
CN101984422A
Scene interaction optimization method and system for virtual digital human
CN119248142A
Method and system for providing semantics based technical support
US20170132210A1
Cited By
Intelligent document information real-time retrieval method and system based on RAG technology
CN121210624A
Software process document generation method and electronic equipment
CN121455538A
Method for generating software process document and electronic device
CN121455538B