Search method, system, and electronic device

CN122173690BActive Publication Date: 2026-09-25HUA DATA TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610645361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-09-25
Estimated Expiration
2046-05-12

AI Technical Summary

Technical Problem

[0005]本公开要解决的技术问题是为了克服现有技术中检索策略僵化,无法进行针对性检索的缺陷,提供一种检索方法、系统以及电子设备

Benefits of technology

[0079]在符合本领域常识的基础上,上述各优选条件,可任意组合,即得本公开各较佳实例。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173690B_ABST
    Figure CN122173690B_ABST
Patent Text Reader

Abstract

The present disclosure provides a retrieval method, system and electronic device, the retrieval method comprising: obtaining a user query question; reading a target retrieval strategy matching the semantics of the user query question from a local experience library; wherein the local experience library comprises a local retrieval strategy extracted from a local historical reasoning track; decomposing the user query question into a plurality of sub-query questions according to the target retrieval strategy; generating corresponding keywords according to the plurality of sub-query questions; calling a hybrid retrieval service for retrieval by using the semantic vector corresponding to the sub-query question and the keywords to generate a corresponding target reasoning track; thereby achieving targeted retrieval, avoiding rigid retrieval strategies, and improving the accuracy of retrieving complex problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a retrieval method, system, and electronic device. Background Technology

[0002] With the breakthroughs in Large Language Models (LLM), Retrieval Augmentation (RAG) has become the standard paradigm for addressing the "knowledge illusion" of models and supplementing private domain knowledge and real-time information. Traditional RAG systems typically follow a linear process of "retrieval-ranking-generation": first, relevant chunks are retrieved from an external knowledge base using vector similarity or keyword matching, and then these chunks are used as contextual input to the LLM to generate the answer.

[0003] However, when faced with complex business scenarios requiring multi-step reasoning and in-depth analysis, such as financial auditing, legal consulting, and scientific research reviews, existing RAG technologies suffer from severe "signal-to-noise ratio bottlenecks" and "strategy rigidity," making it difficult for technical solutions to form an efficient closed loop. Specifically, existing RAG technologies suffer from static and rigid retrieval strategies and lack the ability to dynamically program complex problems.

[0004] Existing RAG systems typically employ predefined static retrieval logic, such as fixed Top-K (top K most relevant entries in the search results list) truncation and single query rewriting templates. When faced with complex user queries, such as "compare the R&D investment growth rates of Company A and Company B over the past three years," the system often performs a fuzzy search directly, resulting in a large amount of irrelevant and noisy data, lacking the "planning-execution-reflection" capabilities of human experts. While agent-driven retrieval augmentation generation (Agentic RAG) attempts to introduce planning capabilities, its planning strategies often rely on extremely expensive cue word engineering or large-scale supervised fine-tuning (SFT), lacking self-adaptive and self-evolving capabilities. Summary of the Invention

[0005] The technical problem to be solved by this disclosure is to overcome the shortcomings of rigid retrieval strategies in the prior art, which make it impossible to perform targeted retrieval, and to provide a retrieval method, system and electronic device.

[0006] This disclosure solves the above-mentioned technical problems through the following technical solution:

[0007] Firstly, a retrieval method is provided, the retrieval method comprising the following steps:

[0008] Get the user's query question;

[0009] Read target retrieval strategies that match the semantics of the user's query from a local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories;

[0010] The user query question is broken down into multiple sub-query questions according to the target retrieval strategy;

[0011] Generate corresponding keywords based on the multiple sub-queries;

[0012] Using the semantic vector corresponding to the subquery question and the keywords, a hybrid retrieval service is invoked to perform a retrieval, thereby generating the corresponding target reasoning trajectory.

[0013] Optionally, the target retrieval strategy includes the execution order and dependencies between subqueries; the user query is decomposed into multiple subqueries according to the target retrieval strategy, including:

[0014] Perform semantic analysis on the user query to determine the user's query intent;

[0015] The query intent is divided into initial query sub-problems with the aforementioned dependencies;

[0016] The initial query subproblems are sorted according to the execution order to obtain multiple ordered subquery subproblems.

[0017] Optionally, the retrieval method further includes:

[0018] If the correlation between the knowledge slice contained in the target inference trajectory and the user query question is less than the correlation threshold, the step of decomposing the user query question into multiple sub-query questions according to the target retrieval strategy is returned.

[0019] Optionally, the retrieval method further includes:

[0020] Obtain the query questions corresponding to the user's historical reasoning trajectory for the target type in the historical logs generated based on retrieval enhancement, and obtain a question set; wherein, the user's historical reasoning trajectory also includes the query conclusion;

[0021] For each query question in the question set, a preset number of searches are performed to obtain a group of inference trajectories corresponding to each query question; wherein, each group of inference trajectories contains the inference trajectories of the preset number of searches.

[0022] In response to the presence of a number of target query conclusions in the inference trajectory group that exceed a number threshold, the inference trajectory corresponding to the target query conclusion is determined to be successful, and the inference trajectories corresponding to other query conclusions are determined to be unsuccessful.

[0023] The first retrieval strategy is added to the local experience base; the first retrieval strategy is a retrieval strategy generated based on the difference between successful and unsuccessful inference trajectories.

[0024] And / or,

[0025] Delete the second retrieval strategy from the local experience base; the second retrieval strategy is the retrieval strategy used by multiple failed inference trajectories.

[0026] Optionally, the retrieval method further includes:

[0027] Delete the first knowledge slice from the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is a knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group;

[0028] And / or, increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice that dominates multiple successful reasoning trajectories in the reasoning trajectory group.

[0029] Optionally, the step of deleting the first knowledge slice in the knowledge base corresponding to the hybrid retrieval service specifically includes:

[0030] Obtain the globally unique identifier of the first knowledge slice in the knowledge base;

[0031] Modify the metadata status bit corresponding to the globally unique identifier to an interceptable state. The interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service.

[0032] If the duration during which the first knowledge slice is in the interceptable state exceeds a preset duration, the first knowledge slice is deleted from the knowledge base.

[0033] Optionally, the step of increasing the weight of the second knowledge slice in the knowledge base specifically includes:

[0034] Obtain the current dynamic weight score from the metadata corresponding to the second knowledge slice;

[0035] The dynamic weight increment of the second knowledge slice is calculated based on the number of successful inference trajectories of the target containing the second knowledge slice and / or the confidence score of the successful inference trajectories of the target.

[0036] The dynamic weight score of the second knowledge slice is updated based on the current dynamic weight score and the dynamic weight increment.

[0037] Optionally, the target reasoning trajectory includes at least two target knowledge slices; after the step of calling the hybrid retrieval service to perform retrieval using the semantic vector corresponding to the sub-query question and the keywords, the method further includes:

[0038] In response to a conflict existing in at least two target knowledge slices, a next round of retrieval is performed; wherein the conflict determination method includes at least one of the following:

[0039] Calculate the conflict confidence of any two target knowledge slices, and in response to the conflict confidence exceeding the corresponding conflict confidence threshold, determine that there is a conflict in the at least two target knowledge slices;

[0040] In response to the detection of contradictory numerical values ​​contained in keywords across multiple knowledge slices, it is determined that a conflict exists between at least two target knowledge slices.

[0041] Secondly, a retrieval system is provided, the retrieval system comprising:

[0042] The question retrieval module is used to retrieve user-queried questions.

[0043] The strategy reading module is used to read target retrieval strategies that match the semantics of the user's query question from the local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories;

[0044] The problem decomposition module is used to decompose the user query problem into multiple sub-query problems according to the target retrieval strategy;

[0045] The keyword generation module is used to generate corresponding keywords based on the multiple sub-queries.

[0046] The trajectory generation module is used to call the hybrid retrieval service to perform retrieval using the semantic vector corresponding to the subquery question and the keywords, so as to generate the corresponding target reasoning trajectory.

[0047] Optionally, the target retrieval strategy includes the execution order and dependencies between subqueries; the problem decomposition module includes:

[0048] The intent determination unit is used to perform semantic analysis on the user query question to determine the query intent of the user query question.

[0049] An initial partitioning unit is used to divide the query intent into initial query sub-problems with the aforementioned dependencies;

[0050] The problem sorting unit is used to sort the initial query sub-problems according to the execution order to obtain multiple ordered sub-query problems.

[0051] Optionally, the retrieval system further includes:

[0052] The relevance detection module is used to respond to the fact that the relevance between the knowledge slice contained in the target reasoning trajectory and the user query question is less than the relevance threshold, and return the step of decomposing the user query question into multiple sub-query questions according to the target retrieval strategy.

[0053] Optionally, the retrieval system further includes:

[0054] The question set acquisition module is used to acquire the query questions corresponding to the user's historical reasoning trajectory of the target type in the historical log generated based on retrieval enhancement, and obtain the question set; wherein, the user's historical reasoning trajectory also includes the query conclusion;

[0055] The reasoning trajectory group acquisition module is used to retrieve each query question in the question set a preset number of times to obtain a reasoning trajectory group corresponding to each query question; wherein, each reasoning trajectory group contains the reasoning trajectory of the preset number of times;

[0056] The success / failure determination module is used to determine the reasoning trajectory corresponding to the target query conclusion as successful and the reasoning trajectory corresponding to other query conclusions as unsuccessful when there is a number of target query conclusions in the reasoning trajectory group that is greater than the number threshold.

[0057] The retrieval strategy processing module is used to add a first retrieval strategy to the local experience base; the first retrieval strategy is a retrieval strategy generated based on the difference between successful reasoning trajectories and failed reasoning trajectories;

[0058] And / or,

[0059] The retrieval strategy processing module is used to delete the second retrieval strategy from the local experience base; the second retrieval strategy is the retrieval strategy adopted by multiple failed inference trajectories.

[0060] Optionally, the retrieval system further includes:

[0061] The knowledge slice processing module is used to delete the first knowledge slice in the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is a knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group;

[0062] And / or,

[0063] The knowledge slice processing module is used to increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice that dominates multiple successful reasoning trajectories in the reasoning trajectory group.

[0064] Optionally, the knowledge slicing processing module specifically includes:

[0065] An identifier acquisition unit is used to acquire a globally unique identifier for the first knowledge slice in the knowledge base.

[0066] The status modification unit is used to modify the metadata status bit corresponding to the globally unique identifier to an interceptable state, wherein the interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service;

[0067] The deletion unit is configured to delete the first knowledge slice from the knowledge base in response to the first knowledge slice being in the interceptable state for a period of time exceeding a preset duration.

[0068] Optionally, the knowledge slicing processing module specifically includes:

[0069] The weight acquisition unit is used to acquire the current dynamic weight score in the metadata corresponding to the second knowledge slice;

[0070] An incremental calculation unit is used to calculate the dynamic weight increment of the second knowledge slice based on the number of successful inference trajectories of the target containing the second knowledge slice and / or the confidence score of the successful inference trajectories of the target;

[0071] The weight update unit is used to update the dynamic weight score of the second knowledge slice based on the current dynamic weight score and the dynamic weight increment.

[0072] Optionally, the target reasoning trajectory includes at least two slices of target knowledge; the retrieval system further includes:

[0073] A multi-round retrieval determination module is used to perform a next round of retrieval in response to a conflict existing in the at least two target knowledge slices; wherein the conflict determination method includes at least one of the following:

[0074] Calculate the conflict confidence of any two target knowledge slices, and in response to the conflict confidence exceeding the corresponding conflict confidence threshold, determine that there is a conflict in the at least two target knowledge slices;

[0075] In response to the detection of contradictory numerical values ​​contained in keywords across multiple knowledge slices, it is determined that a conflict exists between at least two target knowledge slices.

[0076] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and for running on the processor, wherein the processor executes the computer program to implement the retrieval method described in the first aspect.

[0077] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the retrieval method described in the first aspect.

[0078] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the retrieval method described in the first aspect.

[0079] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of this disclosure.

[0080] The positive and progressive effects of this disclosure are as follows: To solve the problem of rigid retrieval strategies that cannot perform targeted retrieval, this disclosure reads target retrieval strategies that match the user's query from a local experience base, instead of performing static retrieval based on the user's query. This allows the user's query to be decomposed into multiple sub-queries based on the historical experience of the target retrieval strategies in the local experience base, thus achieving targeted retrieval and solving the problem of decomposing complex problems. Furthermore, based on the semantic vectors corresponding to the sub-queries and the keywords generated for the sub-queries, a hybrid retrieval service is invoked for retrieval, enabling simultaneous vector retrieval and keyword retrieval, significantly improving the precision and efficiency of the retrieval. Attached Figure Description

[0081] Figure 1 A flowchart of a retrieval method provided in Embodiment 1 of this disclosure;

[0082] Figure 2 A detailed flowchart of step S103 provided in Embodiment 1 of this disclosure;

[0083] Figure 3 This is a partial flowchart of a retrieval method provided in Embodiment 1 of this disclosure;

[0084] Figure 4 This is a partial flowchart of a retrieval method provided in Embodiment 1 of this disclosure;

[0085] Figure 5 A detailed flowchart of step S111 provided in Embodiment 1 of this disclosure;

[0086] Figure 6 A detailed flowchart of step S112 provided in Embodiment 1 of this disclosure;

[0087] Figure 7 A partial flowchart of a specific retrieval method provided in Embodiment 1 of this disclosure;

[0088] Figure 8 A detailed flowchart of a retrieval method provided in Embodiment 1 of this disclosure;

[0089] Figure 9 This is an overall architecture diagram of a retrieval method provided in Embodiment 1 of this disclosure;

[0090] Figure 10 This is a schematic diagram of a retrieval system provided in Embodiment 2 of this disclosure;

[0091] Figure 11 This is a unit schematic diagram of a problem decomposition module provided in Embodiment 2 of this disclosure;

[0092] Figure 12 This is a schematic diagram of a retrieval system provided in Embodiment 2 of this disclosure;

[0093] Figure 13 This is a schematic diagram of a knowledge slicing processing module provided in Embodiment 2 of this disclosure;

[0094] Figure 14 This is a schematic diagram of a knowledge slicing processing module provided in Embodiment 2 of this disclosure;

[0095] Figure 15 This is a schematic diagram of a trajectory generation module provided in Embodiment 2 of this disclosure;

[0096] Figure 16 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this disclosure. Detailed Implementation

[0097] The present disclosure is further illustrated below by way of embodiments, but the present disclosure is not limited to the scope of the embodiments described herein.

[0098] The prefixes such as "first" and "second" used in this disclosure are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this disclosure does not constitute a limitation on the described objects. The description of the described objects is given in the claims or the context of the embodiments, and should not be construed as an unnecessary limitation. Furthermore, in the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0099] In this embodiment of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good morals.

[0100] Example 1

[0101] With breakthroughs in large language models, retrieval-enhanced generation (RAG) has become a standard paradigm for addressing model "knowledge illusion," supplementing private domain knowledge, and providing real-time information. Traditional RAG typically follows a linear process of "retrieval-ranking-generation": first, relevant chunks are retrieved from an external knowledge base using vector similarity or keyword matching; then, these chunks are used as contextual input to the large language model to generate answers. However, when facing complex business scenarios requiring multi-step reasoning and in-depth analysis, such as financial auditing, legal consulting, and scientific research reviews, existing RAG technologies face serious "signal-to-noise ratio bottlenecks" and "strategy rigidity" problems, making it difficult for the technical solution to form an efficient closed loop.

[0102] 1. The retrieval strategy is static and rigid, lacking the ability to dynamically program for complex problems.

[0103] Existing search augmentation generation techniques typically employ predefined static search logic (such as fixed Top-K truncation and single query rewriting templates). When faced with complex user queries (such as "compare the R&D investment growth rates of Company A and Company B over the past three years"), the system often performs a fuzzy search directly, resulting in a large amount of irrelevant and noisy data. The system lacks the "plan-execute-reflect" capabilities of human experts, unable to determine when to break down a large problem into sub-problems, when to change keywords for a secondary search, or when to consult specific authoritative data sources. While agent-driven search augmentation generation (Agentic RAG) attempts to introduce planning capabilities, its planning strategies often rely on extremely expensive cue word engineering or large-scale supervised fine-tuning (SFT), lacking self-adaptation and self-evolutionary capabilities.

[0104] 2. Model optimization is costly and difficult to adapt to rapidly changing business needs.

[0105] To enhance an agent's planning and retrieval capabilities, mainstream research directions (such as Agentic Reinforcement Learning) typically rely on reinforcement learning based on human feedback (RLHF) or large-scale PPO (Proximal Policy Optimization) training. These methods require weight updates on massive models with hundreds of billions or even trillions of parameters, necessitating not only the construction of expensive, high-quality labeled datasets but also massive high-performance GPU computing resources (such as H100 / A100 clusters). For most enterprises, this retraining model is neither economically nor timely, resulting in RAG systems being difficult to improve once deployed, unable to learn from historical interaction data, and severely lacking self-evolution capabilities.

[0106] 3. The knowledge base is under-maintained and lacks an automated noise reduction and purification mechanism.

[0107] The knowledge base is the foundation of RAG. In actual operation, as documents accumulate, the knowledge base inevitably generates a large number of redundant, outdated, and even conflicting knowledge chunks. Traditional knowledge base maintenance relies entirely on manual review, which is extremely inefficient. The system lacks an automated attribution mechanism to identify which knowledge chunks played a key role in historical question answering and which led to the model's illusions or incorrect answers. This causes the signal-to-noise ratio of retrieval to continuously decline over time, severely impacting the quality of the final generated answer.

[0108] In summary, there is an urgent need in this field for an innovative technical solution to construct a deep-search agent with deep thinking capabilities, combined with a lightweight, model parameter-updating-free empirical evolution mechanism. This mechanism would enable the system to learn from past successes and failures, dynamically adjust search strategies, and automatically clean up the knowledge base, thereby achieving low-cost, high-precision system self-evolution. The retrieval method disclosed herein can be applied to fields such as artificial intelligence, natural language processing (NLP), and knowledge engineering.

[0109] Specifically, the technical architecture of retrieval enhancement generation consists of an upper-layer intelligent agent and a lower-layer large language model. The upper-layer intelligent agent is responsible for planning and executing steps for the user's query and calling the retrieval layer to retrieve the results. The lower-layer large language model receives the content planned by the upper-layer intelligent agent and the retrieval results, and completes the final answer construction based on its own language understanding and generation capabilities.

[0110] This embodiment provides an improved retrieval method for existing retrieval enhancement generation techniques. It utilizes local retrieval strategies from a local experience base to acquire historical experience, enabling more accurate decomposition of user queries. Based on the decomposed sub-queries, corresponding keywords are generated. Using the semantic vectors of the sub-queries and the keywords, a hybrid retrieval service is invoked for searching, thus obtaining an optimized method for retrieval enhancement generation techniques and enabling the construction of a new upper-layer intelligent agent. Furthermore, all retrieved content is input as context into a large language model to obtain the query conclusion corresponding to the user query, ultimately generating the corresponding target inference trajectory.

[0111] Figure 1 This embodiment provides a flowchart of a retrieval method, which includes the following steps:

[0112] S101. Obtain user query questions.

[0113] S102. Read the target retrieval strategy that matches the semantics of the user's query question from the local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories.

[0114] In this embodiment, instead of directly performing a search when faced with a user query, the system reads a local experience database to obtain the target search strategy. In this embodiment, the local experience database can also be simply referred to as the experience library. The experience database can be a locally maintained experiences.json file. For example, if the user query is "Analyze the performance bottlenecks of the DeepSearch framework when handling long-tail queries," then the locally maintained experiences.json file is read first.

[0115] In one specific embodiment, the local experience base refers to the experience base deployed on the device performing the current search operation. In other methods, if a user switches devices, from using device A to using another device B for searching, the user can migrate the experiences.json file maintained locally on device A to device B according to actual needs. This avoids the situation where the user cannot access the local experience base experiences.json file maintained locally on device A when using the new device B, thus preventing the inability to obtain experience from historical search strategies.

[0116] The experiences.json file stores local retrieval strategies derived from historical evolution. For example, "[Experience 12]: For issues involving performance bottlenecks, prioritize searching architecture design documents over API manuals." This "[Experience 12]" is extracted from the corresponding historical reasoning trajectory. Therefore, based on the user's query, the system first retrieves the target retrieval strategy matching the user's query from the local experience base. Since the local experience base is formed by accumulating corresponding retrieval strategies extracted from the historical reasoning trajectory of the user's past question-and-answer scenarios, it achieves personalized and targeted retrieval of user queries based on the user's historical usage habits. This avoids rigid retrieval strategies, improves retrieval efficiency, and achieves high-efficiency retrieval.

[0117] S103. Decompose the user query into multiple sub-query questions according to the target retrieval strategy.

[0118] In this embodiment, the target retrieval strategy can be injected as context into the prompt words, and semantic analysis can be performed on the prompt words after the target retrieval strategy is injected. Based on the semantic analysis results, the user query question is decomposed into multiple sub-query questions. Each sub-query question is an independently executable atomic retrieval task.

[0119] For example, given the user query "Analyze the root cause of error code 1004 appearing in the Q&A interface after upgrading from V2.0 to V2.1", the agent retrieves the target retrieval strategy from its local experience base based on this query: [Experience 05] (Analyzing error codes requires a clear definition before checking logs). [Experience 05] is injected into the prompt keywords. Based on semantic analysis of [Experience 05] and the user query, the query is broken down into two sub-queries: "Definition of error code 1004" and "V2.1 version change history".

[0120] In one specific embodiment, the DeepSeek-V3 strong inference model is deployed as the core of the upper-layer online agent. When processing user queries, instead of directly searching, the agent first reads the target retrieval strategy from the local experience base and injects this strategy into the prompt words, thus achieving prompt context assembly. Based on the historical experience of these target retrieval strategies, the agent performs multi-hop chain-of-thought planning, decomposing complex user queries into multiple sub-queries and dynamically generating keywords. Unlike traditional single-search methods, this embodiment adopts the ReAct (Reasoning and Acting) paradigm and innovatively introduces an experience injection mechanism to achieve reasoning for complex user queries.

[0121] DeepSeek-V3 is one of the most powerful inference models in the open-source community. It employs a Hybrid Expert (MoE) architecture with a total of 671 bytes of parameters, but only 37 bytes of activation parameters. The DeepSeek-V3 model introduces a load balancing strategy without auxiliary loss and a multi-head latent attention (MLA) mechanism, ensuring extremely high inference efficiency while possessing logical reasoning and code generation capabilities comparable to GPT-4o. In this specific embodiment, DeepSeek-V3 assumes the role of path planning for the upper-layer agent. Leveraging its powerful thought chain reasoning capabilities, it accurately decomposes complex user query problems based on an experience base, avoiding the rigidity of retrieval strategies.

[0122] In one specific embodiment, the acquired target retrieval strategy can be injected into the suggestion words of the DeepSeek-V3 model employing the retrieval method disclosed herein. Suggestion words without injected target retrieval strategies are as follows:

[0123] "You are a deep search expert. Please refer to the following [Search Experience Library] to break down and plan user questions."

[0124] Experience base: {{experiences}}

[0125] ..."

[0126] The specific prompts injected into the target retrieval strategy are as follows:

[0127] "You are a deep search expert. Please refer to the following [Search Experience Library] to break down and plan user questions."

[0128] Experience base: {{experiences}} ...

[0130] The target retrieval strategy for user issues: [Experience 05] (Error codes must be clearly defined before checking logs)

[0131] S104. Generate corresponding keywords based on multiple subquery questions.

[0132] In this embodiment, 1-3 specific keywords can be generated for retrieval based on each subquery question. For example, the keywords "meaning of Q&A interface error code 1004" and "V2.1 update log" can be generated for the two subquery questions "definition of error code 1004" and "V2.1 version change record", respectively.

[0133] S105. Using the semantic vector and keywords corresponding to the subquery question, call the hybrid retrieval service to perform retrieval in order to generate the corresponding target reasoning trajectory.

[0134] In this embodiment, the semantic vector and keywords corresponding to the subquery question are used to call the hybrid retrieval service for retrieval. Specifically, the vector retrieval in the hybrid retrieval service is called according to the semantic vector corresponding to the subquery question, and the keyword retrieval in the hybrid retrieval service is called according to the keywords, thereby realizing the parallel retrieval of two paths and finally obtaining the target reasoning trajectory, thus achieving accurate retrieval of complex questions.

[0135] In this embodiment, to address the problem of rigid retrieval strategies that prevent targeted retrieval, this disclosure reads target retrieval strategies that match the user's query from a local experience base, instead of performing static retrieval based on the user's query. This allows the user's query to be broken down into multiple sub-queries based on historical experience of target retrieval strategies in the local experience base, achieving targeted retrieval and solving the problem of breaking down complex problems. Furthermore, based on the semantic vectors corresponding to the sub-queries and the keywords generated for the sub-queries, a hybrid retrieval service is invoked for retrieval, enabling simultaneous vector retrieval and keyword retrieval, significantly improving the precision of the retrieval.

[0136] In one optional implementation, the hybrid retrieval service includes keyword retrieval and vector retrieval; the keyword index corresponding to keyword retrieval is constructed based on the mapping relationship between the set keywords and knowledge slices; wherein, the knowledge slices are obtained by segmenting the original documents according to the semantic-aware slicing algorithm, and the original documents are relevant documents in the application domain of the hybrid retrieval service, which are used to construct the knowledge base of the hybrid retrieval service; the vector index corresponding to vector retrieval is constructed based on the data structure of the semantic vectors corresponding to the knowledge slices.

[0137] In this embodiment, to construct an industrial-grade hybrid retrieval service that balances semantic understanding and precise matching, thereby providing solid data support for the retrieval method in this embodiment, a high-precision hybrid retrieval service and a corresponding knowledge base were built. Therefore, single vector retrieval was abandoned, and a hybrid retrieval architecture combining an inverted index based on Elasticsearch (keyword matching) and a dense vector index based on a vector database (such as Milvus) (semantic matching) was constructed. A high-performance embedding model was also introduced to achieve multi-granularity, multi-modal feature extraction, providing a high-recall hybrid retrieval service for the upper-layer intelligent agent. The high-performance embedding model can be BGE-M3, a versatile embedding model that supports over 100 languages, has a maximum input length of 8192, and simultaneously supports dense retrieval, sparse retrieval, and multi-vector retrieval. Specifically, the following steps are included:

[0138] First, the original documents undergo preprocessing, specifically including intelligent segmentation using a semantically aware slicing algorithm to obtain multiple knowledge slices, thus constructing the foundation of the knowledge base. The original documents are relevant documents within the application domain of the hybrid retrieval service, used to build the knowledge base for the hybrid retrieval service. The formats of the original documents include Word, Markdown, etc.

[0139] Traditional segmentation algorithms typically perform mechanical segmentation based on a fixed number of characters, such as segmenting every 500 characters. Semantic-aware slicing algorithms, however, differ from traditional mechanical segmentation methods by introducing a sliding window mechanism, enabling more reasonable segmentation. Specifically, the semantic-aware slicing algorithm uses a small encoder model, BGE-Micro, to calculate the semantic vector similarity between adjacent sentences in the original document in real time. When the similarity score between adjacent sentences is lower than a preset semantic threshold, the algorithm determines that a topic shift has occurred and performs physical segmentation at that breakpoint. The semantic threshold can be set according to the actual situation; for example, setting the semantic threshold to 0.6 means that if the similarity score between adjacent sentences is 0.5, then segmentation will occur between those adjacent sentences.

[0140] BGE-Micro is a lightweight encoder model, typically consisting of only 3-4 Transformer layers. Compared to full-size models like BGE-Large, BGE-Micro is more than 10 times faster inference and has extremely low memory usage. In this implementation, BGE-Micro can be used specifically for the slicing stage to quickly determine the semantic coherence of the context, rather than for final retrieval. This application of BGE-Micro ensures that each generated knowledge slice is semantically complete and independent, avoiding the incorrect truncating of a complete logical paragraph, such as a specific legal clause or technical step, thereby significantly improving the recall accuracy of subsequent retrievals.

[0141] Subsequently, the locally deployed BGE-M3 embedding model is invoked to transform these semantically complete knowledge slices into high-dimensional vectors. In this implementation, using BGE-M3 to generate high-dimensional dense vectors can accurately capture the deep semantic relationships of the text, solving the problem that traditional TF-IDF cannot understand synonyms and implicit intentions.

[0142] Second, construct a hybrid index. To compensate for the shortcomings of vector retrieval in proper noun and exact numerical matching, a "sparse-dense" dual-path indexing strategy can be adopted. The sparse index is based on the keyword index, while the dense index is based on the vector index.

[0143] Keyword Indexing: An inverted index is built based on an Elasticsearch cluster. Elasticsearch is a distributed search and analytics engine based on Lucene, supporting real-time retrieval of massive amounts of data. The BM25 algorithm is used for scoring. BM25 is a classic probabilistic retrieval framework that is a deep optimization of TF-IDF (Term Frequency-Inverse Document Frequency). By introducing two parameters—term frequency saturation and document length normalization—it effectively solves the scoring bias problem caused by high-frequency words in long documents, greatly improving the matching ability for proper nouns, error codes, and precise numerical values.

[0144] Vector Indexing: Built on Milvus or Faiss. Milvus is an open-source vector database designed for processing unstructured data, supporting millisecond-level retrieval of billions of vectors; Faiss is a high-efficiency dense vector clustering and similarity search library. Therefore, both are used to store the embedding vectors generated by BGE-M3, and the HNSW indexing algorithm is used to achieve efficient support for semantic fuzzy matching. HNSW is a graph-based Approximate Nearest Neighbor (ANN) search algorithm that mimics the "six degrees of separation" theory, constructing a multi-layered graph structure. The upper-layer graph contains sparse long-range connections for quickly locating the approximate region of the target vector; the lower-layer graph contains dense short-range connections for high-precision local searches. Compared to traditional inverted file (IVF) indexes, HNSW can maintain a recall rate of over 98% while keeping query latency in the millisecond range when processing billions of high-dimensional vectors, perfectly supporting the high-concurrency retrieval requirements of the RAG system.

[0145] Third, weight-based hybrid retrieval and re-ranking logic:

[0146] When the upper-layer agent receives the user's query, it performs keyword retrieval and vector retrieval in parallel, obtaining two Top-N results (the top N most relevant entries in the search result list). Then, the RRF (Reciprocal Rank Fusion) algorithm is used to merge the two results. The RRF algorithm does not rely on specific absolute similarity scores, but rather merges the results based on the ranking of knowledge slices in both the keyword retrieval result set and the vector retrieval result set.

[0147] The core formula of the RRF algorithm is: ,in, It is a constant (usually 60). The document is in the [number]th [section]. Ranking in path retrieval. This RRF algorithm effectively balances the weights of keyword matching and semantic matching, ensuring that knowledge slices ranking higher in any path receive a higher overall score in the final result.

[0148] Finally, the BGE-Reranker-v2-m3 model is introduced to finely score the top 50 results after RRF fusion. Unlike the dual-encoder architecture of the embedding model, BGE-Reranker adopts a cross-encoder architecture, concatenating the "user query question" and "candidate documents" and directly inputting them into the Transformer network for full-attention interactive computation. Although this mechanism has a slightly higher computational cost, it can accurately capture the subtle logical relationships between the query and the documents, thereby outputting highly accurate Top-K results to the upper-layer agent, achieving the output of high-precision hybrid retrieval results.

[0149] In a specific example, during the step of calling the hybrid retrieval service, to fully utilize the corrected retrieval results after updating the dynamic weight score (boost_score) in the metadata, this embodiment improves the traditional RRF algorithm by adopting a weighted inverse ranking fusion algorithm (Weighted-RRF), breaking its limitation of relying solely on retrieval ranking. The calculation formula is as follows:

[0150]

[0151] in, This refers to the ranking position of knowledge slices in keyword retrieval or vector retrieval. This is a preset constant (e.g., 60). For the dynamic weight score read from the metadata, This is a weighting adjustment factor, set according to the actual situation.

[0152] This formula ensures that knowledge slices that are frequently used in historical successful reasoning trajectories, even if they rank slightly lower in single-path retrieval, can still have their final fusion score improved due to compensation from dynamic weight scores, thereby ensuring that knowledge slices, as high-value experiential data, are recalled first.

[0153] In this embodiment, unlike existing technologies that directly retrieve knowledge base information from user queries, it first reads local retrieval strategies from a local experience base to obtain target retrieval strategies that semantically match the user query. Based on these target retrieval strategies, the user query is then broken down into multiple sub-queries, thus leveraging local historical experience. In one optional embodiment, such as... Figure 2 As shown, the target retrieval strategy includes the execution order and dependencies between subqueries; step S103 includes:

[0154] S1031. Perform semantic analysis on the user's query to determine the user's query intent.

[0155] In this implementation, semantic analysis improves the accuracy of query intent identification, avoids semantic confusion, and ensures that the subsequently retrieved content, i.e., the target knowledge slice, truly meets user expectations. In a specific example, semantic analysis of a user's query can distinguish whether "apple" refers to a fruit or a company.

[0156] S1032. Divide the query intent into initial query subproblems with dependencies.

[0157] In this embodiment, the intent type of the user's query question can be identified based on the query intent, such as comparative analysis or multi-hop reasoning, thereby dividing the query into initial query sub-questions with dependencies based on the query intent and avoiding invalid retrieval.

[0158] S1033. Sort the initial query subproblems according to the execution order to obtain multiple ordered subquery problems.

[0159] In this implementation, for user query questions that require multi-step reasoning, the initial query sub-questions are sorted according to the execution order, making the subsequent retrieval path more reasonable.

[0160] In one optional implementation, the retrieval method further includes: in response to the fact that the relevance between the knowledge slice contained in the target reasoning trajectory and the user query question is less than the relevance threshold, returning to step S103.

[0161] In this implementation, after one round of retrieval, if the relevance between the knowledge slices contained in the target inference trajectory and the user's query question is less than the relevance threshold (meaning the knowledge slices in the current target inference trajectory cannot answer the user's query question), the process returns to the step of breaking down the user's query question into multiple sub-queries according to the target retrieval strategy. The user's query question is then re-broken to generate new sub-queries and corresponding keywords. The hybrid retrieval service is then invoked again for a new round of retrieval to ensure that the retrieved knowledge slices can answer the user's query question. The relevance threshold can be set according to the actual situation.

[0162] In a specific example, after the upper-layer agent calls the hybrid retrieval service to perform a round of retrieval, the large language model determines whether the information of the currently obtained knowledge slices is sufficient. For example, whether the relevance of these knowledge slices to the user's query question is less than the relevance threshold. If not, it returns to the question decomposition step S103 and performs the next round of retrieval; if so, it outputs the query conclusion corresponding to the user's query question and obtains the final answer.

[0163] In one alternative implementation, such as Figure 3 As shown, before step S102, the following steps are also included:

[0164] S201. Overview of the knowledge base for obtaining hybrid search services.

[0165] In this embodiment, to ensure that a user's query can be directly answered, it is also necessary to obtain a knowledge base overview of the hybrid retrieval service. This step can be performed simultaneously with the user's query. The knowledge base overview may include basic identifying information such as the knowledge base's name, ID, description, and creation time. Furthermore, the overview covers the knowledge base's scale, such as the number of original documents, the number of knowledge slices, and the knowledge base's domain.

[0166] S202. In response to the determination based on the knowledge base overview of the hybrid retrieval service that the knowledge base cannot directly answer the user's query question, it is determined that the user's query question needs to be broken down.

[0167] In this embodiment, since the knowledge base overview determines that it cannot directly answer the user's query, it is necessary to break down the user's query to achieve a precise answer. This embodiment uses the knowledge base overview to determine when it is necessary to break down a large question into sub-queries.

[0168] In one alternative implementation, such as Figure 4 As shown, the retrieval method also includes:

[0169] S106. Obtain the query questions corresponding to the user's historical reasoning trajectory of the target type in the historical log generated based on retrieval enhancement, and obtain the question set; wherein, the user's historical reasoning trajectory also includes the query conclusion.

[0170] In this embodiment, specific query questions are processed to enable the self-evolution of the retrieval method. These specific query questions are typically "hard problems" with poor user feedback or low model confidence. Specifically, the target type includes those with poor user feedback and model confidence below a confidence threshold, which can be set according to actual circumstances.

[0171] S107. For each query question in the question set, a preset number of searches are performed to obtain a group of inference trajectories corresponding to each query question; wherein, each group of inference trajectories contains a preset number of inference trajectories.

[0172] In this implementation, multiple inference trajectories are triggered in parallel according to a preset number of times. Each inference trajectory may employ different search keywords and different subquery question decomposition logic. Therefore, based on the complete data of the "planning-retrieval-reading-answering" process, a trajectory set containing successful and failed paths is formed, i.e., an inference trajectory group.

[0173] In one specific implementation, the retrieval method of this disclosure is applied to the DeepSeek-V3 model, utilizing the DeepSeek-V3 model to acquire the inference trajectory group in this retrieval method. For each difficult question in the question set, by adjusting the sampling temperature of DeepSeek-V3 (e.g., setting it to 0.7), the DeepSearch process is run in parallel a preset number of times (e.g., n=5) to achieve hybrid retrieval. Due to the randomness of the DeepSeek-V3 model, these 5 runs will generate different search paths: some may find key documents, while others may wander through irrelevant information. Thus, these 5 complete interaction records, including search terms, selected knowledge slices, intermediate inference steps, and the final answer, are saved as a set of inference trajectory groups.

[0174] For example, for the query "Question and Answer Interface Error Code 1004", the sampling temperature was set to 0.7, and the DeepSearch process was run in parallel 5 times, generating differentiated inference trajectory groups:

[0175] Reasoning trajectory A (success): The agent accurately checked the "error code definition" first, clarified that it was a "timeout" problem, then checked the "configuration change", discovered the contradiction of the shortened time threshold, and finally deduced the correct conclusion.

[0176] Inference Trajectory B (Failure): The agent was misled by an irrelevant document in the search results that said "Database connection failed (error code 2003)," and kept searching for database connection pool configuration, completely deviating from the core direction of "timeout."

[0177] Inference trajectory C (failure): Although the agent found the definition of "timeout", it ignored the text paragraph about "time threshold adjustment" when checking the update log and instead searched for network hardware failure, resulting in an attribution error.

[0178] Inference trajectory D and inference trajectory E: other variants.

[0179] These five reasoning trajectories, each containing a complete intermediate step (Thought), search term (Action), and final answer (Response), are packaged together as the original sample for subsequent offline semantic analysis, i.e., the reasoning trajectory group.

[0180] S108. In response to the existence of a number of target query conclusions in the inference trajectory group that are greater than the number threshold, the inference trajectory corresponding to the target query conclusion is judged as successful, and the inference trajectory corresponding to other query conclusions is judged as unsuccessful.

[0181] In this embodiment, the quantity threshold can be set according to actual conditions. Specifically, this embodiment can load a referee model to evaluate multiple inference trajectories in each inference trajectory group. The referee model can achieve empirical distillation of the question set in offline state based on semantic advantages. For example, the referee model can be DeepSeek-R1, thereby realizing the success or failure determination of inference trajectories in the inference trajectory group. Among them, other query conclusions include query conclusions other than the target query conclusion, and also include no conclusion. Specifically, DeepSeek-R1 is different from traditional next word prediction models. DeepSeek-R1 has undergone large-scale pure reinforcement learning training, internalized a very strong chain-of-thought (CoT) capability, and has the characteristic of "slow thinking". It can perform self-game and multi-step logical verification before outputting the final evaluation, thereby being able to keenly detect extremely hidden logical jumps or factual fallacies in the inference trajectory.

[0182] In one specific implementation, a chained prompt word template can be constructed to guide DeepSeek-R1 in determining the success or failure of the final conclusions of multiple inference trajectories within an inference trajectory group. Specifically, this method can be applied to scenarios lacking a standard answer, automatically determining the success or failure of inference trajectories based on a "majority rules" consistency mechanism. The chained prompt word template is as follows:

[0183] 1. Input: The final query results of 5 trajectories (no standard answer required).

[0184] 2. Hint: "Please analyze these 5 tracks to arrive at the final query conclusion."

[0185] Majority voting: Check if there is a significant majority opinion (e.g., more than 3 out of 5 opinions believe that the threshold misconfiguration is the cause).

[0186] Consensus determination: If a majority opinion exists, that opinion is considered a pseudo-standard answer.

[0187] Labeling: The reasoning trajectory that leads to the majority conclusion is labeled 'success'; the reasoning trajectory that leads to the minority conclusion or cannot reach a conclusion is labeled 'failure'.

[0188] If the query results for the five trajectories are extremely scattered and a consensus cannot be reached, then this round of analysis is invalid and no experience will be generated.

[0189] In this specific example, based on the above chained prompt word template, multiple query conclusions corresponding to multiple inference trajectories in an inference trajectory group are input, and the results are analyzed based on the prompt words to determine the success or failure of the final conclusions of multiple inference trajectories in the inference trajectory group.

[0190] In another specific implementation, the DeepSeek-R1 model performs Training-Free GRPO analysis on the inference trajectory groups collected in step S107. DeepSeek-R1 does not calculate numerical gradients, but instead analyzes why successful inference trajectories succeed (e.g., "using specific qualifiers") and why failed inference trajectories fail (e.g., "quoting outdated content") through semantic comparison, thereby extracting structured natural language experience items. Specifically, the locally deployed DeepSeek-R1 model is used as the core evaluator, leveraging its powerful logical backtracking capabilities to perform deep attribution analysis on complex inference trajectories, achieving a significantly higher accuracy than general-purpose models such as DeepSeek-V3. For example, a set of chained prompt word templates is constructed to guide DeepSeek-R1 in performing phased semantic analysis. This semantic analysis can specifically include success / failure determination and difference attribution. The chained prompt word template for success / failure determination is shown above, while the chained prompt word template for difference attribution is as follows:

[0191] 1. Input: A reasoning trajectory marked as "successful" (e.g., reasoning trajectory A) + a reasoning trajectory marked as "failed" (e.g., reasoning trajectory B).

[0192] 2. Hint: "It is known that reasoning trajectory A leads to a consensus-based correct conclusion, while reasoning trajectory B fails. Please compare the differences between the two:"

[0193] At which step in the analysis reasoning trajectory A did it capture the key textual clue (such as '1004 = timeout')?

[0194] Which specific noise slice misled the analytical reasoning trajectory B?

[0195] Lesson Summary: Based on the differences mentioned above, the model generates a general suggestion in natural language. For example, if multiple failures are found to be due to searching overly broad terms, the model will conclude: "[New Lesson]: When analyzing system performance, avoid using broad terms and combine them with specific metrics (such as TPS, Latency) for a combined search."

[0196] In this process, the DeepSeek-R1 model can analyze layer by layer what the successful inference trajectory did right (e.g., "the key knowledge slice A was retrieved by using 'error code' as a keyword in the second round of search"), and what the failed inference trajectory did wrong (e.g., "it was misled by outdated information in knowledge slice B, causing the inference direction to deviate").

[0197] S109. Add the first retrieval strategy to the local experience base; the first retrieval strategy is a retrieval strategy generated based on the difference between successful reasoning trajectories and failed reasoning trajectories.

[0198] In this implementation, a first retrieval strategy, generated based on the differences between successful and failed inference trajectories, is embedded into the local experience base, thus updating the local experience base. Specifically, new first retrieval strategies derived from the experience of the DeepSeek-R1 model are added. If this first retrieval strategy is semantically similar to a local retrieval strategy in the local experience base, the two are merged to generate a more refined retrieval strategy.

[0199] S110. Delete the second retrieval strategy from the local experience base; the second retrieval strategy is the retrieval strategy used by multiple failed inference trajectories.

[0200] In this embodiment, if an old second search strategy is proven to be ineffective or misleading in multiple practices, the second search strategy is removed from the local experience base, thus achieving automatic maintenance of the local experience base.

[0201] Specifically, successful inference paths are identified as high-value retrieval strategies and added to the local experience base (Add). If the added retrieval strategy is semantically similar to an existing local retrieval strategy in the experience base, DeepSeek-R1 is used to merge them into a more refined description (Merge). For retrieval strategies corresponding to failed inference paths, if the strategy has been proven ineffective or misleading in multiple trials, it is removed (Delete). The updated experience base will immediately take effect in the next round of online inference for user queries, without any model training, achieving low-cost, high-precision self-evolution. Furthermore, if the query conclusions of multiple inference paths in a group are extremely scattered and cannot reach a consensus, the analysis in this round is discarded, and no experience is generated. This improves the accuracy of dynamic adjustment of the experience base and avoids interference from inference paths without a clear conclusion bias.

[0202] Therefore, in this implementation, experience is summarized from successful and failed inference trajectories, and the retrieval strategy in the local experience base is dynamically adjusted. The updated local experience base will take effect immediately in the next round of online inference retrieval without any model training, thus realizing a training-free experience purification mechanism and automatic optimization of the retrieval method and automatic updating of the experience base. Specifically, by evaluating multiple inference trajectories in each inference trajectory group, the local experience base can be added, deleted, and modified. During the retrieval process of the upper-layer agent, it becomes increasingly intelligent, learning from historical interaction data and achieving self-evolution. This realizes a lightweight, model parameter-free experience evolution mechanism. The historical interaction data includes historical inference trajectories and local retrieval strategies extracted from them.

[0203] In addition, in a specific implementation, it may include only step S109, only step S110, or both steps S109 and S110.

[0204] In one optional implementation, the retrieval method further includes:

[0205] S111. Delete the first knowledge slice in the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is the knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group.

[0206] In this embodiment, deleting the first knowledge slice from the knowledge base corresponding to the hybrid retrieval service constitutes the purification of the knowledge base, thus maintaining the underlying index of the knowledge base. The first knowledge slice can be a knowledge slice with a specific index number appearing in multiple failed reasoning trajectories within a reasoning trajectory group. Therefore, this first knowledge slice can be deleted from the index logic, achieving noise reduction and purification of the knowledge base. In a specific implementation, a referee model can be used to mark the first knowledge slice as "misleading information," thereby facilitating its deletion.

[0207] In other specific implementations, the first knowledge slice can also be marked as a low-quality slice in the metadata of the knowledge slice, thereby achieving noise reduction and purification of the knowledge base.

[0208] S112. Increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice of multiple successful reasoning trajectories in the dominant reasoning trajectory group.

[0209] In a specific embodiment, a "dynamic weight-aware" mechanism can be adopted: For the inverted index, during the process of calling the hybrid retrieval service, when retrieving the corresponding knowledge slices through keyword retrieval, while calculating the basic BM25 score, the boost_score (dynamic weight score) in the metadata of each knowledge slice is read, and the basic BM25 score and the dynamic weight score are weighted and calculated to obtain the combined score of each knowledge slice. Here, metadata refers to the structured attribute information describing the knowledge slice, including at least one of the following: dynamic weight score, original document source, author, creation time, file type, chapter title, ID of the corresponding original document, and permission tags. Since the dynamic weight score of a knowledge slice changes with subsequent updates to the knowledge base, this means that knowledge slices that have been repeatedly verified as "secondary knowledge slices" in historical successful inference trajectories will receive a higher initial recall ranking; therefore, secondary knowledge slices can also be called "golden slices."

[0210] In this implementation, the second knowledge slice is the one that dominates multiple successful reasoning trajectories. Therefore, if a knowledge slice repeatedly leads to successful answers, its weight is increased. Specifically, the weight of this second knowledge slice corresponds to a dynamic weight score. In hybrid retrieval, due to the "dynamic weight awareness" mechanism, this dynamic weight score participates in the final ranking calculation, making high-quality knowledge slices easier to retrieve. Therefore, through this entire closed-loop process, this implementation achieves a retrieval method with self-purification and self-learning capabilities, completely solving the pain points of high maintenance costs and long optimization cycles in traditional solutions.

[0211] In a specific implementation, it may include only step S111, only step S112, or both steps S111 and S112.

[0212] In this embodiment, by downgrading or soft-deleting the first knowledge slice that caused the reasoning trajectory to fail in the knowledge base, and upgrading the second knowledge slice that led multiple successful reasoning trajectories, automatic noise reduction of the knowledge base is achieved. This dynamic purification of the knowledge base during the retrieval process avoids the generation of a large number of redundant, outdated, or even conflicting knowledge slices in actual operation. The knowledge base processing method in this embodiment solves the problem of the traditional knowledge base maintenance relying entirely on manual review and being inefficient. It adopts an automated attribution mechanism to identify which knowledge slices played a key role in successful historical reasoning trajectories and which knowledge slices led to erroneous query conclusions in failed reasoning trajectories. Based on the automated attribution mechanism, the first and second knowledge slices are processed, avoiding the continuous decline of the signal-to-noise ratio during the retrieval process and improving the quality of the generated query conclusions corresponding to the user's query question.

[0213] In one alternative implementation, directly physically deleting or rebuilding the underlying knowledge base index can lead to significant system overhead and potentially cause lockouts in online hybrid retrieval services. Therefore, this disclosure employs a lightweight, metadata-based soft isolation mechanism for deleting the first knowledge slice, avoiding performance degradation caused by frequent physical modifications. Figure 5 As shown, step S111 specifically includes:

[0214] S1111. Obtain the globally unique identifier of the first knowledge slice in the knowledge base.

[0215] In this implementation, the first knowledge slice extracted from the failed inference trajectory, i.e., the misleading noise chunk, is not immediately erased from the underlying disk. Instead, a globally unique identifier (UUID) corresponding to the first knowledge slice is obtained. The globally unique identifier is the unique identity of the first knowledge slice in the knowledge base, enabling precise positioning of the first knowledge slice.

[0216] S1112. Modify the metadata status bit corresponding to the globally unique identifier to an interceptable state. The interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service.

[0217] In this embodiment, the interceptable states include soft deletion state or isolation state. The metadata status bits of the first knowledge slice can be marked as soft deletion state or isolation state, so that when the hybrid retrieval service is called in the next round, the first knowledge slice in the soft deletion state or isolation state can be directly intercepted in the pre-filtering stage of the hybrid retrieval service.

[0218] In one specific implementation, after obtaining the globally unique identifier of the first knowledge slice, the underlying database's update API can be called to modify the deprecated status bit in the metadata status of the first knowledge slice, such as changing the is_deprecated field from False to True. During the next round of retrieval, the hybrid retrieval service will directly filter out slices with a deprecated status bit of True during the pre-filtering stages of vector approximate nearest neighbor search (ANN) and inverted index matching, thus achieving logical purification of the knowledge base within milliseconds.

[0219] S1113. In response to the first knowledge slice being in the interceptable state for a duration exceeding a preset duration, delete the first knowledge slice from the knowledge base.

[0220] In this implementation, the preset duration can be set as a preset aging period, which can be configured according to the user's actual needs. When the duration for which the first knowledge slice is in an interceptable state exceeds the preset aging period, an asynchronous physical deletion operation is triggered by the underlying storage engine to delete the first knowledge slice and free up storage space. By periodically cleaning up these first knowledge slices through an asynchronous background daemon during off-peak hours, the stability of the system under high concurrency is ensured.

[0221] In one alternative implementation, regarding increasing the weight of the second knowledge slice (i.e., the gold slice), this disclosure designs a dynamic scoring synchronization architecture across dual engines. For example... Figure 6 As shown, step S112 specifically includes:

[0222] S1121. Obtain the current dynamic weight score from the metadata corresponding to the second knowledge slice.

[0223] S1122. Calculate the dynamic weight increment of the second knowledge slice based on the number of successful inference trajectories of the target containing the second knowledge slice and / or the confidence score of the successful inference trajectories of the target.

[0224] In this embodiment, when it is determined that a certain second knowledge slice has repeatedly led to correct reasoning conclusions and successful reasoning trajectories, the dynamic weight increment of the second knowledge slice can be calculated to achieve dynamic weight score updates for the second knowledge slice. The confidence score of the target successful reasoning trajectory is used to represent the contribution of the second knowledge slice to the target successful reasoning trajectory.

[0225] In this implementation, the dynamic weight increment is F(the number of successful inference trajectories and the confidence score of each successful inference trajectory). For example, the dynamic weight increment can be the sum of the confidence scores of each successful inference trajectory. If the second knowledge slice A is used in two successful inference trajectories within a day, with confidence scores of 0.9 and 0.6 respectively, then the dynamic weight increment for the second knowledge slice A is 0.9 + 0.6 = 1.5.

[0226] S1123. Update the dynamic weight score of the second knowledge slice based on the current dynamic weight score and the dynamic weight increment.

[0227] In this implementation, after aggregating the current dynamic weight score and the dynamic weight increment to obtain the updated dynamic weight score, the metadata of the inverted index corresponding to keyword retrieval and the vector database corresponding to vector retrieval are synchronously updated. This ensures that in subsequent retrievals, the updated dynamic weight score is configured to participate in calculating the recall score for single-path retrieval and as a dynamic adjustment coefficient when performing inverse ranking fusion of dual-path hybrid retrieval results.

[0228] Specifically, the dynamic weight score `boost_score` of the metadata corresponding to the second knowledge slice in the inverted index (such as Elasticsearch) is updated, enabling a higher multiplier bonus through the modified BM25 scoring script during subsequent keyword matching. Simultaneously, the dynamic weight score `boost_score` field in the metadata corresponding to the second knowledge slice in the vector database (such as Milvus or Faiss) is updated. In the dual-path fusion stage of hybrid retrieval, such as the RRF inverse ranking fusion stage, this updated `boost_score` can also be directly used as a dynamic adjustment coefficient, breaking the limitation of the traditional RRF algorithm relying solely on the ranking constant. This mechanism ensures that high-value knowledge extracted from historical experience not only remains in the memory of the upper-level agent but is truly solidified in the lowest-level database metadata, completely establishing a feedback loop from cognitive planning to underlying data storage.

[0229] In one optional implementation, step S105 specifically includes:

[0230] Step 1: Concurrently call the inverted index corresponding to keyword retrieval and the vector index corresponding to vector retrieval to obtain the initial Top-N knowledge slice lists for the two paths respectively;

[0231] Step 2: While recalling the two Top-N knowledge slices, simultaneously read the metadata corresponding to each knowledge slice and extract the boost_score field from it;

[0232] Step 3: The extracted boost_score is used as an adjustment coefficient. Combined with the ranking of each knowledge slice in each retrieval path, the weighted inverse ranking fusion algorithm is used to calculate the comprehensive weight score of each slice. The candidate knowledge slices are then reordered according to the comprehensive weight score to generate the target reasoning trajectory, which further improves the refinement of the hybrid retrieval.

[0233] In one specific implementation method Figure 7 This is a flowchart illustrating the process of updating an experience base and a knowledge base based on semantic analysis results in a retrieval method, specifically including:

[0234] S401. Obtain the multi-path trajectory set; this trajectory set is the inference trajectory group; input the multi-path trajectory set into the referee model DeepSeek-R1 for analysis, and execute step S402;

[0235] S402, Success or failure determination stage; The multi-path trajectory set is analyzed using a referee model to determine success or failure. Based on consistent self-consistent voting, step S403 is executed.

[0236] S403, Mark successful / failed trajectories; mark trajectories that lead to the majority conclusion as successful; mark trajectories that lead to other minority conclusions or cannot reach a conclusion as failed;

[0237] S404, Differential Attribution Stage; Based on the differential attribution results, compare and analyze the two different reasoning trajectories, and execute step S405.

[0238] S405. Identify key elements; based on the key elements, analyze the key success actions and the noise that leads to failure, and proceed to step S406.

[0239] S406, Experience Refinement Stage;

[0240] S407, Structured Recommendations: Provide structured recommendations for experience bases and / or knowledge bases based on general, natural language-based recommendations derived from experience.

[0241] S408. Update the execution module; obtain optimized actions based on structured suggestions, and update the execution module to achieve steps S409 and S410.

[0242] S409. Update the experience base; based on the retrieval strategy, implement at least one of the following: adding, merging, and deleting strategies in the experience base;

[0243] S410. Update the knowledge base index; based on the newly added experience, the knowledge slices corresponding to the knowledge base can be downgraded or upgraded to realize the data evolution of the knowledge base.

[0244] In one optional implementation, the target reasoning trajectory includes at least two target knowledge slices, and after step S105, the method further includes: in response to the existence of a conflict in at least two target knowledge slices, performing a next round of retrieval.

[0245] In this embodiment, these target knowledge slices are judged. If a conflict is detected in at least two target knowledge slices, the next round of retrieval is performed. This avoids the conflict in the target knowledge slices from affecting the quality of the query results for the user's query question and avoids generating incorrect query results.

[0246] The conflict determination methods include at least one of the following:

[0247] First, calculate the conflict confidence score between any two target knowledge slices. If the conflict confidence score exceeds the corresponding conflict confidence score threshold, it is determined that there is a conflict between at least two target knowledge slices. The conflict confidence score threshold can be set according to the actual situation.

[0248] In one specific implementation, conflict probability quantification can be achieved based on non-logical implication (NLI), utilizing a cross-encoder architecture to perform pairwise logical verification on multiple retrieved target knowledge slices. The operational logic includes: simultaneously inputting text fragments from two potentially related target knowledge slices into the model, thereby outputting confidence scores in three dimensions: "consistent," "neutral," or "conflicting." A decision threshold can be set; a conflict confidence threshold can be preset, for example, 0.8. When the model's output score for the "conflicting" dimension exceeds this confidence threshold, the two target knowledge slices are automatically determined to have a potential conflict.

[0249] This method of calculating confidence transforms abstract semantic contradictions into comparable probability values, enabling precise capture of logical hard conflicts such as "version A supports this feature" and "version B has deprecated this feature".

[0250] Second, in response to the detection of contradictory values ​​contained in keywords in multiple knowledge slices, it is determined that there is a conflict in at least two target knowledge slices.

[0251] In this implementation, a structured extraction mechanism based on structured attributes can be used to extract target knowledge slices containing parameters, values, or status bits. For example, at least two target knowledge slices can be extracted using an inference model. The inference model can be DeepSeek-R1, and the specific operation logic includes: extracting key indicators (e.g., a timeout threshold set to 3 seconds) from target knowledge slice A, and extracting actual operating indicators (e.g., an average gateway response time of 4.2 seconds) from target knowledge slice B. Logical verification is performed according to preset judgment rules. For example, when the extracted numerical logical relationship violates the logical closed loop of normal system operation, it is directly marked as a "logical conflict point"; such as "actual time consumption" being greater than "set threshold". This numerical comparison method transforms textual descriptions into a comparison of logical facts, which can greatly improve the accuracy of the system in handling rigorous scenarios such as financial audits and technical troubleshooting.

[0252] Third, consistency and self-consistency judgment based on inference trajectories. By simultaneously opening multiple inference trajectories for a user query and observing whether the intermediate conclusions obtained from different paths diverge, the consistency of the inference trajectories can be determined. The specific operational logic includes: in multiple inference trajectories for the same user query, if three out of five independent inference trajectories point to conclusion A, while the other two point to the completely opposite conclusion B, information conflict in the data source can be identified. Furthermore, the voting difference rate can be introduced as a quantitative indicator in this process. When the query conclusion with the highest number of votes fails to reach an absolute majority (e.g., a vote rate below 60%), it is determined that the currently retrieved knowledge base content contains serious noise or conflict, thus triggering a deeper secondary search.

[0253] In practice, these three dimensions do not exist in isolation within the system, but rather work collaboratively using a "layered triggering and progressive combination" logic. In actual operation, they together form a complete closed-loop decision-making process from probability perception to logical confirmation. The specific combination method is as follows:

[0254] 1. Asynchronous triggering and multi-dimensional verification, specifically including:

[0255] a. Preliminary perception of the probability dimension: After obtaining the retrieved target knowledge slices, a coarse screening is first performed through logical implication. If the "conflict" dimension scores of two target knowledge slices exceed the conflict confidence threshold, no conclusion is drawn immediately. Instead, the two target knowledge slices are marked as "potential conflict points".

[0256] b. Deep verification in the logical dimension: For the "potential conflict points" marked above, the agent enters a deep reasoning mode to extract the structured numerical values ​​or factual states. For example, it transforms vague textual descriptions into specific parameter comparisons and verifies whether the conflict truly exists through factual logic, thereby achieving verification in both probabilistic and logical dimensions and improving the accuracy of conflict determination.

[0257] 2. Closed-loop feedback and experience consolidation, specifically including:

[0258] a. Consensus adjudication based on consistency: When ambiguity cannot be completely eliminated even at the logical level, such as when two rules seem reasonable but lead to opposite conclusions, parallel sampling of multiple inference paths is initiated. Utilizing the "majority rules" consistency mechanism, a statistical determination is made as to which inference path better aligns with the current business logic.

[0259] b. Evolutionary storage: The final conflict result will be extracted into a structured experience entry and stored in the experience database experiences.json.

[0260] In this implementation, each dimension can independently identify a certain type of error, such as plain text conflict, numerical logic error, or reasoning divergence. Through the continuous accumulation of experience base, when dealing with similar problems in the future, the system will directly call the solution for the conflict, avoiding repeated triggering of costly deep verification.

[0261] In one specific implementation, based on DeepSeek-V3 inference, the specific workflow of the DeepSearch agent employing the retrieval method disclosed herein and its interaction logic with the hybrid retrieval service are as follows: Figure 8 As shown, it specifically includes: S601, Start: Receive user queries, i.e., user query questions;

[0262] S602, Loading the experience library and system assembly prompts;

[0263] S603, LLM planning; specifically, it uses the DeepSeek-V3 model for execution, thus entering the ReAct iteration loop, thereby realizing the stream_planning loop, dynamically intertwining planning and execution. This is not the traditional "complete planning first and then execution" mode, but emphasizes the closed loop of "planning, executing, and correcting simultaneously" during the reasoning process.

[0264] S604. Determine if the information is sufficient. If not, proceed to step S605; if yes, proceed to step S611. In this stage, the DeepSeek-V3 model retrieves the search results, updates its internal state, and decides whether to proceed to the next search round. This process typically repeats 2-3 rounds until sufficient information is available.

[0265] S605, Problem Decomposition and Keyword Generation: When the DeepSeek V3 model determines that the current information is insufficient, it decomposes the user's query problem into subquery problems and generates corresponding keywords based on the subquery problems.

[0266] S606, Invoke the hybrid search service;

[0267] S607, Parallel execution of dense and sparse retrieval; where dense retrieval is based on dense vector indexes of vector databases (such as Milvus), and sparse retrieval is based on inverted indexes of Elasticsearch and the BM25 algorithm.

[0268] S608 and RRF fusion: The results of dense and sparse searches are fused using the RRF algorithm to merge the two search results and obtain preliminary screening results, making the search results more accurate.

[0269] S609, Refined Reordering: The BGE-Reranker model is used to reorder the initial screening results to ensure that the most relevant content is ranked first, thereby further improving the accuracy of the search.

[0270] S610, Return the Top-K results, read the materials and update the internal status; Return to the next round and execute step S603 again; The Top-K results refer to the top K results in the search results after sorting by relevance. They are used to balance efficiency and quality and are the standard output format of the search system.

[0271] S611. Generate the final answer;

[0272] S612, End: Output the answer.

[0273] In a specific implementation, the DeepSearch agent's workflow and its interaction logic with the hybrid search service are as follows:

[0274] 1. Intent Identification and Decomposition: Based on the knowledge base overview, determine if the current knowledge base has insufficient information. If insufficient, decompose the user's query into subqueries.

[0275] 2. Keyword generation: Generate 1-3 specific search keywords based on the subquery question.

[0276] 3. Tool Invocation: Invoke the hybrid search service interface in step one to obtain the query result SearchResult.

[0277] 4. Deduplication and Reading: Read the retrieved content, including at least two knowledge slices, update the internal state, and decide whether to proceed with the next round of searching. This process is usually repeated 2-3 times until sufficient information is available.

[0278] Specifically, let's take the user's question, "Analyze the root cause of error code 1004 appearing in the Q&A interface after upgrading from V2.0 to V2.1," as an example:

[0279] In the first round of retrieval, the agent, based on experience [Experience 05] (error codes need to be clearly defined before checking logs), decided to break the problem down into two sub-queries: "definition of error code 1004" and "V2.1 version change log," thus achieving intent recognition and problem decomposition. The agent then invoked a tool based on these two sub-queries, calling the hybrid retrieval service to search for the keywords "meaning of Q&A interface error code 1004" and "V2.1 update log." Next, in the reading and reflection phase, the retrieved plain text API documentation showed: "1004 represents an upstream gateway timeout." Simultaneously, the retrieved text update log showed: "Version V2.1 adjusted the upstream service response time from 5 seconds to 3 seconds."

[0280] The agent discovers a potential conflict through textual reasoning: the error code points to "timeout," while the update log shows "wait time shortened." Therefore, a new search task is added: "Average response time of the upstream gateway," initiating a second round of retrieval and implementing dynamic programming. Ultimately, the maintenance monitoring text log obtained through the second round of retrieval shows: "The average response time of the upstream gateway is 4.2s." The agent, combining the pure textual evidence chain, infers that adjusting the timeout threshold (3s) to a level below the average response time (4.2s) is the root cause of the surge in error code 1004, thus arriving at the final query conclusion.

[0281] In one specific implementation method Figure 9 This is a schematic diagram of the specific structure of a retrieval method, showing a dual-loop structure of online reasoning and offline evolution, including an online reasoning loop 81, an offline evolution loop 82, and a core storage layer 83.

[0282] The online reasoning loop 81 includes user query 811, prompt word assembly module 812, DeepSesrch agent core 813, hybrid retrieval service 814, and final answer 815.

[0283] The offline evolution loop 82 includes a trajectory sampling module 821, a semantic advantage analysis module 822, and an evolution and optimization module 823.

[0284] The core storage layer 83 includes an experience base 831, original document sources 832, and a knowledge base index 833.

[0285] User query 811 retrieves user query questions;

[0286] The prompt word assembly module 812 is used to call the experience base 831 in the core storage layer 83 to realize experience injection, specifically including: reading the target retrieval strategy that matches the user's query question from the experience base generated based on retrieval enhancement; and injecting the target retrieval strategy as context into the prompt words;

[0287] The DeepSesrch agent core 813 can use the DeepSeek-V3 model to perform multi-round planning and retrieval based on the prompts of injected experience. The agent performs multi-hop thinking chain planning based on the experience of these target retrieval strategies, decomposes complex user query problems into multiple sub-query problems, dynamically generates keywords, and innovatively introduces an experience injection mechanism to realize reasoning for complex user query problems.

[0288] The hybrid retrieval service 814 is used to call the knowledge base index 833 to provide data support for hybrid retrieval based on multi-round planning and retrieval strategies, obtain Top-K results, and return them to the DeepSesrch agent core 813 to determine whether the current information is sufficient. If so, the final answer is generated; otherwise, the next round of retrieval is performed.

[0289] The final answer 815 is used to output the final answer, which is the query result corresponding to the user's query question.

[0290] The trajectory sampling module 821 is used to obtain the query questions corresponding to the user's historical reasoning trajectory of the target type in the historical log generated based on retrieval enhancement, and obtain the first question set; wherein, the user's historical reasoning trajectory also includes the query conclusion;

[0291] The semantic advantage analysis module 822 is used to perform semantic analysis on each reasoning trajectory group to obtain success or failure judgment results and difference attribution results, and finally obtain semantic attribution and suggestions;

[0292] The evolution and optimization module 823 is used to evolve and optimize the experience base 831 and the knowledge base index 833 based on semantic attribution and suggestions. Specifically, this includes updating historical retrieval strategies in the experience base 831, and adding, deleting, or modifying historical retrieval strategies based on semantic analysis results. It also includes downgrading or upgrading the weight of knowledge slices in the knowledge base index 833 to achieve data purification. The evolution and optimization module 823 enables the online inference loop 81 to learn from past successes and failures, dynamically adjust search strategies, and automatically purify the knowledge base, just like humans do, thereby achieving low-cost, high-precision system self-evolution.

[0293] The experience base 831 is used to provide experience data for the online inference loop 81. The experience base includes local retrieval strategies extracted from the user's historical inference trajectory. The online inference loop 81 can read the target retrieval strategy that matches the user's query question based on the obtained user query question. Specifically, the experience base 831 can be experiences.json. This allows the online inference loop 81 to learn from historical interaction data and achieve self-evolution, thereby realizing a lightweight experience evolution mechanism that does not require updating model parameters. It can summarize experience from historical successes and failures, dynamically adjust the retrieval strategy, and automatically purify the knowledge base, thereby achieving low-cost and high-precision system self-evolution.

[0294] The original document source 832 includes relevant documents in the application field of hybrid retrieval services, which are used to build a knowledge base. The knowledge base index is built by slicing and vectorizing the original documents.

[0295] Knowledge base index 833 includes dense retrieval built on dense vector indexes based on vector databases (such as Milvus) and sparse retrieval built on inverted indexes based on Elasticsearch (ES).

[0296] In this specific implementation, training-free system evolution is achieved. No LLM parameters need to be updated; retrieval performance can be improved simply by maintaining the knowledge base and Prompt contextual hints, significantly lowering the computational barrier and optimization costs. Furthermore, by combining a DeepSearch agent with a hybrid retrieval service, the challenges of decomposing complex user queries and multi-hop queries are solved, significantly improving precision and achieving deep reasoning and accurate retrieval. Moreover, a pioneering knowledge base and experience base purification mechanism based on reasoning feedback is introduced, automatically identifying and removing "toxic" data, including the first knowledge slice in the knowledge base that causes reasoning trajectory failure and the second retrieval strategy used by multiple failed reasoning trajectories in the experience base. This solves the performance degradation problem of the RAG system after long-term operation and achieves self-optimization capabilities for the knowledge base and experience base.

[0297] Example 2

[0298] Corresponding to the aforementioned retrieval method embodiments, this disclosure also provides embodiments of a retrieval system.

[0299] Figure 10 This is a schematic diagram of a retrieval system provided in an embodiment of the present disclosure. The retrieval system 20 includes:

[0300] The question retrieval module 201 is used to retrieve user-queried questions;

[0301] The strategy reading module 202 is used to read target retrieval strategies that match the semantics of the user's query question from the local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories;

[0302] The problem decomposition module 203 is used to decompose the user query problem into multiple sub-query problems according to the target retrieval strategy;

[0303] Keyword generation module 204 is used to generate corresponding keywords based on multiple subquery questions;

[0304] The trajectory generation module 205 is used to call the hybrid retrieval service to perform retrieval using the semantic vector and keywords corresponding to the subquery question, so as to generate the corresponding target reasoning trajectory.

[0305] In this embodiment, to address the problem of rigid retrieval strategies that prevent targeted retrieval, this disclosure reads target retrieval strategies that match the user's query from a local experience base, instead of performing static retrieval based on the user's query. This allows the user's query to be broken down into multiple sub-queries based on historical experience of target retrieval strategies in the local experience base, achieving targeted retrieval and solving the problem of breaking down complex problems. Furthermore, based on the semantic vectors corresponding to the sub-queries and the keywords generated for the sub-queries, a hybrid retrieval service is invoked for retrieval, enabling simultaneous vector retrieval and keyword retrieval, significantly improving the precision of the retrieval.

[0306] In one optional implementation, the target retrieval strategy includes the execution order and dependencies between subqueries; such as... Figure 11 As shown, the problem decomposition module 203 includes:

[0307] The intent determination unit 2031 is used to perform semantic analysis on the user's query question to determine the user's query intent.

[0308] Initial partitioning unit 2032 is used to divide the query intent into initial query sub-problems with dependencies;

[0309] Problem sorting unit 2033 is used to sort the initial query subproblems according to the execution order to obtain multiple ordered subquery problems.

[0310] In this implementation, the intent type of a user's query question can be identified based on the query intent, such as comparative analysis or multi-hop reasoning. This allows for the segmentation of the query into dependent initial sub-questions, avoiding invalid searches. For user queries requiring multi-step reasoning, the initial sub-questions are ordered according to their execution sequence, making subsequent search paths more rational.

[0311] In one optional implementation, the retrieval system further includes:

[0312] The relevance detection module is used to respond to situations where the relevance between the knowledge slices contained in the target reasoning trajectory and the user query question is less than the relevance threshold, and then return the steps of decomposing the user query question into multiple sub-query questions according to the target retrieval strategy.

[0313] In this implementation, after one round of retrieval, if the correlation between the knowledge slice contained in the target inference trajectory and the user's query question is less than the correlation threshold, that is, the knowledge slice in the current target inference trajectory cannot answer the user's query question, the process returns to the step of breaking down the user's query question into multiple sub-queries according to the target retrieval strategy. The user's query question is then re-broken to generate new sub-queries and keywords corresponding to the new sub-queries. The hybrid retrieval service is then invoked again for a new round of retrieval to ensure that the retrieved knowledge slices can answer the user's query question.

[0314] In one alternative implementation, such as Figure 12 As shown, the retrieval system 20 also includes:

[0315] The question set acquisition module 206 is used to acquire the query questions corresponding to the user's historical reasoning trajectory of the target type in the historical log generated based on retrieval enhancement, and obtain the question set; wherein, the user's historical reasoning trajectory also includes the query conclusion;

[0316] The reasoning trajectory group acquisition module 207 is used to retrieve each query question in the question set a preset number of times to obtain the reasoning trajectory group corresponding to each query question; wherein, each reasoning trajectory group contains a preset number of reasoning trajectories;

[0317] The success or failure determination module 208 is used to determine the reasoning trajectory corresponding to the target query conclusion as successful and the reasoning trajectory corresponding to other query conclusions as unsuccessful in response to the existence of a number of target query conclusions in the reasoning trajectory group that are greater than the number threshold.

[0318] The retrieval strategy processing module 209 is used to add a first retrieval strategy to the local experience base; the first retrieval strategy is a retrieval strategy generated based on the difference between successful inference trajectories and failed inference trajectories; and to delete a second retrieval strategy from the local experience base; the second retrieval strategy is a retrieval strategy adopted by multiple failed inference trajectories.

[0319] In this implementation, experience is summarized from successful and failed inference trajectories, and the retrieval strategy in the local experience base is dynamically adjusted. The updated local experience base will take effect immediately in the next round of online inference retrieval without any model training, thus realizing a training-free experience purification mechanism and achieving automatic optimization of the retrieval method and automatic updating of the experience base. Specifically, by evaluating multiple inference trajectories in each inference trajectory group, the local experience base can be added, deleted, and modified. As the upper-layer agent performs retrieval, it becomes increasingly intelligent, learning from historical interaction data and achieving self-evolution. This realizes a lightweight experience evolution mechanism that does not require updating model parameters.

[0320] In one optional implementation, the retrieval system further includes:

[0321] The knowledge slice processing module is used to delete the first knowledge slice in the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is the knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group;

[0322] The knowledge slice processing module is also used to increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice of multiple successful reasoning trajectories in the dominant reasoning trajectory group.

[0323] In this embodiment, by downgrading or soft-deleting the first knowledge slice that caused the reasoning trajectory to fail in the knowledge base, and upgrading the second knowledge slice that led multiple successful reasoning trajectories, automatic noise reduction of the knowledge base is achieved. This dynamic purification of the knowledge base during the retrieval process avoids the generation of a large number of redundant, outdated, or even conflicting knowledge slices in actual operation. The knowledge base processing method in this embodiment solves the problem of the traditional knowledge base maintenance relying entirely on manual review and being inefficient. It adopts an automated attribution mechanism to identify which knowledge slices played a key role in successful historical reasoning trajectories and which knowledge slices led to erroneous query conclusions in failed reasoning trajectories. Based on the automated attribution mechanism, the first and second knowledge slices are processed, avoiding the continuous decline of the signal-to-noise ratio during the retrieval process and improving the quality of the generated query conclusions corresponding to the user's query question.

[0324] In one alternative implementation, such as Figure 13 As shown, the knowledge slicing processing module 210 specifically includes:

[0325] The identifier acquisition unit 2101 is used to acquire the globally unique identifier of the first knowledge slice in the knowledge base;

[0326] The status modification unit 2102 is used to modify the metadata status bit corresponding to the global unique identifier to an interceptable state, wherein the interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service;

[0327] The deletion unit 2103 is used to delete the first knowledge slice in the knowledge base in response to the first knowledge slice being in the interceptable state for a period of time exceeding a preset time.

[0328] In this embodiment, a lightweight soft isolation mechanism based on metadata status bits is used for the deletion operation of the first knowledge slice. This avoids the huge system overhead caused by directly physically deleting or rebuilding the underlying knowledge base index, which may cause the online hybrid retrieval service to experience lockouts.

[0329] In one alternative implementation, such as Figure 14 As shown, the knowledge slicing processing module 210 specifically includes:

[0330] The weight acquisition unit 2104 is used to acquire the current dynamic weight score in the metadata corresponding to the second knowledge slice;

[0331] The incremental calculation unit 2105 is used to calculate the dynamic weight increment of the second knowledge slice based on the number of target successful reasoning trajectories containing the second knowledge slice and / or the confidence score of the target successful reasoning trajectories.

[0332] The weight update unit 2106 is used to update the dynamic weight score of the second knowledge slice based on the current dynamic weight score and the dynamic weight increment.

[0333] In this embodiment, when it is determined that a certain second knowledge slice has repeatedly led to the correct reasoning conclusion and obtained a successful reasoning trajectory, the dynamic weight increment of the second knowledge slice can be calculated to realize the dynamic weight update of the second knowledge slice.

[0334] In one alternative implementation, such as Figure 15 As shown, the trajectory generation module 205 specifically includes:

[0335] The dual-path recall unit 2051 is used to concurrently call the inverted index corresponding to keyword retrieval and the vector index corresponding to vector retrieval to obtain the initial two-path Top-N knowledge slice lists respectively;

[0336] Metadata reading unit 2052 is used to simultaneously read the metadata corresponding to each knowledge slice while recalling two Top-N knowledge slices, and extract the boost_score field from it;

[0337] The dynamic scoring and fusion unit 2053 is used to use the extracted boost_score as an adjustment coefficient, combined with the ranking of each knowledge slice in each retrieval path, to calculate the comprehensive weight score of each slice using a weighted inverse ranking fusion algorithm, and to reorder the candidate knowledge slices according to the comprehensive weight score to generate the target reasoning trajectory, further improving the refinement of hybrid retrieval.

[0338] In one optional implementation, the target reasoning trajectory includes at least two slices of target knowledge; the retrieval system further includes:

[0339] A multi-round retrieval determination module is used to perform a next round of retrieval in response to a conflict between at least two target knowledge slices; wherein the conflict determination method includes at least one of the following:

[0340] Calculate the conflict confidence between any two target knowledge slices. If the conflict confidence exceeds the corresponding conflict confidence threshold, determine that there is a conflict between at least two target knowledge slices.

[0341] In response to the detection of contradictory numerical values ​​contained in keywords across multiple knowledge slices, it is determined that there is a conflict in at least two target knowledge slices.

[0342] In this embodiment, these target knowledge slices are judged. If a conflict is detected in at least two target knowledge slices, the next round of retrieval is performed. This avoids the conflict in the target knowledge slices from affecting the quality of the query results for the user's query question and avoids generating incorrect query results.

[0343] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs.

[0344] Example 3

[0345] Figure 16 This is a schematic diagram of the structure of an electronic device according to an example embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor. When the processor executes the computer program, it implements the retrieval method described in Embodiment 1 above. Figure 16 The electronic device 30 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0346] like Figure 16 As shown, the electronic device 30 can be manifested as a general-purpose computing device, such as a server device. The components of the electronic device 30 may include, but are not limited to: at least one processor 31, at least one memory 32, and a bus 33 connecting different system components (including memory 32 and processor 31).

[0347] Bus 33 includes a data bus, an address bus, and a control bus.

[0348] The memory 32 may include volatile memory, such as random access memory (RAM) 321 and / or cache memory 322, and may further include read-only memory (ROM) 323.

[0349] The memory 32 may also include a program tool 325 (or utility) having a set (at least one) program module 324, such program module 324 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0350] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the retrieval method provided in Embodiment 1 above.

[0351] Electronic device 30 can also communicate with one or more external devices 34 (e.g., keyboard, pointing device, etc.). This communication can be performed via input / output (I / O) interface 35. Furthermore, electronic device 30 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 36. Figure 16 As shown, network adapter 36 communicates with other modules of electronic device 30 via bus 33. It should be understood that, although... Figure 16 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems.

[0352] It should be noted that although several units / modules or sub-units / modules of the electronic device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0353] Example 4

[0354] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the retrieval method provided in Embodiment 1 above.

[0355] The readable storage medium may be more specifically adopted, including but not limited to: portable disk, hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0356] Example 5

[0357] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the retrieval method described in Embodiment 1 above.

[0358] The program code for executing the computer program product of this disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on a user device, partially on a user device, as a stand-alone software package, partially on a user device and partially on a remote device, or entirely on a remote device.

[0359] While specific embodiments of this disclosure have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of this disclosure is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of this disclosure, but all such changes and modifications fall within the scope of protection of this disclosure.

Claims

1. A retrieval method, characterized in that, The retrieval method includes the following steps: Get the user's query question; Read target retrieval strategies that match the semantics of the user's query from a local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories; The user query question is broken down into multiple sub-query questions according to the target retrieval strategy; Generate corresponding keywords based on the multiple sub-queries; Using the semantic vector corresponding to the subquery question and the keywords, a hybrid retrieval service is invoked to perform a retrieval, thereby generating the corresponding target reasoning trajectory; The retrieval method further includes: Obtain the query questions corresponding to the user's historical reasoning trajectory for the target type in the historical logs generated based on retrieval enhancement, and obtain a question set; wherein, the user's historical reasoning trajectory also includes the query conclusion; For each query question in the question set, a preset number of searches are performed to obtain a group of inference trajectories corresponding to each query question; wherein, each group of inference trajectories contains the inference trajectories of the preset number of searches. In response to the presence of a number of target query conclusions in the inference trajectory group that exceed a number threshold, the inference trajectory corresponding to the target query conclusion is determined to be successful, and the inference trajectories corresponding to other query conclusions are determined to be unsuccessful. The retrieval method further includes: Delete the first knowledge slice from the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is a knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group; specifically including: Obtain the globally unique identifier of the first knowledge slice in the knowledge base; Modify the metadata status bit corresponding to the globally unique identifier to an interceptable state. The interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service. If the duration during which the first knowledge slice is in the interceptable state exceeds a preset duration, the first knowledge slice is deleted from the knowledge base. And / or, Increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice that dominates multiple successful reasoning trajectories in the reasoning trajectory group; specifically including: Obtain the current dynamic weight score from the metadata corresponding to the second knowledge slice; The dynamic weight increment of the second knowledge slice is calculated based on the number of successful inference trajectories of the target containing the second knowledge slice and / or the confidence score of the successful inference trajectories of the target. The dynamic weight score of the second knowledge slice is updated based on the current dynamic weight score and the dynamic weight increment.

2. The retrieval method as described in claim 1, characterized in that, The target retrieval strategy includes the execution order and dependencies between subqueries; The user query question is broken down into multiple sub-query questions according to the target retrieval strategy, including: Perform semantic analysis on the user query to determine the user's query intent; The query intent is divided into initial query sub-problems with the aforementioned dependencies; The initial query subproblems are sorted according to the execution order to obtain multiple ordered subquery subproblems.

3. The retrieval method as described in claim 1, characterized in that, The retrieval method further includes: If the correlation between the knowledge slice contained in the target inference trajectory and the user query question is less than the correlation threshold, the step of decomposing the user query question into multiple sub-query questions according to the target retrieval strategy is returned.

4. The retrieval method as described in claim 1, characterized in that, The retrieval method further includes: The first retrieval strategy is added to the local experience base; the first retrieval strategy is a retrieval strategy generated based on the difference between successful and unsuccessful inference trajectories. And / or, Delete the second retrieval strategy from the local experience base; the second retrieval strategy is the retrieval strategy used by multiple failed inference trajectories.

5. The retrieval method according to any one of claims 1-4, characterized in that, The target reasoning trajectory contains at least two slices of target knowledge. After the step of calling the hybrid retrieval service to perform a retrieval using the semantic vector corresponding to the subquery question and the keywords, the method further includes: In response to a conflict existing in at least two target knowledge slices, a next round of retrieval is performed; wherein the conflict determination method includes at least one of the following: Calculate the conflict confidence of any two target knowledge slices, and in response to the conflict confidence exceeding the corresponding conflict confidence threshold, determine that there is a conflict in the at least two target knowledge slices; In response to the detection of contradictory numerical values ​​contained in keywords across multiple knowledge slices, it is determined that a conflict exists between at least two target knowledge slices.

6. A retrieval system, characterized in that, The retrieval system includes: The question retrieval module is used to retrieve user-queried questions. The strategy reading module is used to read target retrieval strategies that match the semantics of the user's query question from the local experience base; wherein, the local experience base includes local retrieval strategies extracted from local historical reasoning trajectories; The problem decomposition module is used to decompose the user query problem into multiple sub-query problems according to the target retrieval strategy; The keyword generation module is used to generate corresponding keywords based on the multiple sub-queries. The trajectory generation module is used to utilize the semantic vector corresponding to the subquery question and the keywords to call the hybrid retrieval service to perform retrieval, so as to generate the corresponding target reasoning trajectory; The question set acquisition module is used to acquire the query questions corresponding to the user's historical reasoning trajectory of the target type in the historical log generated based on retrieval enhancement, and obtain the question set; wherein, the user's historical reasoning trajectory also includes the query conclusion; The reasoning trajectory group acquisition module is used to retrieve each query question in the question set a preset number of times to obtain a reasoning trajectory group corresponding to each query question; wherein, each reasoning trajectory group contains the reasoning trajectory of the preset number of times; The success / failure determination module is used to determine the reasoning trajectory corresponding to the target query conclusion as successful and the reasoning trajectory corresponding to other query conclusions as unsuccessful when there is a number of target query conclusions in the reasoning trajectory group that is greater than the number threshold. The retrieval system also includes: The knowledge slice processing module is used to delete the first knowledge slice in the knowledge base corresponding to the hybrid retrieval service; the first knowledge slice is a knowledge slice that appears in multiple failed reasoning trajectories in the reasoning trajectory group; The knowledge slicing processing module specifically includes: An identifier acquisition unit is used to acquire a globally unique identifier for the first knowledge slice in the knowledge base. A status modification unit is used to modify the metadata status bit corresponding to the globally unique identifier to an interceptable state, wherein the interceptable state is used to indicate that the first knowledge slice will be intercepted when calling the next round of hybrid retrieval service; The deletion unit is configured to delete the first knowledge slice from the knowledge base in response to the first knowledge slice being in the interceptable state for a period of time exceeding a preset time. And / or, The retrieval system also includes: The knowledge slice processing module is used to increase the weight of the second knowledge slice in the knowledge base; the second knowledge slice is the knowledge slice that dominates multiple successful reasoning trajectories in the reasoning trajectory group. The knowledge slicing processing module specifically includes: The weight acquisition unit is used to acquire the current dynamic weight score in the metadata corresponding to the second knowledge slice; An incremental calculation unit is used to calculate the dynamic weight increment of the second knowledge slice based on the number of target successful inference trajectories containing the second knowledge slice and / or the confidence score of the target successful inference trajectories. The weight update unit is used to update the dynamic weight score of the second knowledge slice based on the current dynamic weight score and the dynamic weight increment.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and for running on the processor, characterized in that, When the processor executes the computer program, it implements the retrieval method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the retrieval method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the retrieval method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Automatic construction method of end-to-end agent based on graph structure semantic fusion

    CN120235181A

  • Question and answer generation method and device combined with knowledge search, medium, equipment and product

    CN121480711A