Intelligent Decision-making Method, Computer Device, and Computer-readable Storage Medium
The integration of tree-of-thoughts and ensemble learning enhances large language models' performance in complex scenarios by exploring multiple paths and dynamically refining answers, addressing limitations in coherence and logical reasoning.
Patent Information
- Application Number
- CN202510294959.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing large-scale language models are insufficient when facing complex, multi-step or multi-angle tasks, and it is difficult to conduct long-term dialogue coherence, deep logical reasoning and multi-perspective analysis. The multi-model fusion method lacks dynamic iterative optimization and pruning mechanisms, resulting in unstable results.
A multi-round reasoning and fusion framework combining tree thinking (ToT) and Stacking integrated learning method is adopted. Through the agent, the thinking path is split, the local solution model is used to solve in parallel and perform first-order fusion, and the second-order fusion is carried out in combination with the global meta learner to realize dynamic iterative optimization of multipaths and the optimal output of answers.
It improves the adaptability and robustness of large language models in complex multi-step tasks, can generate more comprehensive and accurate answers, adapt to the needs of complex scenarios, reduce waste of computing resources, and improve inference efficiency.
Smart Images

Figure CN119808966B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of natural language processing, and particularly relates to an intelligent decision-making method, a computer device, and a computer-readable storage medium. Background Art
[0002] In recent years, artificial intelligence technology has developed by leaps and bounds. Especially driven by deep learning and neural networks, multiple core fields have achieved unprecedented breakthroughs. In the field of natural language processing (NLP), the emergence of large language models (LLMs) has set off a technological revolution. There are already a large number of various civilian and commercial language models on the market. Through training on a vast amount of corpus, they demonstrate excellent language understanding and generation capabilities, and are widely used in tasks such as text generation, dialogue systems, machine translation, question-and-answer retrieval, etc. These models not only significantly improve the accuracy of tasks and the fluency of generated content, but also give rise to new application scenarios, such as intelligent customer service, writing assistants, educational platforms, etc.
[0003] However, as these models are applied to more complex and dynamic scenarios, the ability bottleneck of a single model gradually emerges. For example, traditional dialogue models often perform inadequately in scenarios that require long-term dialogue coherence or in-depth logical reasoning; in the context of multi-step task planning or situations that require multi-faceted perspective analysis, a single model is prone to being limited to a certain fixed thinking pattern. Therefore, how to improve the response performance of language models is a technical problem that needs to be urgently solved by those skilled in the art.
[0004] The foregoing description is to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] Based on this, in view of the above problems, an intelligent decision-making method, a computer device, and a computer-readable storage medium are proposed, which can make decisions comprehensively and fully according to requests.
[0006] This application solves its technical problems by adopting the following technical solutions:
[0007] This application provides an intelligent decision-making method, including the following steps: when an intelligent agent obtains request content, the intelligent agent performs an analysis operation on the request content to split out at least one thinking path; uses each thinking path as the input of a local solution model, and solves the thinking path according to the local solution model to obtain corresponding candidate answers; aggregates the candidate answers corresponding to all thinking paths to obtain a candidate answer set; uses the candidate answer set as the input of a global meta-learner, and after performing a fusion processing operation on the candidate answer set according to the global meta-learner, obtains a final answer and outputs it.
[0008] In an alternative embodiment of the present application, the agent performs an analysis operation on the request content, including: determining the content attribute of the request content, where the content attribute includes at least one of text, picture, video, and audio; determining a feature extraction tool according to the content attribute, and using the feature extraction tool to extract key features from the request content; parsing the key features to determine the core intention of the request content; determining at least one reply sub-thought according to the key features and the core intention; and combining the reply sub-thoughts to obtain a thinking path.
[0009] In an alternative embodiment of the present application, the local solution model is composed of multiple solution sub-models; solving the thinking path according to the local solution model includes: the solution sub-model splitting the thinking path into at least one sub-thinking path, where the sub-thinking path includes multiple sub-questions; arranging the sub-questions in the same sub-thinking path in order, and determining the corresponding solution rounds for the sub-questions between different sub-thinking paths according to the arrangement order; the solution sub-model parallelly solving the sub-questions corresponding to the same solution round between multiple sub-thinking paths to obtain the corresponding sub-reply content; summarizing all the sub-reply contents under the same sub-thinking path to obtain the sub-answer of the sub-thinking path; and performing a first-order fusion process on the sub-answers between all the sub-thinking paths to obtain a candidate answer for the thinking path.
[0010] In an alternative embodiment of the present application, the solution sub-model parallelly solves the sub-questions corresponding to the same solution round between multiple sub-thinking paths to obtain the corresponding sub-reply content, including: obtaining the core intention of the request content determined after the analysis operation, and determining a scoring calculation template for the sub-thinking path according to the core intention; after each round of solution, calculating the evaluation score of the sub-reply content according to the scoring calculation template; if the evaluation score is higher than the deep thinking threshold, generating a new sub-question according to the sub-reply content to be added to the parallel solution in the next round; and / or, after each round of solution, calculating the correlation between the sub-reply content of the current round and the sub-questions of the next round; if the correlation is lower than the preset correlation threshold, generating a new sub-question according to the sub-reply content to be added to the parallel solution in the next round.
[0011] In an alternative embodiment of the present application, the solution sub-model solves the sub-problems corresponding to the same solution round among multiple sub-thinking paths in parallel to obtain the corresponding sub-response content, including: obtaining a scoring calculation template; after each round of solution, calculating the evaluation scores of all sub-response contents according to the scoring calculation template; if there is a sub-response content whose evaluation score is lower than the retention threshold, marking the sub-problem of the next round corresponding to the sub-response content as a problem to be deleted; and / or, after each round of solution, judging the number of sub-problems to be solved in parallel in the next round; when the number of sub-problems is greater than the preset upper limit, calculating the evaluation scores of all sub-response contents in the current round according to the scoring calculation template; retaining the top N sub-response contents with the evaluation scores, and marking the sub-problems of the next round corresponding to the un-retained sub-response contents as problems to be deleted; N is a preset integer, and N is less than or equal to the preset upper limit; deleting the sub-thinking path with the problem to be deleted as the root.
[0012] In an alternative embodiment of the present application, performing a first-order fusion process on the sub-answers among all sub-thinking paths includes: obtaining all sub-response contents that make up each sub-answer, as well as the evaluation scores of the sub-response contents; extracting the key features in the sub-response contents, where the key features are used to characterize the core elements within the sub-response; performing weighted scoring according to the evaluation scores, and selecting the key features with the highest scores for combination according to the scoring results to construct a candidate answer; the candidate answer is composed of multiple key features, and the sum of the evaluation scores of the key features that make up the candidate answer is greater than or equal to the sum of the evaluation scores of all sub-response contents under any one sub-thinking path.
[0013] In an alternative embodiment of the present application, after performing a fusion process operation on the candidate answer set according to the global meta-learner, it includes: obtaining the feature information of each candidate answer in the candidate answer set, and using the feature information as the input of the global meta-learner for processing, where the global meta-learner is one or more constituent models among a linear model, a tree model, a neural network, and / or a large language model; the global meta-learner trains and / or infers the feature information to perform weighted fusion on all the feature information, and the processing output is the answer to be output, and the answer to be output is composed of a combination of feature information; judging whether the answer to be output meets the output conditions; if not, using the answer to be output and the request content as the input of the agent for iteration again; if so, marking the answer to be output as the final answer.
[0014] In an alternative embodiment of the present application, when the global meta-learner trains and / or infers the feature information, it includes: retrieving and obtaining additional feature information according to the request content, where the additional feature information is external information related to the request content in the existing database; using the additional feature information and the feature information as the input of the global meta-learner for training and / or inference.
[0015] The present application also provides a computer device, including a processor and a memory: The processor is configured to execute a computer program stored in the memory to implement the method as described above.
[0016] The present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method as described above.
[0017] Adopting the embodiments of the present application has the following beneficial effects:
[0018] The present application can use an agent to obtain the request content, convert it into multiple thinking paths, and hand them over to the local solution model for parallel solution to obtain a diverse set of candidate answers; then use the set of candidate answers as the input of the global meta-learner for global learning, so as to analyze the advantages and disadvantages of the answers under multiple thinking paths, and finally output the optimal and most expected final answer after comprehensive consideration.
[0019] The above description is only an overview of the technical solutions of the present application. In order to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, details are described in detail. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0021] Figure 1 It is a schematic flowchart of an intelligent decision-making method provided by an embodiment.
[0022] Figure 2 It is a schematic diagram of the thinking path of chain thinking provided by an embodiment.
[0023] Figure 3 It is a schematic diagram of the thinking path of tree thinking provided by an embodiment.
[0024] Figure 4 It is a schematic diagram of expanding the thinking depth of tree thinking provided by an embodiment.
[0025] Figure 5 It is a schematic flowchart of a second-order fusion processing provided by an embodiment.
[0026] Figure 6 Schematic diagram of the data processing and transfer path of the intelligent decision-making method provided for an embodiment.
[0027] Figure 7 Schematic block diagram of the structure of a computer device provided for an embodiment. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0029] In the field of natural language processing, the choice of reasoning method directly affects the depth and breadth of the model to solve problems. The traditional Chain of Thought (CoT) reasoning method is a linear process. The model usually gradually unfolds the reasoning centered around a single path, attempting to obtain the final result through staged logical deductions. This method performs well in simple tasks, but has limitations when facing complex problems.
[0030] For example, in scenarios that require multi-angle analysis, the Chain of Thought often has difficulty considering the potential solutions of different paths simultaneously. In tasks with multi-step iterations, once an error occurs or key information is missed in a certain step of reasoning, it is often difficult to correct in subsequent steps, resulting in the quality of the final result being affected. In addition, this linear reasoning method lacks the ability to dynamically adjust and explore diverse solutions. The current common technologies and problems are as follows: 1. Limitations of single-path reasoning; Once a deviation occurs in a certain step of the traditional Chain of Thought or single-model prediction, subsequent steps will be severely affected, making it difficult to correct in a timely manner, and also lacking the ability to explore multiple solutions in parallel. 2. Simple voting or weighting is insufficient to handle high-complexity scenarios; Some systems use simple voting or average weighting methods to fuse the results of multiple models. However, they ignore the potential differences and context features between model outputs, resulting in limited fusion effects. 3. Lack of multi-round iterative evaluation and pruning mechanisms; Currently, many multi-model or multi-path fusion methods terminate after initial fusion, lacking a system for iterative optimization of errors and potential improvement spaces in multiple rounds, and unable to fully explore the solution space of complex problems.
[0031] To overcome the above technical problems, the present application proposes an intelligent decision-making method. To clearly describe the method provided in this embodiment, please refer to Figures 1 - 6 , which includes steps S110 to S130.
[0032] Different from the existing technology that uses the reasoning method of CoT, this application combines the multi-path parallel reasoning idea of the Tree of Thoughts (ToT) with ensemble learning methods such as Stacking to form a multi-round reasoning and fusion framework. By performing preliminary fusion ("first-order fusion") within each thought path and using a global meta-learner ("second-order fusion") at the global level for in-depth integration, the adaptability and robustness of large language models in complex multi-step tasks can be significantly improved. At the same time, combined with heuristic scoring and pruning strategies, it can continuously optimize and screen out redundancies in multiple rounds of iteration, so as to obtain a more comprehensive and accurate final answer while effectively controlling the computational cost.
[0033] Step S110: When the agent obtains the request content, the agent performs an analysis operation on the request content to split out at least one thought path.
[0034] In one embodiment, step S110: The agent performs an analysis operation on the request content, including: determining the content attribute of the request content, where the content attribute includes at least one of text, picture, video, and audio; determining a feature extraction tool according to the content attribute, and using the feature extraction tool to extract key features from the request content; parsing the key features to determine the core intention of the request content; determining at least one reply sub-idea according to the key features and the core intention; and combining the reply sub-ideas to obtain the thought path.
[0035] In one embodiment, in this application, an agent is used as the input interface of the entire system. The agent is different from the input of traditional large language models. The agent can sense the physical / digital environment (such as cameras, APIs) through sensors or interfaces, while traditional large language models can only process text inputs. That is to say, it can be preliminarily understood that the agent is a large model equipped with parsing tools (sensors or interfaces). Therefore, the request content that the agent can accept is no longer limited to text types, and can include various types of content such as text, pictures, videos, and audio. When the agent analyzes the request content, it needs to adapt the corresponding tools for parsing according to the content attribute of the request content. Therefore, after the agent obtains the request content, it can determine the content attribute of the request content. The content attribute is used to characterize what type of data the request content belongs to, and can specifically include one or more of text, pictures, videos, and audio.
[0036] Further, according to the content attribute, the corresponding feature extraction tool is adapted. Different feature extraction tools are applicable to different content attributes, and are used to extract key features from the request content, and determine the core intention of the request content by parsing the key features. Taking text content as an example, the feature extraction tool can specifically be a Bag of Words (BoW), Word Embedding, or a pre-trained language model (BERT, GPT), etc. Taking the Bag of Words as an example specifically, for the input text content, the keywords can be extracted, the word frequencies in the text can be counted, and the grammar and order can be ignored. Thus, the theme of the text content can be judged, that is, the core intention of the request content is determined. For example, if words such as "preferential" and "discount" appear frequently in the text, it points to a promotion intention. Similarly, for image content, the feature extraction tools include deep learning models, multimodal models (CLIP), etc.; for video content, the feature extraction tools include key frame extraction, spatio-temporal feature extraction, etc.; for audio content, the feature extraction tools include acoustic feature extraction, automatic speech recognition (ASR), etc. Moreover, the request content may be a combination of the above content attributes. When analyzing, multiple feature extraction tools can be combined for feature fusion. For example, for multimodal content, the features need to be aligned and fused (such as the vision + audio of a video), and then tools such as SVM, random forest, or neural network are used to classify the core intention of the request content. Thus, the core intention represented by the request content and the extracted key features form small problems that can be solved, that is, at least one reply sub-thought is determined.
[0037] Take a specific example. Suppose the content input by the user is "How to get from Sichuan to Shanghai". The key features are the two geographical locations of "Sichuan" and "Shanghai", and the core intention is "arrival". For this purpose, the means of transportation for arrival can be selected, including but not limited to private cars, tour buses, high-speed trains, airplanes, ferries, etc.; the provinces that can be passed through between Sichuan and Shanghai can also be determined, and the ways of arrival are divergent. One can go up through Shaanxi, go down through Guizhou, or go directly through Chongqing in the shortest way, and then arrive in Shanghai one by one. Suppose the means of transportation selected is a car and the location selected is to go to Chongqing first, then the first step of how to drive a car from Sichuan to Chongqing becomes the first reply sub-thought. With this as a reference, multiple consecutive and combinable reply sub-thoughts can be specifically determined. The reply sub-thoughts are combined to obtain a thinking path, and the combination process needs to meet the requirements that the solution to the problem is continuous and feasible. For example, if the first reply sub-thought is how to drive from Sichuan to Chongqing, then the next thought should be how to start from Chongqing and reach the next sub-goal point. Instead of discontinuously choosing how to start from Guizhou to reach the next sub-goal point. Therefore, this series of coherent thinking steps is the thinking path.
[0038] And it is worth noting that in the face of the above problems, the traditional chain - of - thought reasoning method will center around a single path to obtain the answer. Usually, the following problems exist: 1. Local optimum problem: In multi - step reasoning, once a deviation occurs or key information is missed in a certain step of reasoning, it is difficult to correct in subsequent steps, resulting in the final result deviating from the actual goal. 2. Lack of global vision: The single - path method usually can only unfold along the preset thinking, unable to explore solutions from multiple angles, multiple hypotheses, or multiple strategies in parallel, thus limiting the diversity and comprehensiveness of the reasoning results. 3. Low adaptability: Facing complex tasks that require dynamic adjustment or multi - dimensional analysis, the single - path reasoning method lacks flexibility and is difficult to meet the requirements of complex scenarios.
[0039] Therefore, this application generates thinking paths in a tree - of - thought (ToT) manner. It conducts multi - path parallel exploration of the problem, where each path represents a different hypothesis or thinking angle, overcoming the limitations of single - path reasoning. Inside the path (i.e., the "first - order fusion" stage), it preliminarily screens and merges the outputs of multiple models, which can not only retain multiple ideas but also prevent the explosive growth of the number of candidates within the path, thus achieving richer diverse thinking.
[0040] That is, the thinking paths generated by the agent, in a preferred embodiment, are not just one, but multiple paths are generated. Each path may explore the answer through different logical orders, different hypotheses, or different Prompts. In subsequent stages, each path may further branch, forming a deeper tree - like structure. To illustrate the difference, take the previously mentioned example of how to get from Sichuan to Shanghai to show Figure 2 and Figure 3 the shown chain of thought, where Figure 2 is the thinking path constituted by the chain - of - thought represented by the prior art; while Figure 3 , is the thinking path under the tree - of - thought of this application. Specifically, the whole chain connecting from the Sichuan square to the Shanghai square from start to finish is also a thinking path. It should be understood that Figure 2 、 Figure 3These are all simple examples of thinking paths. In reality, when faced with complex problems, the thinking paths will only be more complex, and there will also be correlations between the reply sub-thinking paths, which will be specifically reflected later in this text. This application adopts a tree-like thinking method, so the implementation process also follows the common process of tree-like thinking: 1. Problem decomposition: Decompose complex problems into several operable sub-problems or steps. 2. Path generation: Based on the language model, for each sub-problem or each stage, generate multiple possible answers or intermediate reasoning results to form different "thinking branches". 3. Path evaluation: Introduce heuristic or evaluation indicators to score the rationality, feasibility or scores of each path, and retain the optimal / most potential path. 4. Expand and repeat: At the new stage or level, use the retained path to continue generating subsequent thinking branches and repeat the process of evaluation and screening. 5. Integrate the answer: At the final stage, summarize the path information obtained through evaluation and expansion to form the final answer. That is to say, 1. Turn the problem into a tree: Each step can give rise to different thinking branches (such as different writing directions, different problem-solving ideas, etc.). 2. Explore multiple possibilities in parallel: Consider multiple alternative solutions simultaneously instead of just following one path. 3. Continuously screen and merge: Regularly evaluate which thinking path is good and which is not suitable, retain the high-quality branches and continue to deepen, and finally integrate into an optimal answer or idea. The advantages of doing this are: more flexible, more comprehensive, reducing the risk of "going down a dead end" single-mindedly, and making thinking become as lush as a "tree". Subsequently, how to implement each step will be revealed separately.
[0041] Step S120: Use each thinking path as the input of the local solution model, and solve the thinking path according to the local solution model to obtain the corresponding candidate answers; aggregate the candidate answers corresponding to all thinking paths to obtain a candidate answer set.
[0042] In one implementation, the local solution model consists of multiple solution sub-models. The solution sub-models can be different types of solution models from each other, or solution models of the same type but with different parameter configurations. Specifically, the local solution model can be matched according to the core intention of the request content. For example, for the request content of traveling from Sichuan to Shanghai exemplified above, the local solution model can be a model related to navigation; similarly, if it is asking about the weather, it can be a weather-related model. And the local solution model can also be composed of multiple sub-models. For example, if it is solved using three different means of transportation, namely cars, trains, and airplanes, they are different sub-models. Taking a car as an example for solution models of the same type but with different parameter configurations, it can be a private car, a bus, or a motorcycle; the road selection can be a national road or a highway, etc., which are also different parameter configurations.
[0043] It should also be noted that although the local solution model provided in this application is composed of multiple solution sub-models, it is different from the simple fusion of existing multiple models. In the prior art, multi-model integration methods often adopt simple voting or weighted average methods, which have the following problems: Insufficient information utilization: There is a lack of in-depth analysis of the characteristics and differences of the outputs of different models, and the potential of multi-model collaboration cannot be tapped. Insufficient robustness: The simple fusion mechanism cannot effectively handle high-complexity tasks and is easily affected by low-quality models or path outputs, resulting in unstable results. Lack of dynamic feedback mechanism: Most methods output results after preliminary fusion, lacking a multi-round optimization strategy based on evaluation and feedback, and it is difficult to continuously improve the results in complex tasks. This application will score and screen the sub-response content obtained after solution, that is, within the sub-thinking path, perform preliminary screening and merging on the outputs of multiple solution sub-models, which can not only retain multiple ideas but also prevent the explosive growth of the number of candidates within the path, thus realizing richer diverse thinking. At the same time, by introducing an evaluation and heuristic pruning strategy after each round of fusion, low-quality or invalid paths are screened out, and at the same time, paths with higher potential are deeply expanded. This "retaining the excellent and discarding the inferior" dynamic iteration mechanism can continuously optimize the thinking tree structure in multiple rounds, improve the inference accuracy and efficiency, and avoid wasting resources on invalid paths. To solve the problem that the inference process lacks means of dynamic iteration and invalid path screening. The above solutions are all included in the processing of step S120, which is collectively called "first-order fusion", and the specific implementation method will be described in detail later.
[0044] In one embodiment, step S120: Solve the thinking path according to the local solution model, including: The solution sub-model splits the thinking path into at least one sub-thinking path, and the sub-thinking path includes multiple sub-questions; The sub-questions within the same sub-thinking path are arranged in order, and the sub-questions between different sub-thinking paths determine the corresponding solution rounds according to the arrangement order; The solution sub-model solves the sub-questions corresponding to the same solution round among multiple sub-thinking paths in parallel to obtain the corresponding sub-response content; Summarize all the sub-response contents under the same sub-thinking path to obtain the sub-answer of the sub-thinking path; Perform first-order fusion processing on the sub-answers between all sub-thinking paths to obtain the candidate answer of the thinking path.
[0045] In one embodiment, taking the previous example as an example, Figure 3The three feasible routes from Sichuan to Shanghai shown are three thinking paths. For each thinking path, the local solution model can be used separately for solution. And according to the description of the local solution model in the previous text, the local solution model consists of multiple sub-models. Each sub-model has its own configuration process. Therefore, when the sub-model solves the thinking path, it will further split into sub-thinking paths. For example, the sub-model is configured to use private cars to take different roads: sub-model A is configured to take the highway, and sub-model B is configured to take the national road. Taking Figure 3 the middle thinking path as an example. Then sub-model A needs to determine how to connect the route from Sichuan → Chongqing → Hubei → Anhui → Jiangsu → Shanghai with the highway; sub-model B determines the situation of connecting this sequence with the national road. Therefore, connecting this path with the highway and connecting this path with the national road further subdivide the thinking path into two sub-thinking paths. At the same time, the sub-thinking path is composed of multiple sub-questions. Taking the sub-thinking path of sub-model A as an example, it can be specifically divided into 5 sub-questions of how to connect Sichuan → Chongqing, Chongqing → Hubei, Hubei → Anhui, Anhui → Jiangsu, and Jiangsu → Shanghai with the highway. These five sub-questions are arranged in order; similarly, for sub-model B, there are also 5 similar sub-questions of how to connect with the national road.
[0046] It can be seen that there is a relevant order among the sub-questions. Therefore, the corresponding solution rounds can be determined according to the arrangement order of the sub-questions. For example, for the sub-question of Sichuan → Chongqing, the sub-questions within sub-model A and the sub-questions within sub-model B can belong to the same solution round. The remaining sub-questions can be arranged according to this rule. It should be noted that the solution round is not forcibly bound to the corresponding order, but represents the corresponding thinking depth. Those with the same thinking depth can be called the same solution round. There is a relevant relationship between the rounds, and their levels are connected to each other. Therefore, during the solution process, the sub-questions corresponding to the same solution round can be solved in parallel to obtain the corresponding sub-answer content, that is, to solve the problems with the same thinking depth in parallel. Thus, the multi-path parallel exploration characteristic of tree-shaped thinking is utilized, not limited to a single linear reasoning chain, and multiple "thinking paths" are unfolded simultaneously.
[0047] In one embodiment, the solving sub-model solves the sub-problems corresponding to the same solving round among multiple sub-thinking paths in parallel to obtain the corresponding sub-response content, including: obtaining the core intention of the request content determined after the analysis operation, and determining the scoring calculation template for the sub-thinking paths according to the core intention; after each round of solving, calculating the evaluation score of the sub-response content according to the scoring calculation template; if the evaluation score is higher than the deep thinking threshold, generating a new sub-problem according to the sub-response content to be added to the parallel solving in the next round; and / or, after each round of solving, calculating the relevance between the sub-response content of the current round and the sub-problems of the next round; if the relevance is lower than the preset relevance threshold, generating a new sub-problem according to the sub-response content to be added to the parallel solving in the next round.
[0048] In one embodiment, the solving sub-model solves the sub-problems corresponding to the same solving round among multiple sub-thinking paths in parallel to obtain the corresponding sub-response content, including: obtaining the scoring calculation template; after each round of solving, calculating the evaluation scores of all sub-response contents according to the scoring calculation template; if the evaluation score of a sub-response content is lower than the retention threshold, marking the sub-problem of the next round corresponding to the sub-response content as a problem to be deleted; and / or, after each round of solving, judging the number of sub-problems to be solved in parallel in the next round; when the number of sub-problems is greater than the preset upper limit, calculating the evaluation scores of all sub-response contents of the current round according to the scoring calculation template; retaining the top N sub-response contents with the evaluation scores, and marking the sub-problems of the next round corresponding to the un-retained sub-response contents as problems to be deleted; N is a preset integer, and N is less than or equal to the preset upper limit; deleting the sub-thinking path with the problem to be deleted as the root.
[0049] In one embodiment, the specific manner and process of the solving sub-model to specifically obtain the sub-response content belong to the scope of the prior art and need to be determined according to specific problems, so it will not be elaborated here. The scoring and screening need to be specifically determined for the sub-response content. Therefore, the core intention of the request content determined after the analysis operation is obtained before scoring, and the scoring calculation template for the sub-thinking paths is determined according to the core intention. For example, taking the example of going from Sichuan to Shanghai cited above, its core intention is navigation, and its scoring template can be the time spent on the journey or the total toll required, etc. Similarly, for other intentions, the corresponding scoring calculation templates are matched. The scoring calculation template can be a requirement set in advance, for example, for the composition intention, the matching scoring calculation template needs to evaluate the fluency of the writing and the rationality of the word usage, etc.; it can also be a specific standard set or proposed separately by the user. For example, in the same scenario of generating a composition, if the user needs to reduce or increase the rhetoric, this standard can also be included in the scoring calculation template.
[0050] Furthermore, for the specific scoring process, not only can the content itself be processed, but also features can be extracted from the sub-responses (such as: text embedding, confidence of the answer, logical coherence score, output length, etc.). For the extracted features, the evaluation scores of the sub-response content are calculated according to the scoring calculation template. For example, if the question is about solving a math problem: sub-solving model A calculates "10", sub-solving model B calculates "8", and sub-solving model C calculates "12", and their intermediate reasoning embeddings or step features are extracted respectively. Note that it is not their results here, but the process of obtaining the results and some intermediate gain information. The calculation process can use heuristic algorithms to calculate the evaluation scores, or it can be calculated by obtaining the true labels, and there is no restriction on this. For example, assuming a task with a standard answer in the request content (such as having true labels or labeled data), the sub-response content and its features can be directly scored. If there are no labels, scoring can be based on heuristics (such as the matching degree with the question keywords, similarity with the existing knowledge base, logical consistency, etc.). Thus, digital indicators are provided for the subsequent fusion link to measure which candidate is "better" or more "reasonable". If the result is only a "generation" task, self-supervised or rules can also be used to evaluate the answer quality.
[0051] After each round of solving, calculate the evaluation score of the sub-response content according to the scoring calculation template, and determine whether to generate new sub-questions for the sub-response content in the current round based on the evaluation score. For this, reference can be made to Figure 4 the schematic diagram shown. And it is worth noting that the newly generated sub-questions can be not only one, but a series of sub-questions, forming a form similar to a sub-thinking path. That is to say, on the sub-response content with high scores, explore and form a new sub-thinking path to conduct deeper thinking and obtain richer answers.
[0052] As Figure 4 shown, the sub-response content obtained by solving each sub-question in the original sub-thinking path represented by the blue squares has gone through a total of four rounds of solving processes from top to bottom. When the evaluation score obtained by scoring the sub-response content in the first round is higher than the preset deep thinking threshold, it means that the sub-response content obtained in the first round can be further thought-expanded to generate new sub-questions to be solved synchronously in the next round of solving. For example, in Figure 4In this process, two new sub-questions are generated. In the second round of the solution round, in addition to the sub-response content on the original blue sub-thinking path, two additional sub-response contents are obtained. However, they are distinguished by green and red respectively, and it can be seen that there is also sub-response content for the third round after the green sub-response content; while the red sub-response content ends here. The green sub-response content corresponds to an evaluation score that is again higher than the deep thinking threshold, so it is possible to conduct in-depth thinking again and dig out new sub-questions to be added to the parallel solution in the third round. The evaluation score corresponding to the red sub-response content is lower than the deep thinking threshold, which means that there is no longer any value in in-depth thinking under this idea. For this, this response content can be discarded, that is, no new sub-questions will be dug out later.
[0053] And it can be seen that by setting the evaluation score and the deep thinking threshold, taking Figure 4 as an example, there were originally only four sub-response contents represented by blue squares. Due to the expansion of in-depth mining, four additional mined sub-response contents are added, realizing the expansion of the number of sub-response contents, thereby generating a more comprehensive and rich answer.
[0054] However, it should be noted that the method provided in this application realizes the expansion of the thinking breadth by using tree-like thinking. But there are still the following challenges in the solution itself: Redundant invalid paths: A large number of generated thinking paths may contain low-quality or invalid solutions. Without an effective screening mechanism, it will lead to waste of computing resources and degradation of system performance. Path explosion problem: As the number of reasoning steps increases, the number of paths may increase exponentially, and it is difficult for traditional methods to control the computational complexity while ensuring the quality of the results. Insufficient interaction between paths: Each path often explores independently, lacking information sharing and comprehensive analysis across paths, and failing to make full use of the potential complementarity between paths.
[0055] At the same time, existing reasoning and fusion methods usually lack a dynamic iteration and feedback mechanism, manifested as: Lack of iterative optimization: In complex reasoning tasks, the results of the initial reasoning may have room for improvement, but existing technologies lack a flexible multi-round iteration mechanism to optimize the answers. Insufficient pruning strategy: Invalid paths will continue to occupy resources in subsequent reasoning, and no effective pruning of low-quality paths is carried out based on heuristic evaluation. Unclear improvement direction: In multi-round reasoning, there is a lack of targeted strategies to preferentially explore high-potential paths, which easily leads to ineffective exploration iterations.
[0056] Therefore, after each round of solution, the relevance between the sub-answer content of the current round and the sub-questions of the next round can be calculated. Relevance represents the degree of association between sub-answer content or sub-questions. If the relevance is lower than the preset relevance threshold, new sub-questions are generated based on the sub-answer content to be solved in parallel in the next round. On the contrary, if the relevance is higher than the threshold, it indicates that there is a correlation between the sub-answer content or sub-questions. In the process of solution or subsequent fusion, the mutual influence is considered to improve the potential complementarity of the data. Through the above setting method, after each round of fusion, an evaluation and heuristic pruning strategy is introduced to screen out low-quality or invalid paths, and at the same time, paths with higher potential are deeply expanded. Through this "retaining the excellent and discarding the inferior" dynamic iteration mechanism, the structure of the thinking tree can be continuously optimized in multiple rounds, improving the reasoning accuracy and efficiency, and avoiding wasting resources on invalid paths.
[0057] In addition to delving deeply into new sub-questions, it is also possible to determine the quality of each obtained sub-answer content. Similarly, the evaluation scores of all sub-answer contents are calculated according to the scoring calculation template; if there are sub-answer contents with evaluation scores lower than the retention threshold, the sub-questions of the next round corresponding to the sub-answer contents are marked as questions to be deleted for deletion later. This ensures the quality of the finally obtained sub-answer content.
[0058] As mentioned before, the tree-like thinking may have a situation where the number of branches explodes. The explosion in the number will lead to an increase in computing power requirements and the generation of a large number of low-quality answers. Therefore, control is needed from the thinking branches. For this, after each round of solution, the number of sub-questions to be solved in parallel in the next round can be judged. When the number of sub-questions is greater than the preset upper limit, the number of branches needs to be controlled to ensure that the sub-questions solved are all of high quality. And it is known that the sub-questions of each round are relevant to the sub-answer content of the previous round. Therefore, whether to retain the subsequent sub-thinking paths can be determined according to the evaluation scores of the sub-answer content of the previous round. For this, the sub-answer content of the current round can be arranged in ascending order of their evaluation scores, and the sub-questions corresponding to the several sub-answer contents with the lowest scores are deleted. Only the first N sub-answer contents with evaluation scores are retained, that is, only the sub-questions corresponding to the subsequent sub-answer contents of these N sub-answer contents are solved in the next round. This ensures that computing power is not wasted and improves the quantity of answers. N is a preset integer, and N is less than or equal to the preset upper limit. By setting N equal to the preset upper limit, it not only ensures that the number of sub-questions to be solved in parallel in the next round will not exceed the system's computing power limit, but also ensures that there is computing power redundancy in subsequent rounds, so that subsequent rounds can continue to think deeply and explore based on high-scoring sub-answer contents.
[0059] For the specific means of deleting sub - problems, it can be to mark the sub - problems to be deleted as problems to be deleted. Before the next round of solving starts, the subsequent sub - problems with the problem to be deleted as the root, that is, the corresponding sub - thinking path branches, are deleted from the tree - like thinking. That is, the sub - problems after the sub - thinking path where the problem to be deleted is located will no longer be solved. Through dynamic iteration and heuristic pruning to ensure efficiency, low - quality or invalid paths are screened out, and at the same time, paths with higher potential are deeply expanded. This "preserving the excellent and discarding the inferior" dynamic iteration mechanism can continuously optimize the structure of the thinking tree in multiple rounds, improve the reasoning accuracy and efficiency, and avoid wasting resources on invalid paths.
[0060] In one embodiment, a first - order fusion process is performed on the sub - answers between all sub - thinking paths, including: obtaining all sub - reply contents that make up each sub - answer, and the evaluation scores of the sub - reply contents; extracting the key features in the sub - reply contents, where the key features are used to represent the core elements within the sub - reply; performing weighted scoring according to the evaluation scores, and selecting the key features with the highest scores for combination according to the scoring results to construct a candidate answer; the candidate answer is composed of multiple key features, and the sum of the evaluation scores of the key features that make up the candidate answer is greater than or equal to the sum of the evaluation scores of all sub - reply contents under any one sub - thinking path.
[0061] In one embodiment, as described above, a thinking path will be split into multiple sub - thinking paths by the solving sub - model in the local solving model; a series of sub - reply contents will be obtained by solving each sub - thinking path through the corresponding solving sub - model. That is, Figure 3 any one thinking path after being solved by the local solving model will obtain Figure 4 sub - reply contents represented by blue and green squares as shown. However, the sub - reply contents correspond to different sub - thinking paths and cannot be simply integrated to output a logical and output - condition - compliant answer. That is, the method of simply fusing multiple models in the prior art cannot be adopted. For this reason, in this application, the sub - reply contents of multiple sub - thinking paths under one thinking path are deeply fused in layers within the path to significantly improve the adaptability to complex multi - step tasks and the reasoning quality. This process is called the first - order fusion process.
[0062] The first-order fusion process can be handled by a locally configured aggregator in the local solution model. The local aggregator is used to perform preliminary fusion / weighting / voting "inside the thinking path": using the previously obtained evaluation scores, along with the sub-response content itself or features extracted therefrom as key inputs, to synthesize multiple sub-response contents generated by multiple models within the same path, or to select the optimal one. For example, in the previous math problem example, we can obtain the following: Final output of Path 1: After local fusion, we guess the answer = 12, confidence = 0.85; Final output of Path 2: After local fusion, the answer = 8, confidence = 0.65; Path N, etc. Thus, for each thinking path, multiple sub-response contents obtained from the local solution model are fused and output as one answer, which is called the candidate answer corresponding to that thinking path.
[0063] Therefore, in step S120, a tree-like thinking processing method is adopted. After processing through processes such as problem decomposition, thinking expansion, and answer integration, candidate answer generation for multiple thinking paths in the local case is achieved. The following improvements can be realized: 1. Parallel thinking effectively reduces the failure probability caused by taking a "dead end" or being confined to a wrong reasoning chain. 2. Strong scalability, each path can be further refined and corrected in subsequent stages, enabling the model to generate more diverse answer candidates. 3. Adapt to complex tasks, such as strategy planning (game games, planning tasks, etc.), multi-step reasoning (logic problems, mathematical proofs), or creative writing (multiple storylines).
[0064] Step S130: Use the candidate answer set as the input of the global meta-learner. After performing a fusion processing operation on the candidate answer set according to the global meta-learner, the final answer is obtained and output.
[0065] In one implementation, as can be seen from the previous description, after the agent thinking decomposition in step S110 and the local solution and first-order fusion of the local solution model in step S120, a candidate answer set aggregated from candidate answers obtained from multiple thinking paths will be obtained. Each candidate answer in the candidate answer set corresponds to a possible answer under different thinking paths. Each answer has its own advantages and disadvantages, but none of them alone is sufficient to constitute the final answer. Therefore, based on the first-order fusion in step S120 of this application, "second-order fusion" is performed by introducing a global meta-learner (such as Stacking). The global meta-learner deeply explores the complementarity between the outputs and features of each path, and comprehensively learns the difference information between different behaviors / models, thereby achieving a more accurate and robust global optimal fusion.
[0066] For a clear description of the second-order fusion process, please refer to Figure 5 .
[0067] Step S210: Obtain the feature information of each candidate answer in the candidate answer set.
[0068] Step S220: Retrieve and obtain additional feature information according to the request content.
[0069] In one embodiment, in the original Stacking technology processing flow, the global meta-learner processes the prediction results of multiple base learners as input. However, it can be seen that each candidate answer in the candidate answer set is actually composed of comprehensive data such as sub-answer content, the reasoning process of the sub-answer content, and scoring. Therefore, it is necessary to extract the feature information therein to facilitate the processing of the global meta-learner. The feature information includes but is not limited to the evaluation score, reasoning process, confidence, etc. in the first-order fusion stage.
[0070] At the same time, additional feature information can also be retrieved and obtained according to the request content. The additional feature information is external information related to the request content in the existing database, such as search engine results, external knowledge base matching degree, dialogue context information, etc., to enrich the feature dimension. For example, the question is "The humidity is 80% tomorrow. Will it rain?" Among them, "humidity 80%" is the known information in the request content. According to the reasoning process in the first-order fusion stage, the connection between humidity and rainfall can be determined. However, in fact, rainfall is also related to other data. For example, the barometric pressure data for tomorrow is mentioned in the chat context, past weather data is retrieved from the weather database, and various meteorological radar images are retrieved. The information related to rainfall other than the humidity itself is additional feature information. It can be obtained according to the relevance of the obtained results. Therefore, the prediction of a single data aspect is limited. If other relevant data can be additionally obtained and processed together, the output result will be more accurate.
[0071] However, it should be noted that referring to Figure 5 , the step S220 of obtaining additional feature information is actually not necessary, but hidden or optional. Because usually, the quality of the additional feature information obtained according to relevance is difficult to guarantee and may affect the final output answer. Therefore, the solution described in this application is only in the case of a preferred embodiment; in other embodiments, the step S220 may not be included. Correspondingly, no additional feature information will be input in the subsequent step S230.
[0072] Step S230: Use the additional feature information and the feature information as the input of the global meta-learner for training and / or inference.
[0073] In one embodiment, the global meta-learner is composed of one or more of a linear model, a tree model, a neural network, and / or a large language model. The global meta-learner can automatically learn how to weight or combine the results of different paths to obtain a more reliable and comprehensive output. The second-order fusion processing process of the global meta-learner can refer to the fusion processing process of the sub-answer contents of multiple sub-thinking paths in the first-order fusion processing process. The global meta-learner can perform secondary training or inference on the feature information to perform weighted fusion on all the feature information. The finally output fused answer, or an answer distribution (if it is a classification / multiple-choice), is called the answer to be output, and the answer to be output is composed of the combined feature information. The answer to be output may also be accompanied by uncertainty estimation (such as a confidence interval, the subjective probability of the model for its own answer), which helps the subsequent link to judge whether to perform further in-depth exploration.
[0074] Step S240: Obtain the answer to be output processed by the global meta-learner; determine whether the answer to be output meets the output conditions.
[0075] In one embodiment, the output conditions are specifically set according to the specific request content. For example, it can be whether the content quality of the answer to be output reaches the set threshold; or whether sufficient in-depth thinking has been carried out during the first-order fusion and second-order fusion processing processes, and whether sufficient depth has been explored; it can also be whether the answer obtained during the second-order fusion process is globally optimal.
[0076] If not, execute step S250: Use the answer to be output and the request content as the input of the agent. And return to step S110.
[0077] In one embodiment, if the quality of the global fusion result is insufficient or has not reached the set threshold, new expansion will be performed based on the existing optimal path, or a new path will be introduced. This will return to step S110 to regenerate the thinking path and perform iteration again. However, it should be noted that when inputting to the agent again during the iteration process, pruning should be performed according to the sub-answer content or thinking path that determines the reason for not meeting the output conditions during the judgment process, so as to directly discard those paths with too low scores in the next round. Or, if there is still room for improvement in the global result or deeper reasoning is required, the optimal path can be retained or new paths can be re-expanded according to the evaluation & pruning strategy, so as to continue the iteration until a satisfactory answer is finally obtained.
[0078] If it is satisfied, execute step S260: Mark the answer to be output as the final answer and output it.
[0079] Taken together, the present application adopts a hierarchical integration approach in its design: the first-order fusion focuses on the integration of multiple models within a single path; the second-order fusion is responsible for the joint learning of multiple global paths. This architecture can be flexibly embedded with various models and heuristic algorithms, and can also be paired with different types of meta-learners for second-order fusion, featuring good scalability; and it can be combined with the mainstream AI technology stack to achieve unified support for tasks of different scales and complexities.
[0080] For better understanding, the present application presents a specific example to describe the implementation process of the solution provided by the present application. For ease of understanding, reference can be made to Figure 6 . Figure 6 The data processing and flow path in the method of the present application is provided, where the rounded rectangles are the input request content and the final answer of the final output; the squares are the intermediate data during the processing; and the circles are the data models utilized in the processing. And Figure 6 the steps are separated by dashed lines to correspond to the processing procedures of steps S110~S130 in the foregoing text.
[0081] A travel planning scenario example is given below to demonstrate the implementation process of the solution provided by the present application. Assume the user inputs: "I want to travel to Japan for 5 days with a budget of $2,000 and want to visit Tokyo, Kyoto, and Hiroshima. Can you give me a detailed itinerary that includes both popular attractions and some local cultural experiences?" After receiving this message, the intelligent agent can start the thinking splitting of step S110. The intelligent agent disassembles this planning requirement and may first generate several different travel ideas from the perspective of the outline to obtain three thinking paths:
[0082] 1. Thinking path 1: Day 1: Sightseeing in Tokyo; Day 2: Tokyo -> Kyoto (by bullet train); Day 3: Cultural experience around Kyoto; Day 4: Kyoto -> Hiroshima, visit the Atomic Bomb Dome and park; Day 5: Return to Tokyo, end of the itinerary.
[0083] 2. Thinking path 2: Day 1: Main attractions in Tokyo (Asakusa, Akihabara); Day 2: Cultural exploration in Tokyo (kimono experience, tea ceremony); Day 3: Tokyo -> Hiroshima, visit Miyajima; Day 4: Hiroshima -> Kyoto; Day 5: Famous temples and Gion historical district in Kyoto, end of the itinerary.
[0084] 3. Thinking path 3: Day 1: First go to Kyoto to experience traditional culture; Day 2: Kyoto -> Osaka or Kobe for transfer; Day 3: Osaka -> Hiroshima; Day 4: Hiroshima -> Tokyo; Day 5: Free travel in Tokyo, end of the itinerary.
[0085] At this time, for the same question "How to reasonably arrange the itinerary?", the intelligent agent generates several parallel solutions with different start and end sequences, transportation methods, attraction selections, etc.
[0086] Afterwards, each thought path may be handed over to the local solution model for processing to enter the first-order fusion processing process corresponding to step S120. Figure 6 It is not shown in the figure, but according to the previous description, the local solution model actually consists of multiple solution sub-models with different configurations and / or types. Assume that solution sub-model A focuses on comprehensive information; solution sub-model B focuses on cultural experience; solution sub-model C focuses on convenient transportation. Therefore, even for the arrangement under thinking path 1, different sub-thinking paths will be generated and corresponding sub-answer contents will be obtained, which can be specifically:
[0087] Solve sub-model A: It arranges detailed information such as hotel recommendations, how to buy subway cards, and a list of popular restaurants. It also gives a budget estimate: about $1,800 (including transportation and accommodation).
[0088] Solution for sub-model B: The description of traditional cultural activities in Kyoto is strengthened, such as copying sutras, kimono dressing, and wagashi making experience. There is less description of modern cultural experiences in Tokyo. Budget estimate: about $1,900.
[0089] Solving sub-model C: Detailed list of money-saving options for JR Pass, night buses / night trains, and a very tight schedule. Budget estimate: about $1,700, but some cultural activities may be sacrificed.
[0090] Similarly, thinking path 2 and thinking path 3 will also generate multiple sub-plans. The sub-responses can be scored later. Based on the request content, the following features can be extracted from the sub-response content for subsequent scoring: budget (numeric value); number of cultural activities (count); whether popular attractions are included (Boolean value); daily itinerary congestion (score); model confidence, etc.
[0091] Because the request content can be determined to be travel planning, and travel planning has no unique correct solution, so this is a task without a standard answer. For this evaluation score calculation template, you can match it to score each plan based on heuristics, such as: feasibility (whether the transportation connection between cities is reasonable); diversity (whether it takes into account modern cities and traditional culture); budget (whether it is close to US$2,000, not overspending or too wasteful). In this way, you can get the score of the response content of each solution sub-model, for example: the solution sub-sub-scheme of solution sub-model A: comprehensive score = 8.5 / 10; the sub-scheme of solution sub-model B: comprehensive score = 7.5 / 10 (high culturality, but lacks the modern experience of Tokyo); the sub-scheme of solution sub-model C: comprehensive score = 7.0 / 10 (transportation is economical but slightly compact).
[0092] After that, the local aggregator in the local solution model can be used to perform first-order fusion output. Within the local solution model corresponding to Thought Path 1, the three sub-response contents, their features, and scores obtained from the solution sub-models A, B, and C are used as inputs for preliminary weighted or voting fusion to finally merge / concatenate into a hybrid solution, which is called the candidate answer for Thought Path 1: During the day, the cultural experience of solution sub-model A / B is combined with well-known scenic spots; for night transportation, refer to the cost-saving solution of solution sub-model C; the budget is between 1700 and 1900 yuan, and the comprehensive score or confidence level is the highest.
[0093] Output the candidate answer: "After local fusion, the five-day itinerary: The first two days in Tokyo (taking into account popular cultural activities), then take the bullet train to Kyoto and stay for 2 nights to experience tea ceremony & kimono dressing, and on the last day go to Hiroshima to feel the historical culture. The estimated cost is 1800 US dollars." Comprehensive score: 0.85.
[0094] Similarly, Thought Path 2 and Thought Path 3 will also perform the same "first-order fusion" within their respective local solution models to generate their respective candidate answers. For example, the candidate answer A2 for Thought Path 2 has a budget of 1900 US dollars and a comprehensive score of 0.78; the candidate answer A3 for Thought Path 3 has a budget of 1700 US dollars and a comprehensive score of 0.80. These are aggregated into a candidate answer set.
[0095] The candidate answer set will be used as the input for the global meta-learner to enter the second-order fusion process corresponding to step S130. The candidate answer set includes not only the above solutions but also the reasoning process, reasoning scores, and corresponding features. As mentioned before, additional feature information can also be obtained, such as: real-time flight ticket prices / accommodation (possibly connecting to a third-party API); seasonal information (e.g., in August in Japan, it is the peak summer season with a large number of tourists); user personal preferences (such as whether they like historical buildings, need foodie check-ins, etc.). And these additional information are packaged into new features and used as the input for the global meta-learner together with the candidate answer set.
[0096] The global meta-learner (such as linear regression or a small model) performs second-order fusion processing, using the feature information of the candidate answer set and the additional feature information as training data to learn: Among the previous user feedback or example data, which features can best indicate whether a travel itinerary is excellent? How to assign weights to these three solutions or directly select the best one? After calculation or training, the global meta-learner may determine that the score of the "candidate answer for Thought Path 1" is the highest, and the external features (convenient transportation, rich scenic spots) are also very good, so it is the best.
[0097] And before output, the global meta-learner can check whether further refinement is needed. If the user is not satisfied with the plan, or wants a lower budget / more cities, he can go back to step S110 again, and perform the next round of expansion based on the well-performing parts of "Thinking Path 1" and "Thinking Path 2", and cut off the poorly performing paths or sub-models. Let different ideas or algorithms complement each other to obtain more robust and accurate results. If the answer to be output already meets the output conditions, a global optimal itinerary can be given directly, and the following will be given: uncertainty score (such as confidence interval 0.85~0.9) The satisfaction of the plan with user needs (budget, location, cultural experience) is such as: "Recommended plan: 5-day itinerary, costing US$1,800~1,900, convenient transportation connection, and both popular attractions and cultural activities; satisfaction score 0.87."
[0098] Finally, the optimal solution selected after the second-order fusion is output to the user and may be presented in the form of text / list / chart on the front end:
[0099] “Travel itinerary: Day 1: Arrive in Tokyo->Visit Sensoji Temple, stroll around Akihabara, and experience izakaya culture at night; Day 2: Deepen Tokyo cultural exploration, choose tea ceremony / kimono / manga cafe, and take the Shinkansen to Kyoto in the evening; Day 3: Full day in Kyoto, Kiyomizu Temple, Arashiyama Bamboo Grove, kimono experience, etc.; Day 4: Kyoto->Hiroshima, visit the atomic bomb site in the morning, go to Miyajima to see the shrine and taste seafood in the afternoon; Day 5: Return to Tokyo, end of the trip.”
[0100] From the above description, we can see that ToT can develop multiple possible routes in the thinking dimension of a problem; Stacking can integrate the insights of multiple different experts or models in a prediction or decision. Putting them together is a "double insurance" way of thinking:
[0101] 1. Use ToT to develop a thinking path: Generate multiple possible solution directions or ideas for each angle or solution step of the problem. Each idea is like "Expert A", "Expert B", "Expert C"... except that these "experts" are candidate solutions formed in different directions of thinking chains.
[0102] 2. Use Stacking to fuse the results of these paths: Summarize the final answer (or intermediate analysis conclusion, confidence score, etc.) of each thinking path into a "fuser". The "fuser" is like a commander-in-chief, who will learn how to optimally refer to and integrate these different ideas to make the final conclusion more robust.
[0103] In other words, ToT+Stacking allows the system to explore widely at the thinking level (avoiding taking only one path) while deeply integrating at the result level (not just a veto or simple averaging, but intelligent weighted integration).
[0104] Therefore, the present application can use an agent to obtain the request content, convert it into multiple thinking paths, and hand them over to the local solution model for parallel solution to obtain a diverse set of candidate answers; then use the set of candidate answers as the input of the global meta-learner for global learning, so as to analyze the advantages and disadvantages of the answers under multiple thinking paths, and finally output the final answer with the best performance and most in line with expectations after comprehensive consideration. The method provided by the present application has high application value in the fields of multi-model collaboration, task planning, creative text generation, policy reasoning, etc., and can handle uncertain and high-dimensional task scenarios, providing a new idea and technical path for solving the limitations of current large language models. Through hierarchical design, when solving high-dimensional and multi-step reasoning problems, this patent can not only make full use of the generation and understanding capabilities brought by large language models, but also combine the advantages of ensemble learning to achieve a more flexible, accurate and efficient problem-solving process. It has broad application potential in various scenarios such as complex question answering, policy planning, text generation, and scientific research reasoning, providing important technical support and innovative ideas for the development of the next generation of intelligent systems.
[0105] Figure 7 shows the internal structure diagram of a computer device in an embodiment. The computer device can specifically be a terminal or a server. As Figure 7 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the intelligent decision-making method. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the intelligent decision-making method. Those skilled in the art can understand that Figure 7 the structure shown in
[0106] In one embodiment, the present application also proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the method described in any of the foregoing embodiments.
[0107] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0109] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An intelligent decision-making method, characterized in that, It includes the following steps: When the intelligent agent obtains the request content, the intelligent agent performs an analysis operation on the request content to split out at least one thinking path; The intelligent agent performing an analysis operation on the request content includes: determining the content attribute of the request content, where the content attribute includes at least one of text, picture, video, and audio; determining a feature extraction tool according to the content attribute, and using the feature extraction tool to extract key features from the request content; parsing the key features to determine the core intention of the request content; determining at least one reply sub-thought according to the key features and the core intention; combining the reply sub-thoughts to obtain the thinking path; Taking each of the thinking paths as the input of the local solution model, and solving the thinking path according to the local solution model to obtain the corresponding candidate answers, where the local solution model is composed of multiple solution sub-models; aggregating all the candidate answers corresponding to the thinking paths to obtain a candidate answer set; The solving the thinking path according to the local solution model includes: the solution sub-model splitting the thinking path into at least one sub-thinking path, where the sub-thinking path includes multiple sub-questions; arranging the sub-questions in the same sub-thinking path in order, and determining the corresponding solving rounds for the sub-questions between different sub-thinking paths according to the arrangement order; the solution sub-model parallelly solving the sub-questions corresponding to the same solving round between multiple sub-thinking paths to obtain the corresponding sub-reply content; aggregating all the sub-reply content under the same sub-thinking path to obtain the sub-answer of the sub-thinking path; performing a first-order fusion process on the sub-answers between all the sub-thinking paths to obtain the candidate answer of the thinking path; The performing a first-order fusion process on the sub-answers between all the sub-thinking paths includes: obtaining all the sub-reply content constituting each sub-answer, and the evaluation scores of the sub-reply content; extracting the key features in the sub-reply content, where the key features are used to represent the core elements in the sub-reply; performing weighted scoring according to the evaluation scores, and selecting and combining the key features with the highest scores according to the scoring results to construct the candidate answer; the candidate answer is composed of multiple key features, and the sum of the evaluation scores of the key features constituting the candidate answer is greater than or equal to the sum of the evaluation scores of all the sub-reply content under any one of the sub-thinking paths; Taking the candidate answer set as the input of the global meta-learner, and after performing a fusion process operation on the candidate answer set according to the global meta-learner, obtaining the final answer and outputting it.
2. The intelligent decision-making method according to claim 1, wherein The solution sub-model parallelly solving the sub-questions corresponding to the same solving round between multiple sub-thinking paths to obtain the corresponding sub-reply content includes: Obtain the core intention of the requested content determined after the analysis operation, and determine the scoring calculation template for the sub-thinking path according to the core intention; after each round of solution, calculate the evaluation score of the sub-response content according to the scoring calculation template; if the evaluation score is higher than the deep thinking threshold, generate a new sub-question according to the sub-response content to be added to the next round of parallel solution; and / or, After each round of solution, calculate the relevance between the sub-response content of the current round and the sub-questions of the next round; if the relevance is lower than the preset relevance threshold, generate a new sub-question according to the sub-response content to be added to the next round of parallel solution.
3. The intelligent decision-making method according to claim 1, wherein The solution sub-model parallelly solves the sub-questions corresponding to the same solution round among multiple sub-thinking paths to obtain the corresponding sub-response content, including: Obtain the scoring calculation template; After each round of solution, calculate the evaluation scores of all the sub-response contents according to the scoring calculation template; if there is a sub-response content whose evaluation score is lower than the retention threshold, mark the sub-questions of the next round corresponding to the sub-response content as questions to be deleted; and / or, After each round of solution, judge the number of sub-questions to be solved in parallel in the next round; when the number of sub-questions is greater than the preset upper limit, calculate the evaluation scores of all the sub-response contents of the current round according to the scoring calculation template; retain the top N sub-response contents with the evaluation scores, and mark the sub-questions of the next round corresponding to the un-retained sub-response contents as the questions to be deleted; N is a preset integer, and N is less than or equal to the preset upper limit; Delete the sub-thinking path with the question to be deleted as the root.
4. The intelligent decision-making method according to claim 1, wherein After performing the fusion processing operation on the candidate answer set according to the global meta-learner, obtain the final answer and output it, including: Obtain the feature information of each candidate answer in the candidate answer set, and use the feature information as the input of the global meta-learner for processing. The global meta-learner is one or more constituent models among a linear model, a tree model, a neural network, and / or a large language model; The global meta-learner trains and / or infers the feature information to perform weighted fusion on all the feature information, and the processing output is the answer to be output, and the answer to be output is composed of the combination of the feature information; Judge whether the answer to be output meets the output conditions; If not, use the answer to be output and the requested content as the input of the intelligent agent for iteration again; If it meets the conditions, mark the answer to be output as the final answer.
5. The intelligent decision-making method according to claim 4, wherein The global meta-learner trains and / or infers the feature information, including: Retrieve and obtain additional feature information according to the requested content, and the additional feature information is external information related to the requested content in the existing database; Use the additional feature information and the feature information as the input of the global meta-learner for training and / or inference.
6. A computer device, characterized in that, Include a processor and a memory; The processor is used to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Thinking tree-based big language model reasoning method and device
CN118709786A
Financial question and answer method, system and equipment based on multi-agent interaction and medium
CN119539095A