Data-aware intelligent analysis method and system, storage medium and computer device

CN122088710BActive Publication Date: 2026-07-21RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-21

Smart Images

  • Figure CN122088710B_ABST
    Figure CN122088710B_ABST
Patent Text Reader

Abstract

The application discloses a data-aware-based intelligent analysis method and system, a storage medium and a computer device. The method comprises the following steps: receiving a target problem and initializing state information, wherein the state information comprises key sub-conclusions, a reasoning state logic diagram and key data descriptions; through an intelligent agent based on a language model, the following operations are executed in a loop: obtaining input information of a current step; updating the state information of the current step according to the action of the previous step and the environment feedback information, and making a decision based on the state information of the current step to obtain the action of the current step and a tool to be called; calling the tool to execute the action of the current step, and obtaining the environment feedback information of the current step after the action is executed; checking whether the target problem is analyzed; if not, entering the next loop; and if yes, generating an analysis result. The method can improve the reasoning ability, response consistency and interpretability of a deep reasoning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a data-aware intelligent analysis method, system, storage medium, and computer device. Background Technology

[0002] With the rapid development of Large Language Model (LLM) technology, a large number of products claiming to possess "deep analysis" capabilities have emerged in the market. However, closer observation reveals that most of these products adopt a technical architecture of "Retrieval-Augmented Generation (RAG) + Inference Model," simulating the "deep analysis" process solely through text-level logical reasoning. While this approach superficially enhances the model's reasoning ability, it remains essentially limited to text understanding and generation, failing to truly achieve the ability to perceive, interact with, and dynamically reason about real data.

[0003] There's a common misconception in the industry that using a "deep reasoning model" equates to having "deep analysis" capabilities. However, this model overlooks the crucial data-driven and task-oriented nature of data analysis. When faced with complex and ever-changing real-world application scenarios, these products often exhibit insufficient reasoning ability, inconsistent responses, and a lack of interpretability, failing to meet enterprises' needs for efficient, accurate, and interpretable intelligent analysis systems. Summary of the Invention

[0004] In view of this, embodiments of this application provide a data-aware intelligent analysis method, system, storage medium, and computer device, the main purpose of which is to solve the technical problems of insufficient reasoning ability, inconsistent response, and lack of interpretability of existing deep reasoning models when facing complex and ever-changing application scenarios.

[0005] According to one aspect of this application, a data-aware intelligent analysis method is provided, the method comprising: Receive the target question and initialize the state information, wherein the state information includes key sub-conclusions, reasoning state logic diagrams, and key data descriptions; Perform the following operations in a loop using a language model-based agent: Obtain the input information for the current step, wherein the input information includes the target problem, the status information of the previous step, the action of the previous step, and the environmental feedback information of the previous step; Based on the actions and environmental feedback information from the previous step, information is extracted, and key sub-conclusions for the current step are generated based on the extracted information. The key sub-conclusions include conclusion content and conclusion type. Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node; Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the state information of the current step; Based on the current state information, make decisions and determine the actions and tools to be used in the current step. The tool is invoked to perform the action of the current step, and after the action is completed, the environmental feedback information of the current step is obtained. Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, the interconnected nodes in the reasoning state logic graph are merged to compress the reasoning state logic graph. Check whether the target problem has been analyzed; if not, proceed to the next loop; if completed, generate analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

[0006] According to another aspect of this application, a data-aware intelligent analysis method is provided, the method comprising: In response to receiving a target question sent by a user, the target question is sent to the server so that the server can analyze the target question based on the above method and obtain the analysis result of the target question; The analysis results of the target problem are presented.

[0007] According to another aspect of this application, a data-aware intelligent analysis system is provided, the system comprising: A data intelligence environment is a set of tools for providing data, acquiring data, and processing data. A language model-based agent receives the target question, initializes state information, and performs the following operations in a loop: Obtain the input information for the current step, wherein the input information includes the target problem and the state information of the previous step, the action of the previous step and the environmental feedback information of the previous step, and the state information includes key sub-conclusions, reasoning state logic diagrams and key data descriptions; Based on the actions and environmental feedback information from the previous step, information is extracted, and key sub-conclusions for the current step are generated based on the extracted information. The key sub-conclusions include conclusion content and conclusion type. Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node; Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the state information of the current step; Based on the current state information, make decisions and determine the actions and tools to be used in the current step. The tool is invoked to perform the action of the current step, and after the action is completed, the environmental feedback information of the current step is obtained. Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, the interconnected nodes in the reasoning state logic graph are merged to compress the reasoning state logic graph. Check whether the target problem has been analyzed; if not, proceed to the next loop; if completed, generate analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

[0008] According to another aspect of this application, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the above-described data-aware intelligent analysis method.

[0009] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described data-aware intelligent analysis method.

[0010] By employing the above technical solutions, the present application provides a data-aware intelligent analysis method, system, storage medium, and computer device. By introducing an intelligent agent as the core execution unit and utilizing this agent to perform closed-loop operations of task planning, state updating, action execution, and feedback learning, complex analysis tasks can be dynamically scheduled and autonomously executed, thereby effectively improving the overall reasoning ability of the model. Furthermore, by performing multi-step reasoning on the task, natural language instructions can be progressively transformed into structured logical steps. Combined with tools, deep perception and interaction with real data can be achieved. This method not only understands the problem but also enables reasoning and operation in a real data environment, thus achieving true "deep analysis" and improving response consistency. In addition, by collecting environmental feedback information in real time during task execution and using state information to structurally record the results of each reasoning step, users can easily trace, interpret, and audit the entire analysis process, thereby improving the interpretability of the analysis results.

[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an intelligent analysis method based on data perception provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating an agent training method provided in an embodiment of this application is shown. Figure 3 The diagram shows a structural schematic of a data-aware intelligent analysis system provided in an embodiment of this application. Detailed Implementation

[0013] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0014] Currently, in the field of artificial intelligence, most products that claim to have "deep analysis" capabilities adopt a technical architecture of "retrieval augmented generation (RAG) + inference model". This architecture usually simulates the "deep analysis" process only through text-level logical reasoning. Although it superficially improves the model's reasoning ability, it is still limited to the understanding and generation of text and does not truly realize the perception, interaction and dynamic reasoning of real business data.

[0015] There's a common misconception in the industry that using a "deep reasoning model" equates to having "deep analysis" capabilities. In reality, this model ignores the crucial data-driven and task-oriented nature of data analysis. When faced with complex and ever-changing real-world application scenarios, these products often exhibit insufficient reasoning ability, inconsistent responses, and a lack of interpretability, failing to meet enterprises' needs for efficient, accurate, and interpretable intelligent analysis systems.

[0016] Furthermore, many products claim their models possess the ability to "continuously learn and evolve," but in reality, most so-called "continuous evolution" is limited to updating static model parameters or periodically updating the knowledge base, lacking the ability to perceive and adapt to dynamic processes. A truly intelligent system should be able to continuously adjust strategies, optimize paths, and iterate itself during task execution, a capability not yet achieved in current technology.

[0017] Furthermore, from a technical perspective, the core of LLM remains "next token prediction," that is, the probabilistic prediction capability based on language models. If it cannot establish effective connections with real-world data, systems, and tools, LLM cannot truly understand the task context, perform logical reasoning, or provide valuable responses to complex problems. Therefore, LLM must be deeply integrated with structured data, executable code, and dynamic environments in the real world to truly achieve "deep analysis" and "intelligent decision-making."

[0018] Against this background, this application proposes a novel intelligent analysis method and system based on data perception. By constructing a data intelligence environment, introducing a multi-step reasoning mechanism, and integrating structured data interaction with code-level logical reasoning, it can achieve true "deep analysis" and a "task-oriented intelligent system".

[0019] In one embodiment, such as Figure 1 As shown, a data-aware intelligent analysis method is provided. Taking the application of this method to computer devices as an example, the method includes the following steps: Step 101: Receive the target question and initialize the state information, which includes key sub-conclusions, reasoning state logic diagrams, and key data descriptions.

[0020] The target problem refers to any information input by the user through the client interface. This information can be a simple question or a complex analysis task, and can include various forms of information such as text, images, and links. State information refers to the information that the agent needs to update at each step of the reasoning process. It can be recorded in structured information to show the logical thread of the agent throughout the reasoning process.

[0021] In this embodiment, key sub-conclusions refer to important intermediate conclusions gradually derived by the agent during the reasoning process, which are used to support the final decision and have traceability and logical coherence; the reasoning state graph refers to the dynamic recording of the agent's reasoning path in the form of a flowchart, including logical structures such as condition judgment, hypothesis verification, and path branching; and key data descriptions refer to the data sources on which the key sub-conclusions are based and their brief descriptions.

[0022] Specifically, upon receiving the target question input by the user, the system first initializes the state information for the current analysis task. In this initial state information, key sub-conclusions are empty, indicating that no intermediate conclusions have been derived; the reasoning state logic diagram is empty, indicating that no reasoning path has been formed; and the key data description is empty, indicating that no relevant data has been acquired. This initialization process establishes a starting point for the agent's analysis flow, enabling subsequent iterative execution to proceed on a standardized basis.

[0023] In a typical LLM-Agent architecture, state information is usually modeled as dialogue history, i.e., a textual record of the system's interactions with the environment. However, this representation of state has significant limitations: as the task progresses, the dialogue history accumulates, easily reaching the upper limit of the LLM context length, forcing subsequent inference to be interrupted. Furthermore, this state modeling based on raw dialogue struggles to support long-term memory and life-long learning, thus limiting the agent's adaptability and scalability in complex, long-cycle tasks.

[0024] To overcome the above problems, this embodiment abstracts the Agent's state information into three core components: key sub-conclusions, reasoning state logic diagrams, and key data descriptions. This can elevate the state information from "raw dialogue" to "structured cognitive summary," thereby greatly alleviating the dilemma of context length limitations and allowing the model to focus more on decision-making. It can also effectively improve the efficiency of sample synthesis. Next, the following steps 102 to 105 are executed cyclically by the language model-based agent: In this context, the "Agent" refers to an AI system that autonomously operates within a data-aware environment. This system is capable of automatically querying data, executing code, performing inference and analysis, and outputting results based on task requirements. The Agent can be trained using various AI models, including Large Language Models (LLMs), Computer Vision (CV) models, Large Multimodal Models (LMMs), and smaller models created through model distillation techniques. The data-aware environment is an environment that integrates real-world data access with procedural reasoning capabilities, supporting the Agent in dynamic data acquisition, logical reasoning, and decision optimization. In this embodiment, the data-aware environment and the Agent are two crucial components of the intelligent analysis system.

[0025] Step 102: Obtain the input information for the current step, which includes the target problem, the status information of the previous step, the action of the previous step, and the environmental feedback information of the previous step.

[0026] Specifically, in each iterative reasoning process, the agent first needs to obtain the input information for the current step. This input information encapsulates the complete context output by the agent and the data intelligence environment during the previous interaction, as well as the target question input by the user. Initially, the agent can obtain the target question and the initial state information. From the second step onwards, the agent can obtain the target question, the updated state information from the previous step, the action performed by the agent in the previous step, and the environmental feedback information returned by the data intelligence environment after the previous action. For example, if the agent's action is to retrieve data, the environmental feedback information is the data table returned after querying data using a domain-specific language; if the agent's action is to process data, the environmental feedback information is the data calculation result obtained after running the code using a code execution tool. This embodiment, by obtaining complete context information before executing each step, allows the agent to fully perceive the historical progress of the task and the response of the external environment, thereby ensuring the consistency of decision-making and information.

[0027] Step 103: Extract information based on the actions and environmental feedback from the previous step, and generate key sub-conclusions for the current step based on the extracted information. Key sub-conclusions include conclusion content and conclusion type.

[0028] Specifically, when updating the status information of the current step, information can first be extracted based on the action of the previous step and the environmental feedback information returned from the data intelligence environment after the action was performed. That is, cognitive content that has a substantial role in advancing the analysis task can be extracted from the raw data or calculation results, and key sub-conclusions of the current step can be generated based on the extracted information.

[0029] Step 104: Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node.

[0030] Specifically, after obtaining the key sub-conclusions, a corresponding node can be added to the current reasoning state logic graph based on the conclusion type of the generated key sub-conclusions. This node is used to represent a key step in this logical reasoning, and the specific content of the key sub-conclusion of the current step is added to this newly added node, so that the entire reasoning path can be extended and recorded in a visual way.

[0031] Step 105: Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the status information of the current step.

[0032] Specifically, based on the environmental feedback information from the previous step and combined with the key sub-conclusions generated in the current step, a structured summary can be created of the data involved or generated in the current step. This summary generates a key data description for the current step, which records the characteristics of the acquired data, such as the data source, relevant indicators, and dimensions, but does not need to include specific numerical values. Through this method, the agent's state information can be systematically updated, thus providing a richer context for the next round of decision-making.

[0033] In this embodiment, by extracting key sub-conclusions from environmental feedback information and adding these key sub-conclusions and their corresponding key data descriptions to the reasoning state logic diagram, the cognitive state of the agent can be comprehensively and structurally updated. This approach not only makes the reasoning path clear and traceable but also provides precise and concise context for subsequent decision-making, avoiding the context length problem caused by redundant original data and improving the efficiency of multi-round reasoning and the interpretability of the conclusions.

[0034] Step 106: Make a decision based on the current state information to obtain the action and tools to be called in the current step.

[0035] Specifically, after obtaining the action and environmental feedback information from the previous step, the agent can first update the state information of the current step. Specifically, in the state update phase, the agent analyzes the environmental feedback information returned after the previous action and extracts new insights from it. For example, if the previous step obtained sales data for a store in the second and third quarters using domain-specific language tools, the state update phase will derive a new key sub-conclusion, such as "the store's sales declined in the second half of the year." Then, key data features are summarized from the obtained data and written into the key data description. Next, this key sub-conclusion and the corresponding key data description are added to the inference state logic graph, thereby updating the agent's inference path.

[0036] Furthermore, after the state information is updated, decision-making can be based on the current step's state information to determine the action for the current step and the tools to be invoked. Specifically, during the decision-making phase, the agent can decide what action to take and which tool to invoke to advance the analysis based on the updated state information—the latest key sub-conclusions, the reasoning path in the reasoning state logic diagram, and the known data in the key data description. For example, should it continue to invoke domain-specific language tools to obtain more granular data, or invoke code execution tools to perform regression analysis on existing data? This step, by transforming environmental feedback information into structured state information and making action decisions based on the latest state information, enables dynamic understanding and autonomous planning of complex tasks, avoiding getting lost in the original text.

[0037] Step 107: Call the tool to execute the action of the current step, and after the action is completed, obtain the environmental feedback information of the current step.

[0038] In this context, intelligent agents can acquire data or perform specific computational and reasoning processes in a data intelligence environment by invoking external tools. In this embodiment, the tools invoked by the intelligent agent may include data query tools and code execution tools. The data query tool can invoke a domain-specific language to send query requests to the underlying data pool to obtain data with the required metrics and dimensions. The code execution tool can write and run code in a specific language to achieve further processing, modeling, and reasoning of the data.

[0039] Specifically, after deciding on the action for the current step and the tools to be invoked, the corresponding tools can be called. For example, if the action decision is to acquire certain data, the agent can construct a query statement conforming to the domain-specific language specification, specifying the required metrics, dimensions, filtering conditions, and time range in the query statement, and then call the domain-specific language tool to extract the corresponding data from the underlying data pool. If the action decision is to perform deeper analysis on known data, the agent can write code and then call a code execution tool to run the code in a secure sandbox environment to perform statistical analysis, modeling, or inference on the data. Furthermore, after the tool completes execution, its execution results can be captured; for example, the agent can obtain the returned data table or the text output by the code and use it as environmental feedback information for the current step. This step, by invoking data query tools and code execution tools based on action decisions, extends the agent's capabilities from text reasoning to data acquisition and code-level computation, enabling the agent to perform practical operations and verifications in a data intelligence environment, thereby providing core execution capabilities for solving complex analytical problems.

[0040] Step 108: Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, merge the interconnected nodes in the reasoning state logic graph to compress the reasoning state logic graph.

[0041] Specifically, when the key sub-conclusions corresponding to multiple nodes are logically related—for example, multiple nodes are supporting conclusions and they collectively support the same new conclusion, or multiple nodes form a continuous and branchless reasoning chain—the system can identify the relationships between these nodes and integrate them into a more general node. The merged node reflects the cognitive outcome or accumulated evidence that multiple steps jointly point to, and the node merging process can be synchronized with the continuous updating of the reasoning state logic diagram. For example, node merging can be performed after each step, or after completing the judgment of a branch, the nodes of the entire branch or a portion of the nodes within the branch can be merged, and so on. In this way, while maintaining the integrity of the core reasoning path, the complexity of the graphical representation of the reasoning state logic diagram can be reduced, thus avoiding the problem of overly long or difficult-to-read visualizations caused by too many steps, making the logical backbone of the entire analysis process clearer and more prominent.

[0042] By merging logically related nodes in the reasoning state logic diagram, the size of the reasoning state logic diagram can be effectively compressed while ensuring that the core reasoning path and key cognitive turning points are not lost. This improves the simplicity and readability of the visualization, enabling users to grasp the main logic of the analysis process more quickly. At the same time, it can also reduce the burden on the system when storing and rendering complex logic diagrams.

[0043] Step 109: Check if the target problem has been analyzed. If not, proceed to the next loop. If it has been analyzed, generate the analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

[0044] Specifically, after each iteration of steps 102 to 108, the agent can check whether the target problem has been analyzed. The checking condition can be that the agent has derived sufficient evidence to support the final conclusion, or that the agent has reached the set maximum steps and invoked a specific termination tool to indicate completion of the analysis. If the analysis is not complete, the process returns to step 102 and enters the next iteration; if the analysis is complete, the agent summarizes the structured information accumulated throughout the analysis process, including key sub-conclusion chains in the reasoning state logic graph to obtain the entire logical derivation process from initial observation to final attribution; a complete reasoning state logic graph that visualizes the decision branches and verification paths at each step; and all key data descriptions that summarize the data characteristics supporting the conclusion at each step. Based on this information, the agent can generate a final analysis report, which includes not only the conclusions but also a complete, traceable chain of evidence and the reasoning process. This step, by integrating the structured cognitive results accumulated over multiple iterations, generates highly interpretable analysis results, enabling users not only to know the analysis conclusions but also to clearly understand how the conclusions were reached step by step, and what data and logic each step relied on.

[0045] In a specific example, a user poses a target question to the system: "Analyze whether the total sales volume of a certain store has decreased over the past week." The agent can then initialize its state information, making the key sub-conclusions, inference state logic graph, and key data description empty. In the first loop, the agent receives the target question and initial state information, and makes a decision based on the initial state. It decides to first obtain the store's overall sales data, so it calls a data query tool to retrieve the daily total sales volume over the past month. After execution, the tool returns a data table containing the daily totals, i.e., the environmental feedback information. The agent checks that the problem analysis is incomplete, so it enters the next loop. At the beginning of the second loop, the input information from the previous step includes the target question, initial state information, the action from the previous call, and the returned data table. The agent updates its state based on this information, extracts the key sub-conclusion "the total sales volume has been continuously decreasing over the past week, with the largest decrease in the last three days" from the data table, adds this conclusion to the key sub-conclusion chain, writes the obtained data features into the key data description, and adds a "Get Overall Trend" node to the inference state logic graph. Based on the updated state information, the agent reconsiders its decision-making process, recognizing the need to further understand which product categories caused the decline. It then decides to invoke the data query tool again and construct query instructions to obtain sales data for different product categories. In the third loop, based on the newly returned category data, the agent updates the key sub-conclusion to "the decline in sales of category A products is the main reason for the overall decline," and adds a branch to the inference state logic graph, showing the process of attributing from the overall trend to specific categories. Subsequently, the agent decides to invoke the code execution tool to write code to calculate the list of specific products in category A with the most severe sales decline and their contribution. In the fourth loop, based on the code execution results, the agent updates the key sub-conclusion to "the sharp drop in sales of products X and Y in category A is the core factor," and refines the attribution path from category to specific product in the inference state logic graph. At this point, the agent has completed the problem analysis. Based on the key sub-conclusion chain accumulated throughout the process, the complete reasoning state logic diagram, and all key data descriptions, a final analysis report is generated, clearly showing the attribution process of the decline in total transaction volume. Each step is supported by data and logic, and users can backtrack and review the entire analysis path.

[0046] The above example, by constructing structured state information containing key sub-conclusions, reasoning state logic diagrams, and key data descriptions, enables the agent to accurately update the current step's state information in each loop based on the previous action and environmental feedback, and make the next decision based on the latest state information. This achieves autonomous planning and dynamic advancement of complex analysis tasks. During the decision-making process, the agent can flexibly call domain-specific language tools to obtain real business data or call code execution tools to perform code-level logical operations, thus breaking through the limitations of traditional models that only perform text-based reasoning. Through multiple rounds of state updates, action decisions, tool execution, and environmental feedback, the agent can gradually accumulate a complete, visualized reasoning path and a traceable chain of evidence. The final result is no longer an isolated answer, but a deep analysis report containing a detailed derivation process. This makes the entire analysis process highly interpretable and auditable, allowing users to clearly trace the reasoning and data sources for each step, thereby establishing full trust in the analysis results. Meanwhile, by transforming environmental feedback information into structured state information instead of raw text, the context length limitation problem in long sequence tasks can be effectively alleviated, laying the foundation for the system's continuous learning and complex task processing capabilities.

[0047] By applying the technical solution of this embodiment, and introducing an intelligent agent as the core execution unit, and utilizing the intelligent agent to perform closed-loop operations of task planning, state updating, action execution, and feedback learning, complex analysis tasks can be dynamically scheduled and autonomously executed, thereby effectively improving the overall reasoning ability of the model. Furthermore, by performing multi-step reasoning on the task, natural language instructions can be progressively transformed into structured logical steps. Combined with tools, this enables deep perception and interaction with real-world data. The above method not only understands the problem but also allows for reasoning and operation in a real-world data environment, achieving true "deep analysis" and improving response consistency. In addition, by collecting environmental feedback information in real time during task execution and using state information to structurally record the results of each reasoning step, users can easily trace back, interpret, and audit the entire analysis process, thereby improving the interpretability of the analysis results.

[0048] In one embodiment, the conclusion types of key sub-conclusions include new conclusions, unchanged conclusions, corroborating conclusions, and refutation conclusions. Based on this, in step 104, the reasoning state logic diagram in the state information can be updated in the following ways: If the conclusion type is a new conclusion, a new node is added below the previous new conclusion node or the initial node of the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node; if the conclusion type is an unchanged conclusion, a new node is added below the new conclusion node associated with the unchanged conclusion in the reasoning state logic diagram, and the key data description of the current step is added to the node; if the conclusion type is a corroborating conclusion, a new node is added below the new conclusion node associated with the corroborating conclusion in the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node; if the conclusion type is a refutation conclusion, a new node is added below the new conclusion node associated with the refutation conclusion in the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node.

[0049] In this embodiment, the conclusion types of key sub-conclusions can be divided into four types: new conclusions, unchanged conclusions, corroborating conclusions, and refutation conclusions. Based on this, the reasoning state logic diagram in the state information can be updated in the following ways: If the conclusion type of the key sub-conclusion generated in the current step is a new conclusion, it indicates that the agent has deduced a completely new cognitive node. At this time, a new node can be added below the node of the previous new conclusion or the initial node in the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description on which the conclusion is based can be added to this new node, thereby extending a new reasoning branch; If the conclusion type is an unchanged conclusion, it means that the agent has not deduced a new cognition in the current step, but has further confirmed the previously existing conclusion through new data. At this time, a new node can be added below the node of the new conclusion associated with this unchanged conclusion in the reasoning state logic diagram. The inference state logic diagram is divided into several nodes. For example, if the conclusion type is "supporting conclusion," it means the agent has provided additional evidence to support an existing new conclusion through new data or analysis. In this case, a new node can be added below the node of the new conclusion associated with this supporting conclusion in the inference state logic diagram, and the key sub-conclusion and key data description of the current step can be added to this node to show that the conclusion has been strengthened. If the conclusion type is "refuting conclusion," it means the agent has refuted a previous new conclusion through new evidence. In this case, a new node can be added below the node of the new conclusion associated with the refuted conclusion in the inference state logic diagram, and the key sub-conclusion and key data description of the current step can be added to this node, thus clearly showing the cognitive correction process on the logic diagram. Through this method, the inference state logic diagram not only records the forward inference path but also accommodates complex logic such as data verification, hypothesis strengthening, and even error correction, making the visualization of the entire analysis process richer and more realistic.

[0050] This embodiment defines four specific types for key sub-conclusions: new, unchanged, corroborating, and refutating. Based on these, it performs differentiated node expansion and information filling on the reasoning state logic diagram. This allows the agent's reasoning path to dynamically reflect the deepening, confirmation, reinforcement, or correction of cognition, thereby constructing a more refined and realistic visual analysis map. This greatly enhances the interpretability and traceability of the analysis process, enabling users to clearly understand the basis and nature of each cognitive change.

[0051] In one embodiment, in step 106, the action of the current step and the tool to be invoked can be determined in the following ways: First, based on the key sub-conclusions and reasoning state logic diagram in the state information of the current step, the reasoning hypothesis of the current step is determined; then, based on the reasoning hypothesis of the current step, the data to be acquired or the processing method of the acquired data is determined; if it is determined that data needs to be acquired, the tool to be invoked is determined to be a data query tool, and the indicator name, aggregation dimension and filtering conditions of the data to be acquired are determined; if it is determined that the acquired data needs to be processed, the tool to be invoked is determined to be a code execution tool, and the processing method of the acquired data is determined.

[0052] In this embodiment, when determining the action and the tool to be invoked in the current step, the inference hypothesis that needs to be verified or explored in the current step can first be determined based on the key sub-conclusions and inference state logic diagram contained in the current step's state information. This inference hypothesis is a specific direction for further decomposing the target problem or deepening the verification of existing conclusions. Then, based on the determined inference hypothesis of the current step, it is analyzed what new data needs to be acquired to verify the hypothesis, or what further processing is needed on the already acquired data. If it is determined that new data needs to be acquired, the tool to be invoked is determined to be a data query tool, and the indicator name, aggregation dimension, and filtering conditions of the data to be acquired are further determined in order to construct an accurate data query request. If it is determined that further processing of the already acquired data is needed, such as statistical analysis or logical judgment, the tool to be invoked is determined to be a code execution tool, and the specific processing method for the acquired data is determined, for example, writing code logic to perform specific calculations or data analysis. In this way, the agent can transform abstract inference hypotheses into specific tool invocation instructions.

[0053] This embodiment guides the agent to form reasoning hypotheses based on the current state information, and then decides whether to call a data query tool to obtain new data or a code execution tool to process existing data according to the specific needs of the hypothesis. This effectively connects the analysis ideas and execution actions, so that each operation has a clear logic and target, avoids blind tool calls, and improves the accuracy and efficiency of the analysis process.

[0054] In one embodiment, the tools that the intelligent agent can invoke include a data query tool and a code execution tool. Based on this, step 107 can be implemented as follows: if the invoked tool is a data query tool, a query instruction is constructed using the data query tool, and action-related data is retrieved from a preset database using the query instruction; if the invoked tool is a code execution tool, action-related code is written using the code execution tool, and the written code is run to complete action-related data processing; finally, data analysis can be performed based on the acquired or processed data to obtain data analysis results, and environmental feedback information for the current step can be obtained based on the data analysis results.

[0055] In this embodiment, if the current step determines that the tool to be invoked is a data query tool, the agent can construct a structured query instruction based on the data indicator name, aggregation dimension, and filtering conditions determined in the action decision, and obtain the original data related to the current inference hypothesis from a preset database through the query instruction. If the current step determines that the tool to be invoked is a code execution tool, the agent can write code to perform specific calculations or logical judgments based on the data processing method determined in the action decision, and run the written code in a secure sandbox environment to further process the existing data. After the tool has finished executing, the agent can perform in-depth analysis based on the acquired data or the processed data results to extract data analysis results directly related to the action of the current step, and format the results into environmental feedback information that can be parsed by the next round of steps. This environmental feedback information may include data tables, calculation results, or execution logs, thereby providing direct input information for subsequent state updates and decision-making.

[0056] In this embodiment, within the agent's interactive environment, the agent can invoke two core tools to complete complex tasks: a data query tool and a code execution tool. The data query tool accesses an underlying data pool to query and extract real data. Its behavior is highly dependent on the structure and content of the external database, simulating the actual process of interacting with data sources in a real system, thus giving the environment a strong sense of realism. The introduction of this tool requires the agent to understand the task requirements and accurately construct query logic to obtain effective information relevant to the current action decision. The code execution tool allows the agent to write and run code in a specific language, such as Python code, to further process, calculate, and reason about the data, thereby completing operations such as numerical analysis, logical judgment, and complex derivation. This mechanism extends traditional text-based reasoning capabilities to the code-level reasoning level, significantly enhancing the agent's abstract thinking and problem-solving abilities. The synergistic use of these two tools enables the environment to support not only information acquisition but also dynamic analysis and logical evolution, thereby constructing a data environment that integrates data access and procedural reasoning. Therefore, the environment designed in this embodiment can be considered a "data-aware environment," whose core feature lies in combining reinforcement learning frameworks with real-world data interaction and code execution capabilities, providing agents with more practical application scenarios for learning and decision-making in complex data-driven tasks.

[0057] This embodiment utilizes data query tools to obtain real data from the database or uses code execution tools to further process the obtained data. This allows the tool execution results to be transformed into structured environmental feedback information, thereby enabling efficient interaction between the intelligent agent and the data intelligence environment. This provides an effective data foundation for subsequent state updates and action decisions.

[0058] In one embodiment, a query instruction can be constructed and data retrieved using the following method: First, extract the indicator name, aggregation dimension, and filtering conditions of the data to be retrieved; then, map the indicator name, aggregation dimension, and filtering conditions of the data to be retrieved to standard data names, and construct a query statement with a preset syntax based on the standard data names; finally, convert the query statement into a structured query instruction, and perform data retrieval in a preset database using the query instruction to obtain data related to the action.

[0059] In this embodiment, when constructing a query instruction, the indicator name, aggregation dimension, and filtering conditions of the data to be retrieved can first be extracted from the data requirements determined by the action in the current step, so as to clarify the data content to be queried and its constraints. Then, the extracted indicator name, aggregation dimension, and filtering conditions can be mapped to standard data names defined in a preset database to ensure the accuracy and consistency of the data query. Based on these standard data names, a query statement conforming to the specifications is constructed according to a preset query syntax. Finally, the constructed query statement can be converted into a structured query instruction that the data query tool can recognize, and data retrieval is performed in the preset database using this query instruction to extract data related to the current action from the corresponding data table, such as a query result set returned in tabular form. In this way, the agent can obtain the required information from a complex underlying data pool more accurately, thereby providing a reliable data foundation for subsequent data analysis and state updates.

[0060] In a specific example, suppose we analyze the reasons for a merchant's declining total sales. In the previous step, the agent, based on the inference hypothesis "investigating which specific products within category A caused the decline," determined the data to be acquired, extracting the indicator name "sales revenue," the aggregation dimension "product identifier," and the filtering condition "product category equals category A." In the current step, the agent first maps this information to standardized data names in the database. For example, "sales revenue" is mapped to the standard field name "gmv," "product identifier" to "item_id," "product category" to "category," and "category" in the filtering condition is mapped to the corresponding category code "cat_01." Then, based on these standard names, a query statement is constructed according to a preset query syntax, such as "select item_id, sum(gmv)from sales_table where category = 'cat_01' and ds between '20260101' and '20260107' group by item_id." Finally, the query statement is converted into a structured query instruction that can be executed by the data query tool, and the retrieval is performed in the preset database to obtain a data table containing each product identifier and its corresponding sales amount, which is returned to the intelligent agent as environmental feedback information.

[0061] This embodiment maps the data query requirements of the intelligent agent's decision-making to standardized data names and constructs structured query instructions, which can perform accurate data retrieval on the preset database, ensuring the accuracy and consistency of data acquisition and providing a reliable data foundation for subsequent data analysis and state updates.

[0062] In one embodiment, in step 109, the following method can be used to determine whether the target problem has been analyzed: after each loop is completed, determine whether the key sub-conclusions obtained from multiple steps can solve the target problem, or whether the number of steps executed has reached the preset maximum step threshold; if the key sub-conclusions obtained from multiple steps can solve the target problem, or the number of steps executed has reached the maximum step threshold, then the target problem is determined to have been analyzed.

[0063] In this embodiment, after each loop execution, the agent can evaluate whether the chain of key sub-conclusions obtained from the accumulated steps can sufficiently and accurately solve the initially received target problem. This can be achieved by determining whether the key sub-conclusions cover the core attribution required for the problem or whether they have met the preset conclusion completeness standard. Simultaneously, the agent can also check whether the number of analysis steps executed from the initialization in step 101 to the current step has reached a preset maximum step threshold. This threshold is used to prevent the analysis process from consuming excessive computing resources due to infinite loops. If the key sub-conclusions obtained from multiple steps are deemed to have solved the target problem after evaluation, or if the number of executed steps has reached the preset maximum step threshold, the agent determines that the target problem has been analyzed, and the process exits the loop and enters the stage of generating the final analysis result. The combination of these two judgment conditions ensures both the validity of the analysis results and the convergence and resource controllability of the entire process.

[0064] This embodiment determines the termination time of the analysis process by judging the completeness of key sub-conclusions and whether the number of executed steps has reached a preset threshold. This ensures the validity of the analysis conclusions and provides a reliable safety boundary for the process, thereby avoiding infinite loops caused by logical complexity or insufficient data. This achieves efficient convergence of the analysis process and reasonable control of resource consumption.

[0065] Furthermore, such as Figure 2 As shown, the agent can be trained using the following methods: Step 201: Generate multiple trajectory samples through a preset multi-data synthesis pipeline. The trajectory samples include a state sequence, an action sequence, and a reward sequence consisting of multiple steps.

[0066] Step 202: Using trajectory samples, train the initial agent in a supervised learning manner to obtain the basic agent.

[0067] Step 203: Using the basic agent as the initial policy for reinforcement learning, the agent interacts with dynamic data in a data intelligence environment, and the policy gradient algorithm is used to iteratively optimize the policy of the basic agent.

[0068] In the process of strategy optimization, a reward is calculated for each step, which includes an immediate reward for evaluating the immediate quality of the current output and a cumulative reward for evaluating the contribution of the current output to the overall task.

[0069] Step 204: When the preset training termination condition is met, the trained agent is obtained.

[0070] Among them, the trajectory sample refers to the complete path of the agent from the initial state to the final state during the task execution process, including all intermediate states, actions and rewards; the immediate reward refers to the feedback reward that the agent receives immediately after performing an action, which is used to evaluate the local quality of the current output; the accumulated reward refers to the long-term contribution of the agent's current output to the overall progress of the task and the achievement of the final goal, emphasizing the long-term impact of the action.

[0071] In this embodiment, before training the agent, multiple trajectory samples can be generated through preset data synthesis pipelines. Each trajectory sample includes a state sequence, action sequence, and reward sequence consisting of multiple steps. These trajectory samples simulate the complete behavioral path of the agent when solving different types of analysis tasks in a data intelligence environment. Then, these trajectory samples can be used to train the initial agent in a supervised learning manner, enabling the initial agent to learn basic analytical decision-making patterns and obtain a basic agent with preliminary analytical capabilities. Next, this basic agent is used as the initial policy for reinforcement learning and deployed in a real data intelligence environment to continuously interact with dynamically changing real data. The policy of the basic agent is iteratively optimized using a policy gradient algorithm. Simultaneously, a reward is calculated for each step during policy optimization. This reward includes an immediate reward for evaluating whether the current output meets task specifications such as format correctness and avoids numerical illusion, and a cumulative reward for evaluating the long-term contribution of the current output to the overall task progress and the achievement of the final goal. Finally, when preset training termination conditions are met, such as the policy performance reaching expectations or the training rounds being completed, the trained agent is obtained. In this way, intelligent agents can not only learn from historical experience, but also continuously evolve themselves in real-world environments through interaction with dynamic data, thereby improving their ability to cope with complex and ever-changing analytical tasks.

[0072] This embodiment utilizes a data synthesis pipeline to generate trajectory samples for supervised learning of the initial agent, establishing basic analytical capabilities for the agent. Furthermore, by deploying the supervised-trained agent into a real data intelligence environment and iteratively optimizing its policy using a policy gradient algorithm with immediate and cumulative rewards, the agent can balance short-term behavioral norms with long-term task objectives. This allows for policy self-evolution through continuous interaction with dynamic data, ultimately resulting in an agent capable of adhering to analytical norms and effectively completing complex tasks.

[0073] In one embodiment, in step 201, multiple trajectory samples can be generated by the following methods: A continuous state update pipeline using a dual-agent framework, where the auxiliary agent provides global guidance and the master agent executes sequential action decisions, generates multiple trajectory samples with coherent states; a random state sampling pipeline using a single-agent framework, randomly sampling intermediate states from trajectory samples with coherent states as starting points, generates multiple trajectory samples with discontinuous states; a boundary state awareness pipeline using a single-agent framework, running on a preset data boundary test set, generates multiple trajectory samples with specific abnormal states and boundary states.

[0074] In this embodiment, the core of model learning lies in state updating and action decision-making. State updating refers to updating the current state information based on the action of the previous step and environmental feedback; action decision-making refers to determining the action to be taken in the next step based on the current state information. In a data intelligence environment, the main challenges of action decision-making include decision coherence, generalization ability, and the ability to perceive boundary states. Directly performing reinforcement learning may lead to convergence failure and wasted data computation. Therefore, to address these issues, this embodiment designs three data synthesis pipelines to generate diverse trajectory samples, covering various analysis scenarios that the agent may encounter.

[0075] First, a continuous state update pipeline is employed, utilizing a dual-agent framework. An auxiliary agent, responsible for a global perspective, provides directional guidance for long-sequence tasks, while the primary agent executes specific sequential action decisions based on this guidance. This approach generates trajectory samples with highly coherent state sequences, simulating the agent's steady progress in analyzing tasks under ideal conditions, thereby enhancing the agent's understanding and execution capabilities for long-sequence tasks.

[0076] Secondly, by using a random state sampling pipeline and a single agent framework, one or more intermediate states are randomly selected from the trajectory samples with coherent states already generated by the continuous state update pipeline as new starting points for analysis. The agent continues to perform analysis from these starting points, thereby generating multiple trajectory samples with discontinuous states. These samples are used to simulate scenarios where the analysis process is interrupted or restarted from any intermediate state, thereby enhancing the robustness and generalization ability of the model when facing discontinuous and atypical states.

[0077] Finally, a boundary state-aware pipeline, also employing a single-agent framework, is run on a pre-built data boundary test set. This set includes various anomalies, missing data, or special cases at the data environment boundary. The trajectory samples generated by the agent in this environment can cover multiple behavioral patterns of the agent when dealing with abnormal and boundary states, thereby effectively improving the agent's ability to identify and respond to anomalies or data environment boundaries, thus enhancing the agent's adaptability in complex environments.

[0078] In this embodiment, a portion of trajectory samples can be extracted from the training set constructed by the three data synthesis pipelines to build a test set, which is used to test the performance of the trained agent. The test set is divided into two types: a basic test set, used to test the agent's continuous data mining and analysis capabilities; and a data boundary test set, used to test the agent's boundary perception capabilities.

[0079] The trajectory samples generated by the three pipelines in this embodiment can cover various scenarios such as routine analysis, interruption restart, and anomaly handling. This can provide a comprehensive and high-quality data foundation for subsequent agent training, thereby improving the robustness and adaptability of the agent in dealing with various complex situations in the real world.

[0080] In one embodiment, in step 203, the reward for each step can be calculated by: obtaining the output of the basic agent in the current step, the output including at least one of state thinking information, key sub-conclusions, reasoning state logic diagrams, action thinking information, and tool call information; calculating the immediate reward for the current step based on the output of the current step, wherein the immediate reward includes a numerical illusion penalty determined based on whether the output contains a value not appearing in the input data and a format specification reward determined based on whether the output format conforms to a preset specification; calculating the cumulative reward for the current step based on the output of the current step and the output of historical steps, wherein the cumulative reward includes a key sub-conclusion gain determined based on the key sub-conclusions for the overall progress of the task, and a thinking length gain determined based on the length of the state thinking information and action thinking information; weighted summing of the immediate reward and cumulative reward for the current step to obtain the total reward for the current step, and updating the total reward for historical steps based on the cumulative reward for the current step.

[0081] Specifically, in reinforcement learning frameworks for large language models, whether in single-turn interactions or multi-turn dialogues, the core objective is to optimize the output quality of the large language model under specific tasks. The interaction logic of single-turn reinforcement learning is relatively intuitive; a well-structured and computationally efficient reward function can be constructed to achieve precise alignment of model actions. However, in multi-turn interaction scenarios, credit assignment becomes a key bottleneck restricting model performance. The core challenge lies in how to accurately identify and quantify the marginal contribution of each action to the overall task objective within a long decision sequence.

[0082] Based on this, this embodiment proposes decoupling the reward signal into two dimensions: Intermediate Reward and Accumulated Reward. Intermediate Reward aims to provide immediate feedback on the current agent's output, focusing on evaluating whether the agent's output conforms to task specifications, such as format conformity and numerical illusion. Intermediate Reward depends only on the quality of the current agent's output and does not involve the long-term impact of state transitions. In contrast, Accumulated Reward focuses on the alignment of macro-level task progress with long-term goals. It measures the substantial contribution of the current decision to subsequent inference paths and the final attribution result. For example, this embodiment defines a "key sub-conclusion gain term" to characterize whether the agent derives core insights that significantly advance the analysis process. This type of reward, by reinforcing actions with "milestone" significance, can guide the model to build an understanding of long-term goals within a complex inference space. Based on the above decoupling mechanism, this embodiment defines the total reward of each step as a weighted sum of immediate reward and cumulative reward, and sets a discount factor for the cumulative reward during the integration process, thereby using the discount factor to update the total reward of historical steps in order to balance the impact of current decision and future returns.

[0083] In this embodiment, when calculating the reward for each step, the output of the basic agent at the current step can be obtained. This output may include at least one of the following: state thinking information for assisting state updates, newly generated key sub-conclusions, updated reasoning state logic diagrams, action thinking information for assisting action decisions, and tool invocation information. Then, based on the output of the current step, the immediate reward for the current step can be calculated. The immediate reward mainly includes two parts: one part is a numerical illusion penalty determined based on whether the output text contains values ​​that do not actually appear in the input data of the current step. For example, if there are illusory values ​​in the output that are not supported by data, a negative penalty will be given; the other part is a format specification reward determined based on whether the output format conforms to a preset specification. For example, if intermediate information is not output according to the preset format requirements, such as intermediate information with missing fields in key sub-conclusions or grammatical errors in tool invocation instructions, a negative penalty will be given.

[0084] Furthermore, after calculating the immediate reward, the contribution of the current step to the overall task can be comprehensively evaluated based on the output of the current step and the outputs of all historical steps from the start of the analysis to the current step, in order to calculate the cumulative reward for the current step. The cumulative reward includes a key sub-conclusion gain term, determined by whether the current step derives a crucial sub-conclusion essential for solving the target problem. This gain term reflects the degree of contribution of the step to task progress and is the core part of the agent's reward system, aiming to quantify the "information increment" generated by the agent in multiple rounds of reasoning. In addition, the cumulative reward also includes a reasoning length gain term, determined by whether the text length of state reasoning information and action reasoning information is within a reasonable range, to encourage the model to maintain a reasonable reasoning length while maintaining reasoning quality. Unlike traditional long inference models, the agent needs to pursue efficiency and stability in reasoning. This reward aims to prevent a sharp collapse in reasoning length due to excessive illusion penalties, thereby guiding the model to maintain within a controllable reasoning bandwidth.

[0085] Finally, the immediate reward and cumulative reward of the current step can be weighted and summed to obtain the total reward of the current step. At the same time, based on the cumulative reward calculated for the current step, the total reward of previous historical steps is retrospectively updated using a preset discount factor to accurately reflect the true contribution of historical steps to the current progress, thereby solving the credit allocation problem in long sequence tasks.

[0086] In this embodiment, during the continuous reasoning and analysis process, the agent outputs different intermediate information at different stages of each step, such as state thinking information, key sub-conclusions, reasoning state logic diagrams, action thinking information, and tool call information. After these intermediate information are output, immediate and cumulative rewards can be given based on the quality of this information. Alternatively, immediate and cumulative rewards can be given for the entire step after its completion. It is understood that the reward method for the agent's output information at each stage can be set according to the actual situation, and this embodiment does not impose specific limitations.

[0087] This embodiment calculates immediate and cumulative rewards for the agent's output at each step, performs a weighted sum of the immediate and cumulative rewards for each step, and backtracks and updates the cumulative rewards for historical steps. This allows the agent to strictly adhere to output specifications and avoid generating false data during training, while also effectively guiding the agent to perform actions that have long-term and critical value for solving complex tasks during task analysis. This addresses the credit allocation problem in multi-step decision-making, thereby significantly improving the efficiency of reinforcement learning training for the agent and the quality of the agent's final strategy.

[0088] In one embodiment, the numerical illusion penalty and format compliance reward can be determined as follows: The input and output data of the current step are obtained; a reference numerical set is extracted from the input data; and an output numerical set and structured content are extracted from the output data. Each numerical value in the output numerical set is compared with the numerical value in the reference numerical set, and the number of illusionary numerical values ​​is counted. A numerical illusion penalty is calculated based on the number of illusionary numerical values. Each structured content is compared with preset format requirements, and it is determined whether it conforms to the preset format requirements. A format compliance score is assigned to each structured content based on the comparison results. The format compliance scores of all structured content are aggregated to obtain the format compliance reward.

[0089] In this embodiment, when determining the numerical illusion penalty, the input data and the agent's output at the current step are first obtained. All truly existing values ​​are extracted from the input data to construct a reference value set. For example, all sales figures can be extracted from a database query table as the reference value set. Then, all generated values ​​are extracted from the agent's output to construct an output value set. Next, each value in the output value set is precisely compared with the values ​​in the reference value set to count the number of values ​​that are completely absent from the reference value set as the number of illusion values. Finally, the numerical illusion penalty is calculated based on the number of illusion values, where a higher number of illusion values ​​results in a stronger penalty, thereby inhibiting the agent from fabricating data.

[0090] Furthermore, when determining the format compliance reward, the structured content organized according to preset specifications in the output of the current step can be identified, such as the field structure of key sub-conclusions or the parameter format of tool calls. Then, each structured content output by the agent is compared item by item with the preset format requirements. For example, it checks whether the key sub-conclusions contain required fields such as number, conclusion content, and conclusion type. Based on the comparison results, a format compliance score is assigned to each structured content; a high score is given for complete compliance, and points are deducted for missing or incorrect content. Finally, the format compliance scores of all structured content are aggregated and calculated, for example, by summing or averaging, to obtain the final format compliance reward, which is used to encourage the agent to consistently output compliant content.

[0091] This embodiment can effectively suppress the agent's behavior of fabricating false data by calculating the numerical illusion penalty item for each step. At the same time, by calculating the format standardization reward item for each step, it can effectively strengthen the agent's ability to follow the standard output, so that the agent's output is both real and credible and standardized and usable.

[0092] In one embodiment, the key sub-conclusion gain term in the cumulative reward can be determined in the following way: the key sub-conclusion generated in the current step is judged by a preset conclusion judge; if the key sub-conclusion is judged to be valid, a positive reward value is allocated according to the conclusion type and importance of the key sub-conclusion; if the key sub-conclusion is judged to be invalid, a preset negative reward value is allocated.

[0093] In this embodiment, when calculating the gain term of key sub-conclusions in the cumulative reward, a pre-built conclusion judge can be used to determine the validity of the key sub-conclusions generated by the agent in the current step. For example, a finely tuned language model or a rule-based logic validator can be used as the conclusion judge. The judgment criteria mainly include whether the newly generated key sub-conclusions are consistent with the existing key data descriptions and reasoning logic in the current state information, and whether the conclusions themselves are logically consistent. If the conclusion evaluator determines that the key sub-conclusion is valid, meaning it considers the current conclusion to have a correct role in advancing the analysis task, then it further assigns a corresponding positive reward value based on the type of the key sub-conclusion (e.g., whether it is a completely new conclusion or a corroborating conclusion of an existing one) and its importance within the overall problem-solving framework (e.g., whether it directly points to the core cause of the problem). Higher positive rewards are given to more important and groundbreaking conclusions. Conversely, if the conclusion determines that the key sub-conclusion is invalid—for example, if the conclusion contradicts existing data, contains logical leaps, or is completely erroneous—it assigns a pre-set, larger negative reward value to penalize the agent for misleading or incorrect reasoning. In this way, the cumulative reward accurately reflects the true value of each cognitive achievement.

[0094] This embodiment utilizes a conclusion judge to determine the validity of key sub-conclusions generated at each step, and allocates positive or negative cumulative reward values ​​differently based on the conclusion type and importance. This allows the agent to be accurately guided during reinforcement learning, while effectively suppressing unfounded guesses or erroneous reasoning, thereby continuously improving the accuracy and reliability of the agent's analysis results.

[0095] In one embodiment, in step 203, the agent's policy can be optimized by: acquiring multiple trajectory samples; for each step in each trajectory sample, calculating the advantage function value, wherein the advantage function value is used to measure the superiority or inferiority of the output of the current step relative to the average level; calculating the policy gradient based on the advantage function value, and updating the agent's parameters with the goal of maximizing the reward, wherein, when calculating the policy gradient, the gradients of steps or trajectory samples that meet preset conditions are masked.

[0096] In this embodiment, during the reinforcement learning phase, the Policy Gradient algorithm can be used for end-to-end training of the model. Specifically, during policy optimization, for each step in each trajectory sample, the advantage function value for that step can be calculated. This advantage function value measures the cumulative reward obtained by the action actually performed by the agent in the current state, relative to the average cumulative reward obtained by all possible actions in the same state. A positive value indicates that the action is better than the average, while a negative value indicates that it is worse than the average. Then, based on the calculated advantage function value, the policy gradient is estimated. This gradient indicates how to adjust the agent's policy parameters so that actions with high advantage values ​​are more likely to be adopted in the future, thereby updating the agent's parameters with the ultimate goal of maximizing the expected cumulative reward. Furthermore, during the calculation of policy gradients, for specific steps or entire trajectory samples that meet preset conditions, their gradients are masked, meaning their contributions are not included in the final parameter update. These preset conditions may include situations where the format specification reward in the immediate reward of that step is below a threshold, the key sub-conclusions generated in that step are deemed invalid by the conclusion judge, or the number of invalid key sub-conclusions in the entire trajectory sample exceeds a preset upper limit. This approach prevents the agent from learning incorrect decision-making patterns from pseudo-positive samples where the results are acceptable but the intermediate steps contain obvious errors.

[0097] This embodiment measures the relative merits of each step by calculating the advantage function value and updates the policy gradient based on the advantage function value. This allows the agent to optimize the decision-making strategy in a direction that is more conducive to obtaining high cumulative rewards. At the same time, by performing gradient masking on steps or trajectories that meet preset conditions such as format errors or invalid conclusions, the agent can be effectively prevented from learning from samples containing noise or errors. This can significantly improve the stability of the reinforcement learning process and the quality of the final policy.

[0098] In one embodiment, the policy gradient of the agent can be masked by the following method: if the advantage function value of the current step is positive and at least one immediate reward value is lower than a preset threshold, then the gradient of the current step is prohibited from participating in the update of model parameters; if the advantage function value of the current step is positive and the key sub-conclusion generated by the current step is determined to be invalid, then the gradient of the current step is prohibited from participating in the update of model parameters; if the number of invalid steps in a trajectory sample exceeds a preset threshold, then the gradient of the trajectory sample is prohibited from participating in the update of model parameters.

[0099] In this embodiment, even if the cumulative reward of the entire trajectory is positive in a multi-round decision-making task, it may still contain some logical errors or formatted noise, leading to reward hacking—the behavior of an agent exploiting loopholes in the reward function to obtain high scores in an unexpected way. To prevent the model from learning these pseudo-positive samples, this embodiment further designs three gradient masking strategies, from step-level to trajectory-level, to perform fine-grained filtering of trajectory samples.

[0100] First, for each step's calculated advantage function value, a positive value indicates that the current step's action is better than average. Simultaneously, if at least one immediate reward value for that step is detected to be below a preset threshold (e.g., the format specification reward fails to meet the acceptable standard or there is a numerical illusion penalty), then it is determined that although the final result of that step is acceptable, the execution process has significant flaws. Therefore, the gradient of that step is prohibited from participating in the update of model parameters. In this way, it is ensured that the agent will not neglect adherence to task specifications due to accidental high scores.

[0101] Secondly, if the advantage function value of the current step is positive, but the key sub-conclusion generated by the step is deemed invalid by the conclusion judge, for example, if the current key sub-conclusion contradicts the existing data or is logically invalid, the gradient of the current step is also prohibited from participating in the update, in order to prevent the model from learning a decision pattern that accidentally obtains high returns but has an incorrect reasoning process, thereby solving the false positive problem in credit allocation.

[0102] Finally, the entire trajectory sample is evaluated. If the number of invalid steps in the trajectory sample exceeds a preset threshold—for example, if the number of invalid key sub-conclusions exceeds a preset threshold—the entire trajectory sample is considered of low quality, and the gradients of all steps in that trajectory sample are prohibited from participating in the update of model parameters to maintain the overall purity of the corpus. Through the above multi-layered masking mechanism, it can be ensured that only steps and trajectories with standardized processes, valid conclusions, and high overall quality can make a positive contribution to the agent's learning.

[0103] In a specific example, in analyzing the reasons for a decline in a merchant's total sales, suppose that during training, a trajectory sample consists of five steps. The first three steps are logically correct and receive positive cumulative rewards, but in the fourth step, the agent generates an invalid key sub-conclusion, even though the final conclusion of the fifth step appears correct. After calculating the advantage function value for each step, when preparing to update the policy gradient, the system detects that the invalid key sub-conclusion generated in the fourth step meets a preset masking condition. Therefore, the gradient corresponding to the fourth step is masked, preventing it from participating in the current parameter update. In this way, the agent's parameter update will mainly rely on the positive gradients provided by the first three logically correct steps, thus avoiding being misled by an intermediate erroneous step masked by a coincidentally correct final result.

[0104] This embodiment sets gradient occlusion conditions at three levels: the standardization of immediate rewards, the validity of conclusions, and the overall quality of the trajectory. This can effectively filter out the interference of samples with flawed processes, invalid conclusions, or low quality on model parameter updates, ensuring that the reinforcement learning process learns the correct decision patterns only from high-quality positive samples. This significantly improves the efficiency of training and the robustness and accuracy of the final policy.

[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. In addition, the labels corresponding to each step in the above embodiments are only for identification purposes and are not intended to limit the execution order of the steps. The execution order of the steps in each embodiment can be set according to the actual situation.

[0106] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this application provides a data-aware intelligent analysis system, such as... Figure 3 As shown, the system includes: A data intelligence environment is a set of tools that can be used to provide data, acquire data, and process data. A language model-based agent can receive a target question, initialize state information, and perform the following operations in a loop: Obtain the input information for the current step, wherein the input information includes the target problem and the state information of the previous step, the action of the previous step and the environmental feedback information of the previous step, and the state information includes key sub-conclusions, reasoning state logic diagrams and key data descriptions; Based on the actions and environmental feedback information from the previous step, information is extracted, and key sub-conclusions for the current step are generated based on the extracted information. The key sub-conclusions include conclusion content and conclusion type. Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node; Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the state information of the current step; Based on the current state information, make decisions and determine the actions and tools to be used in the current step. The tool is invoked to perform the action of the current step, and after the action is completed, the environmental feedback information of the current step is obtained. Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, the interconnected nodes in the reasoning state logic graph are merged to compress the reasoning state logic graph. Check whether the target problem has been analyzed; if not, proceed to the next loop; if completed, generate analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

[0107] In specific application scenarios, the conclusion types include new conclusions, unchanged conclusions, corroborating conclusions, and refutation conclusions. Specifically, when the conclusion type is a new conclusion, the agent can add a node below the previous new conclusion node or the initial node in the reasoning state logic diagram, and add the key sub-conclusions and key data descriptions of the current step to the node; when the conclusion type is an unchanged conclusion, add a node below the new conclusion node associated with the unchanged conclusion in the reasoning state logic diagram, and add the key data descriptions of the current step to the node; when the conclusion type is a corroborating conclusion, add a node below the new conclusion node associated with the corroborating conclusion in the reasoning state logic diagram, and add the key sub-conclusions and key data descriptions of the current step to the node; when the conclusion type is a refutation conclusion, add a node below the new conclusion node associated with the refutation conclusion in the reasoning state logic diagram, and add the key sub-conclusions and key data descriptions of the current step to the node.

[0108] In specific application scenarios, the intelligent agent can be used to determine the reasoning hypothesis of the current step based on the key sub-conclusions and reasoning state logic diagram in the state information of the current step; based on the reasoning hypothesis of the current step, determine the data to be acquired or the processing method of the acquired data; if it is determined that data needs to be acquired, then the tool to be called is determined to be a data query tool, and the indicator name, aggregation dimension and filtering conditions of the data to be acquired are determined; if it is determined that the acquired data needs to be processed, then the tool to be called is determined to be a code execution tool, and the processing method of the acquired data is determined.

[0109] In specific application scenarios, the tool is either a data query tool or a code execution tool. Specifically, when the tool is a data query tool, the agent can construct query instructions using the data query tool and retrieve data related to the action from a preset database using the query instructions. When the tool is a code execution tool, the agent can write code related to the action using the code execution tool and run the code to complete data processing related to the action. Based on the acquired or processed data, the agent can perform data analysis to obtain data analysis results and obtain environmental feedback information for the current step based on the data analysis results.

[0110] In specific application scenarios, the intelligent agent can be used to extract the indicator name, aggregation dimension, and filtering conditions of the data to be acquired; map the indicator name, aggregation dimension, and filtering conditions of the data to be acquired to standard data names, and construct a query statement with a preset syntax based on the standard data names; convert the query statement into a structured query instruction, and perform data retrieval in a preset database through the query instruction to obtain data related to the action.

[0111] In specific application scenarios, the intelligent agent can be used to determine whether the key sub-conclusions obtained from multiple steps can solve the target problem or whether the number of executed steps has reached a preset maximum step threshold; if the key sub-conclusions obtained from multiple steps can solve the target problem or the number of executed steps has reached the maximum step threshold, then the target problem analysis is determined to be complete.

[0112] In specific application scenarios, the system further includes a model training module. This module can generate multiple trajectory samples through a pre-defined multi-data synthesis pipeline. Each trajectory sample includes a state sequence, action sequence, and reward sequence consisting of multiple steps. Using these trajectory samples, an initial agent is trained in a supervised learning manner to obtain a basic agent. This basic agent serves as the initial policy for reinforcement learning, interacting with dynamic data in a data intelligence environment. The policy of the basic agent is iteratively optimized using a policy gradient algorithm until a pre-defined training termination condition is met, resulting in a trained agent. During policy optimization, a reward is calculated for each step. This reward includes an immediate reward for evaluating the immediate quality of the current output and a cumulative reward for evaluating the current output's contribution to the overall task.

[0113] In specific application scenarios, the model training module can be used to generate multiple trajectory samples with coherent states through a continuous state update pipeline, employing a dual-agent framework where the auxiliary agent provides global guidance and the main agent executes sequential action decisions; through a random state sampling pipeline, employing a single-agent framework, where intermediate states are randomly sampled from the coherent trajectory samples as starting points to generate multiple trajectory samples with discontinuous states; and through a boundary state awareness pipeline, employing a single-agent framework, which runs on a preset data boundary test set to generate multiple trajectory samples with specific abnormal states and boundary states.

[0114] In specific application scenarios, the model training module can be used to obtain the output of the basic agent in the current step. The output includes at least one of state thinking information, key sub-conclusions, reasoning state logic diagram, action thinking information, and tool call information. Based on the output of the current step, calculate the immediate reward for the current step, wherein the immediate reward includes a numerical illusion penalty determined by whether the output contains a value not appearing in the input data, and a format specification reward determined by whether the output format conforms to a preset specification; based on the output of the current step and the output of previous steps, calculate the cumulative reward for the current step, wherein the cumulative reward includes a key sub-conclusion gain determined by the key sub-conclusion on the overall progress of the task, and a thinking length gain determined by the length of state thinking information and action thinking information; perform a weighted summation of the immediate reward and the cumulative reward for the current step to obtain the total reward for the current step, and update the total reward for previous steps based on the cumulative reward for the current step.

[0115] In specific application scenarios, the model training module can be used to acquire the input data and output of the current step, extract a reference value set from the input data, and extract an output value set and structured content from the output; compare each value in the output value set with the values ​​in the reference value set, and count the number of hallucination values, and calculate the numerical hallucination penalty item based on the number of hallucination values; compare each structured content with the preset format requirements, and determine whether it meets the preset format requirements, and assign a format compliance score to each structured content according to the comparison results; aggregate the format compliance scores of all structured content to obtain the format standardization reward item.

[0116] In specific application scenarios, the model training module can be used to determine the validity of key sub-conclusions generated in the current step using a preset conclusion evaluator; if the key sub-conclusion is determined to be valid, a positive reward value is assigned according to the conclusion type and importance of the key sub-conclusion; if the key sub-conclusion is determined to be invalid, a preset negative reward value is assigned.

[0117] In specific application scenarios, the model training module can be used to acquire multiple trajectory samples, calculate the advantage function value for each step in each trajectory sample, wherein the advantage function value is used to measure the quality of the output of the current step relative to the average level; calculate the policy gradient based on the advantage function value, and update the parameters of the agent with the goal of maximizing the reward, wherein when calculating the policy gradient, the gradient of the step or trajectory sample that meets the preset conditions is masked.

[0118] In specific application scenarios, the model training module can be used to prevent the gradient of the current step from participating in the update of model parameters when the advantage function value of the current step is positive and at least one instant reward value is lower than a preset threshold; prevent the gradient of the current step from participating in the update of model parameters when the advantage function value of the current step is positive and the key sub-conclusion generated by the current step is determined to be invalid; and prevent the gradient of the trajectory sample from participating in the update of model parameters when the number of invalid steps contained in a trajectory sample exceeds a preset threshold.

[0119] It should be noted that other corresponding descriptions of the functional units involved in the data-aware intelligent analysis device provided in this application embodiment can be found in the following references. Figures 1 to 2 The corresponding descriptions in the method will not be repeated here.

[0120] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0121] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0122] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0123] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0124] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0125] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0126] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data-aware intelligent analysis method, characterized in that, The method includes: Receive the target question and initialize the state information, wherein the state information includes key sub-conclusions, reasoning state logic diagrams, and key data descriptions; Perform the following operations in a loop using a language model-based agent: Obtain the input information for the current step, wherein the input information includes the target problem, the status information of the previous step, the action of the previous step, and the environmental feedback information of the previous step; Based on the actions and environmental feedback information from the previous step, information is extracted, and key sub-conclusions for the current step are generated based on the extracted information. The key sub-conclusions include conclusion content and conclusion type. Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node; Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the state information of the current step; Based on the current state information, make decisions and determine the actions and tools to be used in the current step. The tool is invoked to perform the action of the current step, and after the action is completed, the environmental feedback information of the current step is obtained. Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, the interconnected nodes in the reasoning state logic graph are merged to compress the reasoning state logic graph. Check whether the target problem has been analyzed; if not, proceed to the next loop; if completed, generate analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

2. The method according to claim 1, characterized in that, The conclusion types include new conclusions, unchanged conclusions, corroborating conclusions, and refutation conclusions; therefore, based on the conclusion type of the key sub-conclusion, a new node is added to the reasoning state logic graph, and the key sub-conclusion of the current step is added to the node, including: If the conclusion type is a new conclusion, then a new node is added below the node of the previous new conclusion or the initial node of the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node. If the conclusion type is an unchanged conclusion, then a new node is added below the node of the new conclusion associated with the unchanged conclusion in the reasoning state logic diagram, and the key data description of the current step is added to the node. If the conclusion type is a corroborating conclusion, then a new node is added below the node of the new conclusion associated with the corroborating conclusion in the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node. If the conclusion type is a refutation conclusion, then a new node is added below the node of the new conclusion associated with the refutation conclusion in the reasoning state logic diagram, and the key sub-conclusion of the current step and the key data description of the current step are added to the node.

3. The method according to claim 1, characterized in that, The decision-making process based on the current step's state information, resulting in the action for the current step and the tools to be invoked, includes: Based on the key sub-conclusions and reasoning state logic diagram in the current step's state information, determine the reasoning hypothesis for the current step; Based on the reasoning assumptions of the current step, determine the data that needs to be acquired or the processing method for the acquired data; If it is determined that data needs to be obtained, then the tool to be called is a data query tool, and the indicator name, aggregation dimension and filtering conditions of the data to be obtained are determined; If it is determined that the acquired data needs to be processed, then the tool to be invoked is determined to be a code execution tool, and the processing method for the acquired data is determined.

4. The method according to claim 1 or 3, characterized in that, The tool is a data query tool or a code execution tool; then, the step of calling the tool to execute the action of the current step, and obtaining the environmental feedback information of the current step after the action is completed, includes: If the tool is the data query tool, then a query instruction is constructed using the data query tool, and data related to the action is obtained from a preset database using the query instruction; If the tool is the code execution tool, then the code related to the action is written using the code execution tool, and the code is run to complete the data processing related to the action; Data analysis is performed based on the acquired or processed data to obtain data analysis results, and environmental feedback information for the current step is obtained based on the data analysis results.

5. The method according to claim 4, characterized in that, The step of constructing a query command using the data query tool and retrieving data related to the action from a preset database using the query command includes: Extract the metric name, aggregation dimension, and filtering conditions of the data to be acquired; The indicator names, aggregation dimensions, and filtering conditions of the data to be acquired are mapped to standard data names, and a query statement with a preset syntax is constructed based on the standard data names. The query statement is converted into a structured query instruction, and data retrieval is performed in a preset database using the query instruction to obtain data related to the action.

6. The method according to claim 1, characterized in that, The check to see if the target problem has been analyzed is included: Determine whether the key sub-conclusions obtained from multiple steps can solve the target problem or whether the number of executed steps has reached a preset maximum step threshold. If the key sub-conclusions obtained from multiple steps can solve the target problem or the number of steps executed has reached the maximum step threshold, then the target problem analysis is deemed complete.

7. The method according to claim 1, characterized in that, The training method for the intelligent agent includes: Multiple trajectory samples are generated through a pre-set multi-data synthesis pipeline, wherein the trajectory samples include a state sequence, an action sequence, and a reward sequence consisting of multiple steps; Using the trajectory samples, the initial agent is trained in a supervised learning manner to obtain the basic agent; Using the basic agent as the initial strategy for reinforcement learning, it interacts with dynamic data in a data intelligence environment, and iteratively optimizes the strategy of the basic agent through a policy gradient algorithm until the preset training termination condition is met, thus obtaining a trained agent. In the process of strategy optimization, a reward is calculated for each step, which includes an immediate reward for evaluating the immediate quality of the current output and a cumulative reward for evaluating the contribution of the current output to the overall task.

8. The method according to claim 7, characterized in that, The process of generating multiple trajectory samples through a preset multi-data synthesis pipeline includes: A continuous state update pipeline is used, employing a dual-agent framework. The auxiliary agent provides global guidance, while the master agent executes sequential action decisions, generating multiple trajectory samples with coherent states. By using a random state sampling pipeline and a single agent framework, intermediate states are randomly sampled from multiple trajectory samples with coherent states as the starting point to generate multiple trajectory samples with discontinuous states. By using a boundary state perception pipeline and a single agent framework, the pipeline runs on a pre-defined data boundary test set to generate multiple trajectory samples for specific abnormal states and boundary states.

9. The method according to claim 7, characterized in that, The calculation of rewards for each step in the strategy optimization process includes: Obtain the output of the basic intelligent agent in the current step, the output including at least one of state thinking information, key sub-conclusions, reasoning state logic diagram, action thinking information and tool call information; Based on the output of the current step, calculate the immediate reward of the current step, wherein the immediate reward includes a numerical illusion penalty determined based on whether the output contains a value that does not appear in the input data, and a format specification reward determined based on whether the output format conforms to a preset specification. Based on the output of the current step and the output of the historical steps, the cumulative reward of the current step is calculated, wherein the cumulative reward includes a key sub-conclusion gain term determined by the key sub-conclusion on the overall progress of the task, and a thinking length gain term determined by the length of state thinking information and action thinking information. The immediate reward and cumulative reward of the current step are weighted and summed to obtain the total reward of the current step. Based on the cumulative reward of the current step, the total reward of the historical steps is updated.

10. The method according to claim 9, characterized in that, The method for determining the numerical illusion penalty and the format specification reward includes: Obtain the input data and output of the current step, extract the reference value set from the input data, and extract the output value set and structured content from the output; Each value in the output value set is compared with the value in the reference value set, and the number of hallucination values ​​is counted. The numerical hallucination penalty is calculated based on the number of hallucination values. Each structured content is compared with the preset format requirements, and it is determined whether it meets the preset format requirements. Based on the comparison results, a format compliance score is assigned to each structured content. The format compliance scores of all structured content are aggregated to obtain the format compliance reward item.

11. The method according to claim 9, characterized in that, The method for determining the gain term of the key sub-conclusion includes: The validity of the key sub-conclusions generated in the current step is determined using a pre-defined conclusion evaluator. If the key sub-conclusion is determined to be valid, a positive reward value is assigned based on the conclusion type and importance of the key sub-conclusion. If the key sub-conclusion is determined to be invalid, a preset negative reward value is assigned.

12. The method according to claim 7, characterized in that, The iterative optimization of the policy of the basic agent using the policy gradient algorithm includes: Multiple trajectory samples are acquired. For each step in each trajectory sample, a dominance function value is calculated, wherein the dominance function value is used to measure the quality of the output of the current step relative to the average level. The policy gradient is calculated based on the advantage function value, and the parameters of the agent are updated with the goal of maximizing the reward. When calculating the policy gradient, the gradients of steps or trajectory samples that meet preset conditions are masked.

13. The method according to claim 12, characterized in that, The gradient masking process for steps or trajectory samples that meet preset conditions includes at least one of the following: If the advantage function value of the current step is positive and at least one immediate reward value is lower than a preset threshold, then the gradient of the current step is prohibited from participating in the update of model parameters. If the advantage function value of the current step is positive, and the key sub-conclusion generated in the current step is determined to be invalid, then the gradient of the current step is prohibited from participating in the update of the model parameters. If the number of invalid steps in a trajectory sample exceeds a preset threshold, the gradient of that trajectory sample is prohibited from participating in the update of model parameters.

14. A data-aware intelligent analysis method, characterized in that, The method includes: In response to receiving a target question sent by a user, the target question is sent to a server so that the server analyzes the target question based on the method according to any one of claims 1 to 13 and obtains the analysis result of the target question; The analysis results of the target problem are presented.

15. A data-sensing-based intelligent analysis system, characterized in that, The system includes: A data intelligence environment is a set of tools for providing data, acquiring data, and processing data. A language model-based agent receives the target question, initializes state information, and performs the following operations in a loop: Obtain the input information for the current step, wherein the input information includes the target problem and the state information of the previous step, the action of the previous step and the environmental feedback information of the previous step, and the state information includes key sub-conclusions, reasoning state logic diagrams and key data descriptions. Based on the actions and environmental feedback information from the previous step, information is extracted, and key sub-conclusions for the current step are generated based on the extracted information. The key sub-conclusions include conclusion content and conclusion type. Based on the conclusion type of the key sub-conclusion, add a new node to the reasoning state logic graph and add the key sub-conclusion of the current step to the node; Based on the environmental feedback information from the previous step and the key sub-conclusions of the current step, generate the key data description for the current step to update the state information of the current step; Based on the current state information, make decisions and determine the actions and tools to be used in the current step. The tool is invoked to perform the action of the current step, and after the action is completed, the environmental feedback information of the current step is obtained. Based on the conclusion types of the key sub-conclusions of multiple steps in the reasoning state logic graph, the interconnected nodes in the reasoning state logic graph are merged to compress the reasoning state logic graph. Check whether the target problem has been analyzed; if not, proceed to the next loop; if completed, generate analysis results based on the key sub-conclusions of multiple steps, the reasoning state logic diagram, and the key data description.

16. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 14.

17. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 14.