Data analysis method and device based on progressive task learning, equipment and medium
By constructing an interactive data environment and a progressive task learning method, a set of training trajectories covering different difficulty levels is generated, the agent is initialized and its parameters are updated, which solves the problem of lack of autonomous learning in existing technologies and improves analytical capabilities and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack a unified training mechanism that can learn progressively based on task difficulty and combine demonstration trajectories with quality feedback, making it difficult to achieve autonomous learning and improved analytical capabilities in complex domain tasks.
An interactive data environment is constructed to generate a set of training trajectories covering a sequence of tasks with progressively increasing difficulty levels. The basic language model is initialized, and reward signals are generated by analyzing the operation trajectories and output results. The network parameters of the agent are updated, and the training is repeated until the highest level of verification is completed, and the target analysis results are obtained.
It enables intelligent agents to gradually learn complex analytical capabilities in a progressive task system, possessing stronger task understanding, execution stability, and analytical accuracy, and is suitable for fintech and healthcare business scenarios.
Smart Images

Figure CN121882159A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making technology, and in particular to a data analysis method, apparatus, device, and medium based on progressive task learning. Background Technology
[0002] When dealing with complex domain tasks, existing technologies generally rely on human experience or automated tools based on fixed rules, making it difficult to build sustainably scalable analytical workflows from multi-source heterogeneous data. For scenarios requiring layer-by-layer reasoning, dynamic adjustment of the analytical chain, and autonomous learning, existing systems lack mechanisms for progressive learning based on task difficulty, accumulation of operational trajectories, and adaptive parameter updates, thus failing to support the continuous advancement of highly complex tasks. Furthermore, there is a general lack of a quality feedback system covering the entire task execution process, preventing models from achieving a closed-loop capability from demonstration and evaluation to improvement in different task environments.
[0003] In the fintech sector, due to significant structural differences and time dependencies among market data, transaction flow data, industry indicator data, and news and public opinion data, existing intelligent analysis tools often can only perform static indicator processing and simple pattern recognition. They cannot continuously accumulate task execution experience, nor can they gradually improve the model's analytical capabilities based on task difficulty. Existing systems generally lack a progressive training framework based on task sequences, preventing models from gradually increasing their inference depth in challenging tasks such as risk identification, financial anomaly judgment, or cross-cycle trend inference. Furthermore, existing technologies lack a data feedback mechanism based on operational process records, making it difficult for models to self-correct and improve analytical accuracy when dealing with complex market events.
[0004] In the healthcare field, data from imaging examinations, monitoring equipment, medical records, and laboratory indicators are highly heterogeneous. Traditional systems typically rely on fixed processes or empirical rules for auxiliary analysis, lacking adaptive learning methods to adapt to changes in task complexity. When faced with multi-step tasks such as evidence-based path derivation, diagnostic suggestion generation, or multimodal association analysis, existing systems cannot accumulate analytical trajectories through task execution, cannot form a trajectory learning mechanism based on demonstration data, and cannot dynamically update analytical capabilities based on quality feedback. Furthermore, due to the multi-stage and highly specialized nature of medical tasks, existing technologies lack training frameworks for constructing progressive task sequences, making it difficult for models to continuously advance when handling higher-level diagnostic reasoning tasks. Summary of the Invention
[0005] The main objective of this invention is to provide a data analysis method, apparatus, device, and storage medium based on progressive task learning, aiming to solve the technical problem that the existing technology lacks a unified training mechanism that can learn progressively based on task difficulty, combine demonstration trajectories and quality feedback, and autonomously improve analytical capabilities in an interactive data environment.
[0006] To achieve the above objectives, this invention provides a data analysis method based on progressive task learning, comprising: Construct an interactive data environment and a data-driven trajectory synthesis framework based on a teacher model and a target domain knowledge base; The data-driven trajectory synthesis framework is used to generate a training trajectory set covering a sequence of tasks with progressively increasing difficulty levels. The basic language model is initialized as an autonomous data analysis agent within the domain, and the initial difficulty level task in the progressive difficulty level task sequence is set as the current level task; In the interactive data environment, the current level task is input into the domain data autonomous analysis agent, and the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment are obtained; A quality analysis is performed on the analysis operation trajectory and the analysis result output to generate a reward signal; The network parameters of the domain data autonomous analysis agent are updated based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; The autonomous data analysis agent in the domain is subjected to a task promotion and iterative training process, which involves repeatedly executing the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level of verification in the progressive difficulty level task sequence, thus obtaining the trained autonomous data analysis agent in the domain. The target analysis task to be processed is obtained, and the autonomous analysis agent with the trained domain data is used to perform an autonomous analysis process for the target analysis task in the interactive data environment to obtain the target analysis result.
[0007] Furthermore, to achieve the above objectives, the present invention provides a data analysis apparatus based on progressive task learning, comprising: The interactive environment construction module is used to build an interactive data environment and a data-driven trajectory synthesis framework based on the teacher model and the target domain knowledge base. The trajectory synthesis and generation module is used to generate a set of training trajectories covering a progressively more difficult task sequence using the data-driven trajectory synthesis framework. The agent initialization module is used to initialize the basic language model as an autonomous agent for data analysis within the domain, and to set the initial difficulty level task in the progressive difficulty level task sequence as the current level task. The task execution acquisition module is used to input the current level task into the domain data autonomous analysis agent in the interactive data environment, and to acquire the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment. The quality assessment module is used to perform quality analysis on the analysis operation trajectory and the analysis result output, and generate reward signals; The parameter update module is used to update the network parameters of the domain data autonomous analysis agent based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set. The level promotion training module is used to perform task promotion and iterative training process on the autonomous data analysis agent in the domain. It cyclically executes the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level verification of the progressive difficulty level task sequence, and obtains the trained autonomous data analysis agent in the domain. The target task analysis module is used to acquire the target analysis task to be processed, and to use the trained domain data autonomous analysis agent to perform an autonomous analysis process for the target analysis task to be processed in the interactive data environment to obtain the target analysis result.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a data analysis program based on progressive task learning stored in the memory and executable on the processor, wherein when the data analysis program based on progressive task learning is executed by the processor, it implements the steps of the data analysis method based on progressive task learning as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a non-volatile computer-readable storage medium storing a data analysis program based on progressive task learning, wherein the data analysis program based on progressive task learning, when executed by a processor, implements the steps of the data analysis method based on progressive task learning as described above.
[0010] Beneficial Effects: This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses a data analysis method, apparatus, device, and medium based on progressive task learning, comprising: constructing an interactive data environment and establishing a trajectory synthesis framework based on a teacher model and a target domain knowledge base; generating a training trajectory set covering a sequence of tasks with progressive difficulty levels; initializing a basic language model as an autonomous data analysis agent within the domain; receiving and processing tasks of different difficulty levels to obtain analysis operation trajectories and analysis result outputs; generating reward signals based on analysis quality; updating the agent's network parameters in conjunction with demonstration data; and obtaining a trained autonomous data analysis agent within the domain through multiple rounds of task advancement and iterative training; and using the trained agent to execute target analysis tasks in the interactive data environment to obtain target analysis results. This invention combines a course learning mechanism, demonstration trajectory synthesis, and reward feedback training, enabling the agent to gradually acquire complex analytical capabilities within a progressive task system. This results in stronger task comprehension, execution stability, and analytical accuracy when facing real-world tasks, achieving autonomous analysis output for complex tasks. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a data analysis method based on progressive task learning in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the data analysis method based on progressive task learning according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the data analysis device based on progressive task learning of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The data analysis method based on progressive task learning provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can construct an interactive data environment through the client and establish a trajectory synthesis framework based on the teacher model and the target domain knowledge base. This generates a training trajectory set covering task sequences with progressively increasing difficulty levels, initializes the basic language model as an autonomous agent for domain-specific data analysis, receives and processes tasks of different difficulty levels to obtain analysis operation trajectories and analysis result outputs, generates reward signals based on analysis quality, updates the agent's network parameters using demonstration data, and obtains a trained autonomous agent for domain-specific data analysis through multiple rounds of task advancement and iterative training. The trained agent then performs target analysis tasks in the interactive data environment to obtain target analysis results. This invention combines a course learning mechanism, demonstration trajectory synthesis, and reward feedback training, enabling the agent to gradually acquire complex analytical capabilities within a progressive task system. This results in stronger task comprehension, execution stability, and analytical accuracy when facing real-world tasks, achieving autonomous analysis output for complex tasks. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of specific embodiments further illustrates this invention.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the data analysis method based on progressive task learning provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the data analysis method based on progressive task learning proposed in this invention includes the following steps: S10, Construct an interactive data environment and build a data-driven trajectory synthesis framework based on the teacher model and the target domain knowledge base; In this embodiment, the construction of the interactive data environment involves the unified organization of task inputs, state variables, data access interfaces, and feedback generation units, enabling the agent to obtain a computable environmental state through continuous interaction. The environment is represented by a structured modeling approach, representing the analysis context, including a searchable dataset, a set of executable analysis actions, and corresponding environment update logic. The data access interface originates from a consistent encapsulation of data from different sources, used to access indicator data, textual materials, or time-series information. The environment update logic updates the state content in real time based on the analysis actions, enabling the agent to obtain a traceable analysis context.
[0016] Based on an interactive data environment, a data-driven trajectory synthesis framework is constructed using a teacher model and a target domain knowledge base. The teacher model, trained from domain data, generates reference action sequences with coherent reasoning. The target domain knowledge base contains rule items, structured knowledge entries, entity relationships, or sets of technical terms, used to validate the actions generated by the teacher model, ensuring that each action in the trajectory is consistent with domain knowledge. The trajectory synthesis framework receives the environmental state, uses the teacher model to generate reference behaviors, and then verifies their rationality against the knowledge base, forming a complete demonstration trajectory. The trajectory sequence can generate demonstration sets of varying complexity based on different task difficulties for use in subsequent learning processes.
[0017] This embodiment provides a structured, continuous analysis context through an interactive data environment, and generates semantically consistent, logically coherent demonstration trajectories that cover multiple levels of complexity through the teacher model and the target domain knowledge base. This improves the structure, accuracy, and diversity of the demonstration data, thereby enhancing the ability to express complex analytical behaviors and generalize tasks during subsequent training.
[0018] S20, using the data-driven trajectory synthesis framework to generate a training trajectory set covering a progressive difficulty level task sequence; In this embodiment, a data-driven trajectory synthesis framework is used to generate a training trajectory set covering a progressively more complex task sequence. This set relies on task difficulty modeling, trajectory generation strategies, and multi-level complexity organization to form a learnable task sequence. The progressively more complex task sequence expresses a progression from basic capabilities to higher-level analytical abilities. Each level consists of a task description, a reasoning chain, and a set of actions interacting with the environment. The task difficulty is determined by the breakdown of analytical requirements for the target application domain, such as from initial information understanding to cross-data source correlation analysis and multivariate reasoning integration. Each level of task structure maintains a computable and observable form.
[0019] The data-driven trajectory synthesis framework combines the reasoning ability of the teacher model with the correction capabilities of domain knowledge entries, enabling it to generate trajectory sequences with logical structure, action order, and intermediate feedback for tasks of varying difficulty. Upon receiving the task description, the teacher model generates a task decomposition chain and triggers environmental actions one by one. The framework records these actions and corresponding feedback, then combines them with knowledge entries to structurally filter and refine the trajectory content, ensuring that the trajectory set retains both the reasoning path and adherence to domain rules. Each trajectory connects the task input, reasoning steps, instruction actions, and feedback content through temporal concatenation, making the trajectory trainable, reproducible, and scalable.
[0020] To cover all progressive difficulty levels, the framework generates trajectory sets level by level according to the complexity of the task sequence, ensuring that each level has a matching number of trajectories and content granularity. The basic level emphasizes the executability of single action chains, the intermediate level emphasizes cross-variable correlation analysis, and the advanced level emphasizes structured multi-step reasoning under complex task objectives. The trajectory sets are organized in a unified format, with each trajectory having a consistent field structure, such as task text, reasoning expression, action sequence, and environmental feedback, allowing subsequent learning modules to use them directly.
[0021] This embodiment provides progressive supervision for the subsequent learning process through demonstration trajectories of varying difficulty, enabling training support for analytical behaviors of different complexities, thereby improving the model's generalization ability and complex reasoning ability for real-world tasks.
[0022] S30, initialize the basic language model as an autonomous data analysis agent within the domain, and set the initial difficulty level task in the progressive difficulty level task sequence as the current level task; In this embodiment, the process of initializing the basic language model as an autonomous data analysis agent within the domain includes three parts: model structure loading, parameter pre-configuration, and capability expansion. The basic language model is a general-purpose reasoning model trained on massive corpora, possessing the capabilities of cross-text understanding, logic generation, and multi-turn expression. To enable it to perform domain analysis, the environmental interaction interface needs to be adapted during the initial parameter loading process, allowing the model to trigger actions such as retrieval, tool invocation, or command execution within the data environment. The parameter pre-configuration mechanism introduced during initialization provides behavioral boundaries and expression specifications for the model, ensuring the consistency and stability of its generated reasoning chains and query actions.
[0023] Setting the initial difficulty level task in a progressively challenging task sequence as the current difficulty level task is a crucial task scheduling action in the training process. The initial difficulty level task is characterized by low operational complexity, short inference chains, and few interactions with the environment, providing the agent with a clear and low-burden training starting point. The task scheduling module standardizes the task text, task objectives, and corresponding data input formats, enabling the agent to recognize its target structure and input boundaries when executing the current difficulty level task. After initialization, upon receiving the current difficulty level task, the agent generates an internal implicit state based on the task description, including a task understanding vector, action prediction vector, and a behavioral plan for interacting with the environment, thus building an executable foundation for subsequent inference processes.
[0024] There is a close dependency between agent initialization and the current level of task setting. The agent's initial behavior space comes from the model parameter initialization, while the task setting provides the agent with the first set of observable behavior samples for training, enabling the agent to learn behavior within a clear objective and a controllable range of difficulty. The combination of these two aspects gives the agent a clear starting point in state representation, action generation, and decision path planning, thus ensuring that training does not fall into behavioral deviation or inference breakdown in the early stages due to excessive task complexity.
[0025] This embodiment initializes the basic language model and sets an initial difficulty level task, enabling the agent to have a clear starting behavioral space and controllable task complexity in domain analysis training. This ensures that the reasoning expression, human-computer interaction actions, and behavioral trajectories in the early stages of training remain stable, and that subsequent training to improve skills has coherence and scalability, thereby enhancing the agent's performance in complex application scenarios.
[0026] S40, in the interactive data environment, the current level task is input into the domain data autonomous analysis agent, and the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment are obtained; In this embodiment, the process of inputting the current-level task into the domain data for autonomous analysis by an intelligent agent in an interactive data environment includes three parts: task expression transformation, agent input loading, and environmental interaction triggering. The current-level task is typically described in text form, including the objective, constraints, and the range of data to be analyzed. To enable the agent to execute the task in the environment, the task content needs to be converted into executable structured instructions. These structured instructions may include tool invocation formats, data referencing methods, environmental resource paths, and task boundary descriptions, allowing the agent to internally formulate an executable behavioral plan.
[0027] The task content input into the agent triggers the model's inference mechanism, causing the model to map the task text into internal representation vectors, including task understanding vectors, environmental operation vectors, and action prediction vectors. These vectors form a continuously updated implicit state within the model, enabling the agent to initiate behaviors such as data retrieval, logical deduction, tool invocation, and result reading according to task requirements. The inference actions unfold in a layer-by-layer generation manner within the agent, ensuring that each action prediction is based on the environmental feedback and internal state evolution of the previous step.
[0028] During task execution, the agent generates analytical operation trajectories based on task instructions and environmental feedback. These trajectories may include task breakdown text, round-by-round inference chains, tool call logs, data execution requests, data entries returned by the environment, and descriptions of intermediate states. These trajectories are combined chronologically to form a complete behavioral record, which demonstrates the agent's action logic, decision-making basis, and path changes during task execution.
[0029] The final analysis output is a summary generated by the agent upon task completion, integrating the task reasoning chain, environmental feedback data, and the agent's internal inference results. The analysis results may be presented as structured text, summary conclusions, or formalized expressions, encompassing the agent's final judgments and deliverables in this round of execution.
[0030] This embodiment loads the current level task into the agent and executes it in an interactive data environment, enabling the agent to form observable behavioral paths and complete reasoning records. This allows the training process to have clear behavioral inputs, environmental feedback, and task outputs, enabling the agent to obtain clear behavioral samples in task understanding, logic chain construction, and tool interaction capabilities, thereby improving the reliability and interpretability of the agent in subsequent learning stages.
[0031] S50, perform quality analysis on the analysis operation trajectory and the analysis result output, and generate a reward signal; In this embodiment, the quality analysis process for the analysis operation trajectory and analysis result output includes three consecutive operations: logical chain parsing, key indicator comparison, and semantic content inspection. The analysis operation trajectory records inference nodes, tool invocation processes, and environmental interactions. The inference text needs to be broken down into logical nodes, and the consistency of causal relationships, referencing relationships, and inference order between nodes needs to be verified to check for jumps or contradictions in the inference chain. The analysis result output includes structured indicators, text descriptions, or model-generated content. Key indicators need to be extracted and compared numerically with benchmark data in the data environment to identify calculation biases or trend misjudgments. Semantic content inspection uses sentence-level semantic scanning to identify whether sensitive expressions that do not conform to industry standards are included. Finally, the logical consistency results, numerical comparison biases, and semantic scanning markers are merged to generate a reward signal, enabling a quantitative assessment of trajectory quality and output quality.
[0032] This embodiment achieves quantifiable evaluation of the reasoning chain, numerical results, and textual expression by performing quality analysis on the analysis trajectory and output content and generating reward signals. This enables the agent's behavior process and output content to receive controllable quality feedback, providing clear optimization basis for subsequent learning stages and improving analysis consistency and result reliability.
[0033] S60, update the network parameters of the domain data autonomous analysis agent according to the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; In this embodiment, the process of updating network parameters based on the reward signal and demonstration data in the training trajectory set includes three closely related operations: demonstration data retrieval, difference metric calculation, and advantage quantification and gradient update. Demonstration data originates from the training trajectory set and needs to be mapped and filtered according to the current level task identifier. Then, content consistent with the task behavior chain is combined from the filtering results. Demonstration data includes inference fragments, tool call order, and feedback data, which can serve as a reference for expected behavior. The difference metric, targeting the demonstration data, calculates the deviation between the behavior distribution generated by the agent in the current training round and the behavior distribution corresponding to the demonstration data, making the degree of inference deviation measurable in numerical space. The reward signal, as external feedback, needs to be compared with the value assessment. The difference between the reward signal and the estimated reward is converted into an advantage quantity, reflecting the deviation between the actual feedback and the internal prediction. The deviation quantity and the advantage quantity are jointly converted into a differentiable gradient, used to guide the update of parameter weights in the network, enabling the model to achieve joint convergence between the demonstration behavior direction and the environmental feedback direction.
[0034] This embodiment utilizes both behavioral reference information from the example data and environmental feedback from the reward signals, enabling network parameter updates to not only rely on the structured behavioral patterns of the example trajectories but also to adjust behavioral tendencies in conjunction with dynamic feedback, thereby achieving faster convergence to the task objective and higher decision stability.
[0035] S70, execute the task promotion and iterative training process for the autonomous data analysis agent in the domain, and repeatedly execute the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level verification of the progressive difficulty level task sequence, and obtain the trained autonomous data analysis agent in the domain. In this embodiment, the task promotion and iterative training process comprises closely related components such as task difficulty level management, training loop triggering mechanism, task input execution chain, output result acquisition mechanism, reward generation logic, and parameter update triggering conditions. Task difficulty level management relies on a progressively challenging task sequence, using level identifiers to determine the complexity of the current training task, enabling the training process to gradually transition from low to high complexity. The training loop triggering mechanism continuously schedules the training process; each loop requires inputting the current level task into the domain-independent data analysis agent, ensuring the agent always operates in a structured environment. The task input execution chain involves the access and processing of task text, the encapsulation of environment call instructions, and the agent's inference execution process, ensuring that tasks always enter the model in a unified format. The output result acquisition mechanism includes behavior sequence recording and result content extraction, capturing the operation chain, inference fragments, and final generated content during task execution, enabling the training loop to obtain quantifiable actual performance. The reward generation logic outputs feedback values based on task execution, allowing the training loop to map execution quality into numerical feedback. The parameter update trigger condition is based on the difference between the reward value and the task trajectory performance, determining whether to execute the update, allowing the agent to continuously adjust its internal weights in a continuous loop. The promotion mechanism in the training process relies on level judgment logic; when the performance exceeds the set evaluation range, the task is switched to a higher level, enabling the training process to progressively progress from basic to advanced tasks. Finally, when the highest-level task is consistently passed, the training loop terminates, and the state management module outputs the autonomously analyzing domain data of the trained agent.
[0036] This embodiment combines continuous training cycles with a level promotion mechanism, enabling the agent to gradually transition from low-difficulty tasks to high-difficulty tasks. Through continuous feedback and updates, it forms a capability transfer path from basic reasoning to complex analysis, and ultimately achieves stable task execution capabilities and higher analytical consistency after multiple rounds of training.
[0037] S80: Obtain the target analysis task to be processed, and use the trained domain data autonomous analysis agent to perform the autonomous analysis process for the target analysis task to be processed in the interactive data environment to obtain the target analysis result.
[0038] In this embodiment, acquiring the target analysis task to be processed comprises three closely related components: task source parsing, task description structuring, and task execution constraint extraction. Task source parsing can originate from system triggers, external requests, or batch scheduling, identifying the task objective and data focus scope by parsing input text or parameter descriptions. Task description structuring generates a unified field format based on the input content, enabling the trained domain-independent data analysis agent to correctly identify the task. Task execution constraint extraction extracts the time range, analysis scope, indicator dimensions, or execution depth requirements from the task content, enabling the agent to make inference decisions based on complete task conditions. The trained domain-independent data analysis agent possesses the ability to perform reasoning, data querying, logical inference, and multi-round iterative analysis in an interactive data environment. Internally, it operates collaboratively through a reasoning unit, tool invocation unit, context memory unit, and result generation unit. The interactive data environment provides data query channels, simulation scenario interfaces, computation execution containers, and result return channels, enabling the agent to dynamically acquire data, perform calculations, verify hypotheses, or generate inference chains during task execution. The autonomous analysis process comprises five consecutive stages: task understanding, data acquisition, logical deduction, result integration, and output generation. During each invocation, the agent selects the data type, analysis path, and deduction depth based on task requirements, continuously integrating intermediate inference results into the internal analysis chain. The target analysis results are logically closed by the inference unit, and the generation unit outputs structured or textual results, thus forming the final deliverable content.
[0039] This embodiment enables a trained autonomous analysis agent to construct reasoning paths and integrate multi-source data when faced with complex tasks by executing an autonomous analysis process in an interactive data environment, thereby generating structured and coherent target analysis results and improving the completeness and interpretability of the analysis.
[0040] In one embodiment, step S10 above includes: S101 integrates multi-source heterogeneous data reading interfaces and executable code analysis tool libraries and deploys them to a virtual sandbox container to obtain an interactive data environment; S102, Semantic parsing and domain analysis logic extraction are performed on unstructured industry expert reports and compliance guidelines documents to construct a structured target domain knowledge base; S103, using the domain analysis logic and compliance boundary conditions in the target domain knowledge base to perform system-level prompting configuration on the pre-trained language model in order to instantiate the teacher model; S104, establish a two-way automated communication protocol between the teacher model and the interactive data environment, so that the teacher model configured with system-level prompts can schedule the executable code analysis tool library in the virtual sandbox container, process the data obtained by the multi-source heterogeneous data reading interface, and make decisions based on the domain analysis logic and compliance boundary conditions in the target domain knowledge base through the two-way automated communication protocol; S105, a data-driven trajectory synthesis framework is obtained by encapsulating the bidirectional automated communication protocol.
[0041] In this embodiment, the interactive data environment is constructed based on a multi-source heterogeneous data reading interface and an executable code analysis tool library. The multi-source heterogeneous data reading interface targets different data carriers and protocols, unifying channels such as database queries, object storage access, message queue retrieval, and file directory scanning into a set of configurable reading units. Each reading unit corresponds to a data source type and parsing method, specifying connection parameters, authentication methods, field mappings, and data refresh strategies in the configuration. The executable code analysis tool library consists of a set of callable analysis units, such as data cleaning scripts, indicator calculation scripts, statistical test scripts, and model inference scripts. Each analysis unit has fixed input format, output format, runtime timeout, and resource consumption limits in its metadata. The multi-source heterogeneous data reading interface and the executable code analysis tool library are jointly deployed inside a virtual sandbox container. The virtual sandbox container limits the execution boundaries through process isolation, resource quotas, and network access control, ensuring that the analysis scripts execute in a controlled environment, thereby forming an interactive data environment that can be remotely invoked by upper-layer intelligent agents. The interactive data environment orchestrates data reading and code execution in a unified manner, and provides repeatable data query and analysis capabilities for the upper-level inference process through task queues, execution schedulers, and result caching.
[0042] The construction of the target domain knowledge base focuses on unstructured text resources, including industry expert reports and compliance guidelines documents. Industry expert reports typically present analytical approaches, evaluation conclusions, and indicator combinations in long text format, while compliance guidelines documents describe restrictions and prohibited situations in clause form. The semantic parsing process first employs segmentation, sentence-by-sentence, and clause-by-clause processing to break down the long text into semantic fragments. Then, entity recognition, relation extraction, and syntactic dependency analysis are used to identify business objects, key indicators, constraints, and inference relationships. The domain analysis logic extraction process summarizes these elements into rule units and inference paths, such as "which type of indicator should be prioritized when a certain type of constraint is met" and "which type of alternative judgment logic should be used when data is missing or conflicting." Clauses in the compliance guidelines documents undergo a similar parsing process to form a searchable set of compliance boundary conditions, including a list of prohibited behaviors, threshold boundaries, and prohibited combinations of conditions. The analysis rules, inference paths, and boundary conditions are stored in the target domain knowledge base in a structured form. During storage, indexes are created based on business objects, analysis topics, and clause sources to facilitate rapid retrieval and combination by scenario.
[0043] Building upon this foundation, the teacher model is bound to the target domain knowledge base using system-level prompt configuration. Before configuration, the pre-trained language model only possesses general language capabilities. System-level prompt configuration injects domain roles, reasoning requirements, and compliance constraints into the model, translating the domain analysis logic and compliance boundary conditions in the target domain knowledge base into pre-conversational information that the model can follow long-term. For example, it generates explanatory text clarifying which analytical dimensions the model needs to explore, how to handle different data patterns, and which conclusions require proactive compliance checks. System-level prompt configuration includes not only static text but also logical template placeholders, used to instantiate and populate rules retrieved from the knowledge base at runtime, thus forming teacher model behavioral constraints that adjust with task changes. After configuration, the teacher model possesses the ability to organize reasoning using domain rules and automatically avoid known violation patterns when generating analysis trajectories.
[0044] The teacher model and the interactive data environment interact via a bidirectional automated communication protocol. This protocol defines request message structures, response message structures, error codes, call timeouts, and retry strategies, determining how the teacher model initiates data query requests, issues code to execute tasks, and receives execution results. One side of the protocol runs on the inference service where the teacher model resides, while the other side runs on a server-side component within the virtual sandbox container. When generating analysis trajectories, the teacher model triggers environment calls by embedding call markers conforming to the protocol format in its output. The protocol parsing module extracts the call parameters from the output, encapsulates them into a standard request, and delivers it to the interactive data environment. The interactive data environment calls the multi-source heterogeneous data reading interface and the executable code analysis tool library according to the data source identifier and analysis unit identifier in the request. After execution, it returns results in a unified format. The returned content is then written back to the teacher model's context by the protocol parsing module as the basis for subsequent inference. This allows the teacher model to dynamically drive the virtual sandbox container without directly accessing the underlying execution details, forming a stable closed-loop interaction.
[0045] The data-driven trajectory synthesis framework is encapsulated on a bidirectional automated communication protocol. The framework revolves around a complete task execution process, uniformly recording and recombining the natural language inference content generated by the teacher model, the triggered environment call sequence, the intermediate data returned by the interactive data environment, and the final conclusion. Internally, the framework creates trajectory units for each model call and operation nodes for each protocol call, chaining the inference text, tool calls, and returned results into a continuous trajectory through timestamps and call dependencies. The framework provides an interface to control the teacher model's processes: receiving task descriptions, determining the inference starting point, initiating data calls, receiving results, updating the inference chain, and generating complete analysis output. It also collects all intermediate states in real time during execution. The collected content is stored in a trajectory set in a unified format for subsequent training of the autonomous data analysis agent within the domain, enabling imitation learning and reinforcement learning, thereby achieving the goal of constructing training data from the teacher model's behavior.
[0046] This embodiment integrates multi-source heterogeneous data reading interfaces and executable code analysis tool libraries into a virtual sandbox container. Combined with the target domain knowledge base obtained from semantic parsing, it uses system-level prompts to configure and instantiate teacher models and links them through a two-way automated communication protocol. The data-driven trajectory synthesis framework can completely record the reasoning process, tool call behavior, and data response content of the teacher model within a unified structure, forming a high-quality training trajectory covering a progressively difficult task sequence. This provides a data foundation with a clear structure, explicit compliance constraints, and the ability to reflect the real analysis process for the subsequent training of autonomous data analysis agents within the domain.
[0047] In one embodiment, step S20 above includes: S201, Design a task template that includes basic capability dimension, special analysis dimension and comprehensive research dimension based on the business scenario characteristics in the target domain knowledge base, and generate a progressive difficulty level task sequence based on the task template; S202, control the teacher model in the data-driven trajectory synthesis framework to sequentially execute the simulation analysis tasks in the progressive difficulty level task sequence in the interactive data environment; S203, record the thought chain derivation text generated by the teacher model during the execution of the simulation analysis task, the code execution instructions sent to the interactive data environment, and the intermediate result data fed back by the interactive data environment; S204, the task description of the simulation analysis task, the thought chain derivation text, the code execution instructions, and the intermediate result data are structurally concatenated according to the temporal logic to generate a training trajectory set.
[0048] In this embodiment, the data-driven trajectory synthesis framework first relies on the task design stage when generating the training trajectory set. Task design is based on the business scenario characteristics in the target domain knowledge base, extracting information such as business object types, data source structures, analysis target categories, and compliance restrictions from the knowledge base to form a set of attributes describing the scenario. Based on these attributes, three interrelated dimensions are introduced when constructing task templates: a basic capability dimension to cover data retrieval, single-indicator calculation, and simple comparison judgment; a specialized analysis dimension to cover multi-indicator linkage analysis under a single business theme, such as risk factor decomposition and business process bottleneck identification; and a comprehensive research dimension to cover comprehensive reasoning needs across scenarios, multiple time scales, and cross data sources. Each task template defines the analysis target, input data type, available toolkit, expected intermediate conclusion structure, and final output form through fields, and sets difficulty level markers for each of the three dimensions. Based on these templates, the system generates a progressively increasing sequence of task difficulty levels from low to high, ensuring that each task in the sequence forms a monotonically increasing relationship in terms of complexity, dependent data structures, and compliance requirement coverage, thus providing a clear hierarchical path for subsequent course-based training.
[0049] In the trajectory generation and execution phase, the data-driven trajectory synthesis framework drives the teacher model to process tasks in a progressively more difficult task sequence through a control interface. For each task, the framework instantiates a task template to obtain a specific task description text and a structured task configuration, including data source identifiers, a list of callable analysis components, and necessary compliance check items, and passes this content into the teacher model context. When the teacher model executes the simulation analysis task with the support of the interactive data environment, it generates inference text step by step around the task description and triggers environment calls through embedded call markers. In this process, the framework acts as a scheduler and recorder. On the one hand, it parses the data interfaces or code analysis units that need to be called based on the teacher model output and sends code execution instructions to the interactive data environment. The instructions include the target data source, query parameters, running analysis component identifiers, and execution configurations. On the other hand, it receives intermediate result data from the interactive data environment, including raw query results, derived indicators, diagnostic information, and error messages, and writes these results back to the teacher model context so that subsequent inference can continue.
[0050] The organization of the thought chain derivation text, code execution instructions, and intermediate result data on the timeline is completed by the trajectory synthesis stage. The system maintains a time-series log for each simulation analysis task, attaching a strict, monotonous timestamp and association identifier to each teacher model output, each environment call request, and each environment response. The thought chain derivation text is broken down into multiple inference units in the order of generation, each unit corresponding to a clear analytical intent or decision-making transition, such as proposing a hypothesis, selecting a data source, or interpreting results. The code execution instructions record the tool name, input parameters, and call context in the order of invocation. The intermediate result data records a data content summary, data format description, and association with the previous inference unit in the order of return. During the structured trajectory assembly, the system starts with the task description, connecting the inference units, instruction records, and result records according to time sequence and causal relationships. A mapping is established through reference tags to "which instruction is triggered by a certain inference" and "which inference is supported by a certain result," ultimately generating a unified trajectory representation. This representation can adopt a hierarchical structure, encoding task-level meta-information, sequence-level temporal information, and node-level content information separately, so that the entire trajectory can be directly used as a supervision object in subsequent training phases.
[0051] When generating the training trajectory set, the system does not retain only a single execution record, but collects multiple simulation analysis processes for each task in the progressively more difficult task sequence. Diversity can be introduced by changing the random sampling temperature of the teacher model, adjusting the combination of available tools in the interactive data environment, and switching some data source configurations, so that the same task produces different but reasonable thought chains and operation sequences in different execution rounds. All trajectories undergo basic quality checks, such as whether they meet the task requirements, whether there are obvious logical breaks, and whether they contain incomplete environment calls. Qualified trajectories are written into the training trajectory set. Each record in the set is labeled with a task identifier, difficulty level marker, and task dimension tag, facilitating batch use by difficulty and dimension during the training phase, thereby completing the construction of trajectories covering the progressively more difficult task sequence.
[0052] Through the above steps, this embodiment can obtain a set of training trajectories that closely match the real business analysis process in terms of difficulty level, analysis path, and tool usage. This enables subsequent learning based on these trajectories to be no longer limited to the final conclusion, but to accurately inherit the hierarchical task design, reasoning process organization, and tool calling rhythm, thereby significantly improving the convergence efficiency and generalization ability of the domain-independent data analysis agent on complex tasks.
[0053] In one embodiment, step S40 above includes: S401, the task description text of the current level task is converted into a structured instruction containing tool call specifications, and transmitted to the domain data autonomous analysis agent through the input interface of the interactive data environment; S402, The autonomous analysis agent monitoring the data within the domain generates task decomposition logic and reasoning steps based on the structured instructions; S403, capture and record the tool call request initiated by the autonomous data analysis agent in the domain to the interactive data environment and the execution response data fed back by the interactive data environment; S404, the task decomposition logic, the thought process steps, the tool call request and the execution response data are combined in time sequence to form an analysis operation trajectory; S405, extract the final delivery content generated by the autonomous data analysis agent in the domain when the task execution ends and output it as the analysis result.
[0054] In this embodiment, in the interactive data environment, the current-level task is first introduced in the form of task description text. The task description text can be understood as a natural language or semi-structured description of a specific analysis objective, including task background, analysis objectives, key constraints, priority hints, etc. To enable the autonomous data analysis agent within the domain to call tools and execute a replayable analysis process according to established specifications, the task description text needs to be converted into structured instructions. The structured instructions explicitly provide tool calling specifications, which include identifiers of usable tools, field definitions of input and output parameters for each tool, calling order constraints, error handling options, and tagging information associated with compliance rules of the target domain. The conversion process can be achieved through a combination of template parsing, intent recognition, and slot filling. For example, first, a domain intent classification model is used to identify the task type, then a corresponding structured template is matched based on the task type. The object names, time intervals, indicator names, etc., appearing in the task text are filled into the template fields to generate an instruction object containing the tool calling specifications. The generated structured instructions are transmitted to the domain-independent data autonomous analysis agent through the input interface of the interactive data environment. This input interface can be a unified API endpoint, message queue channel, or session context injection mechanism. During transmission, task numbers and context identifiers are attached to facilitate subsequent tracking.
[0055] Upon receiving structured instructions, the autonomous data analysis agent within the domain generates task decomposition logic and reasoning steps based on the task configuration specified in the instructions. The task decomposition logic can be understood as a hierarchical structure that breaks down the overall analysis task into several sub-tasks. Each sub-task corresponds to a clear analysis objective and a set of required tools; for example, data preprocessing might be completed first, followed by single-indicator evaluation, and then multi-dimensional comprehensive judgment. The reasoning steps are the text of the reasoning process generated around the task decomposition logic, progressively recording the analysis intent, assumptions, reasons for data selection, and transitional statements to conclusions. An interactive data environment monitors this process, parsing and buffering newly generated task decomposition and reasoning fragments in real time by accessing the agent's output stream or intermediate callback interface, providing time-ordered content for subsequent trajectory generation. The monitoring process allows setting a minimum time granularity and a maximum buffer length, segmenting long texts, and adding timestamps and source identifiers to each fragment to distinguish whether it originates from task decomposition or reasoning.
[0056] During execution, the autonomous data analysis agent within the domain initiates a tool invocation request to the interactive data environment according to the tool invocation specifications in the structured instructions. The tool invocation request includes at least the target tool name or identifier, a list of input parameters, references to the output data from the previous step, and a description of the purpose of the invocation. Based on the tool invocation request, the interactive data environment schedules the corresponding data interface or code execution unit to perform the actual calculations, such as running indicator calculation scripts, querying external databases, or triggering risk control model inference, and encapsulates the execution results into execution response data. The execution response data includes the original result data, necessary statistical summaries, error codes and error messages, execution time, and association markers with the task context. The system captures and records each tool invocation request and its corresponding execution response data through a unified listening channel, establishing a one-to-one correspondence between requests and responses during storage and appending sequence numbers and timestamps to ensure accurate reconstruction of the invocation order and dependencies later.
[0057] During the trajectory organization phase, the system combines the task breakdown logic, reasoning steps, tool call requests, and execution response data collected in the previous phase into an analysis operation trajectory according to time sequence and causal relationships. The combination process can begin by establishing trajectory containers at the task number level, then sorting all content by timestamp, and finally creating a graph structure or sequence structure using request identifiers and context references. For each tool call, the preceding and following reasoning segments are bound to it, forming a local sub-chain of "proposing analysis intent – initiating tool call – receiving execution response – updating reasoning conclusion," with the entire trajectory consisting of multiple interconnected sub-chains. The trajectory structure can simultaneously retain text content, structured metadata, and inter-call dependencies to meet the needs of subsequent training at different granularities. After trajectory construction is complete, the system extracts the final delivery content from the domain-specific data autonomous analysis agent's output stream at the end of the task. This content is typically a summary output for the current task level, such as an analysis report, a list of decision recommendations, or a structured conclusion table. This final delivery content is stored in association with the aforementioned analysis operation trajectory and is used as part of the analysis results output, providing a benchmark for subsequent quality analysis.
[0058] Through the above steps, this embodiment can obtain a complete behavioral record covering input, inference, tool calls, and output without changing the actual business analysis process. This allows subsequent training and evaluation stages to utilize both result-level information and fine-grained signals at the process level, thereby improving trajectory supervision quality and enhancing the ability of the domain-independent data autonomous analysis agent to learn complex task execution paths and tool usage strategies.
[0059] In one embodiment, step S50 above includes: S501, parse the logical deduction steps in the analysis operation trajectory, and determine the logical validity of the logical deduction steps according to the preset analysis paradigm of the target domain; S502, extract key indicator data from the analysis results output, and compare the key indicator data with the benchmark fact data in the interactive data environment. S503, Perform text semantic scanning on the analysis results output to detect whether there is sensitive content that violates the compliance requirements of the target domain; S504, based on the result of the logical validity determination, the numerical comparison result of the key indicator data, and the detection result of the sensitive content, a reward signal in scalar numerical form is generated.
[0060] In this embodiment, when the analysis operation trajectory and analysis result output enter the quality analysis stage, the system first extracts the logical deduction steps from the analysis operation trajectory. The analysis operation trajectory has recorded the task breakdown, tool calls, execution responses, and reasoning text in chronological order. The logical deduction steps can be distinguished by tagging information or specific fields, such as in the form of thought chain text, decision node descriptions, and conclusion transition statements. In implementation, a separate trajectory field can be set for logically related content, and type tags can be added when the trajectory is generated. Alternatively, during the quality analysis stage, segments that perform reasoning functions can be filtered out through key phrase patterns, structured tags, or segmented coding results. The extracted logical deduction steps are organized into a directed sequence or tree structure, with each node containing preconditions, referenced data sources, and obtained local conclusions, facilitating subsequent checks on consistency and completeness.
[0061] The pre-defined analytical paradigm for the target domain provides a reference for determining logical validity. This paradigm is derived from domain expert experience and existing business processes. The pre-defined analytical paradigm can be designed as a set of rules, a process template, or a graph structure. For example, in a risk assessment task, it requires first completing a data integrity check, then performing a single-indicator evaluation, and finally executing a comprehensive score and conclusion explanation. In a efficacy assessment task, it requires first confirming the consistency of sample grouping and baseline, then analyzing changes in key indicators, and finally providing an explanation of adverse events. During quality analysis, the system maps the logical deduction steps onto these pre-defined templates, checking whether necessary steps are covered, whether conclusions are used without proper understanding, and whether there are any inconsistencies. In implementation, graph matching, finite state machines, or constraint checkers can be used to encode the type of each inference node as a state, judge the legality of state transitions, record missing nodes or abnormal transitions, generate logical validity determination results, and include quantitative indicators such as coverage and number of conflicts.
[0062] The analysis results output section contains the final conclusions and key quantitative results after task execution. Key indicator data is extracted from the analysis results output, typically including returns, volatility, probability of default, test statistics, confidence interval boundaries, and clinical endpoint values, depending on the target domain definition. To ensure automatic extraction, these indicators can be output in a structured format during the results generation stage, such as using key-value pairs, table fields, or tokenized segments. During quality analysis, numerical values and units are parsed into standard internal representations through field names, indicator labels, or templated sentences, and a connection is established with the benchmark fact data in the interactive data environment. The benchmark fact data in the interactive data environment comes from authoritative data sources or validated data warehouses, such as historical market databases, macroeconomic statistical databases, and registry study data tables. The system locates the corresponding benchmark record based on the task context, and then compares the key indicator data with the benchmark values. It can calculate absolute error, relative error, and interval inclusion relationships according to business needs, such as determining whether the interval estimate in the analysis conclusion covers the benchmark estimate range. During the comparison process, thresholds and weights can be set to differentiate the importance of different indicators, obtaining numerical comparison results and corresponding scores.
[0063] Text semantic scanning performs content security and compliance checks on the natural language portion of the analysis output. Target domain compliance requirements include regulatory clauses, internal risk control standards, privacy protection requirements, and ethical guidelines. This can be achieved by combining dictionary matching, pattern recognition, and deep semantic classification models to identify potentially sensitive content. For example, in the financial sector, it can detect undisclosed investment advice, misleading profit promises, and descriptions of high-risk products without clearly stated risks. In the healthcare sector, it can detect unauthorized patient details, recommendations that do not comply with medication guidelines, and claims of efficacy exceeding the permitted scope. The system can assign weights to each type of sensitive content and record the sensitivity type, severity level, and context of each hit during the scan, thereby generating sensitive content detection results, including both Boolean judgments and quantitative sensitivity scores.
[0064] The results of logical validity assessment, numerical comparison of key indicator data, and sensitive content detection are ultimately aggregated into a scalar numerical reward signal. To achieve this transformation, the system needs to construct a mapping function to unify the quality evaluation across the three dimensions into the same numerical range. For example, normalization can be used to compress the score of each dimension to between zero and one, followed by a linear combination with preset weights. Alternatively, a piecewise function can be used to set strong penalties for serious logical errors and serious compliance violations. Logical validity can be mapped to a combination of coverage score and conflict penalty, numerical comparison can be mapped to the inverse score of the error function, and sensitive content detection can be mapped to a penalty function for sensitive content hit rate. The system can also introduce non-linear compression during the combination process, applying milder penalties for minor issues and setting sharply reduced reward values for significant deviations from facts or violations of high-risk compliance clauses. The resulting scalar reward signal is bound to the current level task instance and the corresponding analysis operation trajectory, providing a unified optimization objective for updating network parameters during subsequent training phases.
[0065] Through the above steps, this embodiment can simultaneously encode the inference quality, reliability of the result data, and compliance level of the content in a single numerical feedback. This allows the subsequent training phase to guide the autonomous analysis agent in the domain data to optimize in three aspects simultaneously: maintaining the integrity of the inference structure, the credibility of the quantification conclusions, and the compliance of the text output. This significantly improves the usability and controllability of the agent in complex analysis tasks.
[0066] In one embodiment, step S60 above includes: S601, retrieve and extract demonstration data that has a mapping relationship with the current level task from the training trajectory set; S602, input the demonstration data into the supervisory loss function to determine the difference loss between the generation probability distribution of the domain data autonomous analysis agent and the demonstration data, and define the difference loss as prediction bias; S603, using a reinforcement learning value network to analyze the domain data, the autonomous analysis agent determines the expected reward in the current state, and determines the policy gradient based on the advantage difference between the reward signal and the expected reward; S604, based on the prediction bias and the policy gradient, perform gradient-based backpropagation update on the network parameters of the autonomous data analysis agent in the domain.
[0067] In this embodiment, when updating the network parameters of the autonomous analysis agent within the domain using reward signals and the training trajectory set, it is first necessary to retrieve demonstration data that has a mapping relationship with the current level task from the training trajectory set. The training trajectory set is a structured trajectory accumulated by the teacher model when executing a progressively more difficult task sequence in an interactive data environment. Each trajectory contains task description, thought process text, tool call records, and intermediate results. The current level task has already written trajectory metadata in the form of task identifier, task type, and business scenario tags when generating the training trajectory set. Therefore, strategies such as task identifier matching, task type filtering, and scenario tag intersection can be used to select several trajectories from the training trajectory set that are consistent with the current level task in terms of objectives, input distribution, and constraints, thus obtaining a set of demonstration data. The demonstration data can exist in the form of complete trajectories, or key segments can be extracted from complete trajectories, such as high-quality thought process text and corresponding tool call sequences, to guide the agent to generate better behavior.
[0068] After retrieving the demonstration data, it needs to be input into the supervised loss function to calculate the difference loss between the generation probability distribution of the autonomous analysis agent within the domain and the demonstration data. Here, the generation probability distribution refers to the probability allocation given by the agent to output units such as the next action, the next label, and the next tool invocation command under the given input conditions of the current task level. For example, in a text generation scenario, the generation probability distribution could be the conditional probability of each word in the vocabulary; in a tool invocation scenario, it could be the probability of candidate tools and their parameter combinations. To align the agent's behavior with the demonstration data, the demonstration data can be considered as the target distribution, and the agent's output as the prediction distribution. A supervised loss function can be constructed using cross-entropy, least squares, or other metrics to make the agent tend to output results consistent with the demonstration at the demonstration location. The difference loss reflects the degree of deviation between the generation probability distribution and the demonstration data. This difference is defined as the prediction bias, which helps distinguish between errors caused by imitation and biases caused by policy exploration when subsequently combined with reinforcement learning signals.
[0069] When relying solely on imitation, agents may become overly dependent on demonstrations and lack the ability to adapt to different task instances. Therefore, reinforcement learning value networks are needed. A reinforcement learning value network takes the current state as input and outputs the agent's expected reward in that state. The current state can be represented as a vector by encoding the task description of the current level task, historical interaction records, tool invocation context, and environmental feedback. This vector is then mapped to a scalar expected reward by the value network through multiple nonlinear transformations. The value network can share some language encoding layers or be built separately to improve its ability to express the reward structure. Based on the previously generated reward signal, the actual reward obtained can be calculated. Comparing the reward signal with the expected reward yields an advantage difference, which characterizes the superiority or inferiority of the current action relative to the value network's estimate. If the advantage difference is positive, it indicates that the current policy performs better than the value network's expectation in that state, requiring positive reinforcement during policy updates. If the advantage difference is negative, it is necessary to suppress the current policy from repeating this type of output in similar states.
[0070] To combine the prediction bias from the supervised direction and the advantage difference from the reinforcement learning direction for updating network parameters, a joint gradient signal needs to be constructed. The prediction bias can be obtained by taking the gradient of the supervised loss function with respect to the network parameters, while the advantage difference can be obtained by multiplying the policy gradient method by the derivative with respect to the log probability of the action. Numerically, these two are weighted and balanced using weight coefficients to control the relative strength of imitation learning and policy optimization. Network parameters include the weights of each layer of the language model, policy head parameters, and value network parameters. These parameters determine the specific values of the generation probability distribution and expected reward in the forward computation. After constructing the joint loss, the system uses an automatic differentiation mechanism to differentiate the entire computation graph, obtaining the gradient-based parameter update direction. The backpropagation update process proceeds from top to bottom according to the network structure, propagating the gradient of the loss with respect to the output layer to the input side layer by layer, and numerically adjusting the weight matrix, bias vector, and normalization coefficients of each layer. To ensure the stability of updates, techniques such as learning rate control, gradient pruning, and parameter regularization can be introduced to ensure that the adjustment of network parameters can respond to the information provided by reward signals and demonstration data, without causing excessive oscillations in a single training round.
[0071] In the aforementioned update process, prediction bias focuses on narrowing the gap between the agent's output and high-quality demonstrations, while advantage difference focuses on amplifying effective exploration using reward feedback from the interactive environment. Both work together on the network parameters, enabling the autonomous data analysis agent within the domain to gradually develop a strategy that both conforms to the teacher's behavioral patterns and yields high rewards in the real-world environment, based on the data distribution corresponding to the current level of the task. By continuously repeating this process of retrieving demonstrations, calculating prediction bias, estimating expected returns, constructing advantage difference, and performing backpropagation updates, the network parameters converge towards a more stable and efficient analytical behavior, driven by training data and interactive feedback.
[0072] This embodiment retrieves demonstration data that maps to the current task level from the training trajectory set, calculates the prediction bias between the generated probability distribution and the demonstration data using a supervised loss function, and then uses a reinforcement learning value network to provide the expected reward and construct an advantage difference based on the reward signal. Both types of information are updated and uniformly applied to the network parameters through gradient-based backpropagation. This allows for the simultaneous imitation of high-quality demonstration behaviors and reinforcement of high-reward strategies during a single training process. This enables the autonomous data analysis agent within the domain to gradually improve its decision-making quality and environmental adaptability on the current task level while maintaining interpretability and behavioral stability in the analysis process, thereby reducing reliance on manual rule adjustments and simple offline fitting.
[0073] In one embodiment, step S70 above includes: S701, During the cyclic execution process, an evaluation time window is set to statistically analyze the average reward score and task success rate of the autonomous analysis agent in the domain when performing the current level task; S702, compare the average reward score and the task success rate with a preset set of promotion thresholds to determine whether the domain data autonomous analysis agent meets the capability assessment conditions. S703, if the ability assessment conditions are met and the current level task is not at the highest level of the progressive difficulty level task sequence, then the next level task in the progressive difficulty level task sequence is retrieved as the new current level task, and the steps of task input, result acquisition, reward generation and parameter update are continuously executed in a loop. S704, if the capability assessment conditions are met and the current level task is at the highest level of the progressive difficulty level task sequence, then the domain data autonomous analysis agent is determined to have passed the highest level verification, the parameter update step is stopped and the network parameters of the domain data autonomous analysis agent are locked, and the trained domain data autonomous analysis agent is obtained.
[0074] In this embodiment, when the autonomous data analysis agent performs task promotion and iterative training, a training loop control logic needs to be established first, running around a sequence of tasks with progressively increasing difficulty levels. During training, the current level task has already been set in the previous stage, and the four processing stages—task input, result acquisition, reward generation, and parameter update—can be repeatedly invoked. To avoid deviations in promotion judgment caused by fluctuations in single task performance, an evaluation time window is introduced within the training loop to aggregate performance data over a period of time. The evaluation time window can be defined in units of the number of completed task rounds, such as statistically analyzing the execution results of several consecutive current level tasks, or it can be configured in units of training rounds or environmental interaction steps, loaded through a parameter configuration file or training control module, ensuring that the evaluation granularity is adjustable under different training environments.
[0075] Within the evaluation time window, for each execution of a task at the current difficulty level, the training control module collects a single reward value from the reward generation stage and a marker indicating whether the task was successfully completed from the result acquisition stage. The reward value can be a scalar reward signal output from the quality analysis stage, and the task success marker can be determined based on whether the analysis results meet pre-defined conditions for correctness, compliance, and completeness. As the number of task executions within the window increases, the system accumulates the sum of reward values and the number of successful tasks, which are then divided by the total number of tasks to calculate the average reward score and the task success rate. The average reward score reflects the overall quality level of the agent's performance at the current difficulty level, while the task success rate reflects the stability of task completion under real-world business constraints; both constitute the two fundamental indicators for capability evaluation.
[0076] After obtaining the average reward score and task success rate, they need to be compared with a preset set of promotion thresholds. The promotion threshold set can be stratified in the training configuration according to difficulty level, with each level corresponding to a set of thresholds, such as the minimum acceptable average reward, minimum acceptable success rate, etc., and other statistical indicator thresholds can also be added. The comparison process can be completed by the evaluation module, which matches the average reward score and task success rate calculated for the current level of task within the evaluation time window with the corresponding threshold intervals. If both indicators reach or exceed the corresponding threshold intervals, a mark indicating that the ability evaluation conditions are met is generated; otherwise, it is considered that the current level of task has not been mastered, and iterative training at that level needs to continue. Through this threshold set design, the training process can apply differentiated promotion standards at different difficulty levels.
[0077] When the capability assessment conditions are met and the current level task is not the highest level in the progressive difficulty level task sequence, the training control module needs to perform task promotion processing. Promotion processing involves finding the index position in the progressive difficulty level task sequence, shifting the current level task one position to the right, reading the next level task from the sequence, and replacing the current level task with this task. After the task replacement is completed, the loop control logic remains unchanged, still sequentially executing training steps such as task input, result acquisition, reward generation, and parameter updates. However, the task content, data distribution, and constraints become more complex as the level increases, prompting the agent to continue learning at the new difficulty level.
[0078] When the capability assessment conditions are met and the current task level is already at the highest level in the progressive difficulty task sequence, it indicates that the agent has met the preset capability requirements for all levels of tasks in the sequence. At this point, the assessment module sends a signal to the training control module indicating that the highest level verification has been passed. The training control module no longer triggers new parameter update processing but instead performs a network parameter locking operation. Network parameter locking can be achieved by freezing the trainable flag of the model weights, stopping optimizer updates, and persisting a snapshot of the current parameters, saving the parameter states of the current language model's main body, policy part, and value estimation part as a stable set of parameters. In this way, the training process naturally terminates, resulting in a trained agent capable of autonomously analyzing in-domain data. This agent can then be directly integrated into the inference environment without further parameter drift during online operation.
[0079] This embodiment introduces an evaluation time window into the iterative training process to statistically analyze the average reward score and task success rate, and compares them with a set of promotion thresholds configured at different levels. This closely links the capability evaluation conditions with a progressively more difficult task sequence. Combined with conditional branch control for switching to the next level of tasks and network parameter locking when the highest level is passed, the autonomous data analysis agent in the domain can automatically complete the gradual promotion from low to high difficulty during the training process. This avoids overtraining on simple tasks and prevents premature promotion to high-difficulty tasks that could lead to insufficient capability. Furthermore, by stopping parameter updates and fixing the current parameter state after passing the highest level verification, a trained agent with stable performance and clear capability boundaries across the entire task sequence is obtained, providing a reliable foundation for subsequent deployment in real-world business scenarios.
[0080] In one embodiment, a data analysis apparatus based on progressive task learning is provided, which corresponds one-to-one with the data analysis method based on progressive task learning described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the data analysis device based on progressive task learning of the present invention. The modules include: interactive environment construction module 10, trajectory synthesis and generation module 20, agent initialization module 30, task execution acquisition module 40, quality assessment module 50, parameter update module 60, level promotion training module 70, and target task analysis module 80. Detailed descriptions of each functional module are as follows: Interactive environment construction module 10 is used to build an interactive data environment and a data-driven trajectory synthesis framework based on the teacher model and the target domain knowledge base; The trajectory synthesis and generation module 20 is used to generate a training trajectory set covering a progressive difficulty level task sequence using the data-driven trajectory synthesis framework. The agent initialization module 30 is used to initialize the basic language model as an autonomous data analysis agent in the domain, and to set the initial difficulty level task in the progressive difficulty level task sequence as the current level task. The task execution acquisition module 40 is used to input the current level task into the domain data autonomous analysis agent in the interactive data environment, and to acquire the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment. Quality assessment module 50 is used to perform quality analysis on the analysis operation trajectory and the analysis result output, and generate a reward signal; Parameter update module 60 is used to update the network parameters of the domain data autonomous analysis agent based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; The level promotion training module 70 is used to perform task promotion and iterative training process on the autonomous data analysis agent in the domain. It cyclically executes the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level verification of the progressive difficulty level task sequence, and obtains the trained autonomous data analysis agent in the domain. The target task analysis module 80 is used to acquire the target analysis task to be processed, and to use the trained domain data autonomous analysis agent to perform an autonomous analysis process for the target analysis task to be processed in the interactive data environment to obtain the target analysis result.
[0081] In one embodiment, the interactive environment construction module 10 is specifically used for: Integrate multi-source heterogeneous data reading interfaces and executable code analysis tool libraries, and deploy them to a virtual sandbox container to obtain an interactive data environment; Semantic parsing and domain analysis logic extraction are performed on unstructured industry expert reports and compliance guidelines documents to construct a structured target domain knowledge base; The pre-trained language model is instantiated by using the domain analysis logic and compliance boundary conditions in the target domain knowledge base to provide system-level prompts for the teacher model. A two-way automated communication protocol is established between the teacher model and the interactive data environment, enabling the teacher model, configured with system-level prompts, to schedule the executable code analysis tool library in the virtual sandbox container, process data obtained from multi-source heterogeneous data reading interfaces, and make decisions based on the domain analysis logic and compliance boundary conditions in the target domain knowledge base through the two-way automated communication protocol; A data-driven trajectory synthesis framework is obtained by encapsulating the bidirectional automated communication protocol.
[0082] In one embodiment, the trajectory synthesis and generation module 20 is specifically used for: Based on the business scenario characteristics in the target domain knowledge base, a task template is designed that includes basic capability dimensions, specialized analysis dimensions, and comprehensive research dimensions, and a progressive difficulty level task sequence is generated based on the task template; The teacher model in the data-driven trajectory synthesis framework is controlled to sequentially execute the simulation analysis tasks in the progressive difficulty level task sequence in the interactive data environment; Record the thought chain derivation text generated by the teacher model during the execution of the simulation analysis task, the code execution instructions sent to the interactive data environment, and the intermediate result data fed back by the interactive data environment; The task description of the simulation analysis task, the thought chain derivation text, the code execution instructions, and the intermediate result data are concatenated in a time-series logical structure to generate a training trajectory set.
[0083] In one embodiment, the task execution acquisition module 40 is specifically used for: The task description text of the current level task is converted into a structured instruction containing tool call specifications, and transmitted to the domain-independent data autonomous analysis agent through the input interface of the interactive data environment; The autonomous analysis agent monitoring the data within the domain generates task decomposition logic and reasoning steps based on the structured instructions; Capture and record tool call requests initiated by the autonomous data analysis agent within the domain to the interactive data environment, as well as the execution response data fed back by the interactive data environment; The task decomposition logic, the thought process steps, the tool call request, and the execution response data are combined in chronological order to form an analysis operation trajectory. Extract the final delivery content generated by the autonomous analysis agent within the domain when it finishes task execution, and output it as the analysis result.
[0084] In one embodiment, the quality assessment module 50 is specifically used for: The logical deduction steps in the analysis operation trajectory are analyzed, and the logical validity of the logical deduction steps is determined according to the preset analysis paradigm of the target domain. Extract key indicator data from the analysis results output, and compare the key indicator data with the baseline fact data in the interactive data environment. The analysis results are subjected to text semantic scanning to detect whether there is any sensitive content that violates the compliance requirements of the target domain; Based on the results of the logical validity determination, the numerical comparison results of the key indicator data, and the detection results of the sensitive content, a reward signal in scalar numerical form is generated.
[0085] In one embodiment, the parameter update module 60 is specifically used for: Retrieve and extract exemplary data that are mapped to the current level task from the training trajectory set; The demonstration data is input into the supervision loss function to determine the difference loss between the generation probability distribution of the autonomous data analysis agent in the domain and the demonstration data, and the difference loss is defined as the prediction bias. The agent autonomously analyzes the expected reward of the domain data using a reinforcement learning value network, and determines the policy gradient based on the advantage difference between the reward signal and the expected reward. Based on the prediction bias and the policy gradient, gradient-based backpropagation updates are performed on the network parameters of the autonomous data analysis agent within the domain.
[0086] In one embodiment, the rank promotion training module 70 is specifically used for: During the cyclic execution process, an evaluation time window is set to statistically analyze the average reward score and task success rate of the autonomous analysis agent in the domain when performing the current level task. The average reward score and the task success rate are compared with a preset set of promotion thresholds to determine whether the autonomous data analysis agent in the domain meets the capability assessment conditions. If the ability assessment conditions are met and the current level task is not at the highest level in the progressive difficulty level task sequence, then the next level task in the progressive difficulty level task sequence is retrieved as the new current level task, and the steps of task input, result acquisition, reward generation and parameter update are continuously executed in a loop. If the capability assessment conditions are met and the current level task is at the highest level in the progressive difficulty level task sequence, then the domain data autonomous analysis agent is determined to have passed the highest level verification, the parameter update step is stopped, and the network parameters of the domain data autonomous analysis agent are locked, thus obtaining the trained domain data autonomous analysis agent.
[0087] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a data analysis method based on progressive task learning.
[0088] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a data analysis method based on progressive task learning.
[0089] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Construct an interactive data environment and a data-driven trajectory synthesis framework based on a teacher model and a target domain knowledge base; The data-driven trajectory synthesis framework is used to generate a training trajectory set covering a sequence of tasks with progressively increasing difficulty levels. The basic language model is initialized as an autonomous data analysis agent within the domain, and the initial difficulty level task in the progressive difficulty level task sequence is set as the current level task; In the interactive data environment, the current level task is input into the domain data autonomous analysis agent, and the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment are obtained; A quality analysis is performed on the analysis operation trajectory and the analysis result output to generate a reward signal; The network parameters of the domain data autonomous analysis agent are updated based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; The autonomous data analysis agent in the domain is subjected to a task promotion and iterative training process, which involves repeatedly executing the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level of verification in the progressive difficulty level task sequence, thus obtaining the trained autonomous data analysis agent in the domain. The target analysis task to be processed is obtained, and the autonomous analysis agent with the trained domain data is used to perform an autonomous analysis process for the target analysis task in the interactive data environment to obtain the target analysis result.
[0090] In one embodiment, a non-volatile computer-readable storage medium is provided, which may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, it performs the following steps: Construct an interactive data environment and a data-driven trajectory synthesis framework based on a teacher model and a target domain knowledge base; The data-driven trajectory synthesis framework is used to generate a training trajectory set covering a sequence of tasks with progressively increasing difficulty levels. The basic language model is initialized as an autonomous data analysis agent within the domain, and the initial difficulty level task in the progressive difficulty level task sequence is set as the current level task; In the interactive data environment, the current level task is input into the domain data autonomous analysis agent, and the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment are obtained; A quality analysis is performed on the analysis operation trajectory and the analysis result output to generate a reward signal; The network parameters of the domain data autonomous analysis agent are updated based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; The autonomous data analysis agent in the domain is subjected to a task promotion and iterative training process, which involves repeatedly executing the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level of verification in the progressive difficulty level task sequence, thus obtaining the trained autonomous data analysis agent in the domain. The target analysis task to be processed is obtained, and the autonomous analysis agent with the trained domain data is used to perform an autonomous analysis process for the target analysis task in the interactive data environment to obtain the target analysis result.
[0091] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0093] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0094] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.
Claims
1. A data analysis method based on progressive task learning, characterized in that, Includes the following steps: Construct an interactive data environment and a data-driven trajectory synthesis framework based on a teacher model and a target domain knowledge base; The data-driven trajectory synthesis framework is used to generate a training trajectory set covering a sequence of tasks with progressively increasing difficulty levels. The basic language model is initialized as an autonomous data analysis agent within the domain, and the initial difficulty level task in the progressive difficulty level task sequence is set as the current level task; In the interactive data environment, the current level task is input into the domain data autonomous analysis agent, and the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment are obtained; A quality analysis is performed on the analysis operation trajectory and the analysis result output to generate a reward signal; The network parameters of the domain data autonomous analysis agent are updated based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set; The autonomous data analysis agent in the domain is subjected to a task promotion and iterative training process, which involves repeatedly executing the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level of verification in the progressive difficulty level task sequence, thus obtaining the trained autonomous data analysis agent in the domain. The target analysis task to be processed is obtained, and the autonomous analysis agent with the trained domain data is used to perform an autonomous analysis process for the target analysis task in the interactive data environment to obtain the target analysis result. 2.The data analysis method based on progressive task learning according to claim 1, wherein, An interactive data environment is constructed, and a data-driven trajectory synthesis framework based on a teacher model and a target domain knowledge base is built, including: Integrate multi-source heterogeneous data reading interfaces and executable code analysis tool libraries, and deploy them to a virtual sandbox container to obtain an interactive data environment; Semantic parsing and domain analysis logic extraction are performed on unstructured industry expert reports and compliance guidelines documents to construct a structured target domain knowledge base; The pre-trained language model is instantiated by using the domain analysis logic and compliance boundary conditions in the target domain knowledge base to provide system-level prompts for the teacher model. A two-way automated communication protocol is established between the teacher model and the interactive data environment, enabling the teacher model, configured with system-level prompts, to schedule the executable code analysis tool library in the virtual sandbox container, process data obtained from multi-source heterogeneous data reading interfaces, and make decisions based on the domain analysis logic and compliance boundary conditions in the target domain knowledge base through the two-way automated communication protocol; A data-driven trajectory synthesis framework is obtained by encapsulating the bidirectional automated communication protocol.
3. The data analysis method based on progressive task learning as described in claim 1, characterized in that, The data-driven trajectory synthesis framework is used to generate a set of training trajectories covering a sequence of tasks with progressively increasing difficulty levels, including: Based on the business scenario characteristics in the target domain knowledge base, a task template is designed that includes basic capability dimensions, specialized analysis dimensions, and comprehensive research dimensions, and a progressive difficulty level task sequence is generated based on the task template; The teacher model in the data-driven trajectory synthesis framework is controlled to sequentially execute the simulation analysis tasks in the progressive difficulty level task sequence in the interactive data environment; Record the thought chain derivation text generated by the teacher model during the execution of the simulation analysis task, the code execution instructions sent to the interactive data environment, and the intermediate result data fed back by the interactive data environment; The task description of the simulation analysis task, the thought chain derivation text, the code execution instructions, and the intermediate result data are concatenated in a time-series logical structure to generate a training trajectory set.
4. The data analysis method based on progressive task learning as described in claim 1, characterized in that, In the interactive data environment, the current-level task is input into the domain-specific autonomous data analysis agent, and the analysis operation trajectory and analysis result output generated by the domain-specific autonomous data analysis agent when executing the current-level task in the interactive data environment are obtained, including: The task description text of the current level task is converted into a structured instruction containing tool call specifications, and transmitted to the domain-independent data autonomous analysis agent through the input interface of the interactive data environment; The autonomous analysis agent monitoring the data within the domain generates task decomposition logic and reasoning steps based on the structured instructions; Capture and record tool call requests initiated by the autonomous data analysis agent within the domain to the interactive data environment, as well as the execution response data fed back by the interactive data environment; The task decomposition logic, the thought process steps, the tool call request, and the execution response data are combined in chronological order to form an analysis operation trajectory. Extract the final delivery content generated by the autonomous analysis agent within the domain when it finishes task execution, and output it as the analysis result.
5. The data analysis method based on progressive task learning as described in claim 1, characterized in that, Perform quality analysis on the analysis operation trajectory and the analysis result output to generate a reward signal, including: The logical deduction steps in the analysis operation trajectory are analyzed, and the logical validity of the logical deduction steps is determined according to the preset analysis paradigm of the target domain. Extract key indicator data from the analysis results output, and compare the key indicator data with the baseline fact data in the interactive data environment. The analysis results are subjected to text semantic scanning to detect whether there is any sensitive content that violates the compliance requirements of the target domain; Based on the results of the logical validity determination, the numerical comparison results of the key indicator data, and the detection results of the sensitive content, a reward signal in scalar numerical form is generated.
6. The data analysis method based on progressive task learning as described in claim 1, characterized in that, The network parameters of the domain-specific data autonomous analysis agent are updated based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set, including: Retrieve and extract exemplary data that are mapped to the current level task from the training trajectory set; The demonstration data is input into the supervision loss function to determine the difference loss between the generation probability distribution of the autonomous data analysis agent in the domain and the demonstration data, and the difference loss is defined as the prediction bias. The agent autonomously analyzes the expected reward of the domain data using a reinforcement learning value network, and determines the policy gradient based on the advantage difference between the reward signal and the expected reward. Based on the prediction bias and the policy gradient, gradient-based backpropagation updates are performed on the network parameters of the autonomous data analysis agent within the domain.
7. The data analysis method based on progressive task learning as described in claim 1, characterized in that, The autonomous agent for in-domain data analysis is subjected to a task promotion and iterative training process, which involves repeatedly executing steps of task input, result acquisition, reward generation, and parameter update until the autonomous agent passes the highest level of verification in the progressively more difficult task sequence. This process yields a fully trained autonomous agent for in-domain data analysis, including: During the cyclic execution process, an evaluation time window is set to statistically analyze the average reward score and task success rate of the autonomous analysis agent in the domain when performing the current level task. The average reward score and the task success rate are compared with a preset set of promotion thresholds to determine whether the autonomous data analysis agent in the domain meets the capability assessment conditions. If the ability assessment conditions are met and the current level task is not at the highest level in the progressive difficulty level task sequence, then the next level task in the progressive difficulty level task sequence is retrieved as the new current level task, and the steps of task input, result acquisition, reward generation and parameter update are continuously executed in a loop. If the capability assessment conditions are met and the current level task is at the highest level in the progressive difficulty level task sequence, then the domain data autonomous analysis agent is determined to have passed the highest level verification, the parameter update step is stopped, and the network parameters of the domain data autonomous analysis agent are locked, thus obtaining the trained domain data autonomous analysis agent.
8. A data analysis device based on progressive task learning, characterized in that, The data analysis device based on progressive task learning includes: The interactive environment construction module is used to build an interactive data environment and a data-driven trajectory synthesis framework based on the teacher model and the target domain knowledge base. The trajectory synthesis and generation module is used to generate a set of training trajectories covering a progressively more difficult task sequence using the data-driven trajectory synthesis framework. The agent initialization module is used to initialize the basic language model as an autonomous agent for data analysis within the domain, and to set the initial difficulty level task in the progressive difficulty level task sequence as the current level task. The task execution acquisition module is used to input the current level task into the domain data autonomous analysis agent in the interactive data environment, and to acquire the analysis operation trajectory and analysis result output generated by the domain data autonomous analysis agent when executing the current level task in the interactive data environment. The quality assessment module is used to perform quality analysis on the analysis operation trajectory and the analysis result output, and generate reward signals; The parameter update module is used to update the network parameters of the domain data autonomous analysis agent based on the reward signal and the demonstration data corresponding to the current level task in the training trajectory set. The level promotion training module is used to perform task promotion and iterative training process on the autonomous data analysis agent in the domain. It cyclically executes the steps of task input, result acquisition, reward generation and parameter update until the autonomous data analysis agent in the domain passes the highest level verification of the progressive difficulty level task sequence, and obtains the trained autonomous data analysis agent in the domain. The target task analysis module is used to acquire the target analysis task to be processed, and to use the trained domain data autonomous analysis agent to perform an autonomous analysis process for the target analysis task to be processed in the interactive data environment to obtain the target analysis result.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a data analysis program based on progressive task learning stored in the memory and executable on the processor. When executed by the processor, the data analysis program based on progressive task learning implements the steps of the data analysis method based on progressive task learning as described in any one of claims 1-7.
10. A non-volatile computer-readable storage medium, characterized in that, The storage medium stores a data analysis program based on progressive task learning, which, when executed by a processor, implements the steps of the data analysis method based on progressive task learning as described in any one of claims 1-7.
Citation Information
Cited By
Intelligent agent-based financial service processing method, device, equipment and storage medium
CN122332665A