Large model-based thinking processing method and system
By constructing a large-scale model atomic agent database and a streaming monitoring simulation environment, the problem of the uninterpretability of traditional large-scale models is solved, achieving transparency and real-time risk prediction of large-scale models, improving user trust and system stability, supporting compliance verification and policy optimization in highly interpretable scenarios, and promoting commercial applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU YATUO INFORMATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional large-scale models suffer from insufficient user trust, lack of dynamic risk monitoring, difficulty in meeting compliance requirements in highly interpretable scenarios, and difficulty in dynamic strategy optimization due to the unexplainable nature of the decision-making process, which seriously restricts their commercial application expansion.
By establishing a large-scale atomic agent database, constructing a streaming monitoring simulation environment, loading streaming monitoring risk levels and setting time sampling windows, and performing streaming simulation of the atomic agent's thinking process, we can achieve transparency of decision-making logic and real-time risk prediction. Combined with deep learning, we can construct multi-dimensional mapping relationships and dynamically optimize the strategy model.
It achieves transparency in the large model answer generation path, enhances user trust, identifies abnormal decision risks in real time, meets compliance verification requirements in highly interpretable scenarios, dynamically adjusts resource allocation, avoids performance degradation, and promotes commercial applications.
Smart Images

Figure CN121901073A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model technology, and in particular to a thinking processing method and system based on large models. Background Technology
[0002] In the field of artificial intelligence, large-scale user question-answering models have been widely applied in various scenarios such as finance, healthcare, and intelligent customer service. However, as a "black box" system, the unexplainable nature of the decision-making process of traditional large-scale models has become a core bottleneck restricting their further commercialization. Specifically, when processing user questions, large-scale models cannot show users the specific path to generate answers, such as data retrieval strategies and the weighting of reasoning steps. Users have difficulty understanding the report retrieval logic and the reliability of the answer source, leading to insufficient trust in the results. Secondly, the dynamic performance of large-scale models in different user interaction scenarios and system resource environments lacks real-time monitoring methods, making it impossible to identify abnormal decision-making risks caused by flawed thinking strategies or resource overload in advance. Furthermore, in fields with extremely high interpretability requirements, such as financial risk control and medical diagnosis, customers and regulatory agencies need to clearly trace the thinking process of large-scale models to verify the compliance of the results. The traditional black-box model, because it cannot provide a transparent display of the decision-making chain, leads to customer skepticism about the model's output, seriously hindering the expansion of commercial applications. In addition, the thinking strategies of large-scale models lack a quantitative evaluation system, making it difficult to dynamically optimize strategies based on real-time operating data, which may lead to the accumulation of uncontrollable performance degradation risks in long-term operation.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide a thinking processing method and system based on a large model, which aims to solve the technical problems that the traditional user question-and-answer large model suffers from insufficient user trust, lack of dynamic risk monitoring, difficulty in meeting compliance requirements in highly interpretable scenarios, and difficulty in dynamic optimization of strategies due to the black box nature of the decision-making process. These problems severely restrict its commercial application and expansion.
[0005] To achieve the above objectives, the present invention provides a thinking processing method based on a large model, the method comprising: Establish a large-scale model atomic agent database, which includes basic parameter information of atomic agents, running scenario information, thinking strategy design information, and performance feedback information after the atomic agents run; The operational scenario information includes user interaction scenario information and system resource scenario information, and the risk level of streaming monitoring is assessed based on the user interaction scenario information and system resource scenario information. Based on the basic parameter information and the thinking strategy design information, a streaming monitoring simulation environment is constructed for the atomic agent to simulate the thinking process of the atomic agent. The streaming monitoring risk level is loaded into the streaming monitoring simulation environment, and a time sampling window is set. Based on the time sampling window, several thinking stages of the atomic agent are simulated in streaming mode, and the monitoring evaluation results are obtained through the stage simulation results.
[0006] Optionally, assessing the streaming monitoring risk level based on the user interaction scenario information and system resource scenario information includes: The user load stages are divided according to the user interaction scenario information and sorted from high to low according to the interaction complexity to generate an interaction load sequence. Based on the system resource scenario information, the resource disturbance stages are divided and sorted from high to low according to the resource fluctuation amplitude to generate a resource disturbance sequence; The interactive load sequence and resource disturbance sequence are arranged and combined, and the arrangement and combination results are sorted according to the level of interactive load and resource disturbance to generate a comprehensive risk sequence. Based on the large model atomic agent database, a comprehensive risk sensitivity is set, and the comprehensive risk sequence is assessed for risk level through streaming monitoring using the comprehensive risk sensitivity.
[0007] Optionally, the step of performing streaming monitoring risk level assessment on the comprehensive risk sequence using the comprehensive risk sensitivity includes: Traverse the comprehensive risk sequence to determine whether there is a risk item corresponding to the comprehensive risk sensitivity; If it exists, the comprehensive risk sensitivity is marked on the corresponding risk item, and the ratio of the comprehensive risk sensitivity to the risk sensitivity to the risk sensitivity to the risk sensitivity to the risk sensitivity to the risk level is determined. The risk level of the streaming monitoring is evaluated based on the ratio. If none exists, then select any of the aforementioned risk items and determine the relationship between the risk borne by the risk item and the overall risk sensitivity. If the risk level is higher than the overall risk sensitivity, the overall risk sequence is marked as a high-risk sequence; if it is lower than the overall risk sensitivity, the overall risk sequence is marked as a low-risk sequence.
[0008] Optionally, assessing the streaming monitoring risk level based on the ratio includes: Obtain the expected performance index of the atomic agent, obtain the continuous performance of the atomic agent on each risk item based on the comprehensive risk sensitivity and ratio, and evaluate the stability performance based on the expected performance index. Set a mutation recovery time, which is the policy reset recovery time of the atomic agent after encountering an abnormal thinking node. Test the atomic agent with a risk item higher than the comprehensive risk sensitivity within a unit time to obtain the anti-interference performance. The risk level of the streaming monitoring is assessed based on the stability and anti-interference performance.
[0009] Optionally, establishing a large model atomic agent database includes: Collect historical operational data of atomic agents, including the design of various atomic agent thinking strategies; The historical running data of the atomic agent is classified according to the basic parameter information, running scenario information, thinking strategy design information and performance feedback information after the atomic agent runs, and a data index entry is established based on the classification results. A subset of thinking strategies is designed and constructed according to the thinking strategy of each atomic agent, and a historical trajectory tracking system is established for each thinking strategy scheme to form a large model atomic agent database, which is used by the streaming monitoring simulation environment to call data.
[0010] Optionally, the data to be accessed by the streaming monitoring simulation environment includes: The current user interaction scenario information and system resource scenario information are determined. The user interaction scenario information and system resource scenario information are matched by the index of the large model atomic agent database, and simulated scenario factors are constructed. The simulated scenario factors are digital representations of the running scenario information in the streaming monitoring simulation environment. The basic parameter information and thinking strategy design information of the current atomic agent are determined. The basic parameter information and thinking strategy design information are matched by indexing the large model atomic agent database. The time sampling window is divided into stages, and different time sampling factors are set for each stage according to the thinking complexity. Based on the determined current user interaction scenario information, system resource scenario information, and the current atomic agent's basic parameter information and thinking strategy design information, the large model atomic agent database is matched, and the performance feedback information of the corresponding atomic agent after running is tracked to provide data support for the stage simulation results.
[0011] Optionally, the streaming simulation of several thinking stages of the atomic agent based on the time sampling window includes: Set a time sampling factor to define the conversion ratio between the sampling interval and the actual running time during the simulated thinking process, thereby accelerating the simulation process to quickly obtain long-term thinking performance; Based on the actual operating scenario information and thinking strategy design information of the atomic agent, several specific thinking stages are defined. Each stage represents a core link in the thinking process of the atomic agent, and each thinking stage corresponds to a time sampling factor. At each defined thinking stage, the corresponding user interaction parameters and system resource parameters are loaded, and a streaming simulation is performed. By synthesizing the simulation results of all the aforementioned thinking stages, the thinking performance of the atomic agent in the hypothetical full lifecycle interaction is evaluated, and the simulation results of the aforementioned stages are obtained.
[0012] Optionally, obtaining monitoring and evaluation results through stage simulation results includes: Deep learning is performed on the large model atomic agent database to obtain the mapping relationship between the basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs, as well as the mapping relationship between the fitted basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs. A data update window is set up to collect real-time data on the thinking performance of the online atomic agent, and the collected data is entered into the large model atomic agent database. The large model atomic agent database is dynamically updated according to the data update window, and the mapping relationship is iteratively optimized. The monitoring reliability of the stage simulation results is evaluated based on the mapping relationship to obtain the monitoring evaluation results.
[0013] Furthermore, to achieve the above objectives, the present invention also provides a large-model-based thinking processing system, the system comprising: a memory, a processor, and a large-model-based thinking processing program stored in the memory and executable on the processor, the large-model-based thinking processing program being configured to implement the steps of the large-model-based thinking processing method as described in any one of the above descriptions.
[0014] In addition, to achieve the above objectives, the present invention also provides a medium storing a large-model-based thinking processing program, which, when executed by a processor, implements the steps of the large-model-based thinking processing method as described above.
[0015] This invention provides a large-scale model-based thinking processing method. By constructing an atomic agent database and performing streaming simulation of the thinking process, the method decomposes the complex decision-making logic of the large model into traceable atomic thinking stages, making the answer generation path transparent. This solves the core problem of the uninterpretable results of traditional black-box models, significantly improving user trust in report retrieval logic and answer reliability. Furthermore, by constructing a comprehensive risk sequence based on user interaction scenarios and system resource scenarios, and setting comprehensive risk sensitivity and time sampling windows, the method achieves real-time risk prediction of the large model under different loads, proactively identifying abnormal decision risks and dynamically adjusting resource allocation, thus improving system stability in high-concurrency and resource-fluctuation scenarios. In fields with strict interpretability requirements, such as financial risk control and medical diagnosis, the solution simulates the staged thinking output of the environment. Trajectory provides clients and regulatory agencies with complete decision-making traceability capabilities, meeting compliance verification needs, breaking down trust barriers in traditional black-box models, and promoting the commercial application of large-scale models in vertical industries. Based on performance feedback information after atomic agent operation, combined with multi-dimensional mapping relationships constructed by deep learning, the effectiveness of thinking strategies is analyzed in real time. The strategy model is dynamically iterated through data update windows to avoid performance degradation caused by long-term operation, forming a closed loop of simulation-monitoring-optimization, and continuously improving the response efficiency and decision accuracy of large-scale models. By matching real-time scenario parameters through database indexes, adaptive simulation scenario factors and time sampling factors are dynamically generated, enabling large-scale models to quickly adapt to diverse application environments. While ensuring monitoring accuracy, simulation efficiency is optimized, computing power consumption under high-load scenarios is reduced, and resource utilization is improved. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating an embodiment of the large-model-based thinking processing method of the present invention.
[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] Reference Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the large-model-based thinking processing method of the present invention, which presents an embodiment of the large-model-based thinking processing method of the present invention.
[0020] In one embodiment, the large-model-based thinking processing method includes: Step S100: Establish a large-scale model atomic agent database. The large-scale model atomic agent database includes basic parameter information of atomic agents, running scenario information, thinking strategy design information, and performance feedback information after the atomic agents run.
[0021] The large-model atomic agent database can be a structured data set storing the basic parameters, operating scenarios, thinking strategies, and performance feedback of atomic thinking units during the decision-making process of a large model. It can be used to achieve the discretized expression of the large model's decision-making logic and the structured retention of historical operating data. In this embodiment, the large-model atomic agent database can capture the intermediate inference states of the large model in historical tasks through a log collection module, slice and encode them at the atomic granularity, and store them in the database, supporting key-value indexing and timestamp marking. Furthermore, the large-model atomic agent database can include, but is not limited to, one or more of the following: basic parameter information, thinking strategy design information, and performance feedback information. The large-model atomic agent database can provide the configuration basis for atomic units in the streaming monitoring simulation environment and provide the input source for risk level and scenario parameters for the time sampling window.
[0022] Step S200: The running scenario information includes user interaction scenario information and system resource scenario information, and the streaming monitoring risk level is assessed based on the user interaction scenario information and system resource scenario information.
[0023] The streaming monitoring risk level can be a dynamically evaluated label calculated based on real-time parameters of user interaction scenarios and system resource scenarios, used to guide the triggering strength of monitoring strategies. In this embodiment, the streaming monitoring risk level can trigger a highly sensitive sampling strategy by matching scenario labels (such as high concurrency requests plus low memory margin) through a rule engine, or by training a risk-sampling mapping model based on historical performance feedback data to dynamically generate sampling window parameters. Furthermore, the streaming monitoring risk level can be linked with the time sampling window to affect the evaluation frequency and granularity.
[0024] Step S300: Based on the basic parameter information and thinking strategy design information, construct a streaming monitoring simulation environment for the atomic agent and perform streaming simulation of the atomic agent's thinking process.
[0025] The streaming monitoring simulation environment can be a dynamic execution container that reproduces the thought process of atomic agents in a virtual space. It supports the stage-by-stage deduction of decision paths according to time series and can be used to transform the invisible internal reasoning process of a large model into an observable and operable simulation trajectory. In this embodiment, the streaming monitoring simulation environment can call a lightweight inference engine to load atomic units stage by stage based on the parameters and strategies in the atomic agent database, simulating its input-processing-output flow without relying on the computing power of the original large model. Furthermore, the streaming monitoring simulation environment can simulate the conditional jump logic of atomic agents through a state machine model, or replace the original model layer through function injection to simulate the input and output behavior of atomic thinking with proxy functions, thereby reconstructing the observable execution flow of its thought path without calling the original large model.
[0026] Step S400: Load the streaming monitoring risk level in the streaming monitoring simulation environment, set the time sampling window, perform streaming simulation on several thinking stages of the atomic agent based on the time sampling window, and obtain the monitoring evaluation results through the stage simulation results.
[0027] The time sampling window can be a discrete time interval unit defined during the streaming simulation process. It is used to capture the thinking state of the atomic agent at a specific moment and trigger risk assessment. This allows for high-precision timing capture of key nodes in the thinking process, reducing the computational overhead of end-to-end monitoring. In this embodiment, the time sampling window can dynamically calculate the sampling interval and triggering conditions based on real-time parameters of user interaction scenarios and system resource scenarios, matching a preset sampling factor table through a database index. Furthermore, the time sampling window can include, but is not limited to, high-interaction-density sampling windows, low-resource-margin sampling windows, and high-compliance-requirement sampling windows. In the streaming monitoring simulation environment, based on the interval points set by the time sampling window, the simulation can be paused, a snapshot of the current atomic agent's state can be extracted, and its input, intermediate reasoning, and output features can be recorded. Through lightweight semantic consistency verification or state vector embedding analysis, a traceable, phased thinking trajectory can be generated, supporting the replay and anomaly localization of the decision-making chain.
[0028] Taking a financial risk control Q&A system as an example, the large-model-based thinking processing method in this embodiment can be as follows: When a user inquires whether there is a fraud risk in the loan application, the system loads financial-specific atomic units from the atomic agent database to construct a simulation environment; based on the current resource scenario of high concurrent requests and database response delay, a low resource reserve sampling window is matched, and the thinking state is captured at key inference nodes (such as credit score calculation and historical transaction comparison); the stage trajectory output by the simulation environment shows abnormal offset of risk factor weights, triggering the resource scheduling module to temporarily increase the calculation priority, and at the same time outputting the complete decision chain to the regulatory interface for auditing.
[0029] This embodiment provides a large-model-based thinking processing method. It achieves atomic and structured storage of decision logic by establishing a large-model atomic agent database. Based on basic parameters and strategies, it constructs a reproducible streaming monitoring simulation environment. It dynamically loads risk levels and sets time sampling windows according to user interaction and system resource scenarios. It extracts and analyzes the state of the thinking stage by sampling points. It can achieve the technical effects of making black-box reasoning transparent through structural decoupling, completely reproducing the decision path without relying on original computing power, accurately controlling monitoring overhead with an adaptive sampling mechanism, and forming a simulation-monitoring-optimization closed-loop operation combined with performance feedback. It supports the traceability and verifiability requirements in high compliance scenarios.
[0030] In one embodiment, the risk level of streaming monitoring is assessed based on user interaction scenario information and system resource scenario information, including: The user load stages are divided based on user interaction scenario information and sorted from high to low according to interaction complexity to generate an interaction load sequence. The interaction load sequence can be a sequence of user request types sorted from high to low complexity based on user interaction scenario information. It can be used to characterize the semantic depth, multi-hop inference requirements, and contextual dependency strength of user queries. In this embodiment, the interaction load sequence can be based on an intent recognition model to score the multi-hop inference depth of the questions and sort them in descending order of score. Alternatively, it can construct a complexity index using semantic tree depth and entity reference count, and then cluster and sort the requests. This allows for the structured quantification of user interaction complexity, providing a foundational input for multi-dimensional risk modeling. For example, the interaction load sequence can include, but is not limited to, one or more of the following: multi-condition cross-validation requests, open-ended inference requests, and high contextual dependency requests.
[0031] Based on the system resource scenario information, the resource disturbance stages are divided and sorted from high to low according to the resource fluctuation amplitude to generate a resource disturbance sequence; The resource disturbance sequence can be a sequence of runtime environment disturbance types sorted from high to low based on system resource scenario information, and can be used to characterize the dynamic instability of infrastructure such as computing resources, memory, and network bandwidth. In this embodiment, the resource disturbance sequence can collect CPU utilization variance, memory reclamation frequency, and network RTT standard deviation, and sort them by comprehensive disturbance index, or based on historical resource anomaly event clustering, label typical disturbance patterns and sort them by occurrence intensity, thereby abstracting resource fluctuations into a comparable disturbance type sequence, supporting coupled analysis with interactive load. For example, the resource disturbance sequence can include, but is not limited to, one or more of the following: sudden memory exhaustion, GPU scheduling latency, and cross-node communication jitter.
[0032] The interactive load sequence and resource disturbance sequence are arranged and combined, and the results of the arrangement and combination are sorted according to the level of interactive load and resource disturbance to generate a comprehensive risk sequence. The comprehensive risk sequence can be a multidimensional risk ranking list formed by permuting and combining interactive load sequences and resource disturbance sequences. It reflects the priority of potential decision anomalies under different coupled scenarios, providing a tiered risk triggering basis for the streaming monitoring simulation environment and supporting dynamic allocation of monitoring resources according to priority. In this embodiment, the comprehensive risk sequence can combine each item of the interactive load sequence and resource disturbance sequence using a Cartesian product, employing a weighted product method: interaction complexity weight × resource fluctuation intensity weight, to generate a comprehensive score and rank it. Alternatively, a two-dimensional risk matrix can be constructed, dividing the combination type by quadrant and then ranking it by the frequency of historical anomalies. This allows for the first-ever collaborative risk modeling of interactive behavior and system state, breaking through the limitations of single-dimensional risk assessment. The comprehensive risk sequence can include, but is not limited to, one or more of the following: high interaction-high disturbance combination, high interaction-low disturbance combination, and low interaction-high disturbance combination.
[0033] Based on a large model atomic agent database, a comprehensive risk sensitivity is set, and the risk level of the comprehensive risk sequence is assessed through streaming monitoring using the comprehensive risk sensitivity. The comprehensive risk sensitivity can be a dynamically configured weight adjustment parameter based on historical performance feedback information in the atomic agent database. This parameter is used to weight and score each item in the comprehensive risk sequence, enabling adaptive adjustment of risk level assessment and avoiding misjudgments or omissions in heterogeneous scenarios using fixed thresholds. In this embodiment, the comprehensive risk sensitivity can analyze the correlation between historical simulation results and performance degradation events using a deep learning model, outputting weight coefficients for different combinations, and periodically updated by performance feedback data. This allows it to read the permutations and combinations of the comprehensive risk sequence, outputting normalized risk level labels, and driving the dynamic configuration of the time sampling window. In this embodiment, the comprehensive risk sensitivity can also input each combination of items in the comprehensive risk sequence into a neural network for embedding encoding, outputting risk probability values and categorizing them, or using a rule engine to match predefined combination patterns and combine them with dynamic weight coefficients to output level ranges. This enables risk assessment to have scenario-adaptive capabilities, achieving an evolution from static thresholds to dynamic weighted assessment.
[0034] Taking a medical diagnostic question-answering system as an example, the large-model-based thinking processing method in this embodiment can be as follows: when the system simultaneously encounters a large number of multi-hop diagnostic reasoning requests (high interaction load) and response delays caused by backups in the backend database (high resource disturbance), the interaction load sequence and the resource disturbance sequence are respectively sorted; a high-interaction-high-disturbance high-priority risk item is formed by combining them through Cartesian products; the comprehensive risk sensitivity model determines based on historical data that this combination has led to omissions in the diagnostic reasoning path, and thus outputs the highest risk level; the streaming monitoring simulation environment accordingly enables a dense sampling window to block potential erroneous reasoning chains in advance, and at the same time triggers a cache preloading mechanism.
[0035] This embodiment provides a large-scale model-based thinking and processing method. It generates an interaction load sequence based on user interaction scenario information and a resource disturbance sequence based on system resource scenario information. The two are then combined to generate a comprehensive risk sequence. Based on the comprehensive risk sensitivity, the comprehensive risk sequence is subjected to streaming monitoring and risk level assessment. This allows user interaction complexity and system resource fluctuations to be quantified into sortable sequences, achieving risk modeling through their synergistic coupling. Through a dynamic weighting mechanism, the combined items are adaptively weighted and scored, accurately identifying potential decision anomalies in high-interaction and high-resource-disturbance collaborative scenarios. This enables early triggering of simulation enhancement and resource scheduling, providing quantifiable risk warning basis for high-compliance scenarios, and supporting the technical effect of real-time risk prediction and dynamic optimization closed loop.
[0036] In one embodiment, the risk level assessment of the comprehensive risk sequence is performed by streaming monitoring based on the comprehensive risk sensitivity, including: traversing the comprehensive risk sequence and determining whether there is a risk item corresponding to the comprehensive risk sensitivity.
[0037] The comprehensive risk sequence can be a multidimensional risk ranking list formed by arranging and combining interaction load sequences and resource disturbance sequences. It reflects the priority of potential decision anomalies under different scenario coupling states and can be used as an input source for risk level assessment, providing a rankable set of scenario coupling risk items. In this embodiment, the comprehensive risk sequence can include, but is not limited to, one or more of the following: high interaction-high disturbance combination, high interaction-low disturbance combination, and low interaction-high disturbance combination. The comprehensive risk sensitivity can be a dynamically configured weight adjustment parameter based on historical performance feedback information in the atomic agent database. It is used to weight and score each item in the comprehensive risk sequence and can be used as a dynamic benchmark for risk level assessment, supporting adaptive grading and default inference mechanisms.
[0038] Furthermore, the comprehensive risk sensitivity can be analyzed using a deep learning model to correlate historical simulation results with performance degradation events, outputting weight coefficients for different combinations, and periodically updated by performance feedback data. In this embodiment, the comprehensive risk sensitivity serves as an evaluation threshold benchmark, matching and judging its relative relationship with each item in the comprehensive risk sequence to determine the output logic path for the risk level. Traversing the comprehensive risk sequence to determine if there is a risk item corresponding to the comprehensive risk sensitivity can be done by comparing each item in the sorted order of the comprehensive risk sequence to detect if there is a risk item that perfectly matches the comprehensive risk sensitivity value. Further, traversing the comprehensive risk sequence to determine if there is a risk item corresponding to the comprehensive risk sensitivity can be achieved by using hash mapping to quickly locate the exact match of the sensitivity value in the sequence, or by judging an approximate match using floating-point tolerance ranges, avoiding missed detections due to precision errors. This allows for precise alignment of risk items with the sensitivity benchmark, supporting subsequent proportional calculations or default inference branch selection.
[0039] If it exists, the overall risk sensitivity will be marked on the corresponding risk item, and the ratio of the overall risk sensitivity to the overall risk sensitivity to the ratio of the overall risk sensitivity to the risk sensitivity to be determined. The risk level of the streaming monitoring will be assessed based on the ratio.
[0040] One approach to assessing the risk level of streaming monitoring involves labeling the overall risk sensitivity on corresponding risk items and determining the proportion of items above or below that sensitivity in the statistical sequence after matching the corresponding risk items. This proportion serves as an indicator of risk distribution skewness. Furthermore, assessing the risk level of streaming monitoring can be achieved by calculating the cumulative probability of items exceeding the sensitivity threshold; if this probability exceeds 70%, it is marked as a high-risk trend. Alternatively, a risk distribution histogram can be constructed, and the direction of anomaly clustering can be determined based on the skewness coefficient. This allows for the quantification of the overall risk tendency of the system through distribution skewness, avoiding misjudgments caused by single-point anomalies and improving assessment robustness.
[0041] If none exists, then select any risk item and determine the relationship between the risk borne by the risk item and the overall risk sensitivity.
[0042] The process involves selecting any risk item, determining its relationship with the overall risk sensitivity, and labeling the overall risk sequence as high-risk or low-risk. This can be achieved by selecting the first or most representative risk item in the sequence when no matching item exists, comparing its risk value with the overall risk sensitivity, and outputting a high / low risk label. Furthermore, this process can be implemented by selecting the highest-ranked risk item as the benchmark, as it represents the highest potential risk, or by selecting the median risk item to reduce interference from extreme values. This allows for the construction of an inferential assessment mechanism when no precise match is found, ensuring the system can still output an operational monitoring level under new or rare risk modes.
[0043] Taking a financial anti-fraud question-and-answer system as an example, the large-model-based thinking and processing method in this embodiment can be as follows: when the system encounters a new type of user question pattern (such as nested conditional queries), the corresponding risk item has not appeared in the historical data; the comprehensive risk sensitivity is 0.72, and there is no completely matching item in the comprehensive risk sequence; the system selects the highest risk item (0.68) in the sequence for comparison, determines that it is lower than the sensitivity, marks it as a low-risk sequence, but starts enhanced sampling; subsequent simulations found that this pattern causes the inference path to skip key compliance verification nodes, triggers the performance feedback mechanism to update the database, and adds the risk item with a weight of 0.75 in the next iteration, completing the self-discovery and self-adaptation of new risks.
[0044] This embodiment provides a large-scale model-based thinking and processing method. It traverses the comprehensive risk sequence and determines whether a risk item corresponding to the comprehensive risk sensitivity exists. If it exists, the sensitivity is marked and the risk level is assessed based on the higher / lower ratio. If it does not exist, representative risk items are selected for relative relationship judgment and the sequence level is marked. Accurate trend identification is achieved through dynamic benchmark matching and distribution skewness analysis. The continuity of assessment under new models is ensured through a default inference mechanism. By introducing comprehensive risk sensitivity as a dynamic benchmark, two parallel assessment paths can be constructed. When a matching item exists, the overall trend is judged based on the risk distribution ratio; when no matching item exists, inferential labeling is performed using the nearest neighbor item. This breaks through the rigidity of traditional fixed thresholds, enabling risk assessment to have context awareness and asymmetric response capabilities. It can identify edge risks when resources fluctuate drastically or interaction patterns change abruptly. Combined with multi-dimensional coupling modeling of the comprehensive risk sequence, this process can dynamically identify new risk patterns without relying on historical abnormal samples. It provides interpretable, auditable, and credible risk judgment basis for high-compliance scenarios such as financial risk control and medical diagnosis, and is the core decision engine supporting real-time risk prediction and dynamic resource optimization closed loop.
[0045] In one embodiment, assessing the risk level of streaming monitoring based on proportion includes: obtaining the expected performance index of the atomic agent, obtaining the continuous performance of the atomic agent on each risk item based on the comprehensive risk sensitivity and proportion, and evaluating the stability performance based on the expected performance index.
[0046] The expected performance metric can be the standard value of inference efficiency, accuracy, or response consistency that the atomic agent should achieve under ideal, perturbation-free conditions. It can be used as a benchmark for stability assessment, quantifying the deviation between actual performance and the ideal state. In this embodiment, the expected performance metric can be generated from the average performance data of stable operation under historical low-risk scenarios extracted from the atomic agent database, and then smoothed and filtered to produce a benchmark curve. For example, the expected performance metric can include, but is not limited to, one or more of the following: average inference latency threshold, answer consistency confidence, and step integrity coverage. The comprehensive risk sensitivity can be a dynamically configured weighted parameter based on historical performance feedback information in the atomic agent database. It is used to weight and score each item in the comprehensive risk sequence, serving as a dynamic benchmark for performance evaluation and driving continuous performance degradation analysis and anti-interference test triggering conditions. In this embodiment, the comprehensive risk sensitivity analyzes the correlation between historical simulation results and performance degradation events using a deep learning model, outputting weight coefficients for different combinations, and is periodically updated by performance feedback data. For example, the comprehensive risk sensitivity can be combined with the expected performance index to form an evaluation threshold system, which serves as a screening condition for triggering anti-interference testing. Its value determines the disturbance intensity of the selected risk item.
[0047] To obtain the sustained performance of the atomic agent on each risk item based on the comprehensive risk sensitivity and proportion, one approach is to sort the data by comprehensive risk sequence in streaming simulation, continuously run multiple rounds of inference for each risk item with a sensitivity higher than the expected value, and record the performance degradation curve of its performance index relative to the expected value. Furthermore, obtaining the sustained performance of the atomic agent on each risk item based on the comprehensive risk sensitivity and proportion can be achieved by using a sliding window to statistically analyze the latency fluctuation variance of five consecutive inferences for each risk item, calculating the deviation rate relative to the expected value, or by constructing a performance degradation function model to fit the nonlinear relationship between risk intensity and the decrease in inference accuracy. This allows for continuous quantification of performance degradation in high-risk scenarios, supporting refined modeling for stability assessment.
[0048] Set a mutation recovery time, which is the policy reset recovery time of the atomic agent after encountering an abnormal thinking node. Test the atomic agent with a risk item with a higher overall risk sensitivity within a unit time to obtain the anti-interference performance.
[0049] The mutation recovery time, or mutation recovery time, can be defined as the time window required for an atomic agent to reset its policy and return to normal inference after encountering an abnormal thinking node. It quantifies the model's self-healing ability under sudden disturbances and serves as a core evaluation dimension for anti-interference performance. In this embodiment, a disturbance term higher than the overall risk sensitivity is artificially injected into the streaming monitoring simulation environment, and the time stamp difference from the abnormal trigger to state convergence is recorded. For example, mutation recovery time may include, but is not limited to, policy reload recovery time, cache reconstruction recovery time, and inference path reconstruction time. The mutation recovery time can be used as an output indicator of anti-interference performance, input together with stability performance into the risk level assessment module, and its value is affected by the test intensity controlled by the overall risk sensitivity.
[0050] Testing atomic agents with risk terms that have a higher sensitivity than the overall risk level to obtain their anti-interference performance can be achieved by forcibly activating specified risk terms with higher sensitivity in the overall risk sequence within a unit time, and observing the time required for the atomic agent to recover from an abnormal state to a stable output. Furthermore, testing atomic agents with risk terms that have a higher sensitivity than the overall risk level to obtain their anti-interference performance can be achieved by injecting the highest priority risk term and resetting its internal state, recording the complete time from input perturbation to output convergence, or by applying perturbations of different intensities in multiple parallel simulation instances and statistically analyzing the median distribution of recovery times. This allows for the establishment of a resilience measure for the model under sudden failures, shifting the evaluation from passive monitoring to active stress testing.
[0051] The risk level of streaming monitoring is assessed based on its stability and anti-interference performance.
[0052] Assessing the risk level of streaming monitoring based on stability and anti-interference performance can be achieved by weighting and fusing stability performance (performance degradation rate) and anti-interference performance (mutation recovery time) to output a comprehensive risk level label. Furthermore, assessing the risk level of streaming monitoring based on stability and anti-interference performance can be accomplished by employing a multi-criteria decision model, assigning weights to degradation rate and recovery time respectively, calculating a comprehensive risk score and categorizing it, or by constructing a two-dimensional assessment matrix with stability on the horizontal axis and resilience on the vertical axis, dividing the risk level into four quadrants. This allows for a leap from a single risk distribution to a two-dimensional capability assessment, enabling the risk level to reflect the model's true operational resilience.
[0053] Taking a medical diagnostic question-and-answer system as an example, the large-model-based thinking processing method in this embodiment can be as follows: when the system triggers a high-interaction-high-disturbance risk item under high-concurrency consultation requests, the overall risk sensitivity is 0.75; the system detects that the inference latency under this risk item increases by 42% compared to the expected value, and determines it to be moderately unstable; then the risk item is forcibly injected and timed, and it is found that the atomic agent completes the policy reset and restores output consistency within 1.2 seconds; combined with the stability decay rate and recovery time, it is assessed as a medium-high risk level, triggering cache preloading and inference priority improvement to avoid subsequent requests from causing diagnostic delays due to latency accumulation.
[0054] This embodiment provides a large-model-based thinking and processing method. By obtaining the expected performance index of the atomic agent and extracting the continuous performance degradation curve under high-risk items in combination with the comprehensive risk sensitivity, it quantifies the mutation recovery time by forcibly triggering a disturbance higher than the sensitivity and measuring the policy reset response. By integrating the dual-dimensional performance of stability and anti-interference into the comprehensive risk level assessment basis, it can achieve fine-grained evaluation of model stability and dynamic resilience, breaking through the traditional static risk distribution judgment mode. This enables the system to predict the vulnerability and resilience of the model in real high-load scenarios, thereby providing quantifiable, comparable, and verifiable evidence of decision credibility for scenarios such as financial risk control and medical diagnosis. It supports dynamic resource scheduling and policy self-optimization closed loop, fundamentally solving the technical bottleneck of black box models being unmeasurable and unadjustable.
[0055] In some embodiments, a large-scale atomic agent database is established, including: collecting historical operational data of atomic agents, including the design of various atomic agent thinking strategies.
[0056] Collecting historical operational data of atomic agents can involve extracting the inputs, intermediate inferences, outputs, and resource consumption records of atomic agents in various tasks from unstructured operational logs. Furthermore, this historical operational data can be collected by intercepting intermediate state outputs during the inference process at the agent layer and encapsulating them into structured event streams at the atomic unit level. This allows for the construction of a raw data foundation covering multiple strategies and scenarios, supporting subsequent classification and indexing.
[0057] The historical running data of the atomic agent is classified according to the basic parameter information, running scenario information, thinking strategy design information and performance feedback information after the atomic agent runs, and a data index entry is established based on the classification results. A subset of thinking strategies is designed and constructed according to the thinking strategy of each atomic agent, and a historical trajectory tracking system is established for each thinking strategy scheme to form a large model atomic agent database, which is used by the streaming monitoring simulation environment to call data.
[0058] The thinking strategy subset can be an independent data group composed of several strategy categories, divided into historical operational data based on the differences in the thinking strategy design of the atomic agents. This can be used to achieve isolated storage and rapid retrieval of different inference logic modes, avoiding simulation deviations caused by mixed strategies. In this embodiment, the thinking strategy subset can be clustered and grouped based on the logical rules, conditional branches, and output constraint features in the thinking strategy design information, forming a data subset bound to strategy labels. Furthermore, the thinking strategy subset can serve as a matching target for data indexing, allowing the streaming simulation environment to select and adapt the strategy subset for initialization based on the current scenario. For example, the thinking strategy subset can adopt a retrieval-priority strategy subset, an inference-depth strategy subset, a resource-conservative strategy subset, etc.
[0059] Historical trajectory tracking can be a trajectory dataset that fully records and times-anchors the phased state sequences of each thinking strategy scheme during execution in different operating scenarios. It can be used to provide a replayable temporal evidence chain for strategy effectiveness evaluation and abnormal pattern recognition. In this embodiment, historical trajectory tracking can capture intermediate state snapshots by time sampling windows each time the atomic agent runs, and associate input parameters, resource load, and output results to form a timestamped chain trajectory. Furthermore, historical trajectory tracking can be stored in the corresponding thinking strategy subset as empirical data for the execution of the strategy subset, supporting the trajectory reproduction function in the simulated environment. For example, historical trajectory tracking can include, but is not limited to, high-accuracy trajectory sets, high-latency trajectory sets, and resource fluctuation response trajectory sets.
[0060] The data index entry point can be a query interface used to quickly locate and match the optimal subset of thinking strategies and their historical trajectory based on real-time scenario parameters. It can be used to dynamically load the most suitable decision-making mode at startup of the simulation environment, reducing strategy search time and the risk of mismatches. In this embodiment, the data index entry point can construct a composite index key using multi-dimensional fields (such as user interaction type, system resource utilization, and compliance level) to match pre-built strategy subset labels in the database. Furthermore, the data index entry point can connect real-time parameters and thinking strategy subsets of user interaction scenarios and system resource scenarios, serving as the data scheduling hub for the initialization of the streaming simulation environment. For example, the data index entry point can employ scenario-matching indexes, performance-optimized indexes, and compliance-oriented indexes.
[0061] Classifying the historical operational data of atomic agents based on basic parameter information, operational scenario information, strategic design information, and performance feedback information after atomic agent execution can be achieved by labeling and grouping the raw operational data according to four information dimensions to form an identifiable set of categories. Furthermore, this classification of historical operational data can be automated using unsupervised clustering algorithms (such as DBSCAN) to automatically categorize strategic design features, thereby enabling semantic organization of operational data and providing structured input for strategy subset construction.
[0062] Constructing a subset of thinking strategies based on the thinking strategy design of each atomic agent can be achieved by aggregating historical data of atomic agents belonging to the same strategy design paradigm into independent data units. Furthermore, this construction of a subset of thinking strategies based on the thinking strategy design of each atomic agent can use the logical rules in the strategy design document as keys to package matching runtime data into strategy templates, thereby forming a reusable and comparable strategy pattern library to support the strategy selection mechanism in the simulation environment.
[0063] Establishing a historical trajectory tracking system for each thinking strategy can involve recording the state sequence and time-related information of each execution instance within each strategy subset, representing the complete thinking phase. Furthermore, this historical trajectory tracking can be enhanced by inserting state snapshot identifiers at each sampling point, forming a chained trajectory log. This provides traceable evidence of the strategy execution process, supporting anomaly detection and strategy performance evaluation.
[0064] Taking a medical diagnostic question-and-answer system as an example, the large model-based thinking processing method in this embodiment can be as follows: when the system receives a consultation on whether to recommend a biopsy for a patient whose CT image suggests lung nodules, the data index entry matches a compliance-priority strategy subset based on high compliance requirements and low response latency scenario parameters, and loads its corresponding historical trajectory tracking; the simulation environment then reproduces the reasoning path that has successfully passed multiple rounds of expert verification under this strategy, ensuring that the output conforms to clinical guidelines, and at the same time, the current simulation trajectory is stored in the corresponding trajectory set to continuously enrich the empirical data of the strategy subset.
[0065] This embodiment provides a large-model-based thinking processing method. By collecting historical operational data of atomic agents, the data is classified according to basic parameters, operational scenarios, thinking strategy design, and performance feedback. A strategy subset is constructed according to the thinking strategy design, and a historical trajectory is established for each strategy scheme. Through the data index entry, efficient matching between scenario parameters and strategy subsets is achieved, enabling the streaming monitoring simulation environment to dynamically load the decision path and historical experience that best fits the current environment. Thus, without increasing the inference load, it achieves precise strategy selection, traceability of the execution process, and long-term self-evolution, enhancing the system's credibility and adaptability in highly compliant scenarios.
[0066] In some embodiments, the data called by the streaming monitoring simulation environment includes: determining the current user interaction scenario information and system resource scenario information, matching the user interaction scenario information and system resource scenario information through the large model atomic agent database index, and constructing simulation scenario factors.
[0067] The large-scale model atomic agent database can be a structured collection storing multi-dimensional operational data of atomic agents, including basic parameters, operational scenarios, strategic design, and performance feedback information. It can serve as a data source for constructing simulation scenario factors, setting time sampling factors, and tracking performance feedback. In this embodiment, the large-scale model atomic agent database can support low-latency queries and matching in high-concurrency scenarios through multi-dimensional index keys (user interaction type, resource load level, strategy label). For example, the large-scale model atomic agent database can provide scenario parameter mapping for simulation scenario factors, provide strategy complexity references for dividing time sampling window stages, and provide historical operational records for performance feedback tracking. Furthermore, the large-scale model atomic agent database can include, but is not limited to, one or more of the following: a set of basic parameter information, a set of operational scenario parameters, and a set of performance feedback indicators.
[0068] Simulated scenario factors can transform real-time user interaction scenarios and system resource scenarios into computable digital representation variables within a streaming monitoring simulation environment. These factors can provide quantifiable and reusable simulated inputs to the abstract runtime environment, supporting dynamic initialization of the simulation environment. In an exemplary embodiment, simulated scenario factors can match current scenario tags through a database index entry point and extract predefined scenario feature vectors (such as request frequency, memory usage, and compliance level coding) as simulation environment parameters. For example, simulated scenario factors can serve as startup configuration parameters for the streaming monitoring simulation environment, working in conjunction with time sampling factors to determine sampling density and evaluation weights. Furthermore, simulated scenario factors can include, but are not limited to, one or more of the following: high-concurrency interaction scenario factors, low resource reserve scenario factors, and strong compliance constraint scenario factors.
[0069] By indexing and matching user interaction scenario information and system resource scenario information in a large-scale atomic agent database, and constructing simulated scenario factors, this can be achieved by encoding the current user request type and system resource status into structured tags, querying pre-stored scenario feature vectors in the database, and injecting them into the simulated environment. Furthermore, this operation can be implemented by using a word embedding model to map natural language interaction descriptions into scenario semantic vectors, matching them with pre-trained scenario templates in the database, or by collecting CPU / memory / IO metrics in real time through a system monitoring interface and mapping them to predefined resource load level codes. This allows for digital modeling of the runtime environment, enabling the simulated environment to respond to real load changes.
[0070] The basic parameter information and thinking strategy design information of the current atomic agent are determined. The basic parameter information and thinking strategy design information are matched by indexing the large model atomic agent database, and the time sampling window is divided into stages.
[0071] The time sampling window can be a discrete time interval unit defined during the streaming simulation process. It is used to capture the thinking state of the atomic agent at a specific moment and trigger risk assessment. This can be used to capture the timing of key nodes in the thinking process, reducing the computational overhead of end-to-end monitoring. In one specific embodiment, the time sampling window can be dynamically divided into multiple sub-windows based on the logical branch complexity and stage dependencies in the thinking strategy design information. Each sub-window corresponds to an independent sampling stage. For example, the time sampling window can receive load intensity indications from simulation scenario factors and adjust the sampling frequency and granularity of each sub-window according to the time sampling factor. Furthermore, the time sampling window can include, but is not limited to, one or more of the following: high interaction density sampling window, low resource margin sampling window, and high compliance requirement sampling window.
[0072] The time sampling factor can be a sampling density control parameter allocated to each stage within the time sampling window. It determines the collection frequency and information granularity of the state snapshot for that stage, enabling high-precision sampling of stages with high complexity and low-frequency sampling of simpler stages, thus optimizing resource utilization efficiency. In this embodiment, the time sampling factor can generate stage-level sampling weights based on the anomaly rate and computation time of each strategy stage in historical performance feedback information, mapped to the reciprocal of the sampling interval. For example, the time sampling factor can be jointly determined by the thinking strategy design information and performance feedback information, affecting the stage division result of the time sampling window and directly influencing the resolution and overhead of the simulated trajectory. Furthermore, the time sampling factor can include, but is not limited to, one or more of the following: high-complexity stage sampling factors, low-latency requirement sampling factors, and high-fluctuation response sampling factors.
[0073] By indexing and matching basic parameter information and strategic design information from a large-scale atomic agent database, and dividing the time sampling window into stages, this can be achieved by retrieving the corresponding stage division rules from the database based on the current atomic agent's strategy type and structural complexity, thus splitting the time sampling window into multiple sub-intervals. Furthermore, this operation can automatically define stage boundaries based on the number of conditional branches and nesting depth in the strategy design document, or by using the frequency of state changes in historical trajectory tracking to cluster and identify high-activity stages as sampling sub-windows. This aligns the sampling rhythm with the inference logic structure, avoiding wasting sampling resources in low-value stages.
[0074] Based on the determined current user interaction scenario information, system resource scenario information, and the basic parameter information and thinking strategy design information of the current atomic agent, the large model atomic agent database is matched and the performance feedback information of the corresponding atomic agent after running is tracked to provide data support for the phase simulation results.
[0075] By matching the current user interaction scenario information, system resource scenario information, and the basic parameter information and strategic design information of the current atomic agent with the large model atomic agent database and tracking the performance feedback information of the corresponding atomic agents after execution, this can be achieved by combining the current scenario, parameters, and strategy labels, and querying the performance feedback records of historical similar execution instances in the database as the evaluation benchmark for the current simulation phase. Furthermore, this operation can be implemented by using cosine similarity to match historical execution trajectories, extracting the anomaly rate and average response latency of the corresponding phase, or by using an online learning model to predict the expected performance deviation under the current combination, and feeding this information back to the risk assessment module. This provides historical behavioral evidence for each simulation phase, improving the contextual sensitivity of anomaly identification and the accuracy of assessment.
[0076] Taking real-time financial risk control Q&A as an example, the large-model-based thinking processing method in this embodiment can be as follows: when the system receives an inquiry about whether abnormal fluctuations in a company's financial statements over the past three years constitute fraud risk, the database matches simulated scenario factors with high compliance and low response latency; based on its thinking strategy being a multi-round cross-validation type, the time sampling window is divided into four sub-stages: data extraction, correlation analysis, expert rule comparison, and confidence weighting; a high sampling factor (snapshot every 100ms) is allocated to the expert rule comparison stage, and 500ms is allocated to the other stages; at the same time, the performance feedback of this strategy in similar requests in history is retrieved, and it is found that the output drift often occurs in the "confidence weighting" stage. Based on this, the simulation environment increases the risk weight of this stage and triggers resource priority scheduling in advance.
[0077] In one embodiment, streaming simulation of several thinking stages of the atomic agent is performed based on a time sampling window, including: setting a time sampling factor to define the conversion ratio between the sampling interval and the actual running time during the simulated thinking process, thereby accelerating the simulation process to quickly obtain long-cycle thinking performance.
[0078] The time sampling factor can be a scaling parameter used to map simulated time to actual running time, determining the time density of sampling points during the simulation. It can be used to compress simulation duration while maintaining the semantic integrity of the thinking phase, improving the efficiency of long-cycle behavior evaluation. In this embodiment, the time sampling factor can be determined by querying a preset factor mapping table from a database or dynamically calculating its optimal value based on historical performance feedback, according to the complexity of the atomic agent's thinking strategy and the resource constraints of the running scenario. In this embodiment, the time sampling factor is loaded into the streaming monitoring simulation environment as a configuration parameter for the atomic agent's thinking phase, controlling the temporal rhythm of the simulation within this phase. For example, the time sampling factor can include, but is not limited to, one or more of the following: high compression ratio factor, medium fidelity factor, and no compression factor.
[0079] Based on the actual operating scenario information and thinking strategy design information of the atomic agent, several specific thinking stages are defined. Each stage represents a core link in the atomic agent's thinking process, and each thinking stage corresponds to a time sampling factor.
[0080] The atomic agent thinking stage can be a process where the complete thinking process of the atomic agent is divided into several sub-units with independent input, processing logic, and output characteristics according to semantic function. This can be used to achieve semantic segmentation of the decision path of a large model, supporting scenario-specific customized simulation and evaluation. In this embodiment, the atomic agent thinking stage can cluster and label historical inference logs based on the conditional branches and functional module divisions in the thinking strategy design information, combined with typical interaction patterns in the running scenario information. In this embodiment, the atomic agent thinking stage is bound to specific time sampling factors and scenario parameters, serving as the execution unit of the streaming monitoring simulation environment, and its output constitutes the input of the stage simulation results. For example, the atomic agent thinking stage can include, but is not limited to, one or more of the following: information retrieval stage, multi-source inference stage, and compliance verification stage.
[0081] At each defined thinking stage, the corresponding user interaction parameters and system resource parameters are loaded, and a streaming simulation is performed.
[0082] The streaming monitoring simulation environment can be a dynamic execution container that reproduces the thinking process of atomic agents in a virtual space. It supports the stage-by-stage deduction of decision paths according to time series and can be used to transform the originally invisible internal reasoning process of a large model into an observable and operable simulation trajectory. In this embodiment, the streaming monitoring simulation environment supports dynamically switching sampling factors and scene parameters according to the boundaries of the atomic agent thinking stages, realizing segmented heterogeneous simulation, rather than continuous deduction under a single fixed sampling strategy. In this embodiment, the streaming monitoring simulation environment loads the configuration information and corresponding time sampling factors of the atomic agent thinking stages, triggers parameter reloading at the boundary of each stage, and completes continuous simulation across stages. Furthermore, loading the corresponding user interaction parameters and system resource parameters for streaming simulation at each atomic agent thinking stage can be achieved by dynamically updating the input environment variables and restarting the deduction when the simulation environment executes to a certain stage boundary, based on the scene parameter set bound to that stage. Furthermore, this operation can replace environmental variables in the simulation environment through parameter injectors, simulate the impact of different user intentions on the inference path, or simulate resource disturbances such as insufficient system memory and network latency, and observe the stability changes of the stage output, thereby achieving high-fidelity reproduction and differentiated evaluation of behavior at each stage under multi-scenario pressure.
[0083] By synthesizing the simulation results of all thinking stages, the thinking performance of the atomic agent in the hypothetical full lifecycle interaction is evaluated, and the stage simulation results are obtained.
[0084] Furthermore, by comprehensively evaluating the simulation results of all thinking stages to assess the atomic agent's thinking performance throughout the hypothetical lifecycle interactions, this can be achieved by sequentially piecing together indicators such as output status, risk level, and response latency at each stage to form a complete profile of the thinking trajectory. Further, this operation can identify logical breaks or policy drift points by constructing inter-stage state transition matrices, or perform consistent clustering of outputs at each stage to mark potential accumulated deviation patterns over long-term operation. This allows for the generation of thinking behavior assessment reports covering multiple scenarios and long periods, supporting strategy iteration and compliance verification.
[0085] Taking a medical diagnostic question-and-answer system as an example, the thinking process based on a large model in this embodiment can be as follows: In a long-cycle interaction simulating three consecutive patient consultations, the system divides the thinking process into three atomic stages: symptom retrieval, medical record comparison, and contraindication drug screening. A high compression factor is set for the first two stages to accelerate the simulation, while no compression factor is set for the contraindication drug screening stage to ensure compliance accuracy. The simulation environment dynamically loads the corresponding patient medical history and system load parameters at the boundaries of each stage. Finally, by combining the outputs of the three stages, it is found that the misjudgment rate of the contraindication drug screening stage increases under high concurrency, triggering a strategy optimization instruction.
[0086] This embodiment defines the conversion ratio between the simulated sampling interval and the actual running time by setting a time sampling factor. Based on the running scenario information and thinking strategy design information, it defines several atomic agent thinking stages. In each thinking stage, corresponding user interaction parameters and system resource parameters are loaded for streaming simulation. The simulation results of all thinking stages are combined to evaluate the thinking performance of the atomic agents in the assumed full life cycle interaction. This allows the streaming monitoring simulation environment to adopt differentiated simulation speeds and parameter environments at different stages, while maintaining the integrity of the inference logic. It significantly reduces the computational overhead of the full life cycle simulation. Through staged parameter loading and result splicing, it achieves for the first time a segmented high-fidelity evaluation of the dynamic behavior of a large model in a real complex interaction environment, supporting the collaborative realization of risk prediction, strategy optimization, and compliance traceability.
[0087] In one embodiment, monitoring and evaluation results are obtained through stage simulation results, including: performing deep learning on a large model atomic agent database to obtain the mapping relationship between basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs, as well as the mapping relationship between the fitted basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs.
[0088] The large-model atomic agent database can be a structured data set storing the basic parameters, operating scenarios, thinking strategies, and performance feedback of atomic thinking units during the decision-making process of a large model. It can be used to achieve the discretized expression of the decision-making logic of a large model and the structured retention of historical operating data. In this embodiment, the large-model atomic agent database can serve as both an input source and an output target for a deep learning model, supporting the continuous injection of online operating data through a data update window to drive dynamic iteration of mapping relationships. The large-model atomic agent database can provide training and inference data sources for the multi-dimensional mapping relationships constructed by deep learning, accepting real-time feedback data written through the data update window to maintain data freshness.
[0089] The multi-dimensional mapping relationship constructed by deep learning can be a non-linear correlation function between basic parameters, operating scenarios, strategic design information, and atomic agent performance feedback information modeled by deep neural networks. This can be used to transform multi-dimensional input variables into quantitative predictions of performance feedback, enabling data-driven evaluation of the effectiveness of strategic thinking. In this embodiment, the multi-dimensional mapping relationship constructed by deep learning can employ a multilayer perceptron or graph neural network architecture. The inputs are encoded basic parameters, scenario labels, and policy vectors, and the output is the predicted value of multi-dimensional performance feedback indicators. The prediction error is minimized through training with historical data. The multi-dimensional mapping relationship constructed by deep learning can be trained using historical data from a large model atomic agent database. The output results are used to evaluate the reliability of the simulation results during the evaluation phase, and its parameters are dynamically optimized by an incremental learning mechanism triggered by a data update window. The multi-dimensional mapping relationship constructed by deep learning can include, but is not limited to, parameter-feedback mapping models, scenario-feedback mapping models, and policy combination-feedback mapping models.
[0090] Deep learning is applied to a large-scale atomic agent database to obtain mapping relationships between basic parameter information, operational scenario information, and strategic decision design information, and the performance feedback information after the atomic agents run. This also includes mapping relationships between the fitted basic parameter information, operational scenario information, and strategic decision design information, and the performance feedback information after the atomic agents run. This can be achieved by using structured records in the database as training samples, with encoded basic parameters, scenario information, and strategy features as inputs, and corresponding performance feedback metrics as outputs. A neural network is then trained to establish a nonlinear mapping function. Furthermore, this operation can be achieved by using a multi-task learning framework to simultaneously predict multi-dimensional feedback metrics such as response latency, accuracy, and compliance deviation, thereby transforming the assessment of strategic decision effectiveness from a qualitative description to a quantifiable prediction.
[0091] A data update window is set up to collect real-time data on the thinking performance of the online atomic agents, and the collected data is entered into the large model atomic agent database. The large model atomic agent database is dynamically updated according to the data update window, and the mapping relationship is iteratively optimized.
[0092] The data update window can define the time interval and data filtering rules for collecting the actual thinking performance of atomic agents during online operation and writing it into the database. This can be used to achieve a data loop between the simulated environment and the real operating environment, ensuring that the mapping relationship is continuously updated as the environment evolves. In this embodiment, the data update window can dynamically adjust the sampling period based on system load and interaction frequency, retaining only thinking trajectory data with significant performance differences or abnormal patterns, which are then added to the database after verification. The data update window can inject online operating data into the large model atomic agent database, triggering incremental training of the mapping relationship and ensuring that the model remains synchronized with real-world behavior. The data update window can include, but is not limited to, high-frequency update windows, anomaly-triggered update windows, and periodic batch update windows.
[0093] A data update window is set up to collect real-time data on the thinking performance of online atomic agents and dynamically update the database. This can be achieved by capturing intermediate states and final outputs of the atomic agent's thought process during actual execution, filtering them according to the data update window rules, and writing them into the database. Furthermore, this operation can be implemented by capturing the input / output and execution time of the atomic agents through a lightweight agent module, uploading only data samples that deviate from historical distributions. This ensures that the database continuously reflects the actual operating status and prevents mapping relationships from becoming invalid due to data expiration.
[0094] The reliability of monitoring the phase simulation results is evaluated based on the mapping relationship, and monitoring evaluation results are obtained.
[0095] Evaluating the reliability of monitoring the stage simulation results based on the mapping relationship can be achieved by inputting the predicted performance indicators of the stage simulation output into the mapping relationship model, comparing its consistency with the distribution of actual historical feedback, and outputting a confidence score. Furthermore, this operation can be implemented by calculating the KL divergence between the simulated predicted values and the historical predicted mean of the mapping model as a reliability metric. This allows for assigning a quantifiable credibility label to the simulation output, making the monitoring results verifiable rather than merely observational records.
[0096] Taking a financial risk control Q&A system as an example, the large model-based thinking and processing method in this embodiment can be as follows: When simulating a user continuously querying loan risks, the system generates a stage simulation trajectory and outputs the expected compliance deviation rate; the deep learning mapping model predicts a typical deviation rate of 2.3% in this scenario based on historical data, the simulation result is 3.1%, and the confidence level is less than 70% after KL divergence calculation. The system marks the simulation path as low reliability and triggers a strategy retraining instruction; at the same time, such queries that actually occur during online operation are collected and written into the database, and the data update window triggers incremental learning of the mapping model to gradually correct the prediction deviation.
[0097] This embodiment constructs a multi-dimensional mapping relationship through deep learning, transforming the relationship between the input parameters of the atomic agent and the performance feedback into a predictable quantization function. By setting a data update window, an online feedback-driven closed-loop mechanism is formed, enabling the mapping relationship to continuously adapt to the evolution of real-world operation. Based on this mapping relationship, the confidence level of the stage simulation results is evaluated, achieving for the first time the verifiability and reliability calibration of the simulation process. This transforms the simulation trajectory, which was originally only used for playback, into a dynamic evaluation system with predictive capabilities and credibility measurement, supporting the dual improvement of strategy self-optimization and monitoring credibility. This overcomes the limitations of traditional models that cannot quantify evaluation and verify the effectiveness of simulations.
[0098] Furthermore, to achieve the above objectives, the present invention also provides a large-model-based thinking processing system, the system comprising: a memory, a processor, and a large-model-based thinking processing program stored in the memory and executable on the processor, the large-model-based thinking processing program being configured to implement the steps of the large-model-based thinking processing method as described in any one of the above descriptions.
[0099] In addition, to achieve the above objectives, the present invention also provides a medium storing a large-model-based thinking processing program, which, when executed by a processor, implements the steps of the large-model-based thinking processing method as described above.
[0100] Other embodiments or specific implementations of the large-model-based thinking processing system described in this invention can be found in the above-described method embodiments, and will not be repeated here.
[0101] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A thinking and processing method based on a large model, characterized in that, The method includes: Establish a large-scale model atomic agent database, which includes basic parameter information of atomic agents, running scenario information, thinking strategy design information, and performance feedback information after the atomic agents run; The operational scenario information includes user interaction scenario information and system resource scenario information, and the risk level of streaming monitoring is assessed based on the user interaction scenario information and system resource scenario information. Based on the basic parameter information and the thinking strategy design information, a streaming monitoring simulation environment is constructed for the atomic agent to simulate the thinking process of the atomic agent. The streaming monitoring risk level is loaded within the streaming monitoring simulation environment, and a time sampling window is set. Based on the time sampling window, streaming simulation is performed on several thinking stages of the atomic agent, and monitoring evaluation results are obtained through the stage simulation results.
2. The large-model-based thinking processing method as described in claim 1, characterized in that, The assessment of the streaming monitoring risk level based on the user interaction scenario information and system resource scenario information includes: The user load stages are divided according to the user interaction scenario information and sorted from high to low according to the interaction complexity to generate an interaction load sequence. Based on the system resource scenario information, the resource disturbance stages are divided and sorted from high to low according to the resource fluctuation amplitude to generate a resource disturbance sequence; The interactive load sequence and resource disturbance sequence are arranged and combined, and the arrangement and combination results are sorted according to the level of interactive load and resource disturbance to generate a comprehensive risk sequence. Based on the large model atomic agent database, a comprehensive risk sensitivity is set, and the comprehensive risk sequence is assessed for risk level through streaming monitoring using the comprehensive risk sensitivity.
3. The large-model-based thinking processing method as described in claim 2, characterized in that, The step of assessing the risk level of the comprehensive risk sequence through streaming monitoring using the comprehensive risk sensitivity includes: Traverse the comprehensive risk sequence to determine whether there is a risk item corresponding to the comprehensive risk sensitivity; If it exists, the comprehensive risk sensitivity is marked on the corresponding risk item, and the ratio of the comprehensive risk sensitivity to the risk sensitivity to the risk sensitivity to the risk sensitivity to the risk sensitivity to the risk level is determined. The risk level of the streaming monitoring is evaluated based on the ratio. If none exists, then select any of the aforementioned risk items and determine the relationship between the risk borne by the risk item and the overall risk sensitivity. If the risk level is higher than the overall risk sensitivity, the overall risk sequence is marked as a high-risk sequence; if it is lower than the overall risk sensitivity, the overall risk sequence is marked as a low-risk sequence.
4. The large-model-based thinking processing method as described in claim 3, characterized in that, The assessment of the streaming monitoring risk level based on the ratio includes: Obtain the expected performance index of the atomic agent, obtain the continuous performance of the atomic agent on each risk item based on the comprehensive risk sensitivity and ratio, and evaluate the stability performance based on the expected performance index. Set a mutation recovery time, which is the policy reset recovery time of the atomic agent after encountering an abnormal thinking node. Test the atomic agent with a risk item higher than the comprehensive risk sensitivity within a unit time to obtain the anti-interference performance. The risk level of the streaming monitoring is assessed based on the stability and anti-interference performance.
5. The large-model-based thinking processing method as described in claim 1, characterized in that, The establishment of the large-scale model atomic agent database includes: Collect historical operational data of atomic agents, including the design of various atomic agent thinking strategies; The historical running data of the atomic agent is classified according to the basic parameter information, running scenario information, thinking strategy design information and performance feedback information after the atomic agent runs, and a data index entry is established based on the classification results. A subset of thinking strategies is designed and constructed according to the thinking strategy of each atomic agent, and a historical trajectory tracking system is established for each thinking strategy scheme to form a large model atomic agent database, which is used by the streaming monitoring simulation environment to call data.
6. The large-model-based thinking processing method as described in claim 5, characterized in that, The data provided for the streaming monitoring simulation environment to access includes: The current user interaction scenario information and system resource scenario information are determined. The user interaction scenario information and system resource scenario information are matched by the index of the large model atomic agent database, and simulated scenario factors are constructed. The simulated scenario factors are digital representations of the running scenario information in the streaming monitoring simulation environment. The basic parameter information and thinking strategy design information of the current atomic agent are determined. The basic parameter information and thinking strategy design information are matched by indexing the large model atomic agent database. The time sampling window is divided into stages, and different time sampling factors are set for each stage according to the thinking complexity. Based on the determined current user interaction scenario information, system resource scenario information, and the basic parameter information and thinking strategy design information of the current atomic agent, the large model atomic agent database is matched, and the performance feedback information of the corresponding atomic agent after running is tracked to provide data support for the stage simulation results.
7. The large-model-based thinking processing method as described in claim 1, characterized in that, The streaming simulation of several thinking stages of the atomic agent based on the time sampling window includes: Set a time sampling factor to define the conversion ratio between the sampling interval and the actual running time during the simulated thinking process, thereby accelerating the simulation process to quickly obtain long-term thinking performance; Based on the actual operating scenario information and thinking strategy design information of the atomic agent, several specific thinking stages are defined. Each stage represents a core link in the thinking process of the atomic agent, and each thinking stage corresponds to a time sampling factor. At each defined thinking stage, the corresponding user interaction parameters and system resource parameters are loaded, and a streaming simulation is performed. By synthesizing the simulation results of all the aforementioned thinking stages, the thinking performance of the atomic agent in the hypothetical full lifecycle interaction is evaluated, and the simulation results of the aforementioned stages are obtained.
8. The large-model-based thinking processing method as described in claim 7, characterized in that, The process of obtaining monitoring and evaluation results through phased simulation results includes: Deep learning is performed on the large model atomic agent database to obtain the mapping relationship between the basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs, as well as the mapping relationship between the fitted basic parameter information, running scenario information, and thinking strategy design information and the performance feedback information after the atomic agent runs. A data update window is set up to collect real-time data on the thinking performance of the online atomic agent, and the collected data is entered into the large model atomic agent database. The large model atomic agent database is dynamically updated according to the data update window, and the mapping relationship is iteratively optimized. The monitoring reliability of the stage simulation results is evaluated based on the mapping relationship to obtain the monitoring evaluation results.
9. A thinking processing system based on a large model, characterized in that, The system includes: a memory, a processor, and a large-model-based thinking processing program stored in the memory and executable on the processor, the large-model-based thinking processing program being configured to implement the steps of the large-model-based thinking processing method as described in any one of claims 1 to 8.
10. A medium, characterized in that, The medium stores a large-model-based thinking processing program, which, when executed by a processor, implements the steps of the large-model-based thinking processing method as described in any one of claims 1 to 8.