Abnormity analysis method and device based on multi-agent cooperation, equipment and medium
By employing a multi-agent collaborative anomaly analysis method, the problems of incomplete data coverage and rigid decision-making were solved, enabling efficient integration and in-depth analysis of multi-source heterogeneous data, thereby improving the accuracy and adaptability of risk assessment.
Patent Information
- Application Number
- CN202511185257.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies in the fintech and healthcare sectors lack multi-agent collaborative closed-loop optimization mechanisms, resulting in incomplete data coverage, insufficient analytical depth, rigid decision-making, and an inability to dynamically adapt to changing environments, thus affecting the comprehensiveness and accuracy of risk assessment.
Multiple data acquisition agents collect data from various sources in parallel to generate multi-source heterogeneous data; a data fusion agent processes the data into fused data in a unified format; a data analysis agent extracts abnormal feature information; a decision agent generates abnormal analysis results and decision instructions; and a feedback optimization agent collects and compares feedback data after the execution of decision instructions to generate optimization instructions.
It enables efficient integration and in-depth analysis of multi-source heterogeneous data, improves the timeliness and reliability of anomaly identification and response, and enhances the accuracy and adaptability of analysis and decision-making.
Smart Images

Figure CN120929133A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an anomaly analysis method, apparatus, device, and storage medium based on multi-agent collaboration. Background Technology
[0002] In the fintech sector, particularly in risk control within property and casualty insurance, existing technologies exhibit significant shortcomings in data collection. Property and casualty insurance involves various insurance types and policyholder information from different regions and industries, coupled with a volatile market environment and complex claims records. Traditional technologies often only extract partial information from limited data sources, lacking the comprehensive ability to acquire open network information, structured business data, and real-time dynamic data. This insufficient data coverage results in incomplete information upon which risk assessment relies, thus hindering the comprehensiveness and accuracy of risk analysis.
[0003] In the healthcare sector, risk assessment and monitoring also rely on multi-dimensional and multi-source data, including patients' historical health records, real-time monitoring data, medical research information, and changes in the external environment. Current technologies are typically limited to data from a single system or a limited number of healthcare institutions, making it difficult to integrate heterogeneous data from different sources in a timely and comprehensive manner. This not only affects a comprehensive understanding of a patient's risk status but may also lead to delays in identifying key health risks.
[0004] In cross-domain data processing and analysis, existing technologies generally lack efficient data fusion capabilities. Data from different sources vary significantly in structure, format, and update frequency. Existing processing methods are inefficient in terms of format unification, data alignment, and feature extraction, making it difficult to quickly integrate multi-source heterogeneous data into a unified format that can be directly used for analysis. Furthermore, in the data analysis process, traditional technologies often employ single analytical methods, making it difficult to perform targeted mining and in-depth correlation analysis of structured data, unstructured text data, and time-series data, thus limiting the ability to identify potential risk patterns and correlations.
[0005] In the decision-making process, existing technologies often rely on fixed rules or simple models to generate risk response strategies, lacking the ability to dynamically weigh and comprehensively assess the analysis results. When external environments or risk factors change, these fixed-rule models struggle to adjust decisions in a timely manner, leading to lagging response strategies and increasing the likelihood of missed or misjudged risks. This rigidity of the decision-making mechanism directly impacts business security and service quality in both the financial and healthcare sectors. Summary of the Invention
[0006] The main objective of this invention is to provide an anomaly analysis method, apparatus, device, and storage medium based on multi-agent collaboration, aiming to solve the technical problems of incomplete data coverage, insufficient analysis depth, rigid decision-making, and inability to dynamically adapt to changing environments caused by the lack of a multi-agent collaborative closed-loop optimization mechanism in existing technologies.
[0007] To achieve the above objectives, this invention provides an anomaly analysis method based on multi-agent cooperation, comprising:
[0008] Multiple data acquisition agents collect data from various sources in parallel to generate multi-source heterogeneous data.
[0009] The multi-source heterogeneous data is processed by a data fusion intelligent agent to generate fused data in a unified format;
[0010] The data analysis intelligent agent mines the fused data to extract abnormal feature information;
[0011] Based on the aforementioned anomaly characteristic information, the decision-making agent generates anomaly analysis results and decision instructions.
[0012] The feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent.
[0013] Furthermore, to achieve the above objectives, the present invention provides an anomaly analysis device based on multi-agent collaboration, comprising:
[0014] The data acquisition module is used to collect data from multiple sources in parallel through multiple data acquisition agents, generating multi-source heterogeneous data;
[0015] The data fusion module is used to process the multi-source heterogeneous data through a data fusion intelligent agent to generate fused data in a unified format;
[0016] The feature mining module is used to mine the fused data through a data analysis intelligent agent and extract abnormal feature information;
[0017] The decision generation module is used to generate anomaly analysis results and decision instructions based on the anomaly feature information by a decision-making intelligent agent;
[0018] The feedback optimization module is used to collect feedback data after the decision instruction is executed through the feedback optimization agent, compare and analyze the feedback data with the decision instruction, and generate optimization instructions for optimizing the data analysis agent and the decision agent.
[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multi-agent cooperative anomaly analysis program stored in the memory and executable on the processor, wherein when the multi-agent cooperative anomaly analysis program is executed by the processor, it implements the steps of the multi-agent cooperative anomaly analysis method as described above.
[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an anomaly analysis program based on multi-agent cooperation, wherein when the anomaly analysis program based on multi-agent cooperation is executed by a processor, it implements the steps of the anomaly analysis method based on multi-agent cooperation as described above.
[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for anomaly analysis based on multi-agent collaboration, comprising: multiple data acquisition agents collecting data from various sources in parallel to generate multi-source heterogeneous data; a data fusion agent processing the multi-source heterogeneous data into fused data in a unified format; a data analysis agent mining and extracting anomaly feature information from the fused data; a decision-making agent generating anomaly analysis results and decision instructions based on the anomaly feature information; and a feedback optimization agent collecting feedback data after the execution of the decision instructions and comparing it with the decision instructions to generate optimization instructions for optimizing the data analysis agent and the decision-making agent. This invention achieves closed-loop processing of data acquisition, fusion, analysis, decision-making, and feedback optimization through multi-agent collaboration. Different agents cooperate and divide tasks, enabling efficient integration and in-depth analysis of multi-source heterogeneous data. Combined with a feedback optimization mechanism, it continuously improves the accuracy and adaptability of analysis and decision-making, thereby significantly improving the timeliness and reliability of anomaly identification and response. Attached Figure Description
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0023] Figure 1 This is a schematic diagram of an application environment for an anomaly analysis method based on multi-agent collaboration in one embodiment of the present invention;
[0024] Figure 2 This is a flowchart illustrating an embodiment of the anomaly analysis method based on multi-agent collaboration of the present invention.
[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly analysis device based on multi-agent collaboration of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0029] The anomaly analysis method based on multi-agent cooperation provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can generate multi-source heterogeneous data by having multiple data acquisition agents on the user terminal collect data from various sources in parallel. A data fusion agent processes this multi-source heterogeneous data into fused data in a unified format. A data analysis agent mines and extracts abnormal feature information from the fused data. A decision-making agent generates anomaly analysis results and decision instructions based on the anomaly feature information. A feedback optimization agent collects feedback data after the decision instructions are executed and compares it with the decision instructions to generate optimization instructions for optimizing the data analysis agent and the decision-making agent. This invention achieves a closed-loop processing of data acquisition, fusion, analysis, decision-making, and feedback optimization through multi-agent collaboration. Different agents cooperate, enabling efficient integration and in-depth analysis of multi-source heterogeneous data. Combined with a feedback optimization mechanism, the accuracy and adaptability of analysis and decision-making are continuously improved, thereby significantly enhancing the timeliness and reliability of anomaly identification and response. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.
[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multi-agent collaborative anomaly analysis method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0031] like Figure 2 As shown, the anomaly analysis method based on multi-agent cooperation proposed in this invention includes the following steps:
[0032] S10 generates multi-source heterogeneous data by having multiple data acquisition agents collect data from multiple sources in parallel.
[0033] In this embodiment, multiple data acquisition agents collect data from various sources in parallel. This requires establishing independent data acquisition channels for different data sources, with each acquisition channel corresponding to an agent capable of autonomous operation. The data acquisition agents can access open internet data platforms, enterprise internal structured databases, and high-speed, real-time data streaming systems using different data access technologies.
[0034] Parallel data acquisition refers to the simultaneous, independent operation of these intelligent agents in time. Through asynchronous data scheduling and independent thread or process execution mechanisms, data from different sources is simultaneously acquired and stored in the initial local or cloud storage area. Data from various sources includes structured, semi-structured, and unstructured data, such as tabular transaction logs, tagged business records, text-based application information, image-based claims materials, and time series data generated by high-frequency sensors.
[0035] In the process of generating multi-source heterogeneous data, it is necessary to preserve the original data from different sources in their native format and attach metadata such as source identifiers, timestamps, and acquisition channel identifiers to ensure that subsequent processing stages can identify the data's source background and acquisition timeline. The data acquisition agent should have automatic error detection and abnormal data removal functions, and use a data prediction mechanism to determine whether data packets are complete, whether fields are missing, and whether encoding is correct, ensuring that the data entering the fusion stage has high availability and accuracy.
[0036] The parallel architecture relies on a distributed task scheduling and load balancing mechanism to dynamically allocate the priority of collection tasks based on factors such as network latency, data source response time, and bandwidth usage, preventing the blocking of a single data source from affecting the overall collection efficiency.
[0037] In the fintech business, multiple virtualized data acquisition agents can be deployed in a cloud computing environment. Each agent is bound to different API interfaces, database connections, or web crawler modules, and parallel operation is achieved through multithreading or container orchestration. For example, one agent monitors incremental data changes in a bank's transaction history database, another agent crawls real-time index data from a market information interface, and a third agent crawls data from news media and policy announcement web pages. After different types of data are collected, they are directly written to a distributed storage system, generating a unified metadata index for easy subsequent retrieval and analysis.
[0038] In the healthcare sector, multiple intelligent data acquisition units can be configured to establish long-term connections with medical institution HIS systems, regional health information platforms, and wearable device data centers. This enables the synchronous acquisition of electronic medical records, laboratory test results, and real-time vital sign monitoring data. After initial data cleaning at edge computing nodes, the data is uploaded to central storage. In high-throughput environments, a stream processing framework can be introduced, allowing real-time data to directly enter a message queue system after acquisition, providing low-latency access for subsequent integrated intelligent agents. The acquisition frequency can also be adjusted according to business scenarios; for example, increasing the data acquisition rate for financial transactions during peak trading periods and increasing the data acquisition frequency from wearable devices during major public health events.
[0039] This embodiment deploys multiple data acquisition agents in parallel and establishes independent acquisition mechanisms for different sources, enabling simultaneous coverage of various heterogeneous data sources in both time and space, thus improving the comprehensiveness and timeliness of data acquisition. The distributed parallel architecture avoids the performance bottleneck of a single channel and possesses high scalability and fault tolerance. Adding source identifiers and metadata during the acquisition process provides rich contextual support for subsequent fusion and analysis, helping to improve the accuracy and flexibility of risk identification and trend analysis.
[0040] S20, the multi-source heterogeneous data is processed by the data fusion intelligent agent to generate fused data in a unified format;
[0041] In this embodiment, when the data fusion agent processes multi-source heterogeneous data, it needs to design a multi-stage preprocessing and transformation process to address the differences in data source, data type, and data structure. Multi-source heterogeneous data refers to a collection of information from multiple independent channels, using different encoding rules, and with varying data structures, such as structured business database tables, semi-structured log files, and unstructured text and multimedia content. The processing first classifies and sorts the data from different sources according to source identifiers and timestamps to ensure that different data within the same time interval can semantically match and correspond.
[0042] In text data processing, steps such as word segmentation, stop word removal, and synonym normalization can be used to standardize natural language information. Then, vectorization methods convert the text into numerical vectors, facilitating direct use by subsequent computer algorithms. In structured data processing, field mapping can be performed to unify fields with the same meaning but different names from different databases or tables into a standard field set; for example, mapping "Customer ID" and "User Number" to the same field identifier. In time series data processing, timestamp alignment is required to ensure that data from different sampling frequencies can be compared and analyzed using a unified time base; for example, synchronizing millisecond-level sensor readings with minute-level market data to the same time scale.
[0043] During the integration phase, the fusion agent reorganizes the vectorized, field-mapped, and time-aligned data according to a unified data model, generating fused data with consistent structure and uniform encoding. This fused data should have complete field descriptions, data type definitions, and consistent encoding rules, facilitating direct reading by subsequent analysis models.
[0044] In the fintech business, multi-source heterogeneous data can be input into a distributed data fusion framework. The fusion agent calls the text processing module to perform word segmentation and vectorization processing on the news and public opinion data, calls the mapping module to standardize customer information fields in different banking systems, and calls the time alignment module to align real-time transaction data and market index data to a second-level time scale. Finally, a unified transaction analysis data table is generated and stored in a central data warehouse.
[0045] In the healthcare sector, fusion agents can be used to standardize and vectorize electronic medical record text using medical terminology, map test item fields from different medical institution laboratory systems to a unified test code table, and align wearable device monitoring data such as heart rate and blood pressure with the timeline of inpatient medical records. This ultimately generates a fusion dataset encompassing structured indicators, text descriptions, and continuous monitoring data for clinical auxiliary analysis or health risk prediction. Furthermore, fusion rules can be adjusted for specific business scenarios; for example, in emergency medical events, real-time monitoring data can be prioritized, while the fusion of historical records can be delayed to improve timeliness.
[0046] This embodiment transforms multi-source heterogeneous data into a unified format through a fusion agent, eliminating differences in data type, structure, and time base, and achieving seamless integration of cross-source data. This unified format of fused data facilitates direct access by various types of analysis models, improving data utilization and analysis efficiency, reducing the workload of manual processing and cleaning, and enhancing the accuracy and scalability of risk analysis and pattern recognition.
[0047] S30, the data analysis agent mines the fused data and extracts abnormal feature information;
[0048] In this embodiment, when the data analysis agent mines the fused data, it needs to adopt differentiated analysis strategies for different types of content in the fused data and form abnormal feature information through result correlation. The fused data is a cross-source information collection that has undergone format unification and structure standardization, and usually contains multi-dimensional content such as structured numerical values, vectorized text, and time series. During the analysis process, the fused data is first loaded from the central data storage location, and the processing path is divided according to the data type.
[0049] For structured numerical data, cluster analysis can be used to divide data samples into different clusters using metrics such as Euclidean distance and cosine similarity, identifying outliers or anomalous groups that differ significantly from the majority of the data. For example, in financial transactions, multidimensional clustering of customer transaction characteristics can identify customer groups whose transaction patterns differ significantly from similar groups. In healthcare, clustering of patient test results can mark anomalous groups whose results differ excessively from those of the general population.
[0050] For vectorized text data, deep semantic analysis models can be used to extract the implicit topic distribution and semantic features in the text, calculate the deviation of the topic probability distribution from historical benchmarks, and identify semantically anomalous patterns. For example, in financial sentiment analysis, sudden negative topics can be detected; in medical record analysis, rare disease terms appearing in case descriptions can be identified.
[0051] For time series data, volatility pattern analysis can be performed. This involves applying methods such as trend decomposition, cycle analysis, and mutation detection to the time-aligned data to identify time points or periods outside the normal fluctuation range. Examples include unexpected and drastic price fluctuations in financial markets, or a patient's vital signs suddenly exceeding the normal range within a short period.
[0052] During the results integration phase, the analytical intelligence will cross-dimensionally correlate the abnormal data groups obtained from cluster analysis, the abnormal topics obtained from deep semantic analysis, and the abnormal fluctuation points obtained from fluctuation pattern analysis to form abnormal feature information. This information contains both specific abnormal content and retains its location identifier in the original data, so that the subsequent decision-making module can directly call it.
[0053] In the fintech business, fused data can be input into a multi-task analysis framework. The analytical agent calls the clustering analysis module to process a unified transaction data table and identify high-risk customer groups; the semantic analysis module to process public opinion text vectors and detect negative keywords with significantly increased frequency; and the volatility analysis module to process market time series and capture price changes exceeding volatility tolerance. The analysis results are then used by an association engine to generate anomaly feature information including high-risk customer groups, negative public opinion themes, and abnormal fluctuation times, which serves as input for subsequent risk control.
[0054] In the healthcare field, fused data can be input into a clinical analysis platform. The analytical agent then uses a clustering analysis module to group patient test result data, identifying patient groups at potential risk; a semantic analysis module to process medical record text vectors, discovering symptom combinations with abnormal frequency; and a fluctuation analysis module to assess trends in continuous monitoring data, marking time points of acute changes. The analysis results, combined with patient ID, timestamp, and test item number, form complete abnormal feature information for use by subsequent diagnostic support systems.
[0055] The analysis parameters can also be adjusted according to the scenario. For example, the cluster radius and fluctuation detection threshold can be reduced when high sensitivity detection is required, and these thresholds can be increased when false alarms need to be reduced. A rule filtering mechanism can also be introduced to remove abnormal patterns that are acceptable to the business.
[0056] This embodiment analyzes the multi-path mining of fused data by the intelligent agent, enabling the simultaneous discovery of potential anomalies across numerical, semantic, and temporal dimensions. These anomaly information is then correlated and integrated to form cross-dimensional anomaly feature information. This processing method significantly improves the comprehensiveness and accuracy of anomaly detection, providing sufficient and precise data for subsequent risk assessment or intervention strategy formulation.
[0057] S40, the decision-making agent generates anomaly analysis results and decision instructions based on the anomaly feature information;
[0058] In this embodiment, when generating anomaly analysis results and decision instructions, the decision-making agent needs to analyze, quantify, and comprehensively evaluate multiple types of anomalous factors in the anomaly feature information, and select appropriate decision actions based on preset judgment criteria. Anomaly feature information typically includes anomalous data groups, anomalous themes, and anomalous fluctuation points output by the analysis agent; each type of factor carries risk signals of different dimensions.
[0059] First, the decision-making agent receives and parses anomalous feature information, treating anomalous data groups as a set of structured indicators, including attributes such as the number of group members and the degree of feature deviation; anomalous topics as semantic risk labels, including topic keywords, frequency of occurrence, and degree of difference from historical baselines; and anomalous fluctuation points as abrupt events in a time series, including attributes such as time location, amplitude, and duration. The parsing process requires the establishment of a data index to ensure that detailed information for each anomalous factor can be quickly accessed during subsequent processing.
[0060] After analysis, the decision intelligence will assign weights to each anomaly. Weight assignment can be based on historical risk event statistics, business expert experience, or weight distributions trained by machine learning models. For example, higher weights can be assigned to anomaly topics that are highly correlated with risk occurrences in historical events, medium weights can be assigned to anomalous data groups with a wide impact but small deviations, and higher weights can be assigned to sudden and far-reaching anomalous fluctuations.
[0061] Subsequently, the decision-making agent combines the weights of each anomalous factor with their corresponding quantified values to generate a comprehensive anomaly index. The quantified values can represent the degree of anomaly, such as the feature deviation ratio of anomaly groups, the semantic similarity shift of anomalous topics, or the standard deviation multiple of the amplitude of anomalous fluctuation points. The combination of weights and quantified values can be achieved through weighted summation, nonlinear function mapping, etc., to reflect the overall risk level.
[0062] After the comprehensive anomaly indicator is generated, it is compared with the preset decision threshold. When the comprehensive anomaly indicator exceeds the threshold, it indicates that the current risk level is unacceptable and proactive intervention measures are required, such as freezing the account, triggering manual review, or suspending certain types of transactions. When the comprehensive anomaly indicator does not exceed the threshold, a monitoring instruction is generated, the object is added to the continuous monitoring list, and the next review time is set.
[0063] When generating anomaly analysis results, the decision-making intelligence will record information such as comprehensive anomaly indicator values, judgment categories (intervention or monitoring), and a list of involved anomaly factors, forming a standardized structure that can be used in subsequent stages. The anomaly analysis results and decision instructions are output synchronously, providing directly usable operational basis for the feedback optimization module or execution system.
[0064] In the fintech business, anomaly information can be input into a decision-making AI, which analyzes high-risk customer groups within anomalous data clusters, negative public opinion events within anomalous themes, and sudden market fluctuations within anomalous volatility points. By analyzing historical default rates, it was found that events where negative public opinion and sudden volatility occurred simultaneously resulted in significant losses in 90% of cases; therefore, these two factors are given higher weight in the weighting process. When the combined anomaly indicators exceed a threshold, an intervention instruction to freeze high-risk accounts is generated, and the indicator values and triggering reasons are recorded.
[0065] In the healthcare field, abnormal characteristic information can be input into a decision-making intelligent agent, which analyzes high-risk patient groups within abnormal data clusters, rare symptom combinations within abnormal themes, and rapid changes in vital signs at abnormal fluctuation points. Historical case analysis reveals that the simultaneous occurrence of rare symptom combinations and rapid changes in vital signs often foreshadows rapid deterioration of the condition; therefore, these two types of factors are given higher weighting coefficients in the weighting allocation. When the comprehensive abnormal indicators exceed a threshold, an intervention instruction for immediate transfer to intensive care is generated, and detailed analysis results are pushed to the medical command system.
[0066] It can also dynamically adjust decision thresholds based on different business scenarios. For example, it can appropriately raise the threshold during periods of frequent market fluctuations to reduce misjudgments, and lower the threshold during periods of critical medical monitoring to improve sensitivity. At the same time, it can continuously update the weight allocation strategy through reinforcement learning models to adapt to new risk patterns.
[0067] This embodiment transforms multi-source, multi-dimensional anomaly information into quantifiable comprehensive risk indicators and automatically generates intervention or monitoring instructions under a unified judgment mechanism. By introducing a combination of weight allocation and quantification, it can maintain decision-making consistency while taking into account the differences in the impact of different risk factors, thereby significantly improving the accuracy and timeliness of risk response.
[0068] S50, the feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent.
[0069] In this embodiment, when the feedback optimization agent performs this processing step, it needs to continuously monitor the actual execution results after the decision instruction is issued and transform them into analyzable feedback data. This feedback data may come from multiple systems or external interfaces, such as business execution systems, monitoring sensors, transaction logs, external data sources, etc., covering the final state of the decision object, key event records during the execution process, and external condition information that affects the execution results.
[0070] After data collection is complete, the feedback optimization AI will match and analyze the feedback data based on the content of the original decision-making instructions and the execution objectives. The matching process first identifies key information such as the specific object of the decision-making instruction, the execution action, and the expected effect, and then locates the corresponding actual execution performance in the feedback data. This process requires the establishment of a unified identification system, such as achieving data alignment through a mapping relationship between decision-making instruction IDs and execution record IDs.
[0071] Comparative analysis involves calculating the differences between the actual execution results and the expected results in the decision-making instructions, obtaining numerical or categorical difference information. These differences may manifest as numerical deviations (such as the gap between actual completion time and planned time), event deviations (such as operations that should have been performed not being performed), or effect deviations (such as risk reduction being less than expected). During the difference calculation process, different comparison algorithms can be applied based on different types of indicators, such as using Euclidean distance to calculate continuous numerical differences, Boolean logic to compare execution status, and correlation coefficients to measure the consistency of results.
[0072] The results of variance analysis are correlated with information on dynamic environmental changes to identify external factors contributing to the deviations. This information can be provided through a separate environmental monitoring module, such as market fluctuations, policy adjustments, and changes in the state of external systems. By establishing a correlation model between variances and environmental variables, it is possible to determine whether the deviations are caused by changes in external conditions, thus providing a basis for subsequent optimization.
[0073] After the difference and cause analysis is completed, the feedback optimization agent will generate optimization instructions. These instructions contain two types of content: one type is for the data analysis agent, used to adjust the structure, parameters, or data processing logic of its analysis model, thereby improving its anomaly detection and feature extraction capabilities; the other type is for the decision-making agent, used to adjust its strategy selection, parameter weight allocation, or decision model structure to improve the accuracy and adaptability of subsequent decisions. Optimization instructions must adopt a standardized format so that the target agent can directly parse and execute them, for example, by issuing parameter update lists, strategy modification descriptions, or model replacement files.
[0074] In the fintech business, this process can be applied to credit risk management. After the decision-making agent issues a loan rejection instruction to the execution system, the feedback optimization agent tracks subsequent customer behavior data, market credit information, and changes in the bad debt rate to collect actual results. If it is found that a significant proportion of rejected customers have successfully obtained loans on other platforms and repaid on schedule, it indicates that the original analysis model may have overestimated the risk. The feedback optimization agent then issues optimization instructions to the data analysis agent to adjust the risk scoring model parameters and to the decision-making agent to correct the loan rejection threshold.
[0075] In the healthcare field, this process can be applied to remote patient monitoring. When the decision-making agent issues a transfer order to intensive care, the feedback optimization agent collects information on the patient's treatment effectiveness, changes in vital signs, and relevant environmental information (such as equipment operating status and nursing resource allocation) after transfer, comparing this data with the expected recovery curve. If it is found that the recovery speed of some patients is significantly lower than expected, and the difference is highly correlated with the instability of a certain type of equipment, the feedback optimization agent will issue optimization instructions to the data analysis agent to adjust the weight of the equipment data, and will also issue optimization instructions to the decision-making agent to improve the sensitivity to the status of that type of equipment.
[0076] It can also be applied to equipment maintenance decision-making in the manufacturing industry, providing feedback to optimize the intelligent agent to track equipment operation data after the execution of maintenance instructions, identify cases that have not reached the expected service life, and analyze the causes in combination with environmental factors, thereby optimizing subsequent maintenance strategies and fault prediction models.
[0077] This embodiment establishes a closed-loop optimization mechanism that traces back from execution results to the analysis and decision-making modules, enabling the system to automatically correct its analysis methods and decision-making strategies in the face of environmental changes or model deviations. By combining actual feedback with external environmental factors to generate targeted optimization instructions, the system can continuously improve analysis accuracy and decision-making effectiveness, and reduce the risk of misjudgments or omissions caused by fixed models.
[0078] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for anomaly analysis based on multi-agent collaboration, comprising: multiple data acquisition agents collecting data from various sources in parallel to generate multi-source heterogeneous data; a data fusion agent processing the multi-source heterogeneous data into fused data in a unified format; a data analysis agent mining and extracting anomaly feature information from the fused data; a decision-making agent generating anomaly analysis results and decision instructions based on the anomaly feature information; and a feedback optimization agent collecting feedback data after the execution of the decision instructions and comparing it with the decision instructions to generate optimization instructions for optimizing the data analysis agent and the decision-making agent. This invention achieves a closed-loop processing of data acquisition, fusion, analysis, decision-making, and feedback optimization through multi-agent collaboration. Different agents cooperate and divide tasks, enabling efficient integration and in-depth analysis of multi-source heterogeneous data. Combined with a feedback optimization mechanism, it continuously improves the accuracy and adaptability of analysis and decision-making, thereby significantly improving the timeliness and reliability of anomaly identification and response.
[0079] In one embodiment, step S10 above includes:
[0080] S101, the first data acquisition agent uses web crawling technology to obtain entity attribute data from open network data sources and generate open network raw datasets;
[0081] S102, the second data acquisition agent obtains abnormal feature data from the structured database based on interface call technology to generate a structured raw dataset;
[0082] S103, through the third data acquisition agent based on long connection monitoring technology, acquires environmental dynamic data from real-time data stream and generates real-time data stream raw dataset;
[0083] S104, Perform data pre-validation on the open network original dataset, the structured original dataset, and the real-time data stream original dataset to generate the open network valid dataset, the structured valid dataset, and the real-time data stream valid dataset, respectively;
[0084] S105, integrate the open network valid dataset, structured valid dataset, and real-time data stream valid dataset to generate multi-source heterogeneous data.
[0085] In this embodiment, the goal of parallel data acquisition by multiple data acquisition agents from various sources addresses the issues of insufficient information coverage and timeliness by employing a task decomposition and parallel execution approach. Multiple data acquisition agents refer to acquisition units operating independently within the same time window. Each acquisition unit is responsible for stable data source connections, data capture, initial screening, caching, status reporting, and failure retries. Parallel acquisition emphasizes simultaneously advancing the capture process from different sources under controlled network bandwidth, connection count, and acquisition quotas, maintaining balance through queues and rate limiters. The data from various sources includes three types of input channels: open network data sources, structured databases, and real-time data streams. These three types differ significantly in protocol, timing, and structure, constituting heterogeneous input. Multi-source heterogeneous data refers to a unified carrier after the convergence of the three types of channels, retaining source markers, time markers, and field alignment information to provide a traceable data foundation for subsequent processing.
[0086] The first data acquisition agent uses web crawling technology to obtain entity attribute data from open web data sources and generate a raw open web dataset. Web crawling technology here refers to the process of sending requests to publicly accessible web pages, parsing pages, extracting content, and handling anti-blocking measures, including both static page parsing and script-rendered page parsing. Open web data sources cover portal sites, industry news sites, corporate information disclosure sites, and public forums. Entity attribute data is structured or semi-structured content formed around object identification and relationship descriptions, including names, timestamps, geographic tags, industry tags, organizational relationships, event descriptions, and evidence fragments. During implementation, a crawl list, entry links, and pagination strategies are established for each target domain. Selectors or structure trees are used to locate target areas, extracting text, table, and image metadata. The extracted fields are mapped to unified key names to generate a raw open web dataset, which is then written into a metadata header containing the source domain and crawl time. Abnormal situations such as redirection, CAPTCHAs, and rate limiting are handled by retry strategies, waiting windows, and proxy pools. The crawling status is written to the monitoring channel for subsequent scheduling.
[0087] The second data acquisition agent uses API call technology to retrieve abnormal feature data from a structured database and generate a structured raw dataset. API call technology refers to the process of establishing a session, submitting query requests, and extracting records in pages through a controlled interface gateway, supporting both batch and incremental retrieval modes. The structured database includes internal business databases and authorized external data services. Abnormal feature data consists of numerical and enumerable fields accumulated around risk identification, such as account behavior fields, transaction marker fields, profile tag fields, and derived ratio fields. During implementation, authentication credentials and access whitelists are configured, field lists and field mapping relationships are defined, and retrieval is performed by time watermark or primary key cursor to ensure record order and version consistency. The returned results are written into the structured raw dataset according to the table structure, along with the source database name, table name, extraction time, and field version information for easy traceability and playback.
[0088] The third data acquisition agent uses long-connection monitoring technology to acquire environmental dynamic data from real-time data streams and generate a raw dataset for the real-time data stream. Long-connection monitoring technology here refers to establishing continuous sessions to receive continuously pushed data segments, maintaining a heartbeat, reconnection after disconnection, and offset confirmation. Real-time data streams include market data broadcasts, alarm subscriptions, sensor channels, and message queues. Environmental dynamic data reflects the time-varying state of the external environment, such as price fluctuations, public opinion intensity, equipment operating conditions, and policy event triggers. In implementation, a subscription is established for each topic, with a starting offset and maximum lag window set. Received data segments are timestamped and partitioned, and are indexed into the database according to both arrival order and time order to generate the raw dataset for the real-time data stream. To reduce the impact of jitter, window aggregation and a jitter threshold are introduced to convert sub-second noise into records with stable granularity.
[0089] The raw datasets from the open network, structured datasets, and real-time data streams undergo a data pre-validation process. The goal is to filter out unusable records and supplement necessary metadata. Data pre-validation includes field integrity checks, type consistency checks, value range validity checks, primary key or composite key uniqueness comparisons, timestamp monotonicity checks, and source availability checks. Field integrity checks determine if required fields are missing based on the field list for each dataset. Type consistency checks verify that text, numerical, and enumerated data formats match the field definitions. Value range validity checks use whitelists or range limits to identify outliers. Uniqueness comparisons identify duplicates at the primary key or composite key level. Timestamp monotonicity checks identify out-of-order and reverse-order records. Source availability checks record concatenation fluctuations and error rates. Records that pass pre-validation are written to the valid open network dataset, valid structured dataset, and valid real-time data stream dataset. Records that fail are isolated with reasons for failure for later review.
[0090] The valid datasets enter the integration phase to generate multi-source heterogeneous data. The integration phase performs source tag preservation, field alignment, and time-based alignment. Source tag preservation involves writing the source type and source identifier into the record header. Field alignment uses a mapping table to collapse synonymous fields to a unified key name, while preserving the original key names and mapping relationships to avoid semantic loss. Time-based alignment converts the timestamps of each record to a unified time zone and precision, maintaining two timelines: event time and arrival time. The three types of valid datasets are appended to a unified carrier. The carrier appends a version number and batch number to each record, forming replayable multi-source heterogeneous data. To support subsequent processing, the carrier stores word segmentation indexes for text fields, distribution summaries for numerical fields, and window statistics for time-series fields, all organized with appended indexes without altering the original records.
[0091] This embodiment continuously incorporates open network data sources, structured databases, and real-time data streams within the same time window through parallel acquisition, simultaneously improving coverage and arrival timeliness. Web crawling and API call technologies establish stable crawling paths for semi-structured and structured inputs respectively, while long-connection monitoring technology provides dynamic input in a continuous environment. These three channels work together to reduce blind spots. Data pre-validation removes invalid and abnormal records at the entry point, reducing the burden of subsequent processing and minimizing sources of misjudgment. The integration process, based on source tag preservation, field alignment, and time benchmark alignment, forms traceable multi-source heterogeneous data, maintaining its original form while providing a unified access format. The resulting input base is simultaneously improved in coverage, quality, and consistency, providing a stable starting point and a replayable chain of evidence for subsequent fusion, analysis, decision-making, and feedback optimization.
[0092] In one embodiment, step S20 above includes:
[0093] S201, the data fusion agent receives the open network effective dataset, structured effective dataset, and real-time data stream effective dataset from the multi-source heterogeneous data;
[0094] S202, Perform text vectorization transformation on the open network effective dataset to generate a vectorized feature dataset;
[0095] S203, Perform field mapping transformation operation on the structured valid dataset to generate a structured mapped dataset;
[0096] S204, Perform a timestamp alignment operation on the effective dataset of the real-time data stream to generate a time-aligned dataset;
[0097] S205, integrate the vectorized feature dataset, structured mapping dataset, and time-aligned dataset to generate fused data in a unified format and store it in a central data repository.
[0098] In this embodiment, the data fusion agent first establishes a stable data channel session with the multi-source input channels, receiving open network valid datasets, structured valid datasets, and real-time data stream valid datasets in batches. Each batch carries a source tag, acquisition time, batch number, and version number. After entering the receiving buffer, it undergoes idempotency verification and arrival order registration to avoid subsequent deviations caused by duplicate writing and out-of-order delivery. During the receiving process, a unified record identifier and field visibility list are established for the three types of data, clarifying the processing paths for text fields, numerical fields, time fields, and enumerated fields. The data body and metadata are stored separately for subsequent parallel processing.
[0099] Valid datasets from open networks are incorporated into a text vectorization pipeline. This pipeline comprises five sub-stages: cleaning, segmentation, normalization, representation construction, and compression. The cleaning stage removes footnotes, advertising remnants, control characters, and uninformative markers, retaining entity names, time expressions, place names, and organizational terms, which are then solidified using regular expressions or dictionaries. The segmentation stage employs a hybrid strategy of backtracking and sub-word segmentation for mixed Chinese and English text, generating a sequence of terms with positional numbers. The normalization stage performs case unification, synonym folding, stop word weakening, and numerical standardization to ensure that synonym expressions converge to a consistent label space. The representation construction stage maps the term sequence to vector representations, employing three paths: bag-of-words statistics, term weighting, or contextual representation. All outputs are fixed-dimensional vectors, with alignment indices from text paragraphs to vectors appended at the record level. The compression stage reduces storage and computational overhead through dimensionality compression or quantization, while maintaining an upper bound on the reconstruction error. The output is a vectorized feature dataset, where each record maintains a one-to-one or one-to-many mapping relationship with the original text fragment, and the source tag and time tag are preserved to ensure replayability.
[0100] The structured, valid dataset enters the field mapping and transformation pipeline. This pipeline relies on the field mapping table and code table to complete semantic and unit alignment. Semantic alignment standardizes field names through synonym field folding; for example, the upper limit of the amount and the upper limit of the insured amount are merged into a unified key name. Unit alignment unifies amounts, proportions, durations, frequencies, etc., into agreed units and records the original units and conversion coefficients to preserve reversibility. Type normalization converts strings into numerical values, encodes discrete text, and parses time text into timestamps, strictly handling null values, placeholder values, and outliers, using clearly defined placeholders and valid flags to avoid silent discarding. Cross-table fields are concatenated at the row level or filled with dimension tables using primary keys or composite keys to form wide tables or maintain narrow tables with foreign key indexes; both forms record lineage information. The output is a structured mapped dataset containing standardized fields, numerical values with consistent units, enumerated code results, and strict primary key consistency.
[0101] The valid dataset of real-time data streams enters the timestamp alignment pipeline. This pipeline operates simultaneously on both the event time and arrival time timelines, first performing timezone unification and clock skew correction, then reordering and filling gaps based on event time. The alignment window employs a sliding or flipping strategy, setting a maximum out-of-order tolerance time and a delay arrival level. Late data exceeding the level enters the fill-in channel and is marked with a fill-in flag. Multiple sources are resampled according to a unified reference clock to obtain a sequence of segments with a unified sampling interval, while retaining the original granularity of snapshot indexes to avoid loss of detail. The same entity across sources is time-aligned using entity identifiers and time proximity rules, outputting a time-series aligned dataset to ensure the comparability and aggregability of features within the same time slice.
[0102] The three types of transformation products enter the integration channel, where joint integration of field alignment, time alignment, and entity alignment is performed. Field alignment is based on a unified columnar dictionary to merge fields with the same name and maintain a namespace strategy that avoids field conflicts, thus preventing overwriting. Time alignment uses the time-series aligned dataset as an anchor point, aggregating or aligning the vectorized feature dataset and the structured mapping dataset within the same time slice. The window strategy supports scrolling and segmentation, ensuring that information from different sampling frequencies can be observed synchronously. Entity alignment resolves entity ambiguities in multi-source datasets through entity key and similarity matching, and maintains the merge log and confidence score. Finally, a unified format of fused data is generated, using a hybrid layout of columnar storage and row-based indexing, supporting both batch analysis and random access by entity and time. The fused data, along with version numbers, build batches, and processing pipeline parameters, is written to a central data repository. A partitioning strategy is used to bucket the data by date and source, version tags are used to support incremental overwrite and historical playback, and atomic writes and verification digests are used to ensure consistency and verifiability.
[0103] This embodiment completes the most suitable standardization path for each of the multiple input sources before entering the fusion process. Text vectorization transformation enables semi-structured text to have a computable representation, field mapping transformation aligns the semantics of structured fields with units, and timestamp alignment unifies streaming data with different rhythms to a comparable time base. After the three pipelines converge, consistency is achieved in the three dimensions of fields, time, and entities. The unified format of the fused data improves expressiveness, alignment, and traceability at the same time.
[0104] In one embodiment, step S30 above includes:
[0105] S301, through a data analysis agent, obtains fused data in a unified format from a central data repository;
[0106] S302, perform clustering analysis on the structured mapping dataset in the fused data to generate abnormal data groups;
[0107] S303, Perform deep semantic analysis on the vectorized feature dataset in the fused data to generate anomalous topics;
[0108] S304, Perform fluctuation pattern analysis on the time-aligned dataset in the fused data to generate abnormal fluctuation points;
[0109] S305, associate the abnormal data group, abnormal topic and abnormal fluctuation point to generate abnormal feature information.
[0110] In this embodiment, the data analysis agent retrieves fused data in a unified format from the central data repository via pull or subscription. During reading, it simultaneously loads metadata lists, version tags, and partition information, and establishes temporary indexes by entity dimension and time sharding to ensure that subsequent numerical, textual, and time-series sub-pipelines can run in parallel on a consistent sample set. The unified format fused data undergoes a read-only integrity check and field visibility check upon entry, removing columns with incompatible permissions and recording masking mappings to prevent downstream processing interruptions.
[0111] The structured mapping dataset is fed into the distributed clustering pipeline. First, scaling and robust outlier standardization are performed. Continuous variables are scaled using quantiles, enumerated variables are encoded using target encoding or binary expansion, and missing data is preserved using a backtrackable mask. Then, a similarity metric space is constructed, allowing for the selection of Euclidean distance, Mahalanobis distance, or a learning-based metric. Clustering operators are chosen based on data density and cluster morphology. Density-based methods identify high-density regions and mark outliers by combining neighborhood radius and minimum number of points. Graph partitioning methods minimize cutting costs based on k-nearest neighbor graphs or shared nearest neighbor graphs. Hierarchical methods construct cluster trees layer by layer using agglomeration or splitting strategies. To avoid splitting caused by random initial values, multiple initializations and stability voting are used to achieve consistent allocation, and compactness, separation, and representativeness are calculated for each cluster. Finally, outlier data groups are output, marking low-density or structurally abrupt regions as outlier candidates, and attaching an impact factor and interpretability score to each candidate to support subsequent cross-modal association.
[0112] Vectorized feature datasets are fed into the deep semantic analysis pipeline. Text vectors first capture contextual relationships through sequence modeling or graph-text hybrid modeling. Sequence pathways can employ attention structures to obtain long-range dependencies, while graph pathways can perform information propagation on entity co-occurrence graphs to extract relational semantics. Subsequently, topic or semantic aggregation is performed, using probabilistic topic decomposition or subspace clustering methods to obtain a set of semantic topics, characterized by topic word distribution, representative sentences, and topic confidence. To enhance anomaly sensitivity, a comparison target is constructed, comparing the semantic distribution of the target period with that of the reference period or reference population to highlight topics with sudden increases or mutations, forming a set of anomalous topics. For multi-source texts of the same entity, semantic merging and conflict reconciliation strategies within a time window are used to retain contradictory evidence rather than simple overwriting, in order to explain the anomaly formation process in subsequent time dimensions.
[0113] The time-aligned dataset is then incorporated into the fluctuation pattern analysis pipeline. Noise suppression and band decomposition are performed first, and robust filtering or sparse decomposition yields trend, seasonal, and residual components. Anomaly detection is then performed on the residuals and local trends. Short-term abrupt changes can be identified using sliding statistics and threshold adaptive strategies; slow drifts can be identified through online change point detection or hidden Markov state switching; and periodic disruptions can be determined through spectral peak shift and phase disorder. Detection results at different granularities are consistently fused along the time axis, providing timestamps, amplitudes, durations, and local causal clues for anomalous fluctuation points, such as the count of external events or semantic intensity changes of the same entity within the anomalous time window. To reduce false alarms, a multi-window cross-validation and late data re-checking mechanism is used to confirm initial anomaly detections.
[0114] Cross-modal association is performed on a unified entity key and time slice. First, an alignment view of three types of outputs is constructed: anomalous data groups are mapped to the entity dimension via cluster centers and entity member sets; anomalous topics are mapped to the same dimension via text records and entity mapping tables; and anomalous fluctuation points are already located on the time axis. Then, association metrics are calculated on each time slice and entity set, including co-occurrence strength, conditional mutual information, cross-modal consistency score, and causal lag correlation, identifying synchronous or leading relationships. For cases with time lag, a variable lag sliding window is used to search for the optimal alignment offset within a finite lag range, and overfitting is suppressed using Bayesian or information criteria. Association outputs are expressed as triples or graph structures, with nodes representing anomalous data groups, anomalous topics, and anomalous fluctuation points, and edge weights representing association strength and directionality indicators, along with interpretable evidence. The data analysis agent compresses the above graph structure or summary features into vector or tabular forms, forming anomalous feature information, including entity identifiers, time slices, modality labels, strength, evidence citations, and version numbers, ensuring that downstream decision-making and feedback optimization can directly consume and are traceable.
[0115] This embodiment splits fused data into three highly adaptable pipelines—numerical, semantic, and temporal—in a unified format. Each pipeline completes the most suitable modeling and anomaly extraction, and then aligns and correlates them along the two common dimensions of entity and time. Anomaly evidence is transformed from single-point signals into multimodal synthetic signals. This approach yields quantifiable benefits. First, the joint occurrence of anomaly data clusters, anomaly themes, and anomaly fluctuations significantly increases confidence and reduces misjudgments caused by single-modal noise. Second, robust preprocessing and stability voting mechanisms in the sub-pipelines reduce sensitivity to initial values and noise, ensuring consistent output across multiple batches of data and different time windows, facilitating strategy reuse and monitoring. Third, the correlation process incorporates time-delay search and causal orientation evaluation, enabling anomaly feature information to not only describe static anomalies but also provide possible trigger sequences and impact paths. This provides a structured basis for subsequent weight allocation and threshold decisions, thereby shortening the closed-loop time from analysis to decision and improving overall recognition accuracy.
[0116] In one embodiment, step S40 above includes:
[0117] S401, the decision-making intelligent agent analyzes the abnormal data groups, abnormal topics and abnormal fluctuation points in the abnormal feature information;
[0118] S402, assign weights to the abnormal data group, the abnormal topic, and the abnormal fluctuation point respectively, and generate weight allocation results;
[0119] S403, combine the weight allocation result with the quantitative values of the abnormal data group, the abnormal topic and the abnormal fluctuation point to determine the comprehensive anomaly index;
[0120] S404, compare the comprehensive anomaly index with the preset decision threshold. When the comprehensive anomaly index exceeds the preset decision threshold, generate an intervention instruction; otherwise, generate a monitoring instruction.
[0121] S405, the intervention command or monitoring command is used as a decision command, and an anomaly analysis result containing the comprehensive anomaly indicators is generated.
[0122] In this embodiment, after receiving anomalous feature information, the decision-making agent first performs structural analysis, establishing a mapping table between anomalous data groups, anomalous topics, and anomalous fluctuation points based on entity identifiers and time slices. This ensures that multimodal entries under the same entity and time slice participate in the evaluation within the same computational unit. To avoid inconsistencies in dimensions affecting the synthesis calculation, interval normalization or robust standardization is performed on the quantization values of the three types of objects: density-based or cluster confidence is mapped to the zero-to-one interval; the topic intensity of semantic-based objects is converted into quantile scores based on the distribution of the reference corpus; and the fluctuation amplitude of temporal-based objects is converted into a dimensionless index based on the ratio of historical noise baseline to local fluctuations. The normalization strategy and the corresponding baseline window, reference population, or reference period are registered in the metadata to ensure subsequent traceability.
[0123] Weight allocation employs a configurable mechanism to generate results. Static weights are calculated and normalized using an expert pairwise comparison matrix, ensuring that the sum of weights for anomalous data groups, anomalous topics, and anomalous fluctuation points is equal. Adaptive weights are indirectly mapped from weight updates or action value tables passed through the feedback optimization process. Boundary pruning is performed under the constraints of evolving environmental data and business objectives to prevent extreme weighting in a single modality. To enhance robustness, lower and upper bounds are introduced, and exponential smoothing is applied to nearest-neighbor time slices to reduce significant weight oscillations.
[0124] The comprehensive anomaly index is calculated on the same entity and time slice, using the weighted allocation results as coefficients to combine the three types of quantized values into a composite index. Weighted summation or a nonlinear combination of bit-wise maximum and multiplication by a weight overflow penalty term can be used to enhance the alerting capability of single-modal extreme anomalies. When modalities are missing, the visible modal weights are renormalized and a missing mask is recorded to ensure the composite index remains within the zero-to-one range. An uncertainty metric, derived from quantized value variance propagation or Monte Carlo resampling, is appended after synthesis for reference during threshold comparison.
[0125] Threshold comparison employs a threshold band and hysteresis control strategy. Preset decision thresholds can be a single threshold or dual thresholds (entry threshold and exit threshold) to avoid frequent switching between intervention and monitoring. Under the dual threshold configuration, an intervention command is generated when the comprehensive abnormal indicator first exceeds the entry threshold; the intervention state is maintained until the exit threshold is exceeded; when the comprehensive abnormal indicator falls between the two thresholds, a conservative judgment is made based on uncertainty and trend terms. To suppress occasional spikes, a minimum duration or minimum number of hits can be set; an intervention command is only triggered after consecutive hits reach a set threshold.
[0126] Instruction generation follows a parallel output of executable format and interpretable output. Intervention instructions include action type, target entity, effective time range, intensity or level, review window, and reassessment interval; monitoring instructions include continuous observation period, additional sampling requirements, and reassessment threshold. Anomaly analysis results are simultaneously encapsulated with comprehensive anomaly indicators, three types of quantitative values, weight allocation results, threshold hit evidence, lag status, and uncertainty, along with a version number and verification signature for auditing. The generation process records the calculation path and key intermediate quantities, facilitating feedback to the optimization stage for reconstructing and comparing the "expected result data."
[0127] This embodiment unifies the dimensionality of quantified values for anomalous data groups, anomalous themes, and anomalous fluctuation points, imposes robust constraints on the weight allocation results, and completes comparison and decision-making using threshold banding and hysteresis control. Anomalous signals are transformed from local evidence of a single modality into a composite indicator across modalities. In this way, the triggering of intervention commands relies on consistently aligned multi-source evidence and traceable weighting, reducing false alarms caused by sporadic noise or single-modal bias. The dual threshold and minimum duration mechanism reduces command jitter and improves execution stability. Simultaneously, it outputs anomaly analysis results containing comprehensive anomaly indicators and detailed evidence, facilitating subsequent feedback optimization to accurately reconstruct the desired behavior and quantify deviations, thereby forming a stable closed-loop update path and improving overall identification and handling efficiency.
[0128] In one embodiment, step S50 above includes:
[0129] S501 collects actual execution result data after the decision-making instructions are executed through feedback optimization agent, and monitors dynamic changes in the environment to generate environmental evolution data;
[0130] S502, Obtain historical decision-making instructions and their corresponding expected results data;
[0131] S503, Perform a difference analysis on the actual execution result data and the expected result data to generate a difference value;
[0132] S504, when the difference value exceeds the deviation tolerance threshold, the current decision execution cycle is marked as a decision deviation event;
[0133] S505, correlate the decision deviation event with the environmental evolution data to generate deviation correlation analysis results;
[0134] S506, Based on the deviation correlation analysis results, generate a model parameter adjustment instruction for the data analysis agent and a decision strategy update instruction for the decision-making agent;
[0135] S507, the model parameter adjustment instruction is sent to the data analysis agent, and the decision strategy update instruction is sent to the decision agent.
[0136] In this embodiment, when the feedback optimization agent receives the feedback data after decision execution, it categorizes the sources into execution entity logs, business system records, and operational information collected by third-party monitoring interfaces. It then performs multi-dimensional indexing based on instruction identifiers and timestamps, eliminating duplicate and missing entries to form actual execution result data. Based on this, it continuously subscribes to or polls environmental perception channels, including market data streams, external sensor networks, policy event announcements, and public opinion sources. The collected events, indicators, and state variables are integrated in chronological order to obtain environmental evolution data reflecting changes in external conditions over time. The granularity of the environmental evolution data is configurable, allowing for monitoring of high-frequency changes at the minute level or capturing macro trends on a daily or weekly basis.
[0137] Historical decision-making instructions and their corresponding expected results are obtained by querying the instruction archive and simulation backtesting database. Each record contains the weight allocation, state characteristics, threshold settings, and quantitative results predicted by the simulation model at the time the instruction was issued. During the comparison phase, the system aligns the actual execution result data with the corresponding expected result data according to the indicator dimensions and performs difference analysis on each indicator. The difference value can be calculated using simple absolute difference or normalized difference, or it can use methods such as Mahalanobis distance or dynamic time warping in a multi-dimensional space to capture the overall offset pattern. The difference value is compared with a preset deviation tolerance threshold, which can be dynamically adjusted according to business type and environmental volatility.
[0138] When the discrepancy exceeds the tolerance threshold, the current decision-making execution cycle is marked as a decision deviation event. The deviation event object not only includes the numerical difference but also records the hit indicator, time window, execution conditions, and environmental evolution data fragments. In the association analysis phase, by performing multivariate regression, association rule mining, or causal inference on the deviation event and environmental evolution data, significantly changing external variables at the time of the deviation are identified, and the deviation association analysis results are extracted. This result is structured as a list of influencing factors, correlation coefficients, lag effects, and confidence scores.
[0139] Based on the deviation correlation analysis results, the feedback optimization agent generates two types of instructions. For the data analysis agent, a model parameter adjustment instruction is constructed, which includes the name of the parameter to be adjusted (e.g., clustering threshold, feature weights), the target value or adjustment range, the applicable data domain, and the effective time. For the decision-making agent, a decision strategy update instruction is generated, specifying the strategy parameters to be adjusted (e.g., weight allocation coefficients, action value mapping), the update rules (immediate effect or phased application), and rollback conditions. The instructions are sent through a reliable messaging channel, along with a version number and verification information, ensuring that the recipient correctly applies the update and can revert to a previous version if necessary.
[0140] This embodiment systematically collects actual execution results and environmental change information after decision execution, and performs difference analysis with historical expected results. This allows for timely identification of deviations in strategy and model performance under real-world conditions. Correlation analysis between deviation events and environmental variables provides a basis for parameter and strategy updates, preventing the introduction of new instability factors through untargeted adjustments. Optimization instructions are precisely categorized into model parameter adjustments and decision strategy updates, ensuring that both the data analysis agent and the decision-making agent receive improvement guidance matching their functions. This forms a feedback-optimization-execution closed-loop mechanism, enhancing the overall system's adaptability and long-term stability to environmental changes and data characteristics.
[0141] In one embodiment, after step S50 above, the method further includes:
[0142] S601, receives decision strategy update instructions from feedback optimization agents through decision-making agents;
[0143] S602, the decision-making agent extracts the deviation correlation analysis results from the decision strategy update instruction;
[0144] S603, the decision-making agent updates the action value table of the reinforcement learning decision-making model based on the deviation correlation analysis results;
[0145] S604, the decision-making agent optimizes the weight allocation strategy based on the updated action value table;
[0146] S605, the decision-making agent generates subsequent decision instructions in the reinforcement learning decision model by applying an optimized weight allocation strategy based on the state space composed of environmental evolution data and abnormal feature information in the reinforcement learning decision model.
[0147] S606 receives model parameter adjustment instructions from the feedback optimization agent through the data analysis agent;
[0148] S607, the data analysis agent extracts the deviation correlation analysis results from the model parameter adjustment command;
[0149] S608, the data analysis agent adjusts the similarity threshold of the clustering analysis model based on the deviation correlation analysis results;
[0150] S609, the data analysis agent adjusts the feature weight parameters of the semantic analysis model based on the deviation correlation analysis results;
[0151] S610, the data analysis agent applies the adjusted clustering analysis model and semantic analysis model, and processes the subsequent fused data using the adjusted similarity threshold and feature weight parameters.
[0152] In this embodiment, after the feedback optimization agent generates two types of update information, the decision-making agent first completes instruction reception, verification, and storage. The reception process deduplicates messages, verifies signatures, and compares versions to ensure that duplicate deliveries do not trigger multiple updates, and aligns late messages with the current online model state using a time window. Subsequently, the instruction payload is parsed, and factors related to action selection, influence direction, confidence scores, and applicable scenario labels are extracted from the deviation correlation analysis results and mapped into an internal unified structure to facilitate subsequent update processes by category routing.
[0153] The action value table of the reinforcement learning decision-making model is updated in small steps in online mode. The update logic uses the results of bias correlation analysis as external learning signals, transforming the gains or losses reflected in the actual execution results into utility increments. Local adjustments are made to the neighborhood of states closest to the current environmental conditions and anomalous features to avoid instability caused by a one-time global rewrite. To reduce the impact of noise, a moving average and attenuation coefficient are introduced to integrate old and new information with time weights. To suppress oscillations, a freeze window and upper limit constraint are introduced to limit the magnitude of a single update. To reduce the disturbance of low-quality feedback, a minimum effective threshold is set through confidence scoring, and updates are only written when the evidence strength reaches the threshold.
[0154] The optimized weight allocation strategy is based on the updated action value table, transforming the value mapping from state to action into importance weights for each anomalous factor. Specifically, importance scores are first calculated based on value differences and uncertainty, then normalized using temperature control to map them into a weight vector. The temperature parameter is adjustable; lower temperatures produce sharper selections, while higher temperatures produce smoother weight distributions. To prevent bias caused by a single factor dominating for an extended period, minimum weight guarantees and upper limit pruning are introduced to ensure the effective participation of multi-source features. The strategy object includes a version number and effective range for easy canary deployment and rollback.
[0155] Subsequent decisions are generated within the reinforcement learning decision model. Environmental evolution data and anomaly feature information jointly form a state space vector, which enters a standardized pipeline for time alignment, missing data completion, and dimensional unification. The model performs exploration with a certain probability and utilization with the remaining probability. It reads the value assessment of candidate actions from the action value table and integrates the values according to an optimized weight allocation strategy to generate action scores and confidence levels. Under the premise of meeting safety and business constraints, the system outputs intervention or monitoring instructions, along with a summary of reasons and a snapshot of key weights, to facilitate subsequent auditing and tracking.
[0156] After receiving model parameter adjustment instructions, the data analysis agent completes signature verification, deduplication, and version consistency checks. It also analyzes the abnormal pattern types, affected feature domains, and suggested adjustment ranges related to clustering and semantic modeling in the deviation correlation analysis results. For the clustering analysis model, the similarity threshold is dynamically updated: the threshold is increased to enhance discriminative power when the deviation type is excessive merging, and decreased to enhance aggregation when the deviation type is excessive splitting. Maximum step size and a lookback window are set for threshold changes to avoid frequent jitter. For the semantic analysis model, feature weight parameters are increased or decreased at the channel or term level, using a combination of normalization and sparsity to highlight semantically strong cues while suppressing noisy channels. To reduce catastrophic migration, the new parameters are first applied to the shadow branch for parallel inference on the same batch of fused data, and then the main branch is switched according to the evaluation metric compliance rules.
[0157] The updated clustering and semantic analysis models work together when processing subsequent fused data. After data flows into the analysis pipeline, it first enters the temporal alignment and denoising module. Then, the clustering branch generates group partitions under the new similarity threshold, and the semantic branch extracts topics and clues under the new feature weights. Both perform consistency checks and conflict resolution at the multi-source alignment layer, forming structured intermediate results and writing them to the analysis cache. The system tracks key parameters and outputs, collecting drift metrics, threshold hit rates, and alarm quality in real time as the basis for the next round of feedback. The entire pipeline is linked by version identifiers, allowing for the location, replay, and rollback of any parameter or strategy update.
[0158] Example Description: In the intelligent disease risk management system of a large comprehensive medical institution, multiple data acquisition agents collect information from different medical data sources in parallel. The first data acquisition agent uses web crawling technology to crawl authorized medical literature databases, disease knowledge bases, and public health bulletin platforms, extracting entity attribute information such as clinical manifestations, pathological mechanisms, and epidemic trends of diseases to form an open web raw dataset. The second data acquisition agent, through interface calls with the hospital's electronic medical record system, extracts patient diagnosis codes, laboratory test results, medication records, and surgical records from a structured database to form a structured raw dataset. The third data acquisition agent uses long-connection monitoring technology to connect in real-time with the hospital's wearable device platform, bedside monitoring system, and mobile health terminals to collect dynamic environmental data such as the patient's heart rate, blood pressure, blood oxygen, body temperature, and respiratory rate, forming a real-time data stream raw dataset.
[0159] Before the data enters subsequent analysis, the system performs data pre-validation on the three types of raw datasets, removing data records with abnormal formats, missing key fields, incorrect value ranges, or abnormal timestamps, and generating corresponding open network valid datasets, structured valid datasets, and real-time data stream valid datasets. These valid datasets undergo unified integration processing to form multi-source heterogeneous data, covering patient static information, historical disease course, real-time physiological indicators, and external medical knowledge.
[0160] The data fusion agent performs categorized processing on multi-source heterogeneous data. Open network datasets are converted into vectorized feature datasets using medical natural language processing algorithms, preserving semantic information such as disease keywords, symptom correlations, and treatment plans. Structured datasets undergo field mapping transformation to unify data field standards across different medical institutions, departments, and systems; for example, blood pressure measurements are standardized to millimeter-mercury format. Real-time data stream datasets undergo timestamp alignment, aligning continuously monitored heart rate, blood oxygen, and other data by patient ID and time series to ensure accurate correspondence with medical records. Subsequently, the fusion agent integrates the three processing results to generate unified formatted fused data, which is then stored in a central data repository.
[0161] The data analytics agent retrieves fused data from a central data repository and performs targeted analyses on different subsets of the data. Structured mapping datasets use clustering analysis to identify anomalous data clusters, such as identifying a group of diabetic patients with recent abnormal fluctuations in blood glucose levels. Vectorized feature datasets extract anomalous topics through deep semantic analysis, such as specific drug-resistant strains frequently reported in recent hospital infection cases. Time-aligned datasets identify anomalous fluctuation points through fluctuation pattern analysis, such as periods when patients' heart rates frequently spike late at night. These three types of analytical results are correlated to form anomalous feature information, revealing potential disease risk patterns.
[0162] The decision-making agent analyzes anomalous feature information and assigns weights to anomalous data groups, anomalous themes, and anomalous fluctuation points. For example, it assigns higher weight to the anomalous theme of an outbreak of drug-resistant strains within the hospital, medium weight to blood pressure fluctuations in a single patient, and lower weight to a group of patients with high fever fluctuations across multiple departments. The weighting results are combined with the quantitative values of each anomalous factor to calculate a comprehensive anomalous index. When the comprehensive anomalous index exceeds a preset decision threshold, an intervention instruction is generated, such as immediately initiating infection control measures, isolating the patient, and conducting additional testing; otherwise, a monitoring instruction is generated, such as continuously monitoring blood pressure fluctuation trends over the next week. The final output decision instructions and anomalous analysis results are delivered to the clinical management department and the medical team.
[0163] The feedback optimization agent continuously collects data on the actual effects of decision-making instructions, such as changes in the detection rate of drug-resistant strains after intervention, changes in the secondary infection rate in isolation wards, and improvements in patients' blood glucose fluctuations. Simultaneously, it monitors real-time environmental dynamics such as ward temperature and humidity, air quality, and patient density, generating environmental evolution data. The system retrieves historical decision-making instructions and their expected results, performs difference analysis with the actual results, and generates difference values. When the difference value exceeds the deviation tolerance threshold, the period is marked as a decision deviation event and correlated with the environmental evolution data to obtain deviation correlation analysis results. These analysis results are used to generate two types of optimization instructions: one is a model parameter adjustment instruction for the data analysis agent, and the other is a decision strategy update instruction for the decision-making agent, which are then sent to the corresponding agents.
[0164] The decision-making agent receives policy update instructions, extracts the results of bias correlation analysis, and updates the action value table of the reinforcement learning decision model accordingly. For example, if the system finds that intervention measures for abnormal blood glucose fluctuations are ineffective in the long term, it lowers the value assessment of that action in the corresponding state space, while simultaneously increasing the policy value of adjusting drug dosage. Based on the new action value table, the weight allocation strategy is optimized, and new subsequent decision instructions are generated in the reinforcement learning model by combining the current environmental evolution data and abnormal feature information into the state space, making the balance between intervention and monitoring more consistent with the latest risk landscape.
[0165] The data analysis agent receives instructions to adjust model parameters, extracts the results of bias correlation analysis, and adjusts the similarity threshold of the clustering analysis model based on the prompts. For example, it appropriately increases the threshold to avoid mixing different types of abnormal patient groups together. Simultaneously, it adjusts the feature weight parameters of the semantic analysis model, such as increasing the weight of semantic features related to drug resistance and decreasing the weight of symptoms with low relevance. When processing new fused data, the updated clustering and semantic analysis models can more accurately identify potential risk patterns using the adjusted similarity thresholds and feature weight parameters, thereby forming a rapid early warning capability for nosocomial infections, acute exacerbations of chronic diseases, and major public health events.
[0166] Through this complete multi-agent collaborative closed loop, medical institutions can achieve continuous optimization in disease monitoring, hospital infection control, and chronic disease management, ensuring that multi-source data acquisition, fusion processing, in-depth analysis, intelligent decision-making, and self-improvement continue to operate in a dynamic medical environment.
[0167] In the intelligent risk control platform of a large insurance group, multiple data acquisition agents collect information from different financial data sources in parallel. The first data acquisition agent, based on web crawling technology, scrapes open network data sources such as publicly available financial information platforms, insurance industry association announcements, and enterprise credit information disclosure systems to obtain enterprise registration information, major litigation records, industry risk warnings, and market sentiment, forming an open network raw dataset. The second data acquisition agent, through interface calls with an internal structured database, obtains historical insurance records, claims records, insured property value assessments, and risk level ratings of policyholders and related enterprises, forming a structured raw dataset. The third data acquisition agent employs long-connection monitoring technology to continuously connect with real-time transaction monitoring systems, payment gateways, and risk control sensor nodes, collecting dynamic environmental data such as real-time policy changes, fund flows, and transaction anomaly alarms, forming a real-time data stream raw dataset.
[0168] Before the data enters subsequent analysis, the system performs data pre-validation on three types of raw datasets, removing records with format errors, missing data, abnormal timestamps, and inconsistent data sources, generating open network valid datasets, structured valid datasets, and real-time data stream valid datasets. These three types of valid datasets are then integrated to form multi-source heterogeneous data covering public market sentiment, internal business data, and real-time transaction monitoring information, providing a foundation for subsequent fusion processing and analysis.
[0169] The data fusion agent processes multi-source heterogeneous data according to their source characteristics. Open network effective datasets undergo text vectorization, transforming news headlines, announcement summaries, and company news into computable vector features while preserving semantic relevance information, such as the similarity between "significant debt risk" and "corporate bankruptcy application." Structured effective datasets undergo field mapping transformation, unifying data from different business subsystems with similar field meanings but different names; for example, mapping "compensation amount" and "claims expenditure" to a unified field. Real-time data stream effective datasets undergo timestamp alignment, unifying transaction logs, policy change records, and real-time alarms onto a precise timeline, ensuring synchronous analysis with batch data and external data. The fusion agent integrates the three types of processing results, generating unified formatted fused data, which is stored in a central data repository for use in the analysis phase.
[0170] The data analytics agent retrieves fused data from a central data repository and analyzes it by data type. The structured mapping dataset uses clustering analysis to group customer groups with similar risk characteristics, such as those experiencing a significant increase in surrender rates for high-value insurance policies recently. The vectorized feature dataset uses deep semantic analysis to extract trending topics in public opinion, such as the risk signal of "a natural disaster in a certain region leading to a large number of claims." The time-series aligned dataset uses fluctuation pattern analysis to identify abnormal fluctuations in transaction activity, such as a policyholder making multiple large policy changes within a short period outside of business hours. The analysis results, after correlation operations, generate abnormal feature information, presenting a multi-dimensional risk status.
[0171] The decision-making agent analyzes abnormal feature information and assigns weights to abnormal data groups, abnormal themes, and abnormal fluctuation points. For example, it assigns higher weights to abnormal themes involving high-payout risks, medium weights to groups of policies with concentrated surrenders in a short period, and lower weights to abnormal fluctuation points in individual transactions. Combining the weight allocation results with the quantitative values of each risk factor, a comprehensive anomaly index is calculated. When the comprehensive anomaly index exceeds a preset decision threshold, an intervention instruction is generated, such as freezing some policy operations, initiating manual review, or adjusting underwriting strategies; otherwise, a monitoring instruction is generated, such as continuously tracking the transaction behavior of the customer group. The decision instructions and anomaly analysis results are then distributed to the risk control department and business operations team.
[0172] The feedback optimization agent continuously collects actual result data after the execution of decision-making instructions, such as changes in surrender rates, fraudulent transaction volumes, and claims trends after intervention measures. Simultaneously, it monitors changes in the insurance market environment, such as reinsurance rate adjustments, industry policy changes, and natural disaster warnings, generating environmental evolution data. The system retrieves historical decision-making instructions and expected results, performs difference analysis with actual results, and generates difference values. When the difference value exceeds the deviation tolerance threshold, the period is marked as a decision deviation event and correlated with the environmental evolution data to form a deviation correlation analysis result. This result is used to generate two types of optimization instructions: model parameter adjustment instructions for the data analysis agent and decision strategy update instructions for the decision-making agent, which are then sent to the corresponding agents.
[0173] The decision-making agent receives policy update instructions, extracts the results of bias correlation analysis, and updates the action value table of the reinforcement learning decision model. For example, if historical data shows that the strategy of directly freezing policies for customers with high surrender rates is ineffective, the value of this strategy in the corresponding state space is reduced, while the value of strategies that add a policy cooling-off period or provide additional policy consultation is increased. Based on the updated action value table, the weight allocation strategy is optimized, and new subsequent decision instructions are generated in the reinforcement learning model by combining the current environmental evolution data and anomaly feature information into the state space. This allows risk control decisions to adaptively adjust to changes in market and customer behavior.
[0174] The data analysis agent receives instructions to adjust model parameters, extracts the results of bias correlation analysis, and adjusts the similarity threshold of the clustering analysis model. For example, it increases the threshold under highly volatile market conditions to distinguish between short-term market disturbances and long-term risk trends. Simultaneously, it adjusts the feature weight parameters of the semantic analysis model, such as increasing the weight of features related to natural disasters and policy adjustments, and decreasing the weight of features related to low-risk public opinion topics. When processing new fused data, the updated clustering and semantic analysis models can use the adjusted similarity thresholds and feature weight parameters to more accurately identify potential fraud risks and large payout risks, achieving dynamic optimization of risk identification capabilities.
[0175] Through this closed-loop operation, insurance groups can continuously optimize aspects such as policy management, fraud prevention, and claims risk management, maintaining flexibility and accuracy in risk decision-making when facing complex multi-source data and a rapidly changing market environment.
[0176] This embodiment extends the update chain to both the decision-making and analysis agents, forming a traceable closed loop from feedback data to action value, and from bias evidence to model parameters. In an online environment, it rapidly corrects the direction of value assessment and feature extraction, ensuring that subsequent instructions and detection results continuously align with the latest data distribution and environmental conditions. The small step size, amplitude limit, and confidence threshold of the action value table reduce oscillations caused by noisy feedback. The temperature control and constraints of the weight allocation strategy suppress the long-term dominance of a single factor, improving decision stability and interpretability. Adaptive adjustment of the clustering similarity threshold and semantic feature weights narrows the false positive and false negative range, maintaining a balance between analytical sensitivity and robustness in scenarios of environmental abrupt changes or concept drift. Overall, the coherent feedback-update-application process reduces manual intervention and offline retraining costs, accelerates convergence, and improves generalization performance and business availability across time periods and scenarios.
[0177] In one embodiment, an anomaly analysis device based on multi-agent cooperation is provided, which corresponds one-to-one with the anomaly analysis method based on multi-agent cooperation in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly analysis device based on multi-agent collaboration of the present invention. The modules include a data acquisition module 10, a data fusion module 20, a feature mining module 30, a decision generation module 40, and a feedback optimization module 50. Detailed descriptions of each functional module are as follows:
[0178] Data acquisition module 10 is used to collect data from multiple sources in parallel through multiple data acquisition agents to generate multi-source heterogeneous data;
[0179] The data fusion module 20 is used to process the multi-source heterogeneous data through a data fusion intelligent agent to generate fused data in a unified format;
[0180] Feature mining module 30 is used to mine the fused data through a data analysis intelligent agent and extract abnormal feature information;
[0181] The decision generation module 40 is used to generate anomaly analysis results and decision instructions based on the anomaly feature information by a decision-making intelligent agent;
[0182] The feedback optimization module 50 is used to collect feedback data after the decision instruction is executed through the feedback optimization agent, compare and analyze the feedback data with the decision instruction, and generate optimization instructions for optimizing the data analysis agent and the decision agent.
[0183] In one embodiment, the data acquisition module 10 is specifically used for:
[0184] The first data acquisition agent uses web crawling technology to obtain entity attribute data from open network data sources and generate an open network raw dataset.
[0185] The second data acquisition agent uses interface call technology to obtain abnormal feature data from a structured database and generate a structured raw dataset.
[0186] The third-party data acquisition agent uses long-connection monitoring technology to acquire dynamic environmental data from real-time data streams and generate raw datasets of real-time data streams.
[0187] Data pre-validation is performed on the open network raw dataset, the structured raw dataset, and the real-time data stream raw dataset to generate open network valid dataset, structured valid dataset, and real-time data stream valid dataset, respectively.
[0188] By integrating the open network valid dataset, structured valid dataset, and real-time data stream valid dataset, multi-source heterogeneous data is generated.
[0189] In one embodiment, the data fusion module 20 is specifically used for:
[0190] The data fusion agent receives open network valid datasets, structured valid datasets, and real-time data stream valid datasets from the multi-source heterogeneous data.
[0191] Perform text vectorization transformation on the open network effective dataset to generate a vectorized feature dataset;
[0192] Perform field mapping transformation operations on the structured valid dataset to generate a structured mapped dataset;
[0193] Perform timestamp alignment on the valid dataset of the real-time data stream to generate a time-aligned dataset;
[0194] The vectorized feature dataset, structured mapping dataset, and time-aligned dataset are integrated to generate fused data in a unified format and stored in a central data repository.
[0195] In one embodiment, the feature mining module 30 is specifically used for:
[0196] Data analytics agents acquire fused data in a unified format from a central data repository.
[0197] Cluster analysis is performed on the structured mapping dataset in the fused data to generate outlier data clusters;
[0198] Deep semantic analysis is performed on the vectorized feature dataset in the fused data to generate anomalous topics;
[0199] Perform fluctuation pattern analysis on the time-aligned dataset in the fused data to generate abnormal fluctuation points;
[0200] By associating the abnormal data groups, abnormal topics, and abnormal fluctuation points, abnormal feature information is generated.
[0201] In one embodiment, the decision generation module 40 is specifically used for:
[0202] The decision-making agent analyzes the abnormal data groups, abnormal themes, and abnormal fluctuation points in the abnormal feature information.
[0203] Weights are assigned to the abnormal data group, the abnormal topic, and the abnormal fluctuation point, respectively, and weight assignment results are generated.
[0204] By combining the weight allocation results with the quantitative values of the abnormal data groups, the abnormal themes, and the abnormal fluctuation points, a comprehensive anomaly index is determined.
[0205] The comprehensive anomaly index is compared with a preset decision threshold. If the comprehensive anomaly index exceeds the preset decision threshold, an intervention instruction is generated; otherwise, a monitoring instruction is generated.
[0206] The intervention or monitoring instructions are used as decision instructions, and anomaly analysis results containing the comprehensive anomaly indicators are generated.
[0207] In one embodiment, the feedback optimization module 50 is specifically used for:
[0208] The intelligent agent collects data on the actual execution results after the decision-making instructions are executed through feedback optimization, and monitors dynamic changes in the environment to generate environmental evolution data;
[0209] Obtain historical decision-making instructions and their corresponding expected results data;
[0210] Perform a difference analysis on the actual execution result data and the expected result data to generate a difference value;
[0211] When the difference value exceeds the deviation tolerance threshold, the current decision execution cycle is marked as a decision deviation event;
[0212] By correlating the decision-making bias events with the environmental evolution data, bias correlation analysis results are generated.
[0213] Based on the deviation correlation analysis results, generate model parameter adjustment instructions for the data analysis agent and decision strategy update instructions for the decision-making agent;
[0214] The model parameter adjustment instruction is sent to the data analysis agent, and the decision strategy update instruction is sent to the decision agent.
[0215] In one embodiment, the feedback optimization module 50 is specifically used for:
[0216] The decision-making agent receives decision strategy update instructions from the feedback optimization agent.
[0217] The decision-making agent extracts the deviation correlation analysis results from the decision-making strategy update instructions;
[0218] The decision-making agent updates the action value table of the reinforcement learning decision-making model based on the results of the deviation correlation analysis.
[0219] The decision-making agent optimizes the weight allocation strategy based on the updated action value table.
[0220] In the reinforcement learning decision-making model, the decision-making agent generates subsequent decision instructions by applying an optimized weight allocation strategy based on the state space composed of environmental evolution data and abnormal feature information.
[0221] The data analysis agent receives model parameter adjustment instructions from the feedback optimization agent;
[0222] The data analysis agent extracts deviation correlation analysis results from the model parameter adjustment instructions;
[0223] The data analysis agent adjusts the similarity threshold of the clustering analysis model based on the deviation correlation analysis results;
[0224] The data analysis agent adjusts the feature weight parameters of the semantic analysis model based on the deviation correlation analysis results.
[0225] The data analysis agent applies the adjusted clustering analysis model and semantic analysis model, and uses the adjusted similarity threshold and feature weight parameters to process the subsequent fused data.
[0226] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a multi-agent cooperative anomaly analysis method.
[0227] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements user-side functions or steps of a multi-agent cooperative anomaly analysis method.
[0228] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0229] Multiple data acquisition agents collect data from various sources in parallel to generate multi-source heterogeneous data.
[0230] The multi-source heterogeneous data is processed by a data fusion intelligent agent to generate fused data in a unified format;
[0231] The data analysis intelligent agent mines the fused data to extract abnormal feature information;
[0232] Based on the aforementioned anomaly characteristic information, the decision-making agent generates anomaly analysis results and decision instructions.
[0233] The feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent.
[0234] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0235] Multiple data acquisition agents collect data from various sources in parallel to generate multi-source heterogeneous data.
[0236] The multi-source heterogeneous data is processed by a data fusion intelligent agent to generate fused data in a unified format;
[0237] The data analysis intelligent agent mines the fused data to extract abnormal feature information;
[0238] Based on the aforementioned anomaly characteristic information, the decision-making agent generates anomaly analysis results and decision instructions.
[0239] The feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent.
[0240] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0241] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0242] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0243] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An anomaly analysis method based on multi-agent cooperation, characterized in that, Includes the following steps: Multiple data acquisition agents collect data from various sources in parallel to generate multi-source heterogeneous data. The multi-source heterogeneous data is processed by a data fusion intelligent agent to generate fused data in a unified format; The data analysis intelligent agent mines the fused data to extract abnormal feature information; Based on the aforementioned anomaly characteristic information, the decision-making agent generates anomaly analysis results and decision instructions. The feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent.
2. The anomaly analysis method based on multi-agent collaboration as described in claim 1, characterized in that, Multiple data acquisition agents collect data from various sources in parallel, generating multi-source heterogeneous data, including: The first data acquisition agent uses web crawling technology to obtain entity attribute data from open network data sources and generate an open network raw dataset. The second data acquisition agent uses interface call technology to obtain abnormal feature data from a structured database and generate a structured raw dataset. The third-party data acquisition agent uses long-connection monitoring technology to acquire dynamic environmental data from real-time data streams and generate raw datasets of real-time data streams. Data pre-validation is performed on the open network raw dataset, the structured raw dataset, and the real-time data stream raw dataset to generate open network valid dataset, structured valid dataset, and real-time data stream valid dataset, respectively. By integrating the open network valid dataset, structured valid dataset, and real-time data stream valid dataset, multi-source heterogeneous data is generated.
3. The anomaly analysis method based on multi-agent cooperation as described in claim 1, characterized in that, The multi-source heterogeneous data is processed by a data fusion intelligent agent to generate fused data in a unified format, including: The data fusion agent receives open network valid datasets, structured valid datasets, and real-time data stream valid datasets from the multi-source heterogeneous data. Perform text vectorization transformation on the open network effective dataset to generate a vectorized feature dataset; Perform field mapping transformation operations on the structured valid dataset to generate a structured mapped dataset; Perform timestamp alignment on the valid dataset of the real-time data stream to generate a time-aligned dataset; The vectorized feature dataset, structured mapping dataset, and time-aligned dataset are integrated to generate fused data in a unified format and stored in a central data repository.
4. The anomaly analysis method based on multi-agent cooperation as described in claim 1, characterized in that, The data analysis agent mines the fused data to extract abnormal feature information, including: Data analytics agents acquire fused data in a unified format from a central data repository. Cluster analysis is performed on the structured mapping dataset in the fused data to generate outlier data clusters; Deep semantic analysis is performed on the vectorized feature dataset in the fused data to generate anomalous topics; Perform fluctuation pattern analysis on the time-aligned dataset in the fused data to generate abnormal fluctuation points; By associating the abnormal data groups, abnormal topics, and abnormal fluctuation points, abnormal feature information is generated.
5. The anomaly analysis method based on multi-agent collaboration as described in claim 1, characterized in that, Based on the aforementioned anomaly characteristic information, the decision-making agent generates anomaly analysis results and decision instructions, including: The decision-making agent analyzes the abnormal data groups, abnormal themes, and abnormal fluctuation points in the abnormal feature information. Weights are assigned to the abnormal data group, the abnormal topic, and the abnormal fluctuation point, respectively, and weight assignment results are generated. By combining the weight allocation results with the quantitative values of the abnormal data groups, the abnormal themes, and the abnormal fluctuation points, a comprehensive anomaly index is determined. The comprehensive anomaly index is compared with a preset decision threshold. If the comprehensive anomaly index exceeds the preset decision threshold, an intervention instruction is generated; otherwise, a monitoring instruction is generated. The intervention or monitoring instructions are used as decision instructions, and anomaly analysis results containing the comprehensive anomaly indicators are generated.
6. The anomaly analysis method based on multi-agent cooperation as described in claim 1, characterized in that, The feedback optimization agent collects feedback data after the decision instruction is executed, compares and analyzes the feedback data with the decision instruction, and generates optimization instructions for optimizing the data analysis agent and the decision agent, including: The intelligent agent collects data on the actual execution results after the decision-making instructions are executed through feedback optimization, and monitors dynamic changes in the environment to generate environmental evolution data; Obtain historical decision-making instructions and their corresponding expected results data; Perform a difference analysis on the actual execution result data and the expected result data to generate a difference value; When the difference value exceeds the deviation tolerance threshold, the current decision execution cycle is marked as a decision deviation event; By correlating the decision-making bias events with the environmental evolution data, bias correlation analysis results are generated. Based on the deviation correlation analysis results, generate model parameter adjustment instructions for the data analysis agent and decision strategy update instructions for the decision-making agent; The model parameter adjustment instruction is sent to the data analysis agent, and the decision strategy update instruction is sent to the decision agent.
7. The anomaly analysis method based on multi-agent cooperation as described in claim 1, characterized in that, After collecting feedback data after the execution of the decision instruction by the feedback optimization agent, and comparing and analyzing the feedback data with the decision instruction to generate optimization instructions for optimizing the data analysis agent and the decision agent, the system further includes: The decision-making agent receives decision strategy update instructions from the feedback optimization agent. The decision-making agent extracts the deviation correlation analysis results from the decision-making strategy update instructions; The decision-making agent updates the action value table of the reinforcement learning decision-making model based on the results of the deviation correlation analysis. The decision-making agent optimizes the weight allocation strategy based on the updated action value table. In the reinforcement learning decision-making model, the decision-making agent generates subsequent decision instructions by applying an optimized weight allocation strategy based on the state space composed of environmental evolution data and abnormal feature information. The data analysis agent receives model parameter adjustment instructions from the feedback optimization agent; The data analysis agent extracts deviation correlation analysis results from the model parameter adjustment instructions; The data analysis agent adjusts the similarity threshold of the clustering analysis model based on the deviation correlation analysis results; The data analysis agent adjusts the feature weight parameters of the semantic analysis model based on the deviation correlation analysis results. The data analysis agent applies the adjusted clustering analysis model and semantic analysis model, and uses the adjusted similarity threshold and feature weight parameters to process the subsequent fused data.
8. An anomaly analysis device based on multi-agent collaboration, characterized in that, The anomaly analysis device based on multi-agent collaboration includes: The data acquisition module is used to collect data from multiple sources in parallel through multiple data acquisition agents, generating multi-source heterogeneous data; The data fusion module is used to process the multi-source heterogeneous data through a data fusion intelligent agent to generate fused data in a unified format; The feature mining module is used to mine the fused data through a data analysis intelligent agent and extract abnormal feature information; The decision generation module is used to generate anomaly analysis results and decision instructions based on the anomaly feature information by a decision-making intelligent agent; The feedback optimization module is used to collect feedback data after the decision instruction is executed through the feedback optimization agent, compare and analyze the feedback data with the decision instruction, and generate optimization instructions for optimizing the data analysis agent and the decision agent.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multi-agent cooperative anomaly analysis program stored in the memory and executable on the processor. When executed by the processor, the multi-agent cooperative anomaly analysis program implements the steps of the multi-agent cooperative anomaly analysis method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an anomaly analysis program based on multi-agent cooperation, which, when executed by a processor, implements the steps of the anomaly analysis method based on multi-agent cooperation as described in any one of claims 1-7.
Citation Information
Cited By
Intelligent factory fault prediction self-repairing system based on digital twinning and deep reinforcement learning
CN121143258A
Gateway monitoring method and device, electronic equipment, storage medium and computer product
CN121441792A
Self-adaptive association interaction synchronization method and system for expiration value keeping service data
CN121478883A
Monitoring and early warning method, system and device for inoculation observation period
CN122474378A