Abnormality detection and early warning method, device, equipment and medium

By standardizing and analyzing multi-source business data, generating anomaly scores and operation event sets, the problem of correlation between scattered data in fintech and healthcare businesses is solved, enabling dynamic risk assessment and enhanced security.

CN122023031APending Publication Date: 2026-05-12PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN HEALTH INSURANCE CO LTD
Filing Date
2026-03-03
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to perform unified identification and correlation analysis on fragmented business operation records and behavior logs in the fintech and healthcare sectors, resulting in delayed anomaly identification, insufficient risk assessment, lack of end-to-end traceability, and data security vulnerabilities.

Method used

By acquiring multi-source business data, standardizing the data to generate standardized business datasets and audit trail records, extracting abnormal feature sets and inputting them into anomaly analysis models, generating anomaly scores and operation event sets, and dynamically updating early warning messages and report data.

Benefits of technology

It enables multi-dimensional identification and dynamic risk assessment of abnormal behavior, improving the accuracy, timeliness, and traceability of anomaly identification and enhancing data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023031A_ABST
    Figure CN122023031A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision, and discloses an anomaly detection and early warning method and device, equipment and a medium, and the method comprises the steps: obtaining multi-source business data containing an operation identifier, carrying out the standardization processing, and generating a standardized business data set and an audit tracking record set; extracting an exception feature set and inputting the exception feature set into the exception analysis model to obtain exception scores; executing exception detection to generate an exception operation event set; generating and sending an early warning message and exception report data according to the exception score and the exception operation event set; and continuously updating the risk result based on the incremental business data. The method can be applied to business scenes such as financial science and technology, medical health and the like, and the accuracy, the timeliness and the traceability of anomaly recognition are improved through uniform identification association and standardized processing, combination of anomaly feature extraction, model reasoning and anomaly detection and continuous updating of risk results by utilizing incremental business data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to an anomaly detection and early warning method, device, equipment and medium. Background Technology

[0002] Existing financial risk control technologies generally suffer from problems when handling complex business scenarios, such as fragmented multi-source data, difficulty in establishing a unified correlation between business behavior and fund flows, reliance on static rules for risk identification, and a lack of end-to-end traceability capabilities. Because business operation records, log data, and financial data are stored in a dispersed manner, the system struggles to continuously analyze business behavior under a unified data view. This leads to anomaly identification relying heavily on post-event investigations, resulting in delayed risk discovery and warning information that fails to reflect a complete chain of evidence for anomalies. Furthermore, with the continuous growth of business data, existing technologies also exhibit significant shortcomings in dynamic risk assessment and data security auditing, making it difficult to meet the needs of refined financial risk management.

[0003] In the fintech sector, premium income, claims payments, and refunds for medical insurance policies are scattered across different systems. Business records and logs lack a unified identifier for correlation, making it difficult to establish a complete data chain for the same business activity. Traditional risk control methods often rely on single transactions or static rules for judgment, lacking the ability to comprehensively analyze historical behavioral patterns and multi-dimensional characteristics. This makes it difficult to identify risks such as duplicate refunds, abnormal refund ratios, and refunds during abnormal time periods in a timely manner. Existing systems often generate alerts based solely on a single abnormal result, lacking a mechanism to comprehensively analyze abnormal behavior and its severity, resulting in insufficient support for alert information.

[0004] In the healthcare sector, the diagnosis and treatment process is complex, involving multiple stages such as registration, consultation, settlement, and cost adjustment. Operational data and behavioral logs generated at different stages are stored in a scattered manner, making it difficult to form a complete behavioral trajectory. Existing technologies struggle to continuously record the entire process and establish temporal relationships, making it difficult to pinpoint and trace the source of cost or operational anomalies at the operational trajectory level. Furthermore, healthcare data is large-scale, frequently manipulated, and constantly changing. Traditional systems lack the ability to dynamically adjust risk assessment results in response to data changes, resulting in lagging risk monitoring. Moreover, the lack of robust encryption and auditing mechanisms during long-term data storage and access poses certain security risks. Summary of the Invention

[0005] The main objective of this invention is to provide an anomaly detection and early warning method, apparatus, device, and storage medium, aiming to solve the technical problem that existing technologies are unable to correlate and analyze scattered business operation records and behavior logs under a unified identifier, and thereby achieve multi-dimensional identification of abnormal behavior and dynamic risk assessment that updates with data changes.

[0006] To achieve the above objectives, the present invention provides an anomaly detection and early warning method, comprising: Acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The multi-source business data is standardized to generate a standardized business dataset containing the operation identifier, and an audit trail record set is constructed. Based on the standardized business dataset and the audit trail record set, an abnormal feature set for the operation identifier is extracted, the abnormal feature set is input into a preset abnormal analysis model to output an abnormal score, and the abnormal score is associated with the operation identifier; Anomaly detection is performed on the standardized business dataset to generate a set of abnormal operation events associated with the operation identifier; Based on the anomaly score and the set of abnormal operation events, generate and send early warning messages and anomaly report data; The standardized business dataset and the audit trail record set are updated based on incremental business data, the anomaly score is re-determined and the abnormal operation event set is updated, so as to dynamically update the warning message and the anomaly report data.

[0007] Furthermore, to achieve the above objectives, the present invention provides an anomaly detection and early warning device, comprising: The multi-source data acquisition module is used to acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The data standardization processing module is used to standardize the multi-source business data, generate a standardized business dataset containing the operation identifier, and construct an audit trail record set. An anomaly feature analysis module is used to extract an anomaly feature set for the operation identifier based on the standardized business dataset and the audit trail record set, input the anomaly feature set into a preset anomaly analysis model to output an anomaly score, and associate the anomaly score with the operation identifier; An anomaly detection and judgment module is used to perform anomaly detection on the standardized business dataset and generate a set of abnormal operation events associated with the operation identifier; The early warning and report generation module is used to generate and send early warning messages and abnormal report data based on the abnormal score and the abnormal operation event set. The incremental update scheduling module is used to update the standardized business dataset and the audit trail record set based on incremental business data, redetermine the anomaly score and update the abnormal operation event set, so as to dynamically update the warning message and the anomaly report data.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an anomaly detection and early warning program stored in the memory and executable on the processor, wherein when the anomaly detection and early warning program is executed by the processor, it implements the steps of the anomaly detection and early warning method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an anomaly detection and early warning program, wherein the anomaly detection and early warning program, when executed by a processor, implements the steps of the anomaly detection and early warning method as described above.

[0010] Beneficial Effects: This invention relates to the field of intelligent decision-making technology, and discloses an anomaly detection and early warning method, apparatus, device, and medium, comprising: acquiring multi-source business data containing operation identifiers, the multi-source business data including business operation record data and operation event log data; performing standardization processing on the multi-source business data to generate a standardized business dataset containing operation identifiers, and constructing an audit trail record set; extracting anomaly feature set based on the standardized business dataset and audit trail record set and inputting it into an anomaly analysis model to obtain anomaly scores; performing anomaly detection on the standardized business dataset to generate an abnormal operation event set; generating and sending early warning messages and anomaly report data based on the anomaly scores and abnormal operation event set; updating the standardized business dataset and audit trail record set based on incremental business data, re-determining the anomaly scores, and updating the abnormal operation event set to dynamically update the early warning messages and anomaly report data. This invention can be applied to business scenarios such as fintech and healthcare. By uniformly identifying, standardizing, and constructing time-series tracking of business operation record data and operation event log data, and through the synergistic effect of anomaly feature extraction, anomaly analysis model reasoning, and anomaly detection, it achieves multi-dimensional identification of abnormal behavior. Furthermore, by combining incremental business data to drive the continuous updating of risk assessment results, it ensures that risk warnings and reports remain synchronized with business changes, thereby improving the accuracy, timeliness, and traceability of anomaly identification. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an anomaly detection and early warning method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the anomaly detection and early warning method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly detection and early warning device of the present invention; Figure 4This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The anomaly detection and early warning method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain multi-source business data containing operation identifiers from the client. This multi-source business data includes business operation record data and operation event log data. The server performs standardization processing on the multi-source business data to generate a standardized business dataset containing operation identifiers and constructs an audit trail record set. Based on the standardized business dataset and audit trail record set, it extracts anomaly feature sets and inputs them into anomaly analysis models to obtain anomaly scores. Anomaly detection is performed on the standardized business dataset to generate anomaly operation event sets. Based on the anomaly scores and anomaly operation event sets, the server generates and sends warning messages and anomaly report data. Based on incremental business data, the server updates the standardized business dataset and audit trail record set, redetermines the anomaly scores, and update the anomaly operation event sets to dynamically update warning messages and anomaly report data. This invention can be applied to business scenarios such as fintech and healthcare. By uniformly identifying, standardizing, and constructing time-series tracking for business operation record data and operation event log data, and through the synergistic effects of anomaly feature extraction, anomaly analysis model reasoning, and anomaly detection, it achieves multi-dimensional identification of abnormal behavior. Furthermore, by combining incremental business data to drive continuous updates to risk assessment results, it ensures that risk warnings and reports remain synchronized with business changes, thereby improving the accuracy, timeliness, and traceability of anomaly identification. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the anomaly detection and early warning method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the anomaly detection and early warning method proposed in this invention includes the following steps: S10, acquire multi-source service data containing operation identifiers, the multi-source service data including service operation record data and operation event log data; In this embodiment, the process of acquiring multi-source business data containing operation identifiers is used to form an input data set that can be processed uniformly. Multi-source business data includes business operation record data and operation event log data. Business operation record data represents the result information of business actions, and its fields include monetary values, status markers, processing result markers, and business time fields. Operation event log data represents the execution trajectory information of the same business action within the system, and its fields include event time, event type, trigger source, execution result, and operating environment information. Operation identifiers exist as shared fields in the multi-source business data. Operation identifiers are used to mark the uniqueness of a business action or a financial action. The sources of operation identifiers include serial numbers, acceptance numbers, processing numbers, request identifier values, session identifier values, and other marker fields that can appear simultaneously in records and logs, generated by the business system. The acquisition of multi-source business data requires that the fields in the business operation record data and operation event log data be kept intact during collection to avoid losing field values ​​that can be used for subsequent identification and comparison. Simultaneously, an existence verification is performed on the operation identifiers to remove data records with missing operation identifiers, reducing interference from unrelated data in subsequent processing.

[0016] Business operation record data acquisition targets structured storage media, including relational database tables, columnar storage tables, or file-based detail tables. Acquisition actions include field selection, time range limiting, and incremental flag retention. Time range limiting controls the acquisition window, and incremental flags distinguish between historical and newly arrived records. Operation event log data acquisition targets log or event media, including log files, log aggregation systems, event queues, or streaming log topics. Acquisition actions include log parsing and field extraction. Field extraction extracts key fields such as operation identifiers, event times, and event types from the raw log content. To standardize data input, acquisition results can be encapsulated as a collection of raw data objects containing record bodies, field name sets, and metadata. Metadata includes data source identifiers, acquisition times, and parsing version identifiers. Parsing version identifiers distinguish between versions of different log formats or field mapping rules.

[0017] This embodiment simultaneously acquires business operation record data and operation event log data and retains operation identifiers to form a unified input data set containing result information and execution trajectory information. By confirming the existence of operation identifiers, it reduces unrelated data records and improves the integrity and availability of data association during subsequent processing.

[0018] S20, standardize the multi-source business data to generate a standardized business dataset containing the operation identifier, and construct an audit trail record set; In this embodiment, standardization processing is performed on multi-source business data to eliminate differences in source, format, and field representation, ensuring comparability and composability of business operation record data and operation event log data within the same structural system. Business operation record data originates from the business processing system, with a field structure primarily based on business results. Operation event log data originates from system operation records, with a field structure primarily based on behavioral trajectories. The two types of data exhibit inconsistencies in field naming, time representation, status representation, and encoding methods. Standardization processing unifies fields expressing the same meaning from different sources into the same field name through field mapping rules, unifies fields with different representations into a consistent data type through data type conversion, and converts status fields from different encoding systems into a unified status representation through value retrieval rule conversion.

[0019] Operation identifiers remain unchanged during standardization and participate in field mapping and data reorganization as core related fields. Field mapping rules are derived from the analysis of business table and log structures, forming a field lookup table that specifies the correspondence between original and target field names. Data type conversion is achieved by formatting numeric, time, and text fields; for example, converting timestamps to a unified time format and monetary fields to a uniform precision numeric type. Value retrieval rule conversion is achieved by re-encoding fields such as status codes, behavior codes, and environment flags, ensuring consistent semantics for status flags across different systems within a unified data model.

[0020] The standardized business dataset consists of mapped, transformed, and unified business operation record data. Each record in the standardized business dataset retains an operation identifier as well as business-related numerical fields, status fields, and time fields. The audit trail record set consists of unified operation event log data. After standardization, the operation event log data forms standardized event records. Each standardized event record includes an operation identifier, event time, event type, and operating environment information. Standardized event records are categorized according to operation identifiers and arranged in chronological order to form time-series association sequences. These time-series association sequences describe the complete behavioral trajectory of the same operation identifier within the system. The audit trail record set consists of multiple time-series association sequences, each corresponding to one operation identifier.

[0021] This embodiment unifies fields, types, and values, enabling data from different sources to form a directly processable dataset under the same structural system. It also constructs a complete behavioral trajectory record based on operation identifiers, providing a data foundation with consistent structure and semantics for subsequent correlation analysis.

[0022] S30, Based on the standardized business dataset and the audit trail record set, extract an abnormal feature set for the operation identifier, input the abnormal feature set into a preset abnormal analysis model to output an abnormal score, and associate the abnormal score with the operation identifier; In this embodiment, anomaly feature sets targeting operation identifiers are extracted based on standardized business datasets and audit trail records. These features combine business performance information and behavioral trajectory information scattered across different data carriers into a unified, computable representation. The standardized business dataset contains numerical, status, and time fields directly related to the operation result, while the audit trail record set contains event time, event type, and operating environment information related to the operation process. Around the same operation identifier, the two types of data are logically aggregated, forming a unified data view of result and process information.

[0023] The anomaly feature set is formed by concatenating multiple feature classes. Business numerical features are derived from fields related to amount, frequency, proportion, and intensity in standardized business datasets, reflecting the quantitative performance of operation results. Business status change features are derived from the value changes of status fields at different time points, reflecting the status evolution process. Operation time sequence interval features are derived from the time differences between event times in the audit trail record set, reflecting the operation rhythm and behavioral intensity. Operation environment context features are derived from device information, environment markers, and behavior types in event records, reflecting the operating environment at the time of the operation. After these features are extracted under the same operation identifier, they are vectorized and concatenated to form a unified anomaly feature set, enabling different types of data to be computable in the same dimension.

[0024] An anomaly feature set is input into a pre-built anomaly analysis model to map multidimensional features into risk level expressions. The anomaly analysis model is constructed and applied by pre-training historical data using machine learning methods to obtain a computational function that can calculate and output a score quantifying the degree of anomaly based on the input feature vector. This model is trained and validated before deployment, becoming a pre-built, callable analysis component. The core of the model lies in learning the distinguishing boundary between anomaly and normal patterns in the feature space from historical cases labeled with known risk outcomes (such as normal claims confirmed as fraud).

[0025] In the specific implementation of model construction, when gradient boosting decision trees are used as the basic algorithm, the model is integrated from multiple sequentially generated decision tree models. Each decision tree is a weak learner that recursively selects features and thresholds the feature values ​​to assign input samples to different leaf nodes. The gradient boosting process iteratively adds new decision trees, each aiming to correct the residuals predicted by the previous tree. The hierarchical structure of the model is reflected in the depth of the trees, the number of trees, and the size of the feature subset considered when splitting each tree. The connections are sequentially stacked, with the prediction result of the previous tree serving as the benchmark for the next tree's learning. The final output of the model is the weighted sum of the predictions from all trees. The specific steps necessary for training this model include data preparation, parameter initialization, iterative training and early stopping, and model validation. First, a high-quality training dataset needs to be prepared. This dataset consists of feature vectors obtained by processing historical business operation records and audit event sequences through the same feature engineering process. Each sample is labeled with a true binary label (e.g., 1 represents abnormality or fraud, and 0 represents normal). Key parameters for model training include the learning rate, which controls the contribution weight of each tree to the final result to prevent overfitting; the maximum depth of the decision tree, limiting the complexity of the trees; the total number of trees generated; and the minimum number of samples required for leaf nodes, controlling the granularity of tree growth. The training process aims to minimize a specified loss function (such as log loss), adjusting the parameters of newly added trees in each iteration using gradient descent. Training data is typically divided into training and validation sets. An early stopping strategy is implemented by monitoring performance metrics (such as AUC) on the validation set; training terminates when performance no longer improves, thus determining the final number of trees. The trained model is serialized into a file and loaded into an online inference service.

[0026] In the field of health insurance claims risk control, the training data for the model comes from archived data of historical claims cases. Each dimension in the feature vector corresponds to an indicator highly correlated with fraud risk, extracted from claims data and medical behavior logs. For example, features may include the ratio of a single claim amount to the insured's annual deductible, the matching degree between the level of the hospital visited and the severity of the declared disease, the time density of the same insured's visits to different medical institutions, and specific operation sequences on the self-service platform before and after submitting a claim application. The construction of these features deeply integrates actuarial principles, medical behavior pattern analysis, and anti-fraud investigation experience. In the field of financial transaction risk control, the training data comes from historical transaction records, and features may cover transaction amount, transaction time and location, device fingerprints, and transaction frequency within a short period of time. Its design is closely integrated with payment behavior analysis and money laundering pattern identification. In this way, the decision boundary learned by the model is essentially a mathematical abstraction of abnormal behavior patterns within a specific domain.

[0027] Before inputting raw features into the model, standardization or normalization is typically required. For example, the Z-score method can be used to convert each feature value into a distribution with a mean of 0 and a standard deviation of 1, or min-max scaling can be used to map it to the [0,1] interval to ensure that features with different scales and value ranges have equal importance to the model. The order of features in the vector needs to be consistent during training and inference. The model output data is set by transforming the original output value of the model—the weighted sum of all decision tree predictions—to a preset numerical range that is easier for business to understand through a subsequent mapping function, thus obtaining an anomaly score. For example, the original output value might be converted to a probability value between 0 and 1 using the Sigmoid function, and then linearly mapped to a range of 0 to 100. This final anomaly score is strongly correlated with the operation identifier that triggered the inference (such as a claim application number or transaction serial number). This correlation is achieved by establishing a key-value mapping in memory or a database, where the key is the operation identifier and the value is the calculated anomaly score and an optional timestamp. This connection ensures that every assessed business entity can be traced back to its quantified risk score, providing a direct basis for decision-making in subsequent early warning, report generation, and manual review.

[0028] In essence, the anomaly analysis model receives a set of anomaly features, performs inference operations, and outputs anomaly prediction probability values. These anomaly prediction probability values ​​are converted into anomaly scores through numerical range mapping, giving the results a unified numerical expression. Anomaly scores are associated with operation identifiers and stored through a data index structure, ensuring that each operation identifier corresponds to a directly queryable anomaly score.

[0029] This embodiment aggregates result data and behavioral data along the operation identifier dimension, constructs a multi-dimensional feature representation, and inputs it into the model for reasoning to form a unified and quantified anomaly score, providing a comprehensive judgment basis with process and result information for anomaly identification.

[0030] S40, perform anomaly detection on the standardized business dataset and generate a set of abnormal operation events associated with the operation identifier; In this embodiment, anomaly detection is performed on the standardized business dataset to identify records that deviate from the overall distribution pattern from the data that has undergone unified format conversion, and these records are transformed into a set of abnormal operation events with behavioral meaning. The standardized business dataset already has a unified field structure and time expression format, which can support statistical analysis according to time order and numerical patterns. Data traversal is performed around the operation identifier, so that each record always maintains its association with the corresponding operation identifier during the detection process.

[0031] Anomaly detection relies on dividing the data into multiple consecutive time slices based on temporal sequence. These time slices are derived from the time field in the records, segmenting the data using equal-length or sliding time intervals to ensure comparability within the same time period. The distribution of statistical fields within each time slice, including mean, dispersion, and distribution range, is used to form a statistical threshold. This threshold reflects the normal fluctuation boundaries of the data within that time range.

[0032] Iterate through each record in the standardized business dataset, determining the time slice window to which the record belongs based on the time field. Read the specified numerical field from the record and compare its value with the statistical threshold of the corresponding time slice window. Records with values ​​exceeding the statistical threshold are identified as outliers. Outliers maintain their binding relationship with the operation identifier during identification, ensuring that anomaly detection always revolves around the specific operation.

[0033] Anomaly operation events are generated based on outlier records. Each anomaly operation event includes an anomaly type identifier, time information, and a corresponding operation identifier, indicating why the record was identified as an anomaly. All anomaly operation events are grouped according to their operation identifiers to form an anomaly operation event set, allowing for centralized management and representation of multiple anomalies that may occur under the same operation identifier.

[0034] This embodiment establishes statistical boundaries and identifies deviation records in the time dimension, transforming scattered data anomalies into a set of abnormal operation events with clear operational meaning, thus providing a structured basis for subsequent risk assessment.

[0035] S50, generate and send early warning messages and abnormal report data based on the abnormal score and the abnormal operation event set; In this embodiment, warning messages and anomaly report data are generated and sent based on the anomaly score and the set of anomaly operation events to transform the numerical evaluation results and behavioral anomaly results into readable and transmittable information carriers. The anomaly score comes from the numerical results output by the anomaly analysis model and reflects the overall deviation degree of the behavior corresponding to the operation identifier. The set of anomaly operation events comes from the anomaly detection process and contains multiple anomaly event records associated with the operation identifier. Combining the two is used to construct prompt information and evidence information with judgment basis.

[0036] Anomaly scores are compared with preset anomaly warning thresholds to determine the anomaly warning level. The anomaly warning thresholds are derived from historical data statistics or business risk control strategies, and the anomaly warning level expresses the graded status of risk severity. Warning message content is constructed based on the anomaly warning level and anomaly score. The warning message content includes an anomaly level indication, an anomaly score value, and a brief explanation associated with the operation identifier, enabling the recipient to quickly identify the risk situation.

[0037] The serialization format conversion is performed on the set of abnormal operation events to transform the structured abnormal event data into a unified data representation format. The serialization format is derived from data transmission protocols or storage specifications, such as text structure or key-value structure, enabling the set of abnormal operation events to be directly parsed by external systems. The serialized event data maintains the association between each abnormal operation event and its operation identifier.

[0038] Anomaly report data is generated based on serialized event data, anomaly warning levels, and anomaly scores. This anomaly report data comprehensively displays the evidence of anomalies, including the time information of the anomaly event, anomaly type identifier, numerical deviation, and corresponding operation identifier. The anomaly report data is structurally designed to support storage and transmission, and is traceable.

[0039] Establish a communication channel with the remote control terminal, and distribute early warning messages and anomaly report data through the communication channel. The communication channel originates from a network connection interface or message distribution component, enabling the remote system to receive and process early warning messages and anomaly report data in real time.

[0040] This implementation transforms abnormal scores and sets of abnormal operation events into early warning messages and abnormal report data, enabling a unified expression of numerical evaluation results and evidence of behavioral abnormalities. This facilitates remote systems in quickly identifying risks and obtaining complete evidence of abnormalities.

[0041] S60, based on incremental business data, update the standardized business dataset and the audit trail record set, redetermine the anomaly score and update the abnormal operation event set, so as to dynamically update the warning message and the anomaly report data.

[0042] In this embodiment, updating the standardized business dataset and audit trail record set based on incremental business data maintains data timeliness, ensuring that subsequent anomaly assessment processes reflect the latest business behavior status. Incremental business data originates from continuously generated business operation record data and operation event log data; there are temporal and content differences between incremental business data and existing datasets. Business operation record data from the incremental business data is merged into the standardized business dataset to supplement new business field values ​​and status information. Operation event log data from the incremental business data is merged into the audit trail record set to supplement new behavioral timing information.

[0043] The anomaly score is redefined to reflect the impact of incremental business data on the anomaly assessment results. The anomaly score relies on the combined characteristics of the standardized business dataset and the audit trail record set. After data updates, the anomaly feature set extraction and anomaly analysis model inference need to be re-executed to ensure the anomaly score remains consistent with the latest data state. The composition of the anomaly feature set is still based on business numerical characteristics, business status change characteristics, operation timing interval characteristics, and operational environment context characteristics. These characteristics change after the addition of incremental data, thus affecting the anomaly score.

[0044] The updated abnormal operation event set is used to identify new abnormal behaviors introduced by incremental business data or their impact on existing abnormal behaviors. Anomaly detection relies on the numerical and temporal distributions in the standardized business dataset. After the statistical characteristics of incremental business data change, anomaly detection needs to be re-performed on the records associated with the operation identifier, generating new abnormal operation events and merging or replacing them with the original abnormal operation event set.

[0045] Dynamic updates to early warning messages and anomaly report data reflect the risk status following changes in anomaly scores and the set of anomalous operational events. Early warning messages are reconstructed based on the updated anomaly scores and warning levels, and anomaly report data is regenerated based on the updated set of anomalous operational events, ensuring that the output information remains consistent with the current data status.

[0046] After completing the calculation of anomaly scores, the aggregation of abnormal operation events, and the generation of early warning messages and anomaly reports, the system has formed a highly structured data format with strong correlations. This data not only reflects the business behavior itself, but also the temporal trajectory of the business behavior, the anomaly judgment results, and the risk output content. This type of data needs to be repeatedly accessed in subsequent operating cycles for comparative analysis, model reuse, historical tracing, and compliance verification, and therefore enters the persistent storage stage.

[0047] Once in the persistence phase, data no longer exists merely as input for real-time calculations, but transforms into historical data resources that can be stored long-term and repeatedly accessed. Standardized business datasets are used to retain complete business numerical states, audit trail records are used to retain complete behavioral trajectories, anomaly scores are used to retain anomaly assessment results, sets of abnormal operation events are used to retain the basis for anomaly judgments, and warning messages and anomaly report data are used to retain risk output content. Once this data enters the storage environment, the technical challenges shift from data processing to data security and access control.

[0048] Because these data have operational identification relationships, and these relationships can be used to reconstruct the complete business process and risk assessment process, data content transformation is required before writing to the storage medium, making the data physically unidentifiable. During the writing phase, the data is encrypted and transformed into ciphertext, which is then written to a secure storage area. Encryption processing not only targets the main data content but also covers index and cache information, preventing the inference of the original business meaning from the data structure.

[0049] After data is encrypted and stored, subsequent access must undergo authentication and permission checks. When the system receives a data access command, it parses the identity of the accessing subject and the target object from the command. The target object may be any of the following: standardized business datasets, audit trail records, anomaly scores, sets of abnormal operation events, warning messages, or anomaly report data.

[0050] The identity of the accessing entity and the target object are submitted to the access control policy for matching and determination. The access control policy defines the access scope for different identities to different data objects. The permission determination result is used to decide whether to allow access to encrypted data and trigger the decryption and reading process.

[0051] Every data access action generates an access log. The access log includes the access time, the accessing entity, the target object, and the permission assessment result. These access logs are stored in a structured format as an access audit log for subsequent compliance reviews and tracing of abnormal access behavior.

[0052] This implementation improves the real-time performance and accuracy of risk identification by updating incremental business data-driven data, anomaly assessment, and output information, ensuring that anomaly scores, sets of abnormal operation events, early warning messages, and anomaly report data are always consistent with the latest business behavior.

[0053] In one embodiment, step S10 above includes: S101, Configure a data reading interface for connecting to the business database, and capture business operation record data through the data reading interface according to a preset time slice window; S102, Configure a subscription listener for connecting to the log message queue, and capture operation event log data in real time through the subscription listener; S103, extract the operation identifier from the business operation record data and the operation event log data, and perform a non-empty check on the operation identifier; S104. Based on the result of the non-empty verification, filter out the business operation record data and operation event log data with empty operation identifiers from the business operation record data and the operation event log data, and aggregate the business operation record data and operation event log data with non-empty operation identifiers into the original data buffer pool. S105, the dataset stored in the original data buffer pool is combined into multi-source business data.

[0054] In this embodiment, the configuration of the data reading interface for connecting to the business database reflects the connectivity and data extraction constraints of the data acquisition end. The business database refers to the data storage unit that stores business operation record data. The data reading interface refers to the program interface that carries connection parameters, authentication information, query statement encapsulation, and data reading return. Connection parameters are derived from configuration items such as database address, port, database table identifier, account credentials, and connection pool capacity. Grabbing business operation record data according to a preset time slice window reflects the time boundary control of the business operation record data. The time slice window refers to a continuous interval bounded by a timestamp field. The grabbing action refers to forming query conditions within each time slice window and returning a set of records that meet the conditions. The timestamp field comes from the occurrence time, entry time, or accounting time field in the business operation record data. The time slice window can be set using a fixed duration window or a calendar-aligned window. Fixed duration windows are suitable for high-frequency record scenarios, while calendar-aligned windows are suitable for reconciliation and archiving scenarios. The grabbing process can map the time slice window to a set of query conditions. The set of query conditions includes start time conditions, end time conditions, deduplication conditions, and sorting conditions. The sorting conditions are used to stabilize the output record order so that subsequent processing can form a consistent set boundary.

[0055] Configuring a subscription listener to connect to the log message queue demonstrates the ability to stream operational event log data. The log message queue refers to the message middleware that carries operational event log data, and the subscription listener is the execution unit that maintains the subscription relationship and continuously receives new messages. The subscription relationship is defined by topic identifier, group identifier, and filtering conditions. Real-time capture of operational event log data demonstrates the ability to write to the buffer with low latency. Real-time capture includes two methods: pull reading and push receiving. Pull reading obtains the new message set after the message offset through polling, while push receiving receives messages delivered by the queue through callback functions. Key fields in the operational event log data originate from fields such as event timestamp, event type, behavior code, operator identifier, device identifier, terminal identifier, network address, and request parameter summary. These fields are subject to temporal consistency constraints, which are used to support the temporal reordering of subsequent audit trail record sets.

[0056] Extracting operation identifiers from business operation record data and operation event log data reflects the construction of a unified association key across data sources. An operation identifier is a unique value that can identify a single business operation or a group of business operations with the same affiliation. It originates from fields such as serial number, order number, transaction number, and application number in the business operation record data, and fields such as request identifier, session identifier, and event association identifier in the operation event log data. Extraction actions include field location, formatting, and conflict handling. Field location is determined by a field name mapping table or field path expression. Formatting includes removing whitespace characters, standardizing case, encoding, and length rules. Conflict handling includes priority selection and concatenation when multiple field candidates exist. Performing non-empty validation on operation identifiers reflects data quality control. Non-empty validation includes null value detection and invalid value detection. Null value detection covers empty strings, empty arrays, and empty objects, while invalid value detection covers placeholder strings, strings of all zeros, and illegal character sets. Non-empty validation outputs a validation flag, which serves as the basis for subsequent filtering actions.

[0057] The filtering of business operation records and operation event logs with empty operation identifiers based on the results of non-empty verification reflects a strategy to remove unrelated data. The filtering action is performed at the record granularity, meaning each business operation record and each operation event log is an independent judgment unit. The filtering action removes records marked as invalid from the set. The filtered business operation records and operation event logs are then aggregated into the raw data buffer pool, demonstrating temporary storage and aggregation capabilities. The raw data buffer pool refers to a memory queue, disk queue, or key-value storage area used to hold cross-source aggregated data. The aggregation action includes write partitioning, write order, and write idempotency control. Write partitioning can be based on data type, where business operation records and operation event logs are written to different partitions to maintain structural consistency, or it can be based on operation identifier, where data with the same operation identifier is written to the same partition to improve subsequent association efficiency. Write idempotency control is achieved through deduplication keys, which are derived from combinations of operation identifiers and timestamps or message offsets and event identifiers.

[0058] The dataset stored in the original data buffer pool is converted into multi-source business data to represent the definition of the converged result of multi-source business data. The dataset refers to the summarized result of data after being captured, checked for non-empty data, filtered, and aggregated within a preset time frame. It includes two data types: business operation record data and operation event log data, and maintains the operation identifier as the basic field for cross-source association. The formation of multi-source business data enables subsequent standardization processing to perform field cleaning, format unification, and audit trail reorganization within the same input boundary, avoiding subsequent association failures due to missing operation identifiers in the input data.

[0059] This embodiment captures business operation record data by time-slice window through the data reading interface and captures operation event log data in real time through the subscription listener, forming a multi-source business data input covering existing records and incremental events; by performing non-empty verification on the operation identifier and filtering invalid records, the associatable data is aggregated into the original data buffer pool and output as multi-source business data, so that the subsequent processing stage obtains a stable association key input boundary, thereby reducing the probability of cross-source association failure and improving data processing consistency.

[0060] In one embodiment, step S20 above includes: S201, parse the multi-source business data to distinguish between business operation record data and operation event log data; S202, traverse the business operation record data to identify null fields and non-standard format fields, and use preset cleaning rules to correct null fields and unify non-standard format fields; S203, map the corrected and unified business operation record data to a preset standard data model to generate a standardized business dataset containing the operation identifier; S204, Perform standardization cleaning on the operation event log data, unify the timestamp format and behavior code, and generate standardized event records based on the unified timestamp and behavior code. Each standardized event record in the standardized event record is associated with an operation identifier. S205, the standardized event records are reorganized into a time-series associated sequence according to the operation identifier, and the time-series associated sequence is stored as an audit trail record set.

[0061] In this embodiment, parsing multi-source business data to separate business operation record data and operation event log data reflects the structural decoupling and type merging of the input set. Multi-source business data refers to a data set that converges within the same input boundary, simultaneously containing both business operation record data and operation event log data. The parsing process includes data type determination, field structure identification, and media adaptation. Data type determination is based on source markers, topic markers, table name markers, or message header attributes. Field structure identification is based on field set characteristics, key name sets, or structured markers. Media adaptation covers row-based records, column-based fragments, and message body payloads. The output is divided into a business operation record data set and an operation event log data set, providing input boundaries for subsequent application of different cleaning and reorganization rules for different data formats.

[0062] The system iterates through business operation records to identify null and non-standard format fields, enabling data quality discovery and field-level repair. Traversal refers to scanning each record granularly and executing rules at the field level. Null fields are defined as fields with values ​​of empty strings, empty objects, missing fields, or invalid placeholders. Non-standard format fields are those that do not meet preset format constraints, which are derived from field type definitions, value range definitions, character set definitions, and length definitions. The system outputs a set of field anomaly markers, which drives the selection of preset cleaning rules. Preset cleaning rules are pre-configured sets of correction rules for specific fields or field types. Correcting null fields includes three methods: default value filling, same-record inference filling, and cross-record backfilling. Default value filling is suitable for status and enumerated fields; same-record inference filling is suitable for scenarios where redundant fields exist within the same record and can be mutually verified; and cross-record backfilling is suitable for scenarios where the same operation identifier has a consistent field in adjacent records that can be used for completion. Unifying non-standard format fields includes date and time format standardization, numerical precision standardization, currency or unit standardization, encoding standardization, and string normalization. Date and time format standardization unifies time expressions from different sources to the same time base and representation; numerical precision standardization unifies decimal places and rounding rules; unit standardization unifies fields such as proportions and amounts to the same unit of measurement; encoding standardization unifies different character encodings; and string normalization includes removing invisible characters, correcting full-width and half-width characters, and ensuring consistency between uppercase and lowercase letters. The correction and standardization processing generates corrected and standardized business operation record data, providing field reliability for subsequent mapping.

[0063] The corrected and standardized business operation record data is mapped to a pre-defined standard data model to achieve cross-source field alignment and structured output. The pre-defined standard data model refers to the unified definition of the field set, field types, field semantics, and constraints of the business operation record data. The field set includes operation identifier, timestamp field, business status field, numeric field, and context field, etc. Field semantics defines the meaning and value interpretation of fields, and constraints define mandatory, unique, range, and enumeration constraints. The mapping process includes field name mapping, field type conversion, field value normalization, and derived field generation. Field name mapping converts source field names to standard field names using a mapping table. Field type conversion converts string, integer, floating-point, and time types to standard types. Field value normalization merges synonyms into a unified enumeration set. Derived field generation generates standard fields based on the combination of multiple source fields in the business operation record data; for example, a unified business stage field is generated by combining the business status field and the operation type field. The mapping process maintains the consistent position and semantics of the operation identifier in the standard data model, ensuring that each output record contains the operation identifier. Generating a standardized business dataset containing operation identifiers reflects the formation of a standardized business dataset. A standardized business dataset refers to a set of records structured according to a standard data model, with consistent field structure and consistent format constraints, and can be directly used for subsequent abnormal feature set extraction and anomaly detection.

[0064] Standardizing and cleaning operation event log data ensures temporal and semantic consistency of event-side data. Operation event log data typically includes fields such as event timestamps, behavior codes, event types, request parameter summaries, and terminal environment information. Standardization and cleaning involve unifying the timestamp format and behavior codes. Unifying the timestamp format standardizes event timestamps from different sources to the same time zone and precision level, which can be set to seconds, milliseconds, or microseconds. Standardized actions include parsing the original timestamps, filling in missing time zone information, correcting format differences, and outputting a standardized timestamp. Unifying behavior codes maps behavior codes from different systems to a unified set of behavior codes derived from a behavior dictionary table. This behavior dictionary table defines the correspondence between behavior codes and behavior semantics, as well as the merging relationships of synonymous behaviors. Generating standardized event records based on the unified timestamps and behavior codes reflects event objectification and field convergence. Standardized event records are structured event entries containing standardized timestamps, unified behavior codes, event context fields, and operation identifiers. Each standardized event record is associated with an operation identifier to reflect the traceability of the event and business operation. The associated action is achieved by extracting the operation identifier from the log field, mapping the request identifier to the operation identifier, or finding the operation identifier through the association table within the same session. When the operation identifier is missing, the event record can be marked as unassociatable and removed from the standardized event record output through rule matching to avoid unassociatable events being mixed into the audit trail record set.

[0065] Standardized event records are reorganized into time-series sequences based on operation identifiers to reflect the organizational form and sequential constraints of the audit trail record set. The reorganization process includes grouping, sorting, and sequence encapsulation. Grouping aggregates standardized event records using the operation identifier as the key. Sorting uses the normalized timestamp as the primary sorting key and can be supplemented with message offsets or event sequence numbers as secondary sorting keys to ensure stable sorting of concurrent events with the same timestamp. A time-series sequence refers to a sequence of events arranged chronologically within the same operation identifier range. Sequence encapsulation can use array, list, or event cursor structures and can include sequence statistical fields, such as the number of events, the set of behavior codes, and the set of time differences between adjacent events, for subsequent extraction of operation time interval features and operation environment context features from the audit trail record set. Storing the time-series sequences as an audit trail record set demonstrates persistence and retrieval. Storage can employ a combination of partitioned storage by operation identifier and partitioned storage by time. Partitioned storage by operation identifier supports quick location of sequences with a single operation identifier, while partitioned storage by time supports batch loading by time range. An audit trail record set refers to a collection of time-series related sequences organized using operation identifiers as indexes, providing continuous contextual input from the event side for subsequent extraction of anomaly feature sets.

[0066] This embodiment separates business operation record data and operation event log data by parsing multi-source business data. It identifies, corrects, and unifies null fields and non-standard format fields in the business operation record data, ensuring consistency in field structure and format. By mapping the corrected and unified business operation record data to a preset standard data model, a standardized business dataset containing operation identifiers is formed, ensuring stable field semantics and value constraints for subsequent processing. Furthermore, by unifying the timestamp format and behavior codes of the operation event log data and generating standardized event records associated with operation identifiers, and then reorganizing them into a time-series associated sequence and storing it as an audit trail record set, the event-side data becomes traceable and comparable in both time and semantic dimensions. This improves the data consistency and reusability for subsequent anomaly feature extraction and anomaly detection.

[0067] In one embodiment, step S30 above includes: S301, Extract the business numerical features and business status change features for the operation identifier from the standardized business dataset; S302, extract the operation timing interval feature and operation environment context feature for the operation identifier from the audit trail record set; S303, Perform vectorized concatenation processing on the business numerical features, the business status change features, the operation timing interval features and the operation environment context features to generate an abnormal feature set; S304, Load the set of abnormal features into the anomaly analysis model constructed using the gradient boosting decision tree algorithm for inference; S305, obtain the anomaly prediction probability value output by the anomaly analysis model, and map the anomaly prediction probability value to a preset numerical range to obtain an anomaly score; S306, Establish a key-value pair mapping relationship between the abnormal score and the operation identifier in the data index table.

[0068] In this embodiment, the extraction of anomaly feature sets based on operation identifiers, derived from a standardized business dataset and an audit trail record set, embodies a dual-source feature organization for a single operation object. The standardized business dataset provides a structured record set from the business side, while the audit trail record set provides a time-series related sequence set from the event side. The operation identifier serves as a common index key for both sets, supporting cross-set alignment. The extraction process revolves around the operation identifier, performing record filtering and sequence location. Record filtering utilizes the operation identifier field in the standardized business dataset for equality matching or interval aggregation, while sequence location accesses the corresponding time-series related sequence based on the operation identifier index in the audit trail record set. The anomaly feature set refers to the feature set organized for the input of the anomaly analysis model, containing combinations of numerical features, state features, time features, and contextual features. The combination method is solidified into a unified input format during vectorized concatenation processing.

[0069] Extracting business numerical features and business status change features targeting operation identifiers from standardized business datasets reflects the separation and extraction of measurable information and discrete state evolution information of business records. Business numerical features refer to field-derived results that can be used for quantification, originating from numerical fields, count fields, ratio fields, and monetary fields in the standardized business dataset. Extraction actions include field selection, time window aggregation, and statistical generation. Field selection determines the target field set based on an anomaly correlation configuration table. Time window aggregation constructs an aggregation window based on the timestamp range of the records corresponding to the operation identifier. Statistical generation covers derived features such as mean, maximum, minimum, quantile, difference, and fluctuation range. The difference value is obtained by subtracting the numerical fields of adjacent records, and the fluctuation range is determined by the difference or relative difference between the maximum and minimum values ​​within the window. Business status change features refer to the expression of changes in business status fields over time, originating from state enumeration fields, stage fields, and result fields. Extraction actions include state sequence construction and state transition encoding. The state sequence is constructed by filtering records according to operation identifiers and sorting them by timestamps. State transition encoding maps adjacent state pairs to transition type identifiers and generates features such as transition count, transition direction, dwell time, and backtracking count. Dwell time is calculated from the difference in timestamps between adjacent states, and the number of backtracking counts is obtained by statistically analyzing reverse transitions occurring in the state sequence. Business numerical features and business state change features cover two information sources: intensity anomalies and process anomalies, providing heterogeneous feature inputs for subsequent vectorized concatenation.

[0070] Extracting operation time-series interval features and operation environment context features from the audit trail record set reflects the extraction of the temporal structure and environmental semantics of the event sequence. Operation time-series interval features refer to the measurement of event intervals within a time-series associated sequence corresponding to the same operation identifier. These features originate from the normalized timestamp field in the time-series associated sequence. The extraction process includes calculating adjacent event intervals and summarizing interval distributions. Adjacent event interval calculation involves subtracting adjacent timestamps within the sequence to obtain the interval sequence. Interval distribution summarization generates features such as shortest interval, longest interval, average interval, interval variance, and burstiness. Burstiness is determined by the proportion of short intervals or the clustering index of the interval sequence. Operation environment context features refer to the field expressions related to the event's occurrence environment. These features originate from standardized event records, including terminal type fields, channel fields, region fields, network identifier fields, device fingerprint fields, interaction entry fields, and behavior code fields. The extraction process includes context field normalization, combined encoding, and statistical counting. Context field normalization merges synonyms and unifies the encoding space. Combination encoding combines multiple fields to form composite context keys, such as a combination key for terminal type and channel fields. Statistical counting generates features such as the frequency, percentage, and number of changes of context occurrences. The number of changes is obtained by counting the number of times the composite context key switches in the sequence. Operation timing interval features reflect time behavior characteristics, while operation environment context features reflect environmental consistency and switching characteristics. Together, they supplement the interaction layer information that is difficult to cover in business records.

[0071] Vectorization concatenation is performed on business numerical features, business status change features, operation time-series interval features, and operation environment context features to generate an anomaly feature set that reflects feature space alignment and model input standardization. Vectorization refers to converting different data types into numerical vector representations. Business numerical features can be directly used as continuous numerical dimensions as input. Business status change features and operation environment context features typically contain enumeration or category information. Vectorization processing includes category encoding and sparse representation. Category encoding can use ordinal encoding, one-hot encoding, or target encoding. One-hot encoding maps categories to multiple binary dimensions, while target encoding maps categories to statistical values ​​related to historical anomalies. The encoding method is determined by the input requirements of the anomaly analysis model. Operation time-series interval features are usually continuous numerical values. Vectorization processing includes scale normalization and truncation. Scale normalization maps features with different dimensions to a unified range, while truncation limits extreme values ​​to a preset range to avoid single feature dominance. The concatenation process refers to linking various vectors along dimensions to form a unified vector according to a preset feature order. The preset feature order is defined by a feature dictionary, which records feature names, dimension indices, and value ranges. After concatenation, an abnormal feature set is output. This abnormal feature set is represented here as a group of vector data associated with a single operation identifier. It can take the form of one or multiple vectors. Multiple vectors are used to express the feature evolution across multiple time slices, while a single vector is used to express the overall features after aggregation.

[0072] The process of loading anomaly feature sets into an anomaly analysis model constructed using a gradient boosting decision tree algorithm for inference reflects the model input and inference execution. Loading refers to writing the anomaly feature set into the model input buffer or the input parameter structure of the inference interface. The input parameter structure includes feature vectors and operation identifier index information. The anomaly analysis model refers to the model entity used to output anomaly prediction probability values. Constructed using a gradient boosting decision tree algorithm, the model consists of multiple decision trees combined additively for output. The inference process includes feature missing data handling, tree path traversal, and tree output accumulation. Feature missing data handling follows the missing data strategy fixed during model training. The missing data strategy can be a default branch, missing value imputation, or missing indicator dimension. Tree path traversal selects branches to leaf nodes based on threshold conditions of each tree node. Tree output accumulation additively combines the output scores of each leaf node to obtain the original model output, which is then probabilistically transformed to obtain the anomaly prediction probability value. The anomaly prediction probability value is a quantified result of the probability that the behavior corresponding to the operation identifier exhibits an abnormal tendency, ranging from zero to one or within the transformed probability space.

[0073] The process involves acquiring the anomaly prediction probability value output by the anomaly analysis model and mapping it to a preset numerical range to obtain an anomaly score, reflecting a controllable conversion from probability to rating. The anomaly prediction probability value originates from the inference output of the anomaly analysis model. Mapping to the preset numerical range means mapping the probability space to the rating space, the boundary of which is defined by the preset numerical range. The preset numerical range can be set as a closed interval, and mapping methods include linear mapping, piecewise mapping, and nonlinear mapping. Linear mapping scales the anomaly prediction probability value proportionally to the preset numerical range; piecewise mapping corresponds to different rating gains based on probability threshold ranges; and nonlinear mapping scales the probability after performing an exponential or logarithmic transformation. The mapping process may include a calibration step, which aligns the anomaly prediction probability value with historical observation distributions to improve score comparability. Calibration can be completed through quantile mapping, which uses the quantile points of the historical probability distribution as mapping anchors. The anomaly score refers to the mapped numerical result, used for threshold comparison and level determination in subsequent anomaly detection, early warning message generation, and anomaly report data generation. The correspondence between the anomaly score and the operation identifier is solidified into a searchable structure in the next action.

[0074] A key-value mapping between abnormal scores and operation identifiers is established in the data index table to reflect persistent indexing and fast querying of scores. The data index table is the data structure used to store index records, each containing at least an operation identifier key and an abnormal score value. The key-value mapping relationship refers to the mapping structure where the operation identifier is the key and the abnormal score is the value. The establishment of this mapping includes writes, updates, and consistency control. Writes insert an index record when the operation identifier first appears. Updates overwrite or append versions when new abnormal scores are derived from subsequent inferences using the same operation identifier. Version appending can record the score generation time using a timestamp field, enabling time-based backtracking and incremental updates. Consistency control is used to avoid score mismatches caused by concurrent writes, implemented through mutually exclusive writes at the operation identifier granularity or version number-based comparison updates. The data index table provides the entry point for subsequent stages to read abnormal scores by operation identifier and provides rapid location capabilities for the generation of alert messages and abnormal report data.

[0075] This embodiment achieves information fusion between the business side and the event side around the operation identifier by extracting business numerical features and business status change features from a standardized business dataset and operation time interval features and operation environment context features from an audit trail record set. Through vectorized concatenation processing, heterogeneous features are converted into a unified input expression to form an abnormal feature set, ensuring that the input dimension, order, and value range of the abnormality analysis model remain controllable. The abnormal feature set is loaded into an abnormality analysis model constructed using a gradient boosting decision tree algorithm to perform inference and output anomaly prediction probability values. These anomaly prediction probability values ​​are then mapped to a preset numerical range to obtain anomaly scores, making the degree of anomaly measurable and alignable in numerical form. By establishing a key-value pair mapping relationship between anomaly scores and operation identifiers in a data index table, anomaly scores can be quickly retrieved by operation identifier, supporting the generation and updating of subsequent warning messages and anomaly report data.

[0076] In one embodiment, step S40 above includes: S401, the standardized business dataset is divided into multiple consecutive time slice windows according to the timestamp order; S402, determine the statistical threshold based on the statistical distribution of data within each time slice window; S403, traverse each record in the standardized business dataset, determine the time slice window to which the record belongs based on the record's timestamp, and filter records whose values ​​of specified numerical fields deviate from the statistical threshold of the time slice window as outlier records, and associate each outlier record with an operation identifier; S404, Generate an abnormal operation event containing an anomaly type identifier based on the outlier record; S405, based on the operation identifier, the abnormal operation events are collected into the abnormal operation event set.

[0077] In this embodiment, anomaly detection is performed on the standardized business dataset to identify abnormal behaviors related to operation identifiers from structured business records and form a set of processable objects. Anomaly detection refers to judging deviations of specified numerical fields based on statistical distribution and threshold rules. The output does not terminate with a single record, but generates a set of abnormal operation events to support the construction of subsequent warning messages and anomaly report data. The standardized business dataset comes from standardized business operation record data and includes a unified field format, a unified timestamp expression, and operation identifier fields that can be used for filtering. The anomaly detection stage relies on this uniformity to complete the repeatable execution of window division, threshold determination, and outlier record filtering.

[0078] Standardized business datasets are divided into multiple consecutive time slice windows according to timestamp order to limit the calculation range of statistical distribution and ensure the temporal locality of thresholds. A timestamp is a field representing the time a record occurs; it can be a uniform format such as a millisecond timestamp, a second-level timestamp, or a date-time string. The timestamp order is used to determine the arrangement of window boundaries and record attribution. Multiple consecutive time slice windows refer to a set of windows that are adjacent on the timeline without intervals or continuously covered at preset intervals. The window length is determined by a preset window width, which can be configured as a fixed duration or adaptively determined by data density. The fixed duration method uses minutes, hours, days, etc., as the window width to adapt to periodic and intraday fluctuation scenarios; the adaptive method adjusts the window width based on the number of records per unit time or the intensity of fluctuation in numerical fields, ensuring that each time slice window meets the minimum sample size or minimum statistical stability requirements. Time slice windows are generated by calculating the start and end times of the window and writing them into a window index table. The window index table records the window start time, window end time, and window identifier. The window identifier is used to reference the window range in subsequent threshold determination and record filtering.

[0079] Statistical thresholds are determined based on the statistical distribution of data within each time slice window to transform local data distributions into executable decision boundaries. The statistical distribution refers to the distribution of values ​​for a specified numerical field within the window, and can take the form of mean and standard deviation, quantile sets, box statistics, or robust scaling parameters. The statistical threshold is the numerical boundary derived from the statistical distribution, serving as a reference for deviation judgment. Statistical thresholds can include one-sided or two-sided thresholds. One-sided thresholds only address scenarios with abnormally high or low values, while two-sided thresholds address scenarios with both abnormally high and low values. Statistical thresholds can be generated using various calculation methods and are not limited to a single form. The quantile method uses specified quantiles within the window as thresholds, adapting to non-normal and long-tailed distributions; the mean and standard deviation method uses the mean plus or minus a multiple of the standard deviation to form the threshold, adapting to approximately symmetrical distributions; the box statistics method derives upper and lower bounds from the interquartile range, adapting to data with extreme values. The statistical thresholds are associated with the time slice windows and stored in the threshold table. The threshold table records the window identifier, the identifier of the specified numerical field, and the statistical threshold value. The threshold table provides a window-level search entry for subsequent record filtering.

[0080] Traversing each record in the standardized business dataset and determining the time slice window to which the record belongs based on its timestamp is used to establish a deterministic matching relationship between the record and the window-level statistical threshold. Traversal refers to reading records in batches or by cursor in the standardized business dataset. The reading process can be performed by sorting by timestamp to reduce window positioning overhead, or it can be performed in the original storage order and the positioning can be completed through the window index table. Determining the time slice window to which a record belongs can be achieved by comparing the record timestamp with the start and end times of the window. The comparison methods include interval inclusion judgment and binary search positioning. Interval inclusion judgment is used for scenarios with a small number of windows, while binary search positioning is used for scenarios with a large number of windows to reduce positioning complexity. After the time slice window to which a record belongs is determined, deviation judgment is performed on the value of a specified numerical field in the record and the statistical threshold of the time slice window. Deviation refers to the difference between the value and the statistical threshold that exceeds the allowable range. Deviation judgment can use two types of measures: absolute difference and relative difference. The absolute difference measure is used for scenarios where the numerical field has stable dimensions, and the relative difference measure is used for scenarios where the numerical field changes across magnitudes. Deviation judgment can be combined with tolerance parameters, which are used to control the suppression of slight fluctuations and improve the stability of outlier records. The filtering action outputs a set of outlier records. Outlier records refer to records that meet the deviation criteria. Outlier records retain the original fields and the minimum necessary fields for subsequent event generation, including the operation identifier, timestamp, values ​​of specified numeric fields, and window identifier. Each outlier record is associated with an operation identifier, reflecting its relationship to the business operation object. The association action is completed by reading the operation identifier field from the outlier record, achieving subsequent aggregation by operation identifier without introducing additional mappings.

[0081] Generating anomalous operation events based on outlier records, including anomaly type identifiers, transforms outlier records into aggregateable, transmissible, and traceable event objects. An anomalous operation event is a data object describing the fact that an anomaly occurred, including an operation identifier, anomaly occurrence time, anomaly type identifier, and a set of evidence fields. The anomaly type identifier is a labeling field for anomaly categories, used to distinguish anomalies with different deviation patterns and different field sources. The generation of anomaly type identifiers can be based on a combination of a specified numerical field identifier and a deviation direction. The deviation direction is determined by the relative relationship between the value of the specified numerical field and a statistical threshold; a higher deviation corresponds to a higher-direction deviation, and a lower deviation corresponds to a lower-direction deviation. Anomaly type identifiers can also be generated as graded identifiers based on the degree of deviation, determined by absolute difference, relative difference, or quantile offset, with grading rules defined by a preset interval table. The set of evidence fields for anomaly operation events can include a window identifier, a statistical threshold value, the value of the specified numerical field, the deviation direction, and the degree of deviation. This set of evidence fields provides a structured basis for subsequent anomaly reporting data. The generation process can use event templates. Event templates define the field structure and field filling rules for abnormal operation events. The field filling rules extract operation identifiers and timestamps from outlier records and extract statistical threshold values ​​from the threshold table. Then, they combine the deviation judgment results to fill in the abnormal type identifier and deviation measure.

[0082] Abnormal operation events are aggregated into an abnormal operation event set based on operation identifiers to form an event aggregation result centered on the operation identifier. This supports the subsequent generation of early warning messages and anomaly report data oriented towards a single operation identifier. Aggregation refers to grouping multiple abnormal operation events by operation identifier and summarizing them into a set structure. The set structure can be represented as a mapping table from operation identifiers to an event list, or as an event storage structure partitioned by operation identifier. The aggregation process can construct grouped containers in memory and periodically write them to disk, or it can be directly written to the event storage medium partitioned by operation identifier. The timestamp order of abnormal operation events is preserved during aggregation to express the anomaly evolution trajectory. The timestamp order can be achieved by sorting by the time of anomaly occurrence within each operation identifier group. The abnormal operation event set is the output object of the anomaly detection stage. The subsequent stages of generating and sending early warning messages and anomaly report data based on anomaly scores and the abnormal operation event set can directly read the abnormal operation event set to complete serialization format conversion and anomaly evidence chain organization.

[0083] This embodiment divides the standardized business dataset into multiple consecutive time slice windows based on timestamps and determines a statistical threshold for each time slice window, ensuring that the anomaly judgment benchmark corresponds to the local data distribution. By matching each record to a time slice window based on its timestamp and performing deviation filtering on specified numerical fields, it achieves consistent application of statistical thresholds for outlier records and time slice windows. By generating abnormal operation events containing anomaly type identifiers based on outlier records and aggregating them into an abnormal operation event set based on the operation identifiers, it transforms the anomaly results from record-level output into event-level output that is aggregable, traceable, and directly referenced by subsequent warning messages and anomaly report data.

[0084] In one embodiment, step S50 above includes: S501, compare the anomaly score with the preset anomaly warning threshold to determine the current anomaly warning level; S502, construct an early warning message containing emergency notification text based on the abnormality warning level and the abnormality score; S503, Perform serialization format conversion on the abnormal operation event set to obtain serialized event data; S504, Based on the serialized event data, the anomaly warning level, and the anomaly score, generate anomaly report data containing an anomaly evidence chain; S505, establish a communication channel with the remote control terminal, and distribute the warning message and the anomaly report data through the communication channel.

[0085] In this embodiment, warning messages and anomaly report data are generated and sent based on the anomaly score and the set of anomaly operation events. This integrates the numerical evaluation results and event-based evidence results into a distributable output. The output includes warning messages and anomaly report data. The warning messages carry immediate prompts, while the anomaly report data carries traceable evidence. Both are directed to remote control terminals to support decision-making, review, and data retention. The anomaly score originates from the inference output of the anomaly analysis model and has been associated with the operation identifier. The set of anomaly operation events originates from the aggregation output of the anomaly detection stage and aggregates anomaly operation events with operation identifiers. The generation stage utilizes both in five consecutive actions: determining the anomaly warning level, constructing the warning message, generating serialized event data, organizing anomaly report data, and distributing through the communication channel. The output of the previous action serves as the input for the next action, forming an achievable output flow.

[0086] The abnormal score is compared numerically with a preset abnormal warning threshold to determine the current abnormal warning level, thus mapping continuous scores to discrete levels. The preset abnormal warning threshold refers to a pre-configured set of thresholds, which can contain single or multiple thresholds. Multiple thresholds support multiple levels of abnormal warning. The configuration sources for the abnormal warning thresholds can include historical data distribution statistics, score distribution of labeled abnormal samples, business rule-defined intervals, or compliance requirement-defined intervals. During runtime, the threshold set is a read-only configuration for comparison. Numerical comparison involves determining the magnitude of the abnormal score against the abnormal warning threshold and outputting the level mapping result. The mapping relationship can be implemented using an interval mapping table, which records the correspondence between abnormal warning threshold intervals and abnormal warning level identifiers. The abnormal warning level is a discrete expression of the severity of the abnormality. It can be expressed using enumerated identifiers or numerical identifiers. Enumerated identifiers enhance readability, while numerical identifiers facilitate sorting and aggregation. The process of determining the abnormal warning level can be combined with the hysteresis interval to suppress level fluctuations. The hysteresis interval refers to setting entry and exit thresholds near the level switching boundary so that the level remains stable when the abnormal score fluctuates near the boundary. The hysteresis interval is also formed by expanding the preset abnormal warning threshold and written into the interval mapping table.

[0087] An alert message containing an emergency notification text is constructed based on the anomaly warning level and anomaly score to form an instant notification carrier for remote control terminals. The alert message refers to a message object with a fixed field structure, which at least includes the anomaly warning level, anomaly score, and emergency notification text. The field structure can be expanded to include an operation identifier, trigger time, summary field, and location field. The emergency notification text refers to the text content used to quickly convey the abnormal status. The text content is generated by filling in templates and variables. The templates are sourced from a preset notification template library, and the variables are derived from the anomaly warning level and anomaly score, and can be combined with the identification information of the operation identifier. The construction action includes two sub-actions: message template selection and message content filling. Message template selection locates the corresponding template in the template library based on the anomaly warning level. The template library can be grouped and stored according to the anomaly warning level and supports version numbers for updates. Message content filling formats the anomaly score to a specified precision and fills it into the template placeholder, fills the anomaly warning level identifier into the template placeholder, and generates the emergency notification text. The warning message may also include a message deduplication identifier to avoid duplicate distribution. The message deduplication identifier can be generated by combining the operation identifier, the abnormal warning level and the time slice identifier and written into the warning message field, which is used by the remote management terminal or the sending side to perform deduplication judgment.

[0088] The serialization process converts a set of abnormal operation events into serialized event data, transforming it into a standardized, transmittable, storable, and parsable representation. The abnormal operation event set is an aggregate structure containing multiple abnormal operation events organized by operation identifiers. Serialization converts this aggregate structure into a character or binary sequence. Serialization can employ key-value encoding, structured text encoding, or binary encoding. Key-value encoding represents the mapping between the event list and fields; structured text encoding facilitates parsing by remote control terminals; and binary encoding reduces transmission size and improves parsing efficiency. Serialized event data refers to the output of the serialization conversion, containing a sequence of event entries and necessary metadata. This metadata may include field version numbers, encoding method identifiers, and generation timestamps. A field whitelist can be introduced during the serialization process to control the output range. The field whitelist represents the set of allowed abnormal operation event fields for output, derived from a preset configuration and consistent with the set of evidence fields required for the abnormal report data. The serialization process can introduce field normalization rules to ensure field comparability. These rules include unified timestamp format, unified precision of numerical fields, and unified encoding of anomaly type identifiers. The field normalization rules are consistent with the unified field processing in the anomaly detection stage.

[0089] Anomaly report data, containing anomaly evidence chains, is generated based on serialized event data, anomaly warning levels, and anomaly scores to form a traceable, verifiable, and archiveable data carrier. Anomaly report data refers to a report object with a fixed structure, containing at least the anomaly warning level, anomaly score, serialized event data, and anomaly evidence chain fields. The anomaly evidence chain refers to the set of evidence formed around the criteria for anomaly judgment. This set of evidence includes the time series of the anomalous operation event, a set of anomaly type identifiers, key field values, and corresponding statistical thresholds or deviation measurement information. The source of the anomaly evidence chain includes the mapping information between the set of evidence fields and the anomaly score within the anomalous operation event. The generation process comprises three sub-actions: evidence organization, report field assembly, and report integrity verification. Evidence organization involves parsing serialized event data or directly referencing event entries and sorting them by timestamps to form a time-series segment of the anomaly evidence chain. Report field assembly writes the anomaly warning level and anomaly score into the report header field, the anomaly evidence chain into the report body field, and the serialized event data into the report attachment field or body field. Report integrity verification performs missing and consistency checks on the report fields. Missing checks confirm that the anomaly warning level, anomaly score, serialized event data, and anomaly evidence chain are all filled in. Consistency checks confirm that the set of event entries in the anomaly evidence chain matches the set of event entries in the serialized event data and that the mapping relationship between the anomaly warning level and anomaly score matches the interval mapping table. Anomaly report data can be appended with a report summary field to support quick browsing on remote management terminals. The report summary field extracts the number of key events, the earliest anomaly time, and the latest anomaly time from the anomaly evidence chain and combines them with the anomaly warning level to generate the report.

[0090] A communication channel is established with the remote management terminal, and warning messages and anomaly report data are distributed through this channel to complete output actions and enable the remote side to receive and process them. The remote management terminal refers to the remote system or terminal device that receives warning messages and anomaly report data. The communication channel refers to the data transmission channel between the sending side and the remote management terminal. The communication channel can be a long-connection channel or a short-connection channel; long-connection channels are suitable for continuous push scenarios, while short-connection channels are suitable for on-demand sending scenarios. The establishment action includes endpoint configuration, session parameter negotiation, and connection health checks. Endpoint configuration includes the remote management terminal address, port, protocol identifier, and authentication parameters. Negotiation includes transmission compression options and retry policy options. Health checks include heartbeat sending and response confirmation. The distribution action encapsulates the warning messages and anomaly report data into separate packets and writes them to the communication channel. The encapsulation process may include a message header field and a payload field. The message header field includes a message type identifier and an associated operation identifier, and the payload field contains the warning message or anomaly report data. The distribution action can introduce a sending confirmation mechanism to record the delivery status. The sending confirmation mechanism outputs the confirmation result and can write it to the sending record table so that incremental distribution control can be executed during subsequent dynamic updates. The sending record table can be indexed by operation identifier and timestamp to support fast query.

[0091] This embodiment determines the anomaly warning level by comparing the anomaly score with the anomaly warning threshold and constructs a warning message containing emergency notification text accordingly, so that the abnormal status can be quickly identified by the remote control terminal in a hierarchical and textual form; serialized event data is obtained by performing serialization format conversion on the abnormal operation event set and generating anomaly report data containing anomaly evidence chain by combining the anomaly warning level and the anomaly score, so that the basis for anomaly judgment can be reviewed and retained in the form of a structured evidence set; by establishing a communication channel with the remote control terminal and distributing the warning message and anomaly report data, the output object has a transmittable, acknowledgable and traceable sending process, thereby supporting the timeliness of anomaly handling and the consistency of evidence.

[0092] In one embodiment, step S60 above includes: S601, Receive incremental service data stream, the incremental service data stream includes newly added service operation record data and operation event log data; S602, merge the business operation record data and operation event log data in the incremental business data stream into the standardized business dataset and the audit trail record set respectively, and generate an updated standardized business dataset and an updated audit trail record set; S603, extract the affected operation identifier from the incremental service data stream; S604, Based on the updated standardized business dataset and the updated audit trail record set, perform abnormal feature set extraction and abnormal analysis model inference for the affected operation identifier to generate an updated abnormal score; S605, Based on the updated standardized business dataset, perform anomaly detection on the affected operation identifiers to generate an updated set of abnormal operation events; S606, based on the updated abnormal score and the updated set of abnormal operation events, trigger the generation and sending of warning messages and abnormal report data.

[0093] In this embodiment, incremental service data is used to represent the new portion relative to existing data. The incremental service data stream carries continuously arriving incremental service data and supports writing to the processing queue in chronological order. The incremental service data stream includes newly added service operation record data and operation event log data to simultaneously cover both service-side records and behavior-side records. Receiving the incremental service data stream can be achieved by configuring an access buffer at the data entry point. The access buffer records the batch identifier and the receiving timestamp according to the arrival time and performs a basic integrity check on the incremental service data stream. The integrity check covers field structure consistency, operation identifier existence, and timestamp resolvability. Records that fail the check are written to the isolation area and an abnormal reception flag is output. Records that pass the check enter the merging stage.

[0094] The business operation record data and operation event log data in the incremental business data stream are merged into the standardized business dataset and audit trail record set respectively to keep the two data containers evolving synchronously within the same update cycle. For business operation record data, keyed merging can be implemented. Keyed merging uses the operation identifier and business timestamp as a composite key, and locates records with the same key in the standardized business dataset. When a match is found, field-level update rules are executed. These rules can include three types: overwrite update, maximum value update, and append update. Overwrite update corrects the original field value, maximum value update retains the value corresponding to a larger timestamp, and append update retains multiple changes as an array field. When a match is not found, insertion rules are executed and standard field default values ​​are added. Merging of operation event log data is performed according to the time-series organization of the audit trail record set. During merging, the time-series association sequence is located by the operation identifier, and new log events are inserted into the corresponding positions according to the unified timestamp while maintaining sequence order. After insertion, duplicate event folding rules are executed on adjacent events. Duplicate event folding rules can be determined based on behavior code and timestamp interval thresholds and merged into a single event to reduce redundancy. After the merger is completed, an updated standardized business dataset and an updated audit trail record set are formed. The updated standardized business dataset and the updated audit trail record set can be written with the same version identifier to support the reference of a consistent data view in subsequent processing stages.

[0095] Affected operation identifiers are extracted from the incremental business data stream to limit the scope of subsequent recalculation and reduce redundant processing of irrelevant objects. The extraction of affected operation identifiers can be completed during the receiving or merging phase. During extraction, the operation identifier field is scanned and deduplicated from both the business operation record data and the operation event log data. Deduplication can be maintained using a set structure, and each operation identifier is associated with a change type identifier. The change type identifier distinguishes changes from business operation record data from changes from operation event log data. The extraction results form a set of affected operation identifiers. This set can include a change timestamp range to indicate the time range covered by the incremental business data. The change timestamp range is used to limit the statistical window or the event count range when extracting subsequent anomaly feature sets.

[0096] Based on the updated standardized business dataset and the updated audit trail record set, anomaly feature set extraction and anomaly analysis model inference are performed for affected operation identifiers to update anomaly scores and maintain consistency between the anomaly scores and operation identifiers. Anomaly feature set extraction on the updated standardized business dataset side reads business numerical fields and business status fields to form business feature vectors corresponding to the affected operation identifiers. On the updated audit trail record set side, timestamp sequences and behavior code sequences from the time-series correlation sequence are read to form time-series feature vectors and context feature vectors. Then, the two vectors are aligned and concatenated, with the operation identifier as the primary key to ensure that multiple-source features of the same operation identifier enter the same anomaly feature set. Anomaly analysis model inference takes the anomaly feature set as input and outputs anomaly prediction probability values. These anomaly prediction probability values ​​are mapped to a preset numerical range to obtain the updated anomaly score. The mapping process can use piecewise linear mapping or lookup table mapping. The lookup table entries are configured according to the anomaly prediction probability value range and output a stable score scale. After generating the updated abnormal scores, the corresponding entries for the operation identifiers can be updated and written in the data index table. The update write can retain historical score versions and record update timestamps. Historical score versions are used to support the judgment and backtracking of dynamic update trigger conditions.

[0097] Based on the updated standardized business dataset, anomaly detection is performed on affected operation identifiers to generate an updated set of anomalous operation events, maintaining consistency between the aggregation structure of the anomalous operation event set and the operation identifiers. Within the updated standardized business dataset, anomaly detection can perform time-slice window partitioning on the subset of records corresponding to the affected operation identifiers. The time-slice windows are divided after being sorted by record timestamps, and the partitioning method can use a fixed-duration window or a fixed-record-count window. A statistical threshold is determined based on the statistical distribution of the data within each time-slice window. The statistical threshold can be a quantile threshold, a mean and dispersion threshold, or a moving average threshold, and is bound to a specified numerical field. During the record filtering stage, for each record corresponding to the affected operation identifier, the value of the specified numerical field deviates from the statistical threshold, forming an outlier record set. This outlier record set is maintained... The system retains an operation identifier and a deviation direction identifier, with the deviation direction identifier used to distinguish between excessively high and excessively low values. During the event generation phase, abnormal operation events containing anomaly type identifiers are generated based on outlier records. The anomaly type identifier can be obtained by combining a specified numerical field identifier with the deviation direction identifier. During the event aggregation phase, abnormal operation events are aggregated into an updated abnormal operation event set based on the operation identifier. The updated abnormal operation event set can sort events under the same operation identifier by timestamp and execute event merging rules. The event merging rules can merge adjacent events into composite events based on the anomaly type identifier and a time interval threshold to reduce event fragmentation.

[0098] The generation and transmission of alert messages and anomaly report data are triggered based on the updated anomaly score and the updated set of anomaly operation events to deliver the dynamic update results to subsequent output processes and avoid invalid distribution. Triggering actions can incorporate triggering conditions, which must include at least anomaly score change conditions and event set change conditions. Anomaly score change condition can be determined based on the difference threshold between the updated anomaly score and historical score versions. Event set change condition can be determined based on changes in the number of events in the updated set of anomaly operation events, changes in the anomaly type identifier set, or changes in the latest event timestamp. When the triggering conditions are met, a trigger flag is output, and the affected operation identifier, the corresponding updated anomaly score, and the updated set of anomaly operation events are written to a pending-send queue. This queue is used to interface with the generation and transmission of alert messages and anomaly report data. When the triggering conditions are not met, a suppression flag is output, and the suppression reason is recorded. The suppression reason is used to audit dynamic update behavior and support subsequent threshold adjustments.

[0099] For example, in a health insurance claims intelligent risk control scenario, the system needs to identify and block fraud risks in real time for each claim application. This scenario involves obtaining information from multiple heterogeneous data sources, including the core policy management system, the electronic medical record systems of partner medical institutions, third-party medical data platforms, and customer self-service platforms. Each claim application is assigned a unique claim application number as an operation identifier throughout the entire process. The structured policy and claim information provided by the policy management system, including the insured's identity, policy effective date, scope of insurance liability, date of treatment for this application, diagnosis code, details of medical expenses, and the amount of compensation requested, constitutes business operation record data. At the same time, the operation trajectory of customers submitting materials through mobile applications, the click and approval logs of internal auditors in the processing system, and the call events of data interfaces with medical institutions together constitute continuously generated operation event log data. The system extracts claims business records from the core business database in daily batches of time slice windows by configuring data reading interfaces, and captures the operation event stream from various front-end applications in real time by subscribing to the enterprise's internal message bus.

[0100] For the acquired multi-source business data, the system performs in-depth standardization processing to build a consistent analytical foundation. First, the system parses the raw data stream, accurately separating claim records from different business systems from discrete operation event logs. For claim records, the system iterates through all fields, identifies and automatically fills in missing insured contact information using preset rules, maps diagnostic descriptions and drug names from different medical institutions using varying coding systems to a unified International Classification of Diseases (ICD) code and a unified drug catalog, converts all related expense amounts to a base currency unit, and finally generates a standardized business dataset where each record is firmly linked to a claim application number. For the massive operation event logs, the system performs cleaning and standardization, unifying the format of all timestamps to a standard time, standardizing event codes for various operational behaviors, and generating standardized event records based on the processed timestamps and event codes, where each event is bound to a specific claim application number. Subsequently, the system uses the claim application number as the aggregation key to reassemble all standardized event records related to the same application, such as "client submits application", "uploads medical invoice images", "initial reviewer receives task", "calls medical data verification interface", and "claims adjuster makes preliminary conclusion", into a complete time-series associated sequence, which is persistently stored as an audit trail record set. This allows for the accurate reconstruction of the entire operational process and decision-making process of each claim application from initiation to settlement.

[0101] Based on standardized business datasets and audit trail records, the system extracts a multi-dimensional set of abnormal features for each claim application number. From the standardized claims data, it extracts numerical business features such as the deviation between the claimed amount and the insured's historical average claim amount, the proportion of out-of-pocket expenses in the current medical expenses, the proportion of drug costs in the total expenses, and business status change features such as the frequency of abnormally rapid transitions between "pending review," "supplementary materials," "rejection," and "approved." From the audit trail record sequence, it extracts temporal interval features such as the time difference between submitting the application and the first supplementary materials, the density of applications processed by the same auditor during nighttime hours, and contextual features such as whether the IP address of the device used to submit the application changes frequently, and the distance between the geographical location of the application submission and the applicant's permanent residence. These features are then vectorized and concatenated to form a high-dimensional feature vector, used to comprehensively characterize the behavioral pattern of the claim application. The feature vector is loaded into an anomaly analysis model pre-trained using a gradient boosting decision tree algorithm. The model performs forward inference based on patterns learned from historically confirmed fraud and legitimate claims cases, outputting an anomaly prediction probability value representing the likelihood of fraud. This value is then converted into an intuitive anomaly score from 0 to 100 using a pre-defined mapping function. This anomaly score is immediately strongly correlated with the claim number that generated it and persistently stored as a key-value pair in a risk scoring index table for subsequent real-time querying and correlation analysis.

[0102] Meanwhile, the system performs independent anomaly detection based on statistical regularities on the standardized business dataset itself. The system divides the dataset into continuous time slices on a weekly basis according to the timestamp of the medical visit date in the claim application. For the payout amount data of all claims within each window, the system analyzes its statistical distribution and dynamically calculates the statistical threshold for that window. The system then iterates through each claim record in the dataset, determining its corresponding statistical window based on its medical visit timestamp and judging whether its claim amount significantly deviates from the threshold of that window. For example, records with amounts exceeding the 99.5th percentile within the window are filtered and marked as outliers, and each outlier record clearly has its claim application number. Based on these outliers, the system generates abnormal operation events containing specific anomaly type identifiers such as "abnormally high amount" or "abnormal combination of medical items," and aggregates and organizes all events according to the claim application number to form a structured set of abnormal operation events.

[0103] Next, based on the quantified anomaly score and the identified set of abnormal operation events, the system automatically generates and sends a warning message and detailed anomaly report data. The system compares the anomaly score with preset high, medium, and low-level anomaly warning thresholds to determine the current risk warning level of the claim application. Based on this warning level and the specific anomaly score, the system dynamically assembles and generates a concise warning message, such as "Claim application number CLM20241128001, anomaly score 92, suspected of document forgery, please prioritize manual review." Simultaneously, the system performs serialization format conversion on the set of abnormal operation events to obtain serialized event data that is easy to transmit and parse over the network. This data is then integrated with the anomaly warning level and anomaly score to automatically generate a structured anomaly report. This report not only includes a risk conclusion but also integrates detailed characteristic values ​​that triggered the high-risk determination, a complete timeline of user operations and system review events, related medical document image summaries, and comparison data with the insured's historical claims, forming an irrefutable chain of anomaly evidence. The system then establishes an encrypted communication channel with the investigator's work terminal in the claims and anti-fraud department. Through this channel, real-time warning messages and complete anomaly report data packets are pushed together to ensure that risk information reaches the handling end directly.

[0104] To cope with the continuous influx of claims, the system is designed with a dynamic update mechanism to achieve real-time risk monitoring. When a new claim occurs or the status of an existing claim is updated, an incremental business data stream containing new or changed records is generated. The system receives this data stream and merges the business records and event logs into the existing standardized business dataset and audit trail record set, respectively, forming an updated data set reflecting the latest business landscape. The system quickly extracts the affected claim numbers from the incremental data, and then, based on the updated complete data set, re-executes anomaly feature extraction and model inference calculation only for these affected claim numbers, generating updated anomaly scores to promptly reflect the latest behavioral changes of insured individuals or medical institutions. Simultaneously, based on the updated standardized business dataset, anomaly detection scanning is re-executed for the same batch of affected claim numbers, generating an updated set of abnormal operation events. Finally, based on these updated scores and event sets, the system automatically triggers the generation and sending of early warning messages and anomaly reports again, ensuring that risk situation awareness and business progress remain synchronized, and maintaining continuous pressure on ongoing fraudulent activities.

[0105] During this process, all generated and used data assets, including standardized business datasets, audit trail records, anomaly scores, sets of abnormal operation events, early warning messages, and anomaly report data, are encrypted using national cryptographic algorithms and stored in a secure storage area that meets Level 3 requirements of the Information Security Protection System. When claims auditors, investigators, or the compliance audit system initiate data access requests, the system accurately parses the user's digital identity certificate and the target data object identifier in the request instruction. Based on a pre-configured, role- and attribute-based access control policy, it strictly verifies the user's data access permissions. Regardless of whether the permission verification passes or fails, the system generates an access audit log entry containing a precise timestamp, user identity, type of accessed data object, operation action, and verification result. This entry is securely stored in a tamper-proof log database, ensuring that every step from data collection and analysis to access meets the most stringent data security and compliance audit requirements of the financial and insurance industry. This lays a solid foundation for building a trustworthy, reliable, and transparent intelligent risk control system.

[0106] This embodiment receives incremental business data streams and merges them into a standardized business dataset and an audit trail record set, respectively, to form updated standardized business datasets and updated audit trail record sets, ensuring that the data containers remain consistent within the same update cycle. By extracting affected operation identifiers from the incremental business data streams and limiting the processing scope of anomaly feature set extraction, anomaly analysis model inference, and anomaly detection accordingly, the updated anomaly score and updated set of anomaly operation events can be recalculated and updated within the affected object set. By performing trigger determination and outputting trigger flags based on the updated anomaly score and updated set of anomaly operation events, the generation and transmission of warning messages and anomaly report data are matched with data changes, thereby reducing redundant processing and invalid distribution and improving the controllability and response efficiency of dynamic updates.

[0107] In one embodiment, an anomaly detection and early warning device is provided, which corresponds one-to-one with the anomaly detection and early warning methods described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly detection and early warning device of the present invention. The modules include a multi-source data acquisition module 10, a data standardization processing module 20, an anomaly feature analysis module 30, an anomaly detection and judgment module 40, an early warning and report generation module 50, and an incremental update scheduling module 60. Detailed descriptions of each functional module are as follows: Multi-source data acquisition module 10 is used to acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The data standardization processing module 20 is used to standardize the multi-source business data, generate a standardized business dataset containing the operation identifier, and construct an audit trail record set. The anomaly feature analysis module 30 is used to extract an anomaly feature set for the operation identifier based on the standardized business dataset and the audit trail record set, input the anomaly feature set into a preset anomaly analysis model to output an anomaly score, and associate the anomaly score with the operation identifier; The anomaly detection and judgment module 40 is used to perform anomaly detection on the standardized business dataset and generate a set of abnormal operation events associated with the operation identifier. The early warning and report generation module 50 is used to generate and send early warning messages and abnormal report data based on the abnormal score and the abnormal operation event set. The incremental update scheduling module 60 is used to update the standardized business dataset and the audit trail record set based on incremental business data, redetermine the anomaly score and update the abnormal operation event set, so as to dynamically update the warning message and the anomaly report data.

[0108] Specific limitations regarding anomaly detection and early warning devices can be found in the aforementioned limitations regarding anomaly detection and early warning methods, and will not be repeated here. Each module in the aforementioned anomaly detection and early warning device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0109] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an anomaly detection and early warning method on the server side.

[0110] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the client-side functions or steps of an anomaly detection and early warning method.

[0111] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The multi-source business data is standardized to generate a standardized business dataset containing the operation identifier, and an audit trail record set is constructed. Based on the standardized business dataset and the audit trail record set, an abnormal feature set for the operation identifier is extracted, the abnormal feature set is input into a preset abnormal analysis model to output an abnormal score, and the abnormal score is associated with the operation identifier; Anomaly detection is performed on the standardized business dataset to generate a set of abnormal operation events associated with the operation identifier; Based on the anomaly score and the set of abnormal operation events, generate and send early warning messages and anomaly report data; The standardized business dataset and the audit trail record set are updated based on incremental business data, the anomaly score is re-determined and the abnormal operation event set is updated, so as to dynamically update the warning message and the anomaly report data.

[0112] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The multi-source business data is standardized to generate a standardized business dataset containing the operation identifier, and an audit trail record set is constructed. Based on the standardized business dataset and the audit trail record set, an abnormal feature set for the operation identifier is extracted, the abnormal feature set is input into a preset abnormal analysis model to output an abnormal score, and the abnormal score is associated with the operation identifier; Anomaly detection is performed on the standardized business dataset to generate a set of abnormal operation events associated with the operation identifier; Based on the anomaly score and the set of abnormal operation events, generate and send early warning messages and anomaly report data; The standardized business dataset and the audit trail record set are updated based on incremental business data, the anomaly score is re-determined and the abnormal operation event set is updated, so as to dynamically update the warning message and the anomaly report data.

[0113] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0114] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0115] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0116] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various public, legal, and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

[0117] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An anomaly detection and early warning method, characterized in that, Includes the following steps: Acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The multi-source business data is standardized to generate a standardized business dataset containing the operation identifier, and an audit trail record set is constructed. Based on the standardized business dataset and the audit trail record set, an abnormal feature set for the operation identifier is extracted, the abnormal feature set is input into a preset abnormal analysis model to output an abnormal score, and the abnormal score is associated with the operation identifier; Anomaly detection is performed on the standardized business dataset to generate a set of abnormal operation events associated with the operation identifier; Based on the anomaly score and the set of abnormal operation events, generate and send early warning messages and anomaly report data; The standardized business dataset and the audit trail record set are updated based on incremental business data, the anomaly score is re-determined and the abnormal operation event set is updated, so as to dynamically update the warning message and the anomaly report data.

2. The anomaly detection and early warning method as described in claim 1, characterized in that, Acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data, including: Configure a data reading interface for connecting to the business database, and capture business operation record data through the data reading interface according to a preset time slice window; Configure a subscription listener for connecting to the log message queue, and capture operation event log data in real time through the subscription listener; Extract the operation identifier from the business operation record data and the operation event log data, and perform a non-empty check on the operation identifier; Based on the result of the non-empty verification, the business operation record data and operation event log data with empty operation identifiers are filtered out from the business operation record data and the operation event log data, and the business operation record data and operation event log data with non-empty operation identifiers are aggregated into the original data buffer pool. The datasets stored in the original data buffer pool are combined into multi-source business data.

3. The anomaly detection and early warning method as described in claim 1, characterized in that, The multi-source business data is standardized to generate a standardized business dataset containing the operation identifier, and an audit trail record set is constructed, including: The multi-source business data is analyzed to distinguish between business operation record data and operation event log data; The business operation record data is traversed to identify null fields and non-standard format fields, and null fields are corrected and non-standard format fields are standardized using preset cleaning rules; The corrected and unified business operation record data is mapped to a preset standard data model to generate a standardized business dataset containing the operation identifier; The operation event log data is standardized and cleaned to unify the timestamp format and behavior code. Standardized event records are generated based on the unified timestamps and behavior codes. Each standardized event record is associated with an operation identifier. The standardized event records are reorganized into a time-series associated sequence according to the operation identifier, and the time-series associated sequence is stored as an audit trail record set.

4. The anomaly detection and early warning method as described in claim 1, characterized in that, Based on the standardized business dataset and the audit trail record set, an abnormal feature set for the operation identifier is extracted. The abnormal feature set is input into a preset anomaly analysis model to output an anomaly score, and the anomaly score is associated with the operation identifier, including: Extract business numerical features and business status change features for the operation identifier from the standardized business dataset; Extract the operation timing interval features and operation environment context features for the operation identifier from the audit trail record set; The business numerical features, the business status change features, the operation timing interval features, and the operation environment context features are subjected to vectorized concatenation processing to generate an abnormal feature set. The set of abnormal features is loaded into an anomaly analysis model constructed using the gradient boosting decision tree algorithm for inference; Obtain the anomaly prediction probability value output by the anomaly analysis model, and map the anomaly prediction probability value to a preset numerical range to obtain an anomaly score; Establish a key-value pair mapping relationship between the abnormal score and the operation identifier in the data index table.

5. The anomaly detection and early warning method as described in claim 1, characterized in that, Anomaly detection is performed on the standardized business dataset to generate a set of abnormal operation events associated with the operation identifier, including: The standardized business dataset is divided into multiple consecutive time slice windows according to the timestamp order; Statistical thresholds are determined based on the statistical distribution of data within each time slice window. Iterate through each record in the standardized business dataset, determine the time slice window to which the record belongs based on the record's timestamp, and filter records whose values ​​of specified numerical fields deviate from the statistical threshold of the time slice window as outlier records. Each outlier record is associated with an operation identifier. An abnormal operation event containing an anomaly type identifier is generated based on the outlier record; Based on the operation identifier, the abnormal operation events are grouped into an abnormal operation event set.

6. The anomaly detection and early warning method as described in claim 1, characterized in that, Based on the anomaly score and the set of abnormal operation events, a warning message and anomaly report data are generated and sent, including: The anomaly score is compared with a preset anomaly warning threshold to determine the current anomaly warning level. An early warning message containing an emergency notification text is constructed based on the anomaly warning level and the anomaly score; Perform serialization format conversion on the set of abnormal operation events to obtain serialized event data; Based on the serialized event data, the anomaly warning level, and the anomaly score, anomaly report data containing an anomaly evidence chain is generated; Establish a communication channel with the remote control terminal, and distribute the early warning message and the anomaly report data through the communication channel.

7. The anomaly detection and early warning method as described in claim 1, characterized in that, The standardized business dataset and the audit trail record set are updated based on incremental business data. Anomaly scores are re-determined and the abnormal operation event set is updated to dynamically update the warning message and the anomaly report data, including: Receive incremental service data streams, which include newly added service operation record data and operation event log data; The business operation record data and operation event log data in the incremental business data stream are merged into the standardized business dataset and the audit trail record set, respectively, to generate an updated standardized business dataset and an updated audit trail record set; Extract the affected operation identifier from the incremental service data stream; Based on the updated standardized business dataset and the updated audit trail record set, perform anomaly feature set extraction and anomaly analysis model inference for the affected operation identifier to generate an updated anomaly score; Based on the updated standardized business dataset, anomaly detection is performed on the affected operation identifiers to generate an updated set of abnormal operation events; Based on the updated anomaly score and the updated set of abnormal operation events, the generation and sending of warning messages and anomaly report data are triggered.

8. An anomaly detection and early warning device, characterized in that, The anomaly detection and early warning device includes: The multi-source data acquisition module is used to acquire multi-source business data containing operation identifiers, wherein the multi-source business data includes business operation record data and operation event log data; The data standardization processing module is used to standardize the multi-source business data, generate a standardized business dataset containing the operation identifier, and construct an audit trail record set. An anomaly feature analysis module is used to extract an anomaly feature set for the operation identifier based on the standardized business dataset and the audit trail record set, input the anomaly feature set into a preset anomaly analysis model to output an anomaly score, and associate the anomaly score with the operation identifier; An anomaly detection and judgment module is used to perform anomaly detection on the standardized business dataset and generate a set of abnormal operation events associated with the operation identifier; The early warning and report generation module is used to generate and send early warning messages and abnormal report data based on the abnormal score and the abnormal operation event set. The incremental update scheduling module is used to update the standardized business dataset and the audit trail record set based on incremental business data, redetermine the anomaly score and update the abnormal operation event set, so as to dynamically update the warning message and the anomaly report data.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an anomaly detection and warning program stored in the memory and executable on the processor. When executed by the processor, the anomaly detection and warning program implements the steps of the anomaly detection and warning method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an anomaly detection and early warning program, which, when executed by a processor, implements the steps of the anomaly detection and early warning method as described in any one of claims 1-7.