Financial data anomaly identification method and system based on multi-agent cooperation
By employing a multi-agent collaborative financial data anomaly identification method, this approach utilizes agents for text understanding, data perception, feature reasoning, and anomaly assessment. Combined with dynamic weighted fusion and consistency constraints, it addresses the issues of insufficient multi-source data fusion and inconsistency in module outputs in existing technologies, thereby improving the accuracy and reliability of financial data anomaly identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BANK OF TAIZHOU CO LTD
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-04
AI Technical Summary
Existing methods for identifying financial data anomalies have shortcomings in multi-source data feature extraction, cross-modal information fusion, and multi-module collaborative decision-making, resulting in inconsistent and unstable anomaly identification results, which are difficult to meet the requirements for accuracy, comprehensiveness, and reliability in highly dynamic financial scenarios.
A financial data anomaly identification method based on multi-agent collaboration is adopted. By combining text understanding, data perception, feature reasoning and anomaly evaluation agents with dynamic weighted fusion and consistency constraint mechanism, a deep fusion and accurate evaluation of multi-dimensional anomaly features can be achieved.
It improves the accuracy, comprehensiveness, and reliability of identifying financial data anomalies, providing more efficient and precise technical support for financial risk prevention and control.
Smart Images

Figure CN122508418A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and more specifically, to a method and system for identifying financial data anomalies based on multi-agent collaboration. Background Technology
[0002] Currently, financial data anomaly identification systems serve as core intelligent analysis platforms in the field of financial risk prevention and control. Their multi-source data feature extraction capabilities, cross-modal information fusion capabilities, and multi-module collaborative decision-making capabilities directly determine the accuracy of anomaly identification results and the efficiency of risk warnings. Existing financial data anomaly identification methods mostly model single text or numerical data independently, employing shallow feature extraction to mine anomaly patterns, and deploying each identification module as an independent structure within the analysis process. However, conventional identification schemes are limited by factors such as the lack of deep correlation and fusion between text semantics and numerical features, the inability to capture time-series anomalies across scales, the lack of consistency constraints between module prediction results leading to output conflicts, and the lack of adversarial perturbation verification mechanisms resulting in insufficient model robustness. These existing schemes generally suffer from problems such as one-sided single-dimensional feature representation leading to missed anomaly detections and misjudgments, inconsistent identification results due to lack of collaboration in multi-module decision-making, and poor identification stability in complex perturbation and time-series variation scenarios. They struggle to achieve deep fusion of multi-source heterogeneous data, full-time-scale anomaly mining, and multi-agent collaborative decision-making while ensuring identification efficiency, failing to meet the needs of accurate, comprehensive, and stable anomaly identification in highly dynamic financial scenarios.
[0003] Therefore, improving the accuracy, comprehensiveness, and reliability of financial data anomaly identification is an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method and system for identifying financial data anomalies based on multi-agent collaboration, so as to improve the accuracy, comprehensiveness and reliability of financial data anomaly identification.
[0005] Firstly, this application provides a method for identifying financial data anomalies based on multi-agent collaboration, including: By using the text understanding agent, data perception agent, feature reasoning agent, and anomaly assessment agent in the initial financial data anomaly identification model, data anomaly embedding processing and anomaly loss calculation processing are performed on the financial text training data and financial numerical training data, respectively, to obtain text understanding anomaly loss, data perception anomaly loss, fusion perception anomaly loss, and prediction anomaly loss. The consistency constraint loss among the text prediction anomaly confidence level output by the text understanding agent, the data prediction anomaly confidence level output by the data perception agent, and the fusion prediction anomaly confidence level output by the feature reasoning agent is calculated using the collaborative scheduling agent in the initial financial data anomaly identification model. The total joint loss is obtained by jointly weighting the text understanding anomaly loss, the data perception anomaly loss, the fusion perception anomaly loss, the prediction anomaly loss, and the consistency constraint loss. The internal parameters of the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly assessment agent, and the collaborative scheduling agent in the initial financial data anomaly identification model are updated in reverse based on the total joint loss to obtain the trained financial data anomaly identification model. The financial data to be identified is input into the trained financial data anomaly identification model, and the comprehensive anomaly identification result of the financial data to be identified is output.
[0006] Secondly, this application provides a financial data anomaly identification system based on multi-agent collaboration. The financial data anomaly identification system based on multi-agent collaboration includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the financial data anomaly identification system based on multi-agent collaboration implements the aforementioned financial data anomaly identification method based on multi-agent collaboration.
[0007] The financial data anomaly identification method and system based on multi-agent collaboration provided in this application constructs an intelligent agent architecture that integrates five independent yet collaborative functions: text understanding, data perception, feature reasoning, anomaly assessment, and collaborative scheduling. It utilizes a fine-tuned large language model from the financial domain to process financial text and numerical data. Through feature concatenation, transformer encoder mining of temporal dependencies, and dynamic weighted fusion with consistency constraints, it achieves deep fusion and accurate assessment of multi-dimensional anomaly features. This effectively addresses the problems of one-sided identification of single-type data anomalies, insufficient fusion of multi-source data, and inconsistency in the output of various modules, thereby improving the accuracy, comprehensiveness, and reliability of financial data anomaly identification and providing more efficient and accurate technical support for financial risk prevention and control. Attached Figure Description
[0008] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0009] Figure 1 A flowchart illustrating a method for identifying financial data anomalies based on multi-agent collaboration, provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a financial data anomaly identification system based on multi-agent collaboration, provided in an embodiment of this application.
[0010] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0011] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] Figure 1 This is a flowchart illustrating a method for identifying financial data anomalies based on multi-agent collaboration, provided as an embodiment of this application. It should be understood that in other embodiments, the order of some steps in this method may be shared as needed, or some steps may be omitted or maintained. Figure 1 As shown, the method may include: S110. Using the text understanding agent, data perception agent, feature reasoning agent, and anomaly evaluation agent in the initial financial data anomaly identification model, perform data anomaly embedding processing and anomaly loss calculation processing on the financial text training data and financial numerical training data respectively, to obtain text understanding anomaly loss, data perception anomaly loss, fusion perception anomaly loss, and prediction anomaly loss.
[0013] In this step, the initial financial data anomaly identification model is a model architecture that includes multiple functionally independent but mutually cooperating intelligent agents.
[0014] The text understanding agent focuses on processing financial text data, and its core function is to extract semantic-level anomaly-related information from the text data; the data perception agent targets financial numerical data and is responsible for capturing anomaly patterns in the numerical data; the feature reasoning agent integrates the feature representations from the text understanding agent and the data perception agent to mine deep correlations and temporal dependencies; and the anomaly assessment agent integrates the output results of each agent to perform the final anomaly assessment and loss calculation.
[0015] Data anomaly embedding processing refers to converting raw financial text and numerical data into high-dimensional feature vector representations that can be understood by computers. These vectors can contain anomalous information in the data.
[0016] Anomaly loss calculation involves quantifying the model's prediction error by comparing the model's generated predictions with the actual anomaly labels. Text understanding anomaly loss measures the difference between the predictions made by the text understanding agent after processing financial text training data and the corresponding anomaly labels; data-aware anomaly loss measures the prediction error of the data-aware agent after processing financial numerical training data; fusion-aware anomaly loss measures the prediction error of the feature reasoning agent after fusing textual and numerical anomaly features; and prediction anomaly loss measures the error between the anomaly evaluation agent's comprehensive prediction anomaly score (generated by integrating various confidence levels) and the actual labels.
[0017] S111. The text understanding agent performs semantic anomaly mapping on the financial text training data to generate text anomaly embedding representation and text prediction anomaly confidence, and obtains text understanding anomaly loss based on the text prediction anomaly confidence and the anomaly labeling corresponding to the financial text training data.
[0018] Semantic anomaly mapping is a core operation of text understanding agents, encompassing a series of processes such as semantic analysis of financial text training data, anomaly information extraction, feature vectorization, and anomaly confidence prediction. Text anomaly embedding representation converts semantic anomaly information in financial text into a high-dimensional vector form; this vector set can represent the text's anomalous features in mathematical space.
[0019] The text prediction anomaly confidence score is the probability estimate by which a text understanding agent believes that financial text training data belongs to an anomaly category. For example, it can be a continuous numerical value within a specific range. The anomaly labeling corresponding to the financial text training data is a marker, either manually or through other reliable methods, indicating whether the financial text data contains anomalies. For example, it can be in binary form, such as "1" indicating anomaly and "0" indicating normality. The text understanding anomaly loss is an indicator that measures the difference between the text prediction anomaly confidence score and the anomaly labeling.
[0020] S1111. The financial text training data is processed by rule filtering using a fine-tuned large language model in the financial field and built-in rule prompts. Semantic abnormal text information is extracted from the financial text training data. The semantic abnormal text information is a set of text fragments containing abnormal pointers.
[0021] Fine-tuned large language models in the financial field refer to models obtained by further training and adjusting parameters using large-scale text corpora in the financial field on the basis of general pre-trained large language models, so that they can better understand the professional terminology, grammatical structure and semantic connotation of the financial field.
[0022] Built-in rule prompts are a series of pre-defined text instructions used to guide the model to perform specific rule filtering. These instructions can be based on common anomaly patterns and regulatory requirements in the financial field, such as text snippets describing abnormal revenue recognition, unfair related-party transactions, and abnormal fluctuations in financial indicators.
[0023] Rule-based filtering refers to the model's analysis of the input financial text training data sentence by sentence or paragraph by paragraph based on built-in rule prompts, determining whether it contains semantically anomalous information that conforms to the rules. Semantically anomalous text information refers to text fragments in the financial text training data that can imply or directly indicate the existence of financial anomalies, operational risks, or other problems. These fragments may involve unconventional expressions, contradictory descriptions of financial data, or ambiguous business interpretations. The text fragment set is a collection of multiple such semantically anomalous text fragments.
[0024] S1112. The semantically anomalous text information is encoded and converted using the encoding layer of the fine-tuned large language model in the financial field, and the semantically anomalous text information is converted into a text anomalous embedding representation, which is a set of feature vectors in a high-dimensional vector space.
[0025] In the financial domain, fine-tuning the encoding layer of a large language model can, for example, consist of multiple stacked Transformer encoders, each containing a multi-head self-attention mechanism and a feedforward neural network. Encoding transformation is the process by which the encoding layer converts semantically anomalous textual information (e.g., in the form of word or sub-word sequences) into a fixed-dimensional vector representation.
[0026] In this process, each word or subword in the text segment is first mapped to an initial word embedding vector. Then, a multi-head self-attention mechanism is used to capture the contextual dependencies between words, followed by nonlinear transformation and feature extraction via a feedforward neural network. The text anomaly embedding representation is a high-dimensional vector obtained after encoding and transformation, with dimensions ranging from hundreds to thousands, such as 768 or 1024 dimensions. A high-dimensional vector space refers to the mathematical space in which these vectors reside, where each dimension represents a specific feature of the text. The feature vector set is the collection composed of the high-dimensional vectors corresponding to each text segment in the semantically anomalous text information.
[0027] S1113. The text anomaly embedding representation is subjected to dimension mapping processing through text single-layer linear transformation parameters to generate a text intermediate transformation vector, and the text intermediate transformation vector is subjected to interval compression processing through text nonlinear activation function to generate text prediction anomaly confidence, wherein the text prediction anomaly confidence is a continuous value within a preset interval range.
[0028] The text single-layer linear transformation parameters are a learnable weight matrix and bias vector. The number of rows in the weight matrix is equal to the dimension of the text anomaly embedding representation, while the number of columns is set according to task requirements. It is used to map the high-dimensional text anomaly embedding representation to a new dimensional space.
[0029] Dimension mapping refers to the process of performing matrix multiplication and addition operations on the text anomaly embedding representation and the text single-layer linear transformation parameters to obtain the text intermediate transformation vector. Its mathematical expression can be understood as: text intermediate transformation vector = text anomaly embedding representation × text single-layer linear transformation weight matrix + text single-layer linear transformation bias vector.
[0030] A text nonlinear activation function is a function with nonlinear characteristics, such as the Sigmoid function, Tanh function, ReLU function, etc. Its purpose is to perform nonlinear transformation on the intermediate transformation vector of the text, thereby enhancing the expressive power of the model.
[0031] Interval compression refers to compressing the range of values in the intermediate transformation vector of text into a specific interval using a non-linear activation function. For example, the sigmoid function can compress vector values to between 0 and 1. The text prediction anomaly confidence score is the result obtained after interval compression, representing the model's confidence that the semantically anomalous text information in the input belongs to an anomalous category. The preset interval range can be determined based on the characteristics of the activation function; for example, when using the sigmoid function, the preset interval range is 0 to 1.
[0032] S1114. Obtain the text understanding anomaly loss based on the text prediction anomaly confidence and the anomaly label corresponding to the financial text training data, and calculate the text difference metric between the text prediction anomaly confidence and the anomaly label using the binary cross-entropy loss function to obtain the text understanding anomaly loss.
[0033] The anomaly labeling for the financial text training data is a vector with the same dimension as the text prediction anomaly confidence score, where each element is either 0 or 1, representing whether the corresponding sample is normal or abnormal, respectively. The binary cross-entropy loss function is a loss function commonly used in binary classification tasks, and its calculation formula is based on the difference in probability distribution between the text prediction anomaly confidence score and the anomaly labeling.
[0034] The text difference metric is calculated using a binary cross-entropy loss function to quantify the difference between the text prediction anomaly confidence score and the anomaly label. Specifically, for each sample, the anomaly label is used as the true probability distribution, and the text prediction anomaly confidence score is used as the predicted probability distribution. The cross-entropy between these two distributions is then calculated, and the average of the cross-entropies across all samples yields the text understanding anomaly loss.
[0035] S1115. Based on the text understanding anomaly loss, update the parameters of the financial domain fine-tuned large language model, the parameters of the text single-layer linear transformation, and the parameters of the text nonlinear activation function through the backpropagation algorithm.
[0036] Backpropagation is a common method for training neural networks. Its basic principle is to calculate the gradient of each parameter along the back path of the network based on the calculated loss function value, and then use optimization algorithms such as gradient descent to update the parameters.
[0037] In this step, the gradient of the text understanding anomaly loss with respect to the output of the text nonlinear activation function is first calculated. Then, following the chain rule, the gradients with respect to the intermediate text transformation vector, the text single-layer linear transformation parameters, and the text anomaly embedding representation are calculated sequentially, thus obtaining the gradients for each parameter of the encoding layer of the large language model for fine-tuning in the financial domain. After obtaining the gradients of all parameters, an optimizer (e.g., the Adam optimizer) is used to update the parameters of the large language model for fine-tuning in the financial domain (including the weights and biases of the encoding layer), the text single-layer linear transformation parameters (weight matrix and bias vector), and any learnable parameters that may exist in the text nonlinear activation function (if the activation function has learnable parameters) according to the gradient direction and learning rate, in order to reduce the text understanding anomaly loss.
[0038] S112. The data-aware intelligent agent performs numerical anomaly mapping processing on the financial numerical training data to generate a data anomaly embedding representation and a data prediction anomaly confidence score, and obtains the data-aware anomaly loss based on the data prediction anomaly confidence score and the anomaly label corresponding to the financial numerical training data.
[0039] Numerical anomaly mapping processing is a series of operations performed by a data-aware intelligent agent on financial numerical training data. Its purpose is to extract anomalous features from the numerical data and generate corresponding embedding representations and anomaly confidence scores. Financial numerical training data typically includes various financial indicators, transaction amounts, asset and liability data, and other data presented in numerical form.
[0040] Data anomaly embedding representation transforms anomalous information in financial numerical training data into a high-dimensional vector representation, similar to text anomaly embedding representation, but with a greater emphasis on characterizing numerical features. Data prediction anomaly confidence is the probability estimate by which a data-aware agent believes that financial numerical training data belongs to an anomaly category; similarly, it can be a continuous value within a preset range.
[0041] Anomaly labeling for financial numerical training data identifies whether the numerical data is anomalous, and its format is similar to anomaly labeling for financial text training data. Data-aware anomaly loss is an indicator that measures the difference between the confidence level of data prediction regarding anomalies and the corresponding anomaly labeling.
[0042] S1121. Use a fine-tuned large language model in the financial field to perform missing location identification processing on the financial numerical training data to determine the location of missing data.
[0043] During the actual collection and storage of financial numerical training data, some data may be missing due to various reasons (such as system failures, human error, etc.). Although fine-tuning large language models in the financial field are mainly used for text processing, their pattern recognition capabilities can be transferred to the identification of missing locations in numerical data.
[0044] In practice, financial numerical training data can be converted into text sequences in a specific format (such as combining numerical values and their corresponding indicator names into text), and then input into a large-scale language model for fine-tuning in the financial domain. By analyzing patterns and contextual information in the sequence, the model identifies which positions of data should be present but are actually missing, thereby determining the location of the missing data.
[0045] For example, for a time series of financial indicators that includes multiple quarters, the model can identify a missing value for the "net profit" indicator in a certain quarter by comparing the data distribution of adjacent quarters and the correlation between indicators.
[0046] S1122. Generate a fill value based on the data distribution characteristics around the location of the missing data.
[0047] The distribution characteristics of the data around the location of missing data include statistical features such as the magnitude, trend, variance, and mean of the adjacent data before and after that location. There are various methods to generate imputed values, such as mean imputation, median imputation, linear interpolation, and weighted average based on adjacent data.
[0048] Specifically, the mean or median of several data points before and after the missing data location can be calculated as the filler value; or the value at the missing location can be predicted by fitting a linear function based on the coordinates of the known data points before and after it; or the filler value can be generated by using the moving average method based on the time series characteristics of the data.
[0049] S1123. Fill the missing data positions with the filling values to generate a complete set of numerical data.
[0050] After determining the filler values, they are used to replace the missing markers (such as NaN, null values, etc.) at the locations of the missing data, thus making the originally missing financial numerical training data complete. A complete numerical dataset refers to a dataset where all data locations are filled with valid values. It retains the structure and most of the information of the original data while eliminating the impact of missing data on subsequent processing.
[0051] S1124. The encoding layer of the fine-tuned large language model in the financial field is used to encode and transform the complete numerical data set to obtain a data anomaly embedding representation, which is a set of vector representations under the numerical feature dimension.
[0052] Similar to handling semantically anomalous text information, this approach also utilizes the encoding layer of a fine-tuned large language model in the financial domain to encode the complete numerical dataset. First, the complete numerical dataset is converted into an input format acceptable to the model, for example, by combining each numerical value with its corresponding indicator name, timestamp, and other information into a text sequence. Then, the encoding layer processes these text sequences, using word embeddings, multi-head self-attention mechanisms, and feedforward neural networks to convert the numerical data into high-dimensional vectors. Data anomaly embedding emphasizes the characterization of data anomalies from the perspective of numerical features. Each vector in the vector representation set corresponds to a data item or data segment in the complete numerical dataset, implying the anomalous features of that data item or segment at the numerical level.
[0053] S1125. Perform dimensional mapping processing on the data anomaly embedding representation through data single-layer linear transformation parameters to generate data intermediate transformation vector.
[0054] Similar to the single-level linear transformation parameters for text, the data single-level linear transformation parameters are also learnable weight matrices and bias vectors. The number of rows in the weight matrix equals the dimension of the data anomaly embedding representation, while the number of columns is determined based on subsequent processing requirements. The dimension mapping process is: Data intermediate transformation vector = Data anomaly embedding representation × Data single-level linear transformation weight matrix + Data single-level linear transformation bias vector.
[0055] This process maps the data anomaly embedding representation from the original high-dimensional space to a new dimensional space, preparing for subsequent nonlinear transformations and confidence predictions.
[0056] S1126. The intermediate transformation vector of the data is subjected to interval compression processing by the data nonlinear activation function to generate the data prediction anomaly confidence.
[0057] The nonlinear activation function for the data can also be, for example, commonly used nonlinear functions such as Sigmoid and Tanh. Interval compression compresses the range of values for the intermediate transformation vector of the data into a preset interval, such as between 0 and 1. The data prediction anomaly confidence score is the result obtained after compression processing; it reflects the probability judgment of the data-aware agent regarding the existence of anomalies in the complete numerical data set.
[0058] S1127. Calculate the data difference metric between the data prediction anomaly confidence and the anomaly label corresponding to the financial numerical training data using the binary cross-entropy loss function to obtain the data-perceived anomaly loss.
[0059] The calculation logic in this step is similar to that in S1114. The anomaly labels corresponding to the financial numerical training data are binary vectors, and the anomaly prediction confidence is a continuous probability value. The cross-entropy between the two is calculated using a binary cross-entropy loss function, and then averaged over all samples to obtain a data difference metric, i.e., the data-perceived anomaly loss.
[0060] S1128. Based on the data-perceived abnormal loss, update the parameters of the financial domain fine-tuning large language model, the parameters of the data single-layer linear transformation, and the parameters of the data nonlinear activation function through the backpropagation algorithm.
[0061] This step is similar to the parameter update process in S1115. Based on the data-aware anomaly loss, the gradients of the coding layer parameters, single-layer linear transformation parameters, and nonlinear activation function parameters of the financial domain fine-tuning large language model are calculated using the backpropagation algorithm. Then, the optimizer is used to update the parameters to reduce the data-aware anomaly loss.
[0062] S113. The feature reasoning agent performs time-dependent fusion processing on the text anomaly embedding representation and the data anomaly embedding representation to generate a deep anomaly fusion feature representation and a fusion prediction anomaly confidence score. The fusion-aware anomaly loss is obtained based on the anomaly labeling that corresponds to both the fusion prediction anomaly confidence score and the financial text training data and the financial numerical training data.
[0063] Temporal dependency fusion processing is the core function of feature reasoning agents. It not only needs to fuse two different types of features, namely text anomaly embedding representation and data anomaly embedding representation, but also needs to consider the possible temporal relationships between them, such as the correlation between text descriptions and numerical data at the same point in time, as well as the evolution trend of data at different points in time.
[0064] Deep anomaly fusion feature representation is a high-dimensional feature vector obtained after fusion processing, which comprehensively reflects textual and numerical anomaly information and their temporal dependencies. Fusion prediction anomaly confidence is the probability prediction of whether data is anomalous or not based on the deep anomaly fusion feature representation by the feature reasoning agent.
[0065] Anomaly labeling, which corresponds to both financial text training data and financial numerical training data, is a comprehensive anomaly marker for the same financial entity or event, taking into account both textual and numerical information. The fusion-aware anomaly loss is a measure of the difference between the fusion-predicted anomaly confidence score and this common anomaly label.
[0066] S1131. Perform feature concatenation operation on the text anomaly embedding representation and the data anomaly embedding representation to generate a concatenated anomaly feature sequence.
[0067] Feature concatenation involves joining and combining text anomaly embeddings and data anomaly embeddings along a specific dimensional axis. Assuming the dimensions of the text anomaly embedding are (number of samples, number of time steps, text feature dimension) and the dimensions of the data anomaly embedding are (number of samples, time steps, data feature dimension), then the dimension of the concatenated anomaly feature sequence could be, for example, (number of samples, time steps, text feature dimension + data feature dimension), which means concatenating the text and data feature vectors at the same time step along the feature dimension.
[0068] The concatenated anomaly feature sequence contains both textual and numerical original anomaly feature information, laying the foundation for subsequent temporal dependency mining.
[0069] S1132. The self-attention calculation process is performed on the spliced abnormal feature sequence using the encoder structure of the converter to obtain the self-attention output feature.
[0070] The encoder structure of the transformer consists of multiple identical encoder layers, each containing a multi-head self-attention sub-layer and a feedforward neural network sub-layer. Self-attention computation is the core operation of the multi-head self-attention sub-layer, which enables the model to focus on the dependencies between different positions in the spliced anomalous feature sequence.
[0071] Specifically, for each position in the concatenated anomaly feature sequence, the self-attention mechanism calculates the attention weights between that position and all other positions in the sequence. Then, it performs a weighted sum of the feature vectors of all positions based on these attention weights to obtain the self-attention output feature for that position. Multi-head self-attention, on the other hand, computes multiple different attention heads in parallel to capture different types of dependencies, and then concatenates the outputs of each head. The self-attention output feature integrates information from all positions in the sequence, effectively capturing long-distance temporal dependencies.
[0072] S1133. The encoder structure of the converter is used to perform forward propagation processing on the self-attention output features, and deep anomaly temporal feature representation is extracted from the spliced anomaly feature sequence. The deep anomaly temporal feature representation contains anomaly evolution mode information under different time steps.
[0073] Forward propagation is the operation of a sublayer in a feedforward neural network, which performs nonlinear transformations and feature extraction on the self-attention output features. Feedforward neural networks typically contain two linear transformation layers and a nonlinear activation function (such as ReLU), which can further process the self-attention output features and extract higher-level, more abstract features.
[0074] After processing by the encoder structure of the transformer (a stack of multiple encoder layers), the self-attention output features are transformed into a deep anomaly temporal feature representation. The deep anomaly temporal feature representation not only includes the fusion features of text and numerical values, but also contains the evolution patterns of anomaly features at different time steps, such as how anomaly features change over time and how anomaly features at different time points are related.
[0075] S1134. Calculate the attention energy value corresponding to each time step in the deep anomaly temporal feature representation using learnable attention weight parameters, and perform exponential normalization on the attention energy value to generate the attention weight coefficient corresponding to each time step.
[0076] The learnable attention weight parameters include query, key, and value matrices, which are learned during model training.
[0077] For the temporal feature representation of deep anomalies, it is first multiplied by the query matrix and the key matrix respectively to obtain the query vector and the key vector. Then, the dot product of the query vector and the key vector is calculated to obtain the initial attention energy value, which reflects the correlation between features at different time steps.
[0078] Exponential normalization can be achieved, for example, by using the Softmax function to convert attention energy values into attention weight coefficients, ensuring that the sum of the attention weight coefficients across all time steps is 1. Specifically, the calculation involves taking the exponent of the attention energy value at each time step and then dividing it by the sum of the exponents of all attention energy values at all time steps to obtain the attention weight coefficient for that time step.
[0079] S1135. Multiply the depth anomaly temporal feature representation of each time step by the attention weight coefficient of the corresponding time step to generate a weighted time step feature vector.
[0080] This step involves element-wise multiplying the feature vector of each time step in the deep anomaly temporal feature representation with its corresponding attention weight coefficient to obtain a weighted time step feature vector. The larger the attention weight coefficient of a time step, the greater the contribution of its corresponding feature vector to the subsequent fusion process, thus enabling the model to focus on the time step information that is more important for anomaly identification.
[0081] S1136. Perform a summation and aggregation operation on the weighted time step feature vectors of all the time steps to generate a deep anomaly fusion feature representation, and perform dimensional mapping processing on the deep anomaly fusion feature representation by fusing single-layer linear transformation parameters to generate a fusion intermediate transformation vector.
[0082] The summation and aggregation operation adds the weighted time step feature vectors of all time steps together along the time dimension to obtain a feature vector that integrates information from all time steps, which is the deep anomaly fusion feature representation.
[0083] The fusion single-layer linear transformation parameters are a learnable weight matrix and bias vector used to perform dimensionality mapping on the deep anomaly fusion feature representation. The process is similar to the single-layer linear transformation of text and data, generating a fusion intermediate transformation vector.
[0084] S1137. The fused intermediate transformation vector is subjected to interval compression processing by a fused nonlinear activation function to generate a fused prediction anomaly confidence score.
[0085] Similarly, nonlinear activation functions such as Sigmoid can be used to compress the intermediate transformation vector of the fusion, limiting its value range to a preset range, and obtaining the fusion prediction anomaly confidence. This confidence combines the anomaly information of the text and numerical values as well as the temporal dependency.
[0086] S1138. Based on the fusion prediction anomaly confidence and the anomaly labels corresponding to both the financial text training data and the financial numerical training data, calculate the fusion difference metric between the fusion prediction anomaly confidence and the anomaly labels corresponding to both the financial text training data and the financial numerical training data using the binary classification cross-entropy loss function, and obtain the fusion-aware anomaly loss.
[0087] The anomaly labels corresponding to both financial text training data and financial numerical training data are comprehensive anomaly labels for the same training sample. The cross-entropy between the fusion predicted anomaly confidence and this common label is calculated using a binary cross-entropy loss function. The average of these cross-entropy values yields the fusion difference metric, i.e., the fusion-perceived anomaly loss.
[0088] S1139. Update the parameters of the encoder structure of the transformer, the learnable attention weight parameters, the parameters of the fused single-layer linear transformation, and the parameters of the fused nonlinear activation function through the backpropagation algorithm based on the fused anomalous loss.
[0089] Based on the fusion-aware anomaly loss, the gradients of the parameters of each layer in the transformer encoder structure (e.g., the weights and biases of multi-head self-attention and feedforward neural networks), learnable attention weight parameters (query, key, value matrix, etc.), fusion single-layer linear transformation parameters, and fusion nonlinear activation function parameters are calculated using the backpropagation algorithm. Then, the optimizer is used to update these parameters to reduce the fusion-aware anomaly loss.
[0090] S114. The anomaly assessment agent dynamically weights and fuses the text prediction anomaly confidence, the data prediction anomaly confidence, and the fusion prediction anomaly confidence to generate a comprehensive prediction anomaly score. The prediction anomaly loss is obtained based on the anomaly labeling corresponding to the comprehensive prediction anomaly score and the financial text training data and the financial numerical training data.
[0091] Dynamic weighted fusion refers to an anomaly assessment agent assigning different weight coefficients to the anomaly confidence scores of text prediction, data prediction, and fused prediction based on their reliability or importance, and then performing a weighted summation to obtain a comprehensive anomaly prediction score.
[0092] Unlike fixed weights, dynamic weights can adaptively adjust based on different input data and model states, allowing more reliable confidence levels to carry a larger weight in the overall score. The overall predicted anomaly score is a comprehensive quantitative assessment of the degree of anomaly in financial data. The predicted anomaly loss is a measure of the difference between the overall predicted anomaly score and the common anomaly label.
[0093] S1141. Calculate the text weight coefficient corresponding to the text prediction anomaly confidence level through a learnable dynamic weight generation network, and calculate the data weight coefficient corresponding to the data prediction anomaly confidence level through the learnable dynamic weight generation network.
[0094] A learnable dynamic weight generation network is a small neural network model whose inputs can be, for example, text prediction anomaly confidence, data prediction anomaly confidence, fusion prediction anomaly confidence itself, and some of their statistical characteristics (such as variance, gradient, etc.), and the output is the corresponding weight coefficients.
[0095] Learnable dynamic weight generation networks learn patterns from training data to determine the reliability of various confidence levels under different conditions, thereby generating appropriate weights. For example, when the variance of the confidence level for predicting text anomalies is small, it indicates that the model's judgment of text anomalies is relatively stable, and the dynamic weight generation network can assign it a larger text weight coefficient; conversely, it assigns a smaller coefficient.
[0096] S1142. Calculate the fusion weight coefficient corresponding to the fusion prediction anomaly confidence through the learnable dynamic weight generation network, wherein the sum of the text weight coefficient, the data weight coefficient, and the fusion weight coefficient is a fixed constant.
[0097] After calculating the text weight coefficients and data weight coefficients, the dynamic weight generation network continues to calculate the fusion weight coefficients corresponding to the fusion prediction anomaly confidence levels. To ensure the rationality of the weighted fusion, the weight coefficients are usually normalized so that the sum of the text weight coefficients, data weight coefficients, and fusion weight coefficients equals a fixed constant, such as 1. This can be achieved by using the Softmax function in the output layer of the dynamic weight generation network, which converts the original output weight values into a probability distribution form with a sum of 1.
[0098] S1143. Multiply the text prediction anomaly confidence score by the text weight coefficient to generate a text weighted score.
[0099] The text-weighted score is the component of the text prediction anomaly confidence level in the overall prediction anomaly score. It is calculated by multiplying the text prediction anomaly confidence level by the text weight coefficient.
[0100] S1144. Multiply the data prediction anomaly confidence level by the data weight coefficient to generate a data weighted score.
[0101] The data-weighted score is the proportion of the data prediction anomaly confidence level in the overall prediction anomaly score. It is calculated by multiplying the data prediction anomaly confidence level by the data weight coefficient.
[0102] S1145. Multiply the fusion prediction anomaly confidence by the fusion weight coefficient to generate a fusion weighted score.
[0103] The fusion weighted score is the component of the fusion prediction anomaly confidence level in the overall prediction anomaly score. It is calculated as the product of the fusion prediction anomaly confidence level and the fusion weight coefficient.
[0104] S1146. The text-weighted score, the data-weighted score, and the fusion-weighted score are summed to generate a comprehensive prediction anomaly score.
[0105] The text-weighted score, data-weighted score, and fusion-weighted score are added together to obtain the comprehensive prediction anomaly score. This score integrates the prediction results of the three agents and reflects the relative importance of each result through dynamic weights.
[0106] S1147. Based on the anomaly labels corresponding to the comprehensive predicted anomaly score and the financial text training data and the financial numerical training data, calculate the squared difference between the comprehensive predicted anomaly score and the anomaly labels corresponding to the financial text training data and the financial numerical training data using the mean squared error loss function to obtain the predicted anomaly loss.
[0107] The mean squared error loss function is a commonly used regression loss function. It is calculated as the average of the squared differences between the overall predicted outlier score and the common outlier label (which can be a continuous value between 0 and 1, or a binary value 0 / 1). The squared difference value is an intermediate result in the calculation of the mean squared error loss function. The predicted outlier loss is obtained by averaging the squared differences of all samples.
[0108] S1148. Update the parameters of the learnable dynamic weight generation network using the backpropagation algorithm based on the predicted abnormal loss.
[0109] Based on the predicted anomaly loss, the learnable dynamic weights are calculated using the backpropagation algorithm to generate the parameter gradients of each layer of the network. Then, the optimizer is used to update these parameters, making the allocation of dynamic weights more reasonable, thereby reducing the predicted anomaly loss.
[0110] S120. Calculate the consistency constraint loss among the text prediction anomaly confidence level output by the text understanding agent, the data prediction anomaly confidence level output by the data perception agent, and the fusion prediction anomaly confidence level output by the feature reasoning agent through the collaborative scheduling agent in the initial financial data anomaly identification model.
[0111] The role of the collaborative scheduling agent is to coordinate the outputs of various agents and ensure consistency among them. Text prediction anomaly confidence, data prediction anomaly confidence, and fusion prediction anomaly confidence all determine whether the same financial data is abnormal; ideally, they should have high consistency.
[0112] The consistency constraint loss is used to measure the degree of inconsistency among the three confidence levels. There are several ways to calculate the consistency constraint loss. For example, you can calculate the absolute difference or squared difference between each pair of confidence levels and then sum or average these differences; or you can use more complex consistency measures, such as the Kappa coefficient or correlation coefficient, and convert them into a loss value.
[0113] For example, the absolute difference between the confidence scores of text prediction anomalies and data prediction anomalies, the absolute difference between the confidence scores of text prediction anomalies and fusion prediction anomalies, and the absolute difference between the confidence scores of data prediction anomalies and fusion prediction anomalies can be calculated. Then, these three absolute differences can be added together and averaged as the consistency constraint loss.
[0114] S130. The total joint loss is obtained by jointly weighting the text understanding anomaly loss, the data perception anomaly loss, the fusion perception anomaly loss, the prediction anomaly loss, and the consistency constraint loss.
[0115] The total joint loss is an overall loss function that comprehensively considers the losses of each agent and their consistency. Joint weighting refers to assigning a weight coefficient to each of the text understanding anomaly loss, data perception anomaly loss, fusion perception anomaly loss, prediction anomaly loss, and consistency constraint loss, and then summing the results after multiplying each loss by its corresponding weight coefficient to obtain the total joint loss. These weight coefficients can be, for example, pre-set fixed values or parameters that can be learned during training.
[0116] For example, the weight of the text understanding anomaly loss can be set as w1, the weight of the data perception anomaly loss as w2, the weight of the fusion perception anomaly loss as w3, the weight of the prediction anomaly loss as w4, and the weight of the consistency constraint loss as w5. The total joint loss is calculated as w1 × text understanding anomaly loss + w2 × data perception anomaly loss + w3 × fusion perception anomaly loss + w4 × prediction anomaly loss + w5 × consistency constraint loss.
[0117] S140. Based on the total joint loss, update the internal parameters of the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly assessment agent, and the collaborative scheduling agent in the initial financial data anomaly identification model to obtain the trained financial data anomaly identification model.
[0118] This step is the core of model training, using the total joint loss to guide the update of all agent's internal parameters. The backward update process uses the backpropagation algorithm, starting from the total joint loss and calculating the gradient of each parameter backward along the model's computational path.
[0119] For text understanding agents, the parameters that need to be updated include those for fine-tuning the large language model in the financial domain, the parameters of the single-layer linear transformation of text, and the parameters of the non-linear activation function of text. For data perception agents, these include the parameters for fine-tuning the large language model in the financial domain, the parameters of the single-layer linear transformation of data, and the parameters of the non-linear activation function of data. For feature reasoning agents, these include the parameters of the transformer encoder structure, the learnable attention weight parameters, the parameters of the fused single-layer linear transformation, and the parameters of the fused non-linear activation function. For anomaly assessment agents, the main focus is on the parameters of the learnable dynamic weight generation network. For cooperative scheduling agents, this includes the relevant parameters involved in calculating the consistency constraint loss (such as weight coefficients, if these parameters are learnable).
[0120] The optimizer (such as Adam, SGD, etc.) updates all these parameters based on the calculated gradient, and iterates continuously until the model converges or reaches the preset number of training rounds, finally obtaining the trained financial data anomaly identification model.
[0121] S150. Input the financial data to be identified into the trained financial data anomaly identification model, and output the comprehensive anomaly identification result of the financial data to be identified.
[0122] Financial data to be identified refers to new financial data that needs to be identified for anomaly detection. It can have the same data type and structure as the training data, including financial text data and financial numerical data.
[0123] After the financial data to be identified is input into the trained financial data anomaly identification model, the various agents within the model process the same procedure as during the training phase: the text understanding agent performs semantic anomaly mapping on the financial text data, generating text anomaly embedding representations and text prediction anomaly confidence scores; the data perception agent performs numerical anomaly mapping on the financial numerical data, generating data anomaly embedding representations and data prediction anomaly confidence scores; the feature reasoning agent performs time-dependent fusion processing on the text and data anomaly embedding representations, generating deep anomaly fusion feature representations and fusion prediction anomaly confidence scores; and the anomaly assessment agent performs dynamic weighted fusion of the three confidence scores to generate a comprehensive prediction anomaly score.
[0124] The comprehensive anomaly identification result can be, for example, a comprehensive predicted anomaly score, or a classification result based on the comprehensive predicted anomaly score (e.g., when the score is greater than a certain threshold, it is judged as an anomaly, otherwise it is normal).
[0125] The method provided in this application constructs an intelligent agent architecture that integrates five independent yet collaborative functions: text understanding, data perception, feature reasoning, anomaly assessment, and collaborative scheduling. It utilizes a fine-tuned large language model from the financial field to process financial text and numerical data. Through feature concatenation, transformer encoders, and the mining of temporal dependencies, combined with dynamic weighted fusion and consistency constraint mechanisms, it achieves deep fusion and accurate assessment of multi-dimensional anomaly features. This effectively addresses the problems of one-sided identification of single-type data anomalies, insufficient fusion of multi-source data, and inconsistency in the output of various modules. Consequently, it improves the accuracy, comprehensiveness, and reliability of financial data anomaly identification, providing more efficient and accurate technical support for financial risk prevention and control.
[0126] S210. An adversarial perturbation agent is used to apply perturbation processing to the financial text training data and the financial numerical training data respectively to generate adversarial text training data and adversarial numerical training data.
[0127] Adversarial perturbation generators are agents specifically designed to generate adversarial examples. Their purpose is to improve the robustness of the model by adding carefully designed small perturbations to the original training data to generate adversarial examples that can mislead the model.
[0128] Perturbation refers to adding perturbations to the original features of financial text training data and financial numerical training data. For financial text training data, perturbations could be, for example, replacing, inserting, or deleting certain words in the text. These operations should ensure that the modified text does not change semantically significantly, but can still cause the model to make incorrect predictions. For financial numerical training data, perturbations could be, for example, making small increases or decreases in numerical values, ensuring that the changes are within a reasonable range, but sufficient to affect the model's judgment. Adversarial text training data refers to financial text training data with perturbations added, and adversarial numerical training data refers to financial numerical training data with perturbations added.
[0129] S220. Input the adversarial text training data and the adversarial numerical training data into the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly evaluation agent, and the collaborative scheduling agent in the initial financial data anomaly identification model to obtain the adversarial text prediction anomaly confidence level corresponding to the adversarial text training data and the adversarial data prediction anomaly confidence level corresponding to the adversarial numerical training data.
[0130] The generated adversarial text training data and adversarial numerical training data are input into the initial financial data anomaly detection model in the same way as the original training data. The text understanding agent processes the adversarial text training data to generate adversarial text anomaly embedding representations and adversarial text prediction anomaly confidence scores; the data perception agent processes the adversarial numerical training data to generate adversarial data anomaly embedding representations and adversarial data prediction anomaly confidence scores. This processing flow is consistent with the flow described in S111 and S112, the only difference being that the input data is changed to adversarial samples.
[0131] S230. The feature reasoning agent performs time-dependent fusion processing on the adversarial text anomaly embedding representation corresponding to the adversarial text prediction anomaly confidence and the adversarial data anomaly embedding representation corresponding to the adversarial data prediction anomaly confidence to generate adversarial deep anomaly fusion feature representation and adversarial fusion prediction anomaly confidence.
[0132] This step is similar to the processing flow of S113. The feature reasoning agent receives the adversarial text anomaly embedding representation from the text understanding agent and the adversarial data anomaly embedding representation from the data perception agent. It performs temporal-dependent fusion processing on them, such as feature concatenation, self-attention calculation, forward propagation, and attention-weighted aggregation, and finally generates the adversarial deep anomaly fusion feature representation and the adversarial fusion prediction anomaly confidence.
[0133] S240. The anomaly assessment agent performs dynamic weighted fusion of the anomaly confidence of the adversarial text prediction, the anomaly confidence of the adversarial data prediction, and the anomaly confidence of the adversarial fusion prediction to generate an adversarial comprehensive prediction anomaly score.
[0134] The anomaly assessment agent processes the anomaly confidence scores of adversarial text prediction, adversarial data prediction, and adversarial fusion prediction according to the dynamic weighted fusion method described in S114, and generates an adversarial comprehensive prediction anomaly score.
[0135] S250. Calculate the adversarial prediction anomaly loss based on the anomaly score of the adversarial comprehensive prediction, the anomaly label corresponding to the financial text training data and the financial numerical training data, and calculate the adversarial consistency constraint loss based on the adversarial text prediction anomaly confidence, the adversarial data prediction anomaly confidence, and the adversarial fusion prediction anomaly confidence.
[0136] The calculation method for adversarial prediction anomaly loss is the same as that for prediction anomaly loss in S1147, using the mean squared error loss function to calculate the squared difference between the adversarial comprehensive prediction anomaly score and the common anomaly label. The calculation method for adversarial consistency constraint loss is the same as that for consistency constraint loss in S120, calculating the degree of inconsistency among the adversarial text prediction anomaly confidence, adversarial data prediction anomaly confidence, and adversarial fusion prediction anomaly confidence.
[0137] S260. The adversarial prediction anomaly loss and the adversarial consistency constraint loss are added to the total joint loss, and the internal parameters of all agents in the initial financial data anomaly identification model and the internal parameters of the adversarial perturbation generating agent are updated in reverse according to the total joint loss after addition, so as to obtain the financial data anomaly identification model after adversarial training.
[0138] Based on the original total joint loss, corresponding weight coefficients are assigned to the adversarial prediction anomaly loss and the adversarial consistency constraint loss, and then these are added to the total joint loss to obtain a new total joint loss. The new total joint loss = original total joint loss + w6 × adversarial prediction anomaly loss + w7 × adversarial consistency constraint loss, where w6 and w7 are the weight coefficients of the adversarial prediction anomaly loss and the adversarial consistency constraint loss, respectively. Then, based on the new total joint loss, the internal parameters of all agents (text understanding agent, data perception agent, feature reasoning agent, anomaly assessment agent, and cooperative scheduling agent) in the initial financial data anomaly detection model are updated using the backpropagation algorithm. Simultaneously, the internal parameters of the adversarial perturbation generation agent are also updated, ensuring that the model not only performs well on the original data but also has strong robustness against adversarial examples, ultimately resulting in an adversarially trained financial data anomaly detection model.
[0139] The method provided in this application adds an adversarial perturbation generation agent to apply small and reasonable perturbations to financial text and numerical training data to generate adversarial samples. The adversarial samples are then input into the initial financial data anomaly identification model to complete the loss calculation in the adversarial scenario. The adversarial prediction anomaly loss and adversarial consistency constraint loss are then integrated into the original total joint loss, and the internal parameters of each agent in the model and the adversarial perturbation generation agent are updated synchronously. This adversarial training enhances the model's ability to resist small perturbations, making up for the limitation of the original training which only targets normal samples. This improves the robustness and generalization ability of the financial data anomaly identification model, ensuring that the model can still stably and accurately identify financial data anomalies when facing malicious perturbations or data biases, and further strengthening the reliability of financial risk prevention and control.
[0140] S310. Perform multi-scale time window segmentation processing on the deep anomaly time series feature representation generated by the feature reasoning agent through multi-time scale decomposition agent to generate a set of short-term anomaly time series feature segments and a set of long-term anomaly time series feature segments.
[0141] The role of multi-timescale decomposition agents is to analyze the temporal feature representations of deep anomalies at different time scales to capture anomaly patterns of varying durations. Multi-scale temporal window segmentation refers to using time windows of different sizes to perform sliding segmentation on the temporal feature representations of deep anomalies.
[0142] For example, a smaller time window (such as containing 3 time steps) can be set as a short-term window to capture short-term abnormal fluctuations; a larger time window (such as containing 10 time steps) can be set as a long-term window to capture long-term abnormal trends.
[0143] By sliding these windows, the deep anomaly temporal feature representation can be segmented into multiple short-term and long-term anomaly temporal feature segments. The set of short-term anomaly temporal feature segments consists of the feature segments obtained from all short-term window segmentations, and the set of long-term anomaly temporal feature segments consists of the feature segments obtained from all long-term window segmentations.
[0144] S320. The multi-timescale decomposition agent performs local fluctuation pattern extraction processing on the set of short-term abnormal time series feature segments to generate a short-term abnormal pattern feature vector.
[0145] Local fluctuation pattern extraction is a feature extraction operation performed on a set of short-term anomalous time-series feature segments, aiming to capture the local fluctuation characteristics within each short-term segment. For example, a convolutional neural network (CNN) can be used to process each short-term anomalous time-series feature segment, extracting local features within the segment, such as fluctuation patterns like peaks, valleys, and rates of change, through convolutional kernels.
[0146] The features of each processed short segment are summarized (e.g., through pooling operations) to obtain the local fluctuation pattern features corresponding to each segment. Then, the local fluctuation pattern features of all segments are combined to generate a short-term anomaly pattern feature vector.
[0147] S330. The multi-timescale decomposition agent extracts the trend evolution pattern of the long-term abnormal time series feature fragment set to generate a long-term abnormal trend feature vector.
[0148] Trend evolution pattern extraction is a feature extraction operation performed on a set of long-term anomalous time series feature segments, aiming to capture the long-term trend evolution characteristics within each long-term segment. For example, recurrent neural networks (RNNs) or long short-term memory networks (LSTMs) can be used to process each long-term anomalous time series feature segment. These networks can effectively capture long-term dependencies and trend changes in time series data. By performing sequence modeling on long-term segments, trend evolution pattern features such as upward trends, downward trends, and periodic changes are extracted. Then, the trend evolution pattern features of all long-term segments are combined to generate a long-term anomalous trend feature vector.
[0149] S340. Input the short-term anomaly pattern feature vector and the long-term anomaly trend feature vector into the feature reasoning agent, and perform weighted aggregation processing on the short-term anomaly pattern feature vector and the long-term anomaly trend feature vector through the feature reasoning agent to generate a multi-scale enhanced deep anomaly fusion feature representation.
[0150] After inputting the short-term anomaly pattern feature vector and the long-term anomaly trend feature vector into the feature inference agent, the feature inference agent first assigns learnable weight coefficients to these two vectors, and then multiplies them by their respective weight coefficients to obtain the weighted short-term anomaly pattern feature vector and the weighted long-term anomaly trend feature vector.
[0151] Weighted aggregation can be achieved by concatenating the two weighted vectors, connecting them along the feature dimension to form a new high-dimensional feature vector, i.e., a multi-scale enhanced deep anomaly fusion feature representation. This feature representation integrates anomaly pattern information from different time scales, making it richer and more comprehensive than the original deep anomaly fusion feature representation.
[0152] S350. Generate a multi-scale fusion prediction anomaly confidence score based on the multi-scale enhanced deep anomaly fusion feature representation, and dynamically weight and fuse the multi-scale fusion prediction anomaly confidence score, the text prediction anomaly confidence score, and the data prediction anomaly confidence score through the anomaly evaluation agent to generate a multi-scale comprehensive prediction anomaly score.
[0153] The process of generating multi-scale fusion predicted anomaly confidence based on the multi-scale enhanced deep anomaly fusion feature representation is similar to the process in S1136 and S1137. It involves dimensional mapping through fusing single-layer linear transformation parameters and interval compression through fusing nonlinear activation functions. Then, the anomaly assessment agent dynamically weights and fuses the multi-scale fusion predicted anomaly confidence with the text predicted anomaly confidence and the data predicted anomaly confidence. At this point, the dynamic weight generation network needs to calculate four weight coefficients (text weight coefficient, data weight coefficient, fusion weight coefficient, and multi-scale fusion weight coefficient) and ensure that their sum is a fixed constant. The weighted fusion process is similar to S1143 to S1146, ultimately generating a multi-scale comprehensive predicted anomaly score.
[0154] S360. Calculate the multi-scale prediction anomaly loss based on the anomaly labeling corresponding to the multi-scale comprehensive prediction anomaly score, the financial text training data, and the financial numerical training data, and add the multi-scale prediction anomaly loss to the total joint loss.
[0155] The calculation method for multi-scale anomaly prediction loss is the same as that for prediction anomaly loss, using the mean squared error loss function to calculate the squared difference between the multi-scale comprehensive prediction anomaly score and the common anomaly label. Then, a weighting coefficient w8 is assigned to the multi-scale prediction anomaly loss and added to the total joint loss. The new total joint loss = original total joint loss + w8 × multi-scale prediction anomaly loss.
[0156] S370. Based on the total joint loss after addition, update the internal parameters of all agents in the initial financial data anomaly identification model and the internal parameters of the multi-timescale decomposed agents in reverse to obtain the financial data anomaly identification model trained on multiple timescales.
[0157] Based on the total joint loss after incorporating multi-scale prediction anomaly loss, the internal parameters of all agents in the initial financial data anomaly identification model are updated through the backpropagation algorithm. At the same time, the internal parameters of the multi-timescale decomposition agents are also updated (such as the size of the multi-scale time window, the network parameters involved in the extraction of local fluctuation patterns and trend evolution patterns).
[0158] In this way, the model can learn anomaly patterns at different time scales, thereby improving its ability to identify anomalies in complex financial data and obtaining a financial data anomaly identification model trained at multiple time scales.
[0159] The method provided in this application adds a multi-timescale decomposition agent to segment the deep anomaly time-series feature representation generated by the feature reasoning agent into multi-scale time windows. It extracts short-term anomaly fluctuation patterns and long-term anomaly trend evolution patterns, obtaining short-term anomaly pattern feature vectors and long-term anomaly trend feature vectors. These are then weighted and aggregated by the feature reasoning agent to generate a multi-scale enhanced deep anomaly fusion feature representation. An anomaly assessment agent then performs dynamic weighted fusion of multi-dimensional confidence and calculates the multi-scale anomaly prediction loss, integrating it into the total joint loss to synchronously update the parameters of each agent. This addresses the problem of incomplete anomaly pattern capture and difficulty in considering both short-term fluctuations and long-term trends under a single timescale. This enriches the dimensions of anomaly feature representation, improves the model's accuracy in identifying financial anomalies at different timescales, enhances the model's adaptability to complex and variable financial data anomalies, and further improves the comprehensiveness and reliability of financial data anomaly identification.
[0160] S410. Extract the text anomaly embedding representation output by the text understanding agent as a first feature distribution through the mutual information constraint agent, extract the data anomaly embedding representation output by the data perception agent as a second feature distribution, and extract the deep anomaly fusion feature representation output by the feature reasoning agent as a third feature distribution.
[0161] The role of mutual information-constrained agents is to constrain the relationship between different feature distributions through mutual information, thereby improving the discriminativeness and independence of features.
[0162] Feature distribution refers to the probability distribution of feature vectors in a high-dimensional space. The first feature distribution is the probability distribution followed by the text anomaly embedding representation, reflecting the statistical properties of text anomaly features; the second feature distribution is the probability distribution followed by the data anomaly embedding representation, reflecting the statistical properties of numerical anomaly features; and the third feature distribution is the probability distribution followed by the deep anomaly fusion feature representation, reflecting the statistical properties of fused anomaly features. Mutual information-constrained agents estimate these feature distributions using specific methods (e.g., density estimation methods based on neural networks).
[0163] S420. Calculate the first mutual information estimate between the first feature distribution and the second feature distribution, the second mutual information estimate between the first feature distribution and the third feature distribution, and the third mutual information estimate between the second feature distribution and the third feature distribution through the mutual information constrained agent.
[0164] Mutual information is an indicator that measures the degree of interdependence between two random variables. The larger the mutual information value, the stronger the dependency between the two variables. The first mutual information estimate measures the degree of interdependence between the first and second characteristic distributions; the second mutual information estimate measures the degree of interdependence between the first and third characteristic distributions; and the third mutual information estimate measures the degree of interdependence between the second and third characteristic distributions.
[0165] One method for calculating mutual information estimates is to use a mutual information neural network estimator (MINE), which estimates the lower bound of mutual information between two distributions by training a neural network.
[0166] S430. The mutual information constraint agent performs a summation operation on the first mutual information estimate, the second mutual information estimate, and the third mutual information estimate to generate a total mutual information metric.
[0167] The total mutual information measure is the sum of the first mutual information estimate, the second mutual information estimate, and the third mutual information estimate, which comprehensively reflects the overall degree of interdependence among the three characteristic distributions.
[0168] S440. Generate a mutual information minimization constraint loss based on the total mutual information metric, wherein the mutual information minimization constraint loss is positively correlated with the total mutual information metric.
[0169] The purpose of mutual information minimization constraint loss is to reduce redundancy among different feature distributions and improve feature independence and discriminative ability by minimizing the total mutual information metric. Since mutual information minimization constraint loss is positively correlated with the total mutual information metric, the mutual information minimization constraint loss also decreases when the total mutual information metric decreases.
[0170] For example, the mutual information minimization constraint loss can be directly set as the total mutual information metric, or a linear function of the total mutual information metric (e.g., mutual information minimization constraint loss = α × total mutual information metric, where α is a positive coefficient).
[0171] S450. The mutual information minimization constraint loss is added to the total joint loss, and the internal parameters of the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly evaluation agent, and the cooperative scheduling agent in the initial financial data anomaly identification model are updated in reverse according to the total joint loss after addition. The internal parameters of the mutual information constraint agent are updated in reverse through the mutual information minimization constraint loss to obtain the financial data anomaly identification model after feature decoupling training.
[0172] A weight coefficient w9 is assigned to the mutual information minimization constraint loss and added to the total joint loss. The new total joint loss = original total joint loss + w9 × mutual information minimization constraint loss. Then, based on the total joint loss after adding the mutual information minimization constraint loss, the internal parameters of each agent in the initial financial data anomaly identification model are updated through the backpropagation algorithm. At the same time, the internal parameters of the mutual information constraint agent (such as the neural network parameters used to estimate mutual information) are updated in reverse through the mutual information minimization constraint loss.
[0173] In this way, the model can learn more independent and discriminative features, reduce redundant information between features, thereby improving the performance of anomaly detection and obtaining a financial data anomaly detection model with decoupled feature training.
[0174] The method provided in this application embodiment adds a mutual information constraint agent to extract three types of feature distributions corresponding to text anomaly embedding representation, data anomaly embedding representation, and deep anomaly fusion feature representation. It calculates the mutual information estimates between each type of feature distribution and sums them to obtain the total mutual information metric. Based on this metric, it generates a mutual information minimization constraint loss that is positively correlated with the total mutual information. After incorporating this loss into the total joint loss, it synchronously updates the internal parameters of each agent in the model and the mutual information constraint agent. By minimizing mutual information, it reduces the redundancy between different feature distributions, improves the independence and discriminative ability of features, solves the problems of redundant information interference and insufficient feature discrimination in the process of multi-source feature fusion, thereby optimizing the feature representation quality, improving the accuracy and stability of financial data anomaly identification, and further enhancing the model's adaptability to complex financial anomaly scenarios.
[0175] Figure 2 This is a schematic diagram of the structure of a financial data anomaly identification system 100 based on multi-agent collaboration, provided in an embodiment of this application. Figure 2 As shown, the processor 120 can be used in the financial data anomaly identification system 100 based on multi-agent collaboration, and is used to perform the functions in this invention.
[0176] The financial data anomaly identification system 100 based on multi-agent collaboration can be a general-purpose server or a special-purpose server; both can be used to implement the financial data anomaly identification method based on multi-agent collaboration of the present invention. Although only one server is shown in this invention, for convenience, the functions described in this invention can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0177] For example, a multi-agent collaborative financial data anomaly detection system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the multi-agent collaborative financial data anomaly detection system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present invention can be implemented according to these program instructions. The multi-agent collaborative financial data anomaly detection system 100 also includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0178] For ease of explanation, only one processor is described in the multi-agent collaborative financial data anomaly detection system 100. However, it should be noted that the multi-agent collaborative financial data anomaly detection system 100 of this invention may also include multiple processors. Therefore, the steps executed by one processor described in this invention may also be executed jointly by multiple processors or individually. For example, if the processor of the multi-agent collaborative financial data anomaly detection system 100 executes steps A and B, it should be understood that steps A and B may also be executed jointly by two different processors or individually by one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0179] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0180] Finally, it should be noted that the above-disclosed embodiments are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying financial data anomalies based on multi-agent collaboration, characterized in that, include: By using the text understanding agent, data perception agent, feature reasoning agent, and anomaly assessment agent in the initial financial data anomaly identification model, data anomaly embedding processing and anomaly loss calculation processing are performed on the financial text training data and financial numerical training data, respectively, to obtain text understanding anomaly loss, data perception anomaly loss, fusion perception anomaly loss, and prediction anomaly loss. The consistency constraint loss among the text prediction anomaly confidence level output by the text understanding agent, the data prediction anomaly confidence level output by the data perception agent, and the fusion prediction anomaly confidence level output by the feature reasoning agent is calculated using the collaborative scheduling agent in the initial financial data anomaly identification model. The total joint loss is obtained by jointly weighting the text understanding anomaly loss, the data perception anomaly loss, the fusion perception anomaly loss, the prediction anomaly loss, and the consistency constraint loss. The internal parameters of the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly assessment agent, and the collaborative scheduling agent in the initial financial data anomaly identification model are updated in reverse based on the total joint loss to obtain the trained financial data anomaly identification model. The financial data to be identified is input into the trained financial data anomaly identification model, and the comprehensive anomaly identification result of the financial data to be identified is output.
2. The financial data anomaly identification method based on multi-agent collaboration according to claim 1, characterized in that, The initial financial data anomaly identification model utilizes a text understanding agent, a data perception agent, a feature reasoning agent, and an anomaly assessment agent to perform data anomaly embedding processing and anomaly loss calculation processing on financial text training data and financial numerical training data, respectively, to obtain text understanding anomaly loss, data perception anomaly loss, fusion perception anomaly loss, and prediction anomaly loss, including: The text understanding agent performs semantic anomaly mapping on the financial text training data to generate text anomaly embedding representations and text prediction anomaly confidence scores, and obtains text understanding anomaly loss based on the text prediction anomaly confidence scores and the anomaly labels corresponding to the financial text training data. The data-aware intelligent agent performs numerical anomaly mapping processing on the financial numerical training data to generate data anomaly embedding representation and data prediction anomaly confidence, and obtains data-aware anomaly loss based on the data prediction anomaly confidence and the anomaly labeling corresponding to the financial numerical training data. The feature reasoning agent performs time-dependent fusion processing on the text anomaly embedding representation and the data anomaly embedding representation to generate a deep anomaly fusion feature representation and a fusion prediction anomaly confidence score. The fusion-aware anomaly loss is obtained based on the anomaly labeling that corresponds to both the fusion prediction anomaly confidence score and the financial text training data and the financial numerical training data. The anomaly assessment agent dynamically weights and fuses the text prediction anomaly confidence, the data prediction anomaly confidence, and the fusion prediction anomaly confidence to generate a comprehensive prediction anomaly score. The prediction anomaly loss is then obtained based on the comprehensive prediction anomaly score and the anomaly labels corresponding to both the financial text training data and the financial numerical training data.
3. The financial data anomaly identification method based on multi-agent collaboration according to claim 2, characterized in that, The step of generating text anomaly embeddings and text prediction anomaly confidence scores by performing semantic anomaly mapping on the financial text training data using the text understanding agent, and obtaining text understanding anomaly loss based on the text prediction anomaly confidence scores and the anomaly labels corresponding to the financial text training data, includes: The financial text training data is processed by rule filtering using a fine-tuned large language model in the financial field and built-in rule prompts. Semantic abnormal text information is extracted from the financial text training data. The semantic abnormal text information is a set of text fragments containing abnormal pointers. The semantically anomalous text information is encoded and converted using the encoding layer of the fine-tuned large language model in the financial field, and the semantically anomalous text information is converted into a text anomalous embedding representation, which is a set of feature vectors in a high-dimensional vector space. The anomaly embedding representation of the text is mapped by a single-layer linear transformation parameter to generate an intermediate transformation vector of the text. The intermediate transformation vector of the text is then compressed by a non-linear activation function of the text to generate a text prediction anomaly confidence score. The text prediction anomaly confidence score is a continuous value within a preset interval. The text understanding anomaly loss is obtained by using the text prediction anomaly confidence score and the anomaly label corresponding to the financial text training data, and the text difference metric between the text prediction anomaly confidence score and the anomaly label is calculated by using the binary cross-entropy loss function. Based on the text understanding anomaly loss, the parameters of the fine-tuned large language model in the financial domain, the parameters of the single-layer linear transformation of the text, and the parameters of the nonlinear activation function of the text are updated through the backpropagation algorithm.
4. The financial data anomaly identification method based on multi-agent collaboration according to claim 2, characterized in that, The step of generating a data anomaly embedding representation and a data prediction anomaly confidence score by performing numerical anomaly mapping processing on the financial numerical training data through the data-aware intelligent agent, and obtaining the data-aware anomaly loss based on the data prediction anomaly confidence score and the anomaly labeling corresponding to the financial numerical training data, includes: The missing data locations are identified by using a fine-tuned large language model in the financial field to perform missing data identification processing on the financial numerical training data. Generate filler values based on the data distribution characteristics around the location of the missing data; The missing data positions are filled with the specified values to generate a complete set of numerical data. The complete numerical data set is encoded and transformed using the encoding layer of the fine-tuned large language model in the financial field to obtain a data anomaly embedding representation, which is a set of vector representations under the numerical feature dimension. The data anomaly embedding representation is subjected to dimensional mapping processing by the data single-layer linear transformation parameters to generate an intermediate data transformation vector. The intermediate transformation vector of the data is compressed by a nonlinear activation function to generate an anomaly confidence score for data prediction. The data-perceived anomaly loss is obtained by calculating the data difference metric between the data prediction anomaly confidence and the anomaly labeling corresponding to the financial numerical training data using the binary cross-entropy loss function. Based on the data-aware anomaly loss, the parameters of the financial domain fine-tuning large language model, the parameters of the data single-layer linear transformation, and the parameters of the data nonlinear activation function are updated through the backpropagation algorithm.
5. The financial data anomaly identification method based on multi-agent collaboration according to claim 2, characterized in that, The process involves using the feature reasoning agent to perform time-dependent fusion processing on the text anomaly embedding representation and the data anomaly embedding representation to generate a deep anomaly fusion feature representation and a fusion prediction anomaly confidence score. A fusion-aware anomaly loss is then obtained based on the fusion prediction anomaly confidence score and the anomaly labels corresponding to both the financial text training data and the financial numerical training data. This includes: The text anomaly embedding representation and the data anomaly embedding representation are concatenated to generate a concatenated anomaly feature sequence. The self-attention calculation process is performed on the spliced abnormal feature sequence using the encoder structure of the converter to obtain the self-attention output features; The encoder structure of the converter is used to perform forward propagation processing on the self-attention output features, and deep anomaly temporal feature representation is extracted from the spliced anomaly feature sequence. The deep anomaly temporal feature representation contains anomaly evolution mode information under different time steps. The attention energy value corresponding to each time step in the deep anomaly temporal feature representation is calculated by using learnable attention weight parameters, and the attention energy value is exponentially normalized to generate the attention weight coefficient corresponding to each time step. The deep anomaly temporal feature representation of each time step is multiplied by the attention weight coefficient of the corresponding time step to generate a weighted time step feature vector. The weighted time step feature vectors of all the time steps are summed and aggregated to generate a deep anomaly fusion feature representation. The deep anomaly fusion feature representation is then subjected to dimensionality mapping by fusing single-layer linear transformation parameters to generate a fusion intermediate transformation vector. The fused intermediate transformation vector is subjected to interval compression by fusing a nonlinear activation function to generate a fused prediction anomaly confidence score. Based on the fusion prediction anomaly confidence and the anomaly labels that are jointly corresponding to the financial text training data and the financial numerical training data, the fusion difference metric between the fusion prediction anomaly confidence and the anomaly labels that are jointly corresponding to the financial text training data and the financial numerical training data is calculated using the binary classification cross-entropy loss function to obtain the fusion perception anomaly loss. Based on the fusion-aware anomaly loss, the parameters of the encoder structure of the transformer, the learnable attention weight parameters, the fusion single-layer linear transformation parameters, and the parameters of the fusion nonlinear activation function are updated using a backpropagation algorithm.
6. The financial data anomaly identification method based on multi-agent collaboration according to claim 2, characterized in that, The process involves dynamically weighting and fusing the text prediction anomaly confidence score, the data prediction anomaly confidence score, and the fused prediction anomaly confidence score using the anomaly assessment agent to generate a comprehensive prediction anomaly score. The prediction anomaly loss is then obtained based on the comprehensive prediction anomaly score and the anomaly labels corresponding to both the financial text training data and the financial numerical training data. This includes: The text weight coefficients corresponding to the text prediction anomaly confidence are calculated using a learnable dynamic weight generation network, and the data weight coefficients corresponding to the data prediction anomaly confidence are calculated using the learnable dynamic weight generation network. The fusion weight coefficient corresponding to the fusion prediction anomaly confidence is calculated by the learnable dynamic weight generation network, wherein the sum of the text weight coefficient, the data weight coefficient, and the fusion weight coefficient is a fixed constant. Multiply the text prediction anomaly confidence score by the text weight coefficient to generate a text weighted score; The data prediction anomaly confidence level is multiplied by the data weight coefficient to generate a data weighted score; The fusion prediction anomaly confidence score is multiplied by the fusion weight coefficient to generate a fusion weighted score. The text-weighted score, the data-weighted score, and the fusion-weighted score are summed to generate a comprehensive anomaly prediction score. Based on the anomaly labels that correspond to the comprehensive predicted anomaly score and the financial text training data and the financial numerical training data, the squared difference between the comprehensive predicted anomaly score and the anomaly labels that correspond to the financial text training data and the financial numerical training data is calculated using the mean squared error loss function to obtain the predicted anomaly loss. The parameters of the learnable dynamic weight generation network are updated using the backpropagation algorithm based on the predicted abnormal loss.
7. The financial data anomaly identification method based on multi-agent collaboration according to claim 1, characterized in that, Also includes: Adversarial perturbation-generated intelligent agents apply perturbation processing to the financial text training data and the financial numerical training data respectively to generate adversarial text training data and adversarial numerical training data. The adversarial text training data and the adversarial numerical training data are input into the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly evaluation agent, and the collaborative scheduling agent in the initial financial data anomaly identification model to obtain the adversarial text prediction anomaly confidence level corresponding to the adversarial text training data and the adversarial data prediction anomaly confidence level corresponding to the adversarial numerical training data. The feature reasoning agent performs time-dependent fusion processing on the adversarial text anomaly embedding representation corresponding to the adversarial text prediction anomaly confidence and the adversarial data anomaly embedding representation corresponding to the adversarial data prediction anomaly confidence, generating adversarial deep anomaly fusion feature representation and adversarial fusion prediction anomaly confidence. The anomaly assessment agent dynamically weights and fuses the anomaly confidence scores of the adversarial text prediction, the adversarial data prediction, and the adversarial fusion prediction to generate an adversarial comprehensive prediction anomaly score. The adversarial prediction anomaly loss is calculated based on the anomaly score of the adversarial comprehensive prediction, the anomaly label corresponding to the financial text training data and the financial numerical training data, and the adversarial consistency constraint loss is calculated based on the adversarial text prediction anomaly confidence, the adversarial data prediction anomaly confidence, and the adversarial fusion prediction anomaly confidence. The adversarial prediction anomaly loss and the adversarial consistency constraint loss are added to the total joint loss, and the internal parameters of all agents in the initial financial data anomaly identification model and the internal parameters of the adversarial perturbation generating agent are updated in reverse according to the total joint loss after addition, so as to obtain the financial data anomaly identification model after adversarial training.
8. The financial data anomaly identification method based on multi-agent collaboration according to claim 1, characterized in that, Also includes: By performing multi-scale time window segmentation on the deep anomaly temporal feature representation generated by the feature reasoning agent through multi-time-scale decomposition agent, a set of short-time anomaly temporal feature fragments and a set of long-time anomaly temporal feature fragments are generated. The multi-timescale decomposition agent performs local fluctuation pattern extraction processing on the set of short-term anomaly time series feature segments to generate short-term anomaly pattern feature vectors. The multi-timescale decomposition agent extracts the trend evolution pattern of the long-term abnormal time series feature fragment set to generate a long-term abnormal trend feature vector. The short-term anomaly pattern feature vector and the long-term anomaly trend feature vector are input into the feature reasoning agent, and the feature reasoning agent performs weighted aggregation processing on the short-term anomaly pattern feature vector and the long-term anomaly trend feature vector to generate a multi-scale enhanced deep anomaly fusion feature representation. Based on the multi-scale enhanced deep anomaly fusion feature representation, a multi-scale fusion prediction anomaly confidence score is generated. The anomaly evaluation agent then performs dynamic weighted fusion of the multi-scale fusion prediction anomaly confidence score, the text prediction anomaly confidence score, and the data prediction anomaly confidence score to generate a multi-scale comprehensive prediction anomaly score. The multi-scale prediction anomaly loss is calculated based on the anomaly labeling corresponding to the multi-scale comprehensive prediction anomaly score, the financial text training data, and the financial numerical training data, and the multi-scale prediction anomaly loss is added to the total joint loss. The internal parameters of all agents in the initial financial data anomaly identification model and the internal parameters of the multi-timescale decomposed agents are updated in reverse based on the total joint loss after joining, to obtain the financial data anomaly identification model trained on multiple timescales.
9. The financial data anomaly identification method based on multi-agent collaboration according to claim 1, characterized in that, Also includes: The text anomaly embedding representation output by the text understanding agent is extracted as the first feature distribution through the mutual information constrained agent, the data anomaly embedding representation output by the data perception agent is extracted as the second feature distribution, and the deep anomaly fusion feature representation output by the feature reasoning agent is extracted as the third feature distribution. The mutual information constrained agent calculates the first mutual information estimate between the first feature distribution and the second feature distribution, the second mutual information estimate between the first feature distribution and the third feature distribution, and the third mutual information estimate between the second feature distribution and the third feature distribution; The mutual information constrained agent performs a summation operation on the first mutual information estimate, the second mutual information estimate, and the third mutual information estimate to generate a total mutual information metric. A mutual information minimization constraint loss is generated based on the total mutual information metric, and the mutual information minimization constraint loss is positively correlated with the total mutual information metric. The mutual information minimization constraint loss is added to the total joint loss, and the internal parameters of the text understanding agent, the data perception agent, the feature reasoning agent, the anomaly assessment agent, and the cooperative scheduling agent in the initial financial data anomaly identification model are updated in reverse according to the total joint loss after the addition. The internal parameters of the mutual information constraint agent are updated in reverse through the mutual information minimization constraint loss, so as to obtain the financial data anomaly identification model with feature decoupling training completed.
10. A financial data anomaly identification system based on multi-agent collaboration, characterized in that, The invention includes a processor and a computer-readable storage medium storing machine-executable instructions, which, when executed by a computer, implement the financial data anomaly identification method based on multi-agent collaboration as described in any one of claims 1-9.