Financial data anomaly detection method and system based on time sequence feature decoupling

By decoupling temporal features based on deep separable convolution and attention mechanism, combined with dual-stream residual compensation mechanism and online threshold update, the problem of distinguishing between fluctuations and abnormal signals in financial data anomaly detection is solved, and efficient and accurate anomaly detection is achieved, adapting to changes in data distribution.

CN120634758APending Publication Date: 2025-09-12CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511102210.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing financial data anomaly detection methods have difficulty effectively distinguishing normal cyclical fluctuations from true abnormal signals, and their accuracy decreases when the data distribution changes. They also have high computational complexity and poor model interpretability.

Method used

The proposed method adopts a feature decoupling module based on depthwise separable convolution, a feature separation module based on attention mechanism, an anomaly detection module based on dual-stream residual compensation mechanism, and an online threshold update module. Through temporal feature decoupling and feature separation, the detection threshold is dynamically adjusted, and the residual compensation matrix is ​​constructed in combination with the K-Nearest Neighbors regression algorithm.

Benefits of technology

The accuracy and reliability of anomaly detection have been improved, the false alarm rate has been reduced by about 40%, the detection rate has been increased by about 35%, the computing efficiency has been increased by 8-9 times, and it can adapt to the dynamic changes in the distribution of financial data. The detection accuracy has not decreased by more than 5%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634758A_ABST
    Figure CN120634758A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of financial data analysis, in particular to a financial data anomaly detection method and system based on time sequence feature decoupling, comprising a feature decoupling module, a feature separation module, an anomaly detection module and an on-line threshold updating module, decoupling and separation of three types of time sequence features are realized through depth separable convolution and an attention mechanism, and the detection accuracy is improved. According to the method, feature distribution and a double-flow residual compensation mechanism are constructed and used for fitting probability distribution of prediction value abnormity and judging whether financial data are abnormal or not, an online threshold updating module dynamically updates an offline training threshold of an abnormity detection model, and online updating of the model is achieved. According to the invention, normal seasonal fluctuation and real abnormal signals can be effectively distinguished, and the false alarm rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial data analysis, and in particular to a financial data anomaly detection method and system based on time series feature decoupling. Background Art

[0002] In corporate financial management, timely detection of anomalies in financial data is crucial for preventing financial risks and ensuring stable business operations. Traditional methods for detecting anomalies in financial data are primarily based on static thresholds or simple statistical models, which have significant shortcomings when dealing with complex and volatile financial time series data.

[0003] Most existing financial anomaly detection methods treat time series data as a whole, making it difficult to effectively distinguish normal cyclical fluctuations from true abnormal signals. For example, seasonal sales fluctuations may be mistakenly identified as anomalies, while small fraud activities hidden within normal fluctuations may go unnoticed. Furthermore, traditional methods are poorly adaptable to changes in data distribution. Once the business environment changes, the accuracy of fixed-threshold detection models will significantly decrease.

[0004] With the development of deep learning technology, some neural network-based anomaly detection methods have been introduced into the financial field. However, these methods have high computational complexity, poor model interpretability, and limitations when dealing with the unique properties of financial time series data. In particular, existing methods struggle to perform refined identification and analysis when dealing with the simultaneous presence of cyclical fluctuations, trend changes, and sudden events in financial data.

[0005] Therefore, there is an urgent need for a detection method that can decouple different features of financial time series data and efficiently identify abnormal patterns. Summary of the Invention

[0006] The purpose of the present invention is to provide a financial data anomaly detection method and system based on time series feature decoupling. Through deep separable convolution technology and attention mechanism, three types of time series features in financial data, namely periodic fluctuations, trend changes and burst noise, are separated, and a dual-stream residual compensation mechanism is used for dynamic anomaly detection, thereby overcoming the deficiency of the existing technology in distinguishing normal fluctuations from abnormal signals.

[0007] The present invention proposes a financial data anomaly detection method based on time series feature decoupling, including:

[0008] Collect the company's financial data and build a sample set of the company's financial data;

[0009] Build a financial data anomaly detection model that decouples time series features. The model includes:

[0010] A feature decoupling module based on depthwise separable convolution is used to construct a time series dataset based on the decoupled data of three types of time series features: periodic fluctuations, trend changes, and burst noise;

[0011] A feature separation module based on an attention mechanism, wherein the feature separation module is used to separate the three types of time series features in the time series dataset and construct a feature distribution;

[0012] An anomaly detection module based on a dual-stream residual compensation mechanism, the anomaly detection module being used to fit the probability distribution of predicted value anomalies through the feature distribution and determine whether the data in the company's financial data sample set is abnormal;

[0013] An online threshold updating module is used to dynamically update the offline training threshold of the anomaly detection model based on the error between the predicted value and the true value, and to update the anomaly detection model online.

[0014] As a preferred method, the company's financial data is collected to construct a sample set of the company's financial data, specifically including:

[0015] Data collection: Collect daily data for at least 5 years, including the company's historical financial data and basic financial data;

[0016] Data reconstruction: Statistics are compiled on a monthly basis to build a basic financial data sequence; historical financial data of the company is calculated on a monthly basis, including accounts receivable turnover, accounts payable turnover, and inventory turnover, to build a historical financial data sequence.

[0017] As a preferred approach, a financial data anomaly detection model with decoupled time series features is constructed, specifically including:

[0018] Feature decoupling module: The input dataset of the feature decoupling module is defined as ,in Represents the basic financial data sequence of each month, and m represents the number of basic financial data sequences in the input dataset; the feature decoupling module includes convolutional layers, maximum pooling, batch normalization, and fully connected layers, and uses depthwise separable convolution to extract the two-dimensional feature matrix of the basic financial data sequence of each month. , is the month, and then merged into a multi-scale feature matrix At the same time, we use three types of time series feature extraction modules, namely periodic fluctuation, trend change and burst noise, to reduce the dimension of F, reduce the number of channels of Fi to 1 / 3, and obtain a new data set , Represents the time series characteristics of each month, ∈[1,m];

[0019] Feature separation module based on attention mechanism: The input dataset of the feature separation module is defined as , each Contains the time series features of each month extracted by the feature decoupling module; The data is input into a feature separation module consisting of two residual blocks and one fully connected layer. Depthwise separable convolution is used to extract features and generate an N×N feature matrix V. An attention mechanism is then used to extract features related to periodic fluctuations, trend changes, and sudden noise from the feature matrix V corresponding to different months. Using the feature matrix V, a feature separation dataset corresponding to periodic fluctuations, trend changes, and sudden noise is obtained.

[0020] Anomaly detection module based on dual-stream residual compensation mechanism: using the dataset generated by the feature separation module ; Each data set The input is sent to the feature separation module consisting of 2 residual blocks and 1 fully connected layer. Features are extracted through depthwise separable convolution, and then maximum pooling is performed to generate a new feature matrix. ; Use multi-layer perceptron to transform the feature matrix Projected into the feature space related to the anomaly, the anomaly detection matrix is ​​obtained At the same time, using the data set Construct a data feature matrix related to anomalies ,Will As training data, As training labels, they are input into the anomaly detection module for feature selection, and finally the residual compensation matrix is ​​generated. ; According to the error between the predicted value and the true value, the dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to generate the final anomaly detection threshold.

[0021] Preferably, a depth-wise separable convolution is used to generate a multi-scale feature matrix F, specifically including:

[0022] The input dataset D is batch normalized and the two-dimensional feature matrix of each month is extracted through a depth-wise separable convolution of size 3×3×n×f. ;

[0023] Use pooling to reduce the feature matrix The number of channels is 1 / 3, and then the feature matrix of each month is aggregated into a multi-scale feature matrix F in the time dimension;

[0024] Among them, n represents the feature dimension of each month, and f represents the number of features of each month.

[0025] As a preference, the attention mechanism is used to extract features related to periodic fluctuations, trend changes, and burst noise, specifically including:

[0026] The input features calculated by the attention mechanism are the feature matrix V, and the output features are the feature separation dataset;

[0027] in, is the feature matrix No. Rank The value of the column, is the attention matrix Middle Rank The value of the column, is the weight vector after average pooling;

[0028] The feature separation dataset is generated by the feature matrix V and attention matrix E in the feature separation module.

[0029] As a preference, a dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model based on the error between the predicted value and the true value, thereby generating the final anomaly detection threshold, specifically including:

[0030] Assume the feature matrix Each point in It represents the jth feature of the kth category anomaly detection for the financial data of the i-th week. The anomaly detection module outputs the anomaly detection matrix , k∈[1,3], l∈[1,m], anomaly detection matrix The element representation The probability of being abnormal;

[0031] According to the prediction error of the true value, the anomaly detection weight is adjusted to generate the final anomaly detection model and obtain the final anomaly detection result. ;

[0032] in, represents the j-th feature of the k-th category anomaly detection, Represents the financial data for week i The corresponding true value, is the coefficient of the loss function of the jth feature of the kth category anomaly detection, and is the financial data of the i-th week The probability of being predicted as abnormal, is the residual compensation factor of the jth feature of the kth type of anomaly detection.

[0033] As a preferred method, the K-Nearest Neighbors regression algorithm is used to implement anomaly detection, and the nearest k features are calculated to construct the residual compensation matrix. , specifically including:

[0034] Calculate the distance d between the sample detected for each class of anomaly and other samples in the class, and select the k samples with the smallest distance as the residual compensation matrix ;

[0035] Among them, KNNregk,l( ) is the K-Nearest Neighbors regression algorithm, used to calculate The distance between the kth class and other samples in the anomaly detection.

[0036] As an option, the model training step is also included:

[0037] After the financial data sample set is input into the model, an offline training threshold of an anomaly detection model based on three characteristics: periodic fluctuation, trend change and sudden noise is adopted;

[0038] According to the error between the predicted value and the true value, a dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to complete the training of the financial data anomaly detection model.

[0039] As a preferred method, basic financial data of each month is collected on a monthly basis to construct a basic financial data sequence, specifically including:

[0040] Extract serialized financial indicator time series data through a sliding window with a time window of 5;

[0041] Use depth-wise separable temporal convolution to extract dynamic features of temporal data, and perform feature decoupling through dynamic channel attention;

[0042] The feature matrix is ​​input into the long short-term memory network to obtain the predicted value of accounts receivable turnover rate.

[0043] The financial data anomaly detection system based on time series feature decoupling includes:

[0044] Data collection module, used to collect the company's financial data and build a sample set of the company's financial data;

[0045] The model building module is used to build a financial data anomaly detection model that decouples time series features. The model includes:

[0046] A feature decoupling module based on depthwise separable convolution is used to construct a time series dataset based on the decoupled data of three types of time series features: periodic fluctuations, trend changes, and burst noise;

[0047] A feature separation module based on an attention mechanism, wherein the feature separation module is used to separate the three types of time series features in the time series dataset and construct a feature distribution;

[0048] An anomaly detection module based on a dual-stream residual compensation mechanism, the anomaly detection module being used to fit the probability distribution of predicted value anomalies through the feature distribution and determine whether the data in the company's financial data sample set is abnormal;

[0049] An online threshold updating module, configured to dynamically update an offline training threshold of the anomaly detection model based on an error between a predicted value and a true value, and to update the anomaly detection model online;

[0050] A model training module is used to tune the model in the model building module and train the model based on the company's financial data sample set.

[0051] The present invention has the following beneficial effects:

[0052] 1. By decoupling time series features, we achieve precise separation of cyclical fluctuations, trend changes, and sudden noise in financial data, significantly improving the accuracy and reliability of anomaly detection. Compared with traditional methods, this method can effectively distinguish normal seasonal fluctuations from true anomalies, reducing the false alarm rate by approximately 40% while increasing the anomaly detection rate by approximately 35%.

[0053] 2. Using depthwise separable convolution technology instead of standard convolution significantly reduces computational complexity while maintaining feature extraction effectiveness, increasing computational efficiency by approximately 8-9 times and enabling the system to process larger amounts of financial data.

[0054] 3. Combining the attention mechanism and two-stream residual compensation technology, the system can adaptively adjust the anomaly detection threshold, enabling online model updates and adapting detection results to dynamic changes in the distribution of financial data. Experiments show that under changing operating environments, the detection accuracy of this system decreases by no more than 5%, while the accuracy of traditional fixed threshold methods can drop by over 30%.

[0055] 4. The residual compensation matrix is ​​constructed through the K-Nearest Neighbors regression algorithm, which improves the model's ability to detect different types of anomalies. It performs particularly well in detecting small and persistent anomaly patterns, with detection accuracy increased by approximately 25%. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a flowchart of a method for detecting anomalies in financial data based on time series feature decoupling according to an embodiment of the present invention;

[0057] Figure 2 This is a structural diagram of a timing feature decoupling module according to an embodiment of the present invention;

[0058] Figure 3This is a schematic diagram of the structure of a feature separation module based on an attention mechanism according to an embodiment of the present invention;

[0059] Figure 4 This is a structural diagram of an anomaly detection module based on a dual-stream residual compensation mechanism according to an embodiment of the present invention;

[0060] Figure 5 A schematic structural diagram of an online threshold updating module according to an embodiment of the present invention;

[0061] Figure 6 This is a structural block diagram of a financial data anomaly detection system based on time series feature decoupling according to an embodiment of the present invention. DETAILED DESCRIPTION

[0062] Please refer to the attached Figure 1-6 The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0063] like Figure 1 As shown, the financial data anomaly detection method based on time series feature decoupling provided by the present invention includes the following steps:

[0064] Step S1: Collect the company's financial data and construct a sample set of the company's financial data.

[0065] Step S2: Construct a financial data anomaly detection model with time series feature decoupling, which includes: a feature decoupling module based on depthwise separable convolution, a feature separation module based on attention mechanism, an anomaly detection module based on dual-stream residual compensation mechanism, and an online threshold update module.

[0066] In step S1, collecting the company's financial data and constructing the company's financial data sample set specifically includes: collecting daily data for at least 5 years, including the company's historical financial data and the company's basic financial data; then reconstructing the data, counting the basic financial data of each month on a monthly basis, and constructing a basic financial data sequence. At the same time, calculating the company's historical financial data of each month on a monthly basis, including accounts receivable turnover rate, accounts payable turnover rate and inventory turnover rate, and constructing a historical financial data sequence.

[0067] The present invention preferably collects at least five years of historical data. This is because financial data often exhibits significant seasonal and cyclical characteristics, and data spanning a longer timeframe can more fully capture these characteristics. For example, retail sales typically see a significant increase in the fourth quarter. If only one or two years of data are collected, the model may misinterpret seasonal fluctuations as anomalies. Practice has shown that five years of data can cover the operating cycles of most companies, including annual, quarterly, and monthly fluctuations, providing ample learning samples for the model.

[0068] Specifically, the basic financial data collected by the present invention includes, but is not limited to, basic financial indicators such as operating income, operating costs, gross profit margin, net profit, debt-to-asset ratio, and current ratio. Historical financial data primarily focuses on indicators related to capital flow efficiency, such as accounts receivable turnover, accounts payable turnover, and inventory turnover. These indicators are highly indicative of financial anomalies within a company. For example, a sudden drop in accounts receivable turnover could indicate a change in sales policy, a decrease in customer payment capacity, or even signal fictitious sales.

[0069] In step S2, the process of constructing a financial data anomaly detection model with time series feature decoupling will be described in detail below.

[0070] like Figure 2 As shown, the specific implementation of the feature decoupling module 10 in the present invention is as follows:

[0071] The input dataset of feature decoupling module 10 is defined as ,in Represents the basic financial data sequence of each month, and m represents the number of basic financial data sequences in the input dataset. The feature decoupling module 10 includes a convolution layer 11, a maximum pooling layer 12, a batch normalization layer 13, and a fully connected layer 14. It uses depthwise separable convolution to extract the two-dimensional feature matrix of the basic financial data sequence of each month. , i∈[1,m], i is the month, and then merged into a multi-scale feature matrix .

[0072] Where: D is the input data set, which is an m-dimensional vector; is the basic financial data series of the i-th month, including the values ​​of all financial indicators of that month; m is the total number of financial data series, which usually corresponds to the number of months analyzed; To pass depthwise separable convolution from The extracted two-dimensional feature matrix; F is the multi-scale feature matrix formed by merging the feature matrices of all months.

[0073] At the same time, the feature decoupling module 10 uses the periodic fluctuation extraction submodule 15, the trend change extraction submodule 16 and the burst noise extraction submodule 17 to reduce the dimension of F respectively. The number of channels is reduced to 1 / 3, and a new data set is obtained. , Represents the time series features of each month, i∈[1,m].

[0074] in: is the new data set after processing; The time series features of the i-th month after feature decoupling include three types of decoupled features: periodic fluctuations, trend changes, and burst noise. Each type of feature occupies 1 / 3 of the total number of channels to ensure that the three types of features are balanced in the feature space.

[0075] The specific process of using depth-wise separable convolution to generate the multi-scale feature matrix F is as follows:

[0076] First, the input dataset D is batch normalized, and the two-dimensional feature matrix of each month is extracted by depth-wise separable convolution of size 3×3×n×f. Here, n represents the feature dimension for each month, and f represents the number of features for each month. In practical applications, n is usually set between 10 and 30, corresponding to the number of key financial indicators that a company focuses on; f is usually set between 64 and 256 to capture feature patterns at different scales.

[0077] in: is the intermediate feature matrix; n is the dimension of the monthly financial features, such as 20 dimensions representing 20 financial indicators; f is the number of features, corresponding to the number of feature maps generated by the convolution operation; 3×3 represents the spatial size of the convolution kernel, which is used to capture local feature relationships.

[0078] Take a manufacturing enterprise as an example. The input data contains 20 financial indicators per month (n=20), and the feature number f=128. For the financial data of the enterprise in January 2023 , the depthwise separable convolution first performs depthwise convolution, applies 3×3 convolution kernel to each index independently, and then performs 1×1 pointwise convolution for channel fusion, and finally generates the feature matrix This operation can effectively capture the relationship between different financial indicators, such as the correlation pattern between operating income and profit margin.

[0079] Next, use the pooling operation to reduce the feature matrix The number of channels is reduced to 1 / 3, and the feature matrix of each month is aggregated into a multi-scale feature matrix F in the time dimension. This design enables the model to focus on short-term, medium-term and long-term financial anomaly patterns at the same time.

[0080] Compared with standard convolution, depthwise separable convolution separates spatial convolution and channel convolution, which greatly reduces the computational complexity. The computational complexity of standard convolution is , and the computational complexity of depth-wise separable convolution is reduced to ,in is the convolution kernel size, M and N are the number of input and output channels respectively.

[0081] Where: O represents the big O symbol of algorithm complexity; is the side length of the convolution kernel, which is 3 in this case; M is the number of input channels, corresponding to the number of financial indicators; N is the number of output channels, corresponding to the number of generated features.

[0082] In financial data processing, this optimization enables the model to more efficiently process large amounts of financial indicator data. For example, for a large corporate group with hundreds of financial indicators, traditional convolution may take several hours to process, while depthwise separable convolution can reduce the processing time to just over ten minutes while maintaining feature extraction quality.

[0083] like Figure 3 As shown, the specific implementation of the feature separation module 20 based on the attention mechanism in the present invention is as follows:

[0084] The input dataset of feature separation module 20 is defined as , each Contains the temporal features of each month extracted by the feature decoupling module 10. The feature separation module 20 includes a first residual block 21, a second residual block 22 and a fully connected layer 23. The input is fed into the module, and features are extracted using depthwise separable convolution to generate an N×N feature matrix V, where N is usually set to 64-128, which is sufficient to represent the complex features of financial data.

[0085] in: It is the output of the feature decoupling module and the input of the feature separation module; is the decoupled time series feature of month i; V is an N×N feature matrix, where N represents the feature dimension, typically between 64 and 128. Each element in matrix V represents the strength of association between specific features.

[0086] The feature separation module 20 uses the attention mechanism to extract features related to periodic fluctuations, trend changes, and sudden noise from the feature matrix V corresponding to different months. In specific implementations, the input features calculated by the attention mechanism are the feature matrix V, and the output features are the feature separation dataset. The calculation formula is as follows:

[0087] ,

[0088] ,

[0089] in: is the value of the i-th row and j-th column of the feature matrix V, indicating the degree of association between features; is the value of the i-th row and j-th column in the attention matrix E, which represents the attention weight; is the weight vector after average pooling, which is used to adjust the importance of different features; W is a learnable weight matrix; b is a learnable bias vector; tanh is the hyperbolic tangent activation function, which maps values ​​to the interval [-1,1]; exp is the exponential function; ∑ represents the summation operation, which sums all items of j from 1 to N.

[0090] The feature separation data set is generated by the feature matrix V and the attention matrix E in the feature separation module 20. The specific generation process is as follows:

[0091] ,

[0092] in: is the output feature separation dataset, k∈{1,2,3} corresponds to three types of features: periodic fluctuation, trend change and sudden noise; ∑ represents the summation operation; is the attention weight; is the i-th row of the feature matrix, representing the i-th eigenvector.

[0093] The core idea of ​​the attention mechanism is to calculate the importance weight of each feature in the feature matrix, enabling the model to automatically focus on the most relevant features. This mechanism is particularly important in financial anomaly detection, as anomalies of different periods and types may manifest in different financial indicators.

[0094] Take a commercial retail enterprise as an example. Its seasonal sales fluctuations are obvious. Through the attention mechanism, the model will automatically give higher weights to indicators such as sales revenue and seasonal inventory changes when analyzing periodic fluctuations ( When analyzing sudden noise, the model pays more attention to indicators such as irregular expenses and changes in bad debt ratios. This adaptive feature weighting significantly improves the model's ability to identify different types of financial anomalies.

[0095] Preferably, the attention weight calculation adopts the scaled dot product attention mechanism, and the formula is as follows:

[0096] ,

[0097] Where: Q is the query matrix (Query), which is obtained by linear transformation of the input features; K is the key matrix (Key), which is obtained by linear transformation of the input features; V is the value matrix (Value), which is obtained by linear transformation of the input features; The dimension of the key vector is used to scale the dot product result to prevent the gradient from disappearing; represents the matrix multiplication of the transpose of Q and K, which calculates the similarity between the query and the key; softmax is the normalization function, which converts the similarity into a probability distribution; √ represents the square root operation.

[0098] In practical applications, the attention mechanism is divided into three branches, focusing on three types of time series features: cyclical fluctuations, trend changes, and sudden noise. For example, for a retail company with significant seasonality, the cyclical fluctuation branch may give higher weight to seasonal changes in accounts receivable turnover; while for a rapidly expanding technology company, the trend change branch may be more sensitive to changes in net profit growth rate.

[0099] like Figure 4 As shown, the specific implementation of the anomaly detection module 30 based on the dual-stream residual compensation mechanism in the present invention is as follows:

[0100] The anomaly detection module 30 uses the data set generated by the feature separation module 20 , k∈[1,3],l∈[1,m] as input. The anomaly detection module 30 includes a first residual block 31, a second residual block 32, a fully connected layer 33 and a multi-layer perceptron 34.

[0101] in: is the dataset output by the feature separation module; k is the feature type index, k=1 represents periodic fluctuation features, k=2 represents trend change features, and k=3 represents sudden noise features; is the month index, ∈[1,m]; m is the total number of months.

[0102] First, each dataset The input is sent to the feature separation module consisting of 2 residual blocks and 1 fully connected layer. Features are extracted through depthwise separable convolution, and then maximum pooling is performed to generate a new feature matrix. .

[0103] in: is the feature matrix generated after the residual block and pooling, which contains the k-th feature in the Each residual block adopts the "convolution-batch normalization-activation function-convolution-batch normalization" structure and adds skip connections, which effectively avoids the gradient vanishing problem in deep networks.

[0104] Taking a manufacturing company as an example, when analyzing its financial data for the first quarter of 2023, the data with periodic fluctuation characteristics (k=1) is processed through the residual block. 、 、 Then, we get the feature matrix 、 、 ,These matrices capture the seasonal production rhythm, raw material procurement cycle and other cyclical characteristics of the enterprise.

[0105] Secondly, the feature matrix is ​​transformed into Projected into the feature space related to the anomaly, the anomaly detection matrix is ​​obtained At the same time, the dataset Construct a data feature matrix related to anomalies ,Will As training data, As training labels, they are input into the anomaly detection module 30 for feature selection, and finally the residual compensation matrix is ​​generated. .

[0106] in: is the anomaly detection matrix, which represents the probability distribution of each point being an anomaly; is the data feature matrix, containing the original feature information; is the residual compensation matrix, which is used to adjust the anomaly detection results.

[0107] The structure of the multilayer perceptron 34 consists of two hidden layers. The number of neurons in the first layer is 256, and the ReLU activation function is used. The number of neurons in the second layer is 128, and the Tanh activation function is used. The number of neurons in the output layer is the same as the feature dimension, and the Sigmoid activation function is used to map the output to the interval [0, 1] to represent the abnormality probability.

[0108] According to the error between the predicted value and the true value, the dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to generate the final anomaly detection threshold.

[0109] Specifically, let the feature matrix Each point in It represents the jth feature of the kth category anomaly detection for the financial data of the i-th week. The anomaly detection module 30 outputs the anomaly detection matrix , k∈[1,3], ∈[1,m], anomaly detection matrix The element representation The probability of being abnormal.

[0110] in: is the feature matrix The element in row i and column j in represents the jth financial feature in week i; is the anomaly detection matrix The corresponding element in represents is the probability of anomaly, and its value range is [0,1].

[0111] According to the prediction error of the true value, the anomaly detection weight is adjusted to generate the final anomaly detection model and obtain the final anomaly detection result. :

[0112] ,

[0113] in: is the final anomaly detection result matrix; is the initial anomaly detection matrix; It is an adjustment factor, and its value range is [0,1]. The larger the value, the stronger the adjustment.

[0114] Defined as:

[0115] ,

[0116] in: Indicates the sum of all i and j values; represents the value of the jth feature of the kth category anomaly detection in week i; Represents the financial data for week i The corresponding true value; Indicates the absolute error between the predicted value and the true value; is the loss function coefficient of the jth feature of the kth category anomaly detection, which controls the importance of different features; is the financial data for week i The probability of being predicted as abnormal, i.e. The value of It is the residual compensation factor of the jth feature of the kth type of anomaly detection, which is used to balance the influence of different features.

[0117] In practical applications, and Set according to the importance and volatility of different financial indicators. For example, for profit-related indicators, It is usually set to a higher value (such as 0.8-0.9) because such indicators are more important for anomaly detection; while for auxiliary financial indicators, it can be set to a lower value (such as 0.3-0.5). It is used to balance the detection sensitivity of different types of anomalies. For sudden anomalies that require high sensitivity detection, It can be set lower (such as 0.2-0.4), and for periodic fluctuations, It can be set higher (such as 0.6-0.8).

[0118] Take a trading company as an example. When analyzing its accounts receivable turnover rate, if historical data shows that the normal fluctuation range is 4.5-5.5, and the actual value in a certain month is 3.8, which is obviously too low, the model calculates =4.8 (expected value), =3.8 (actual value), =1.0. If the indicator =0.8 (high importance), =0.4 (high sensitivity), then The calculated result is smaller, making the final Maintaining a high value correctly marks the point as an anomaly. This dynamic adjustment mechanism significantly improves the model's adaptability to different types of financial anomalies.

[0119] The specific implementation method of using the K-Nearest Neighbors regression algorithm to implement anomaly detection in the present invention is as follows:

[0120] Calculate the distance d between the sample detected for each class of anomaly and other samples in the class, and select the k samples with the smallest distance as the residual compensation matrix L_k,l:

[0121] ,

[0122] in: is the residual compensation matrix; Indicates the selection of k samples with the smallest distance; dist() represents the distance calculation function; Indicates the current sample point; 、 、 Etc. represent other sample points in the sample set.

[0123] KNNreg_k,l(x_ij) is the K-Nearest Neighbors regression algorithm used to calculate The core idea of ​​this algorithm is to find the k samples that are most similar to the current sample in the feature space and predict the label of the current sample based on the labels of these samples.

[0124] In practical applications, the choice of parameter k significantly impacts model performance. Experimental results indicate that k is typically set to the square root of the total number of samples. For financial data from small and medium-sized enterprises (e.g., 60 monthly data points), k can be set to 7-8. For large enterprises or high-frequency financial data, k can be increased to 10-15.

[0125] For example, when analyzing anomalies in the R&D expense ratio of a technology company, the KNN algorithm searches the feature space for the eight most similar historical data points to the current month. If the current month's R&D expense ratio is 18%, while the eight most similar historical points average 12%, with fluctuations ranging from 10% to 14%, the current value is likely to be flagged as an anomaly. This similarity-based judgment method is particularly well-suited for handling nonlinear relationships and complex dependencies in financial data.

[0126] The advantage of the K-Nearest Neighbors regression algorithm is that it does not require strong assumptions about data distribution and can adapt to the nonlinear characteristics of financial data. In this paper, the algorithm is used to construct a residual compensation matrix, effectively improving the accuracy of anomaly detection. It is particularly effective in detecting small but persistent anomaly patterns (such as small, persistent fraud).

[0127] The distance calculation method preferably adopts Mahalanobis distance, and the formula is as follows:

[0128] ,

[0129] in: Represents the Mahalanobis distance between samples x and y; x and y are sample points, represented as multidimensional vectors; (xy) represents the difference vector of two samples; (xy)^T represents the transpose of the difference vector; represents the inverse of the sample covariance matrix; √ represents the square root operation.

[0130] The Mahalanobis distance accounts for the correlation between features and is more suitable than the Euclidean distance for highly correlated, multidimensional data such as financial data. For example, for retail companies, sales and marketing expenses are often highly correlated. Using the Mahalanobis distance can more accurately capture the impact of this correlation and avoid distance calculation bias caused by feature correlation.

[0131] The specific implementation of the model training step in the present invention is as follows:

[0132] After inputting the financial data sample set into the model, the model uses offline training thresholds for the anomaly detection model based on three characteristics: cyclical fluctuations, trend changes, and sudden noise. Based on the error between the predicted and true values, a dual-stream residual compensation mechanism dynamically adjusts the anomaly detection weights of the prediction model to complete the training of the financial data anomaly detection model.

[0133] The present invention preferably uses cross-validation to train and evaluate the model. Specifically, five years of collected financial data are divided into training and test sets in an 8:2 ratio. 20% of the training set is then allocated as a validation set. The training process uses an early stopping strategy: if performance metrics (such as the F1 score) on the validation set do not improve for five consecutive epochs, training is terminated to avoid overfitting.

[0134] The training set is used for model parameter optimization; the validation set is used for hyperparameter adjustment and early stopping; and the test set is used for final performance evaluation. An epoch represents a complete pass through the training set during training. The F1 score is the harmonic mean of precision and recall, calculated as F1 = 2 × (precision × recall) / (precision + recall).

[0135] For example, a manufacturing company's five-year monthly financial data consists of 60 sample points. After a 8:2 split, the training set contains 48 sample points and the test set contains 12 sample points. A further 20% of the training set, or approximately 10 sample points, is allocated as the validation set. The model is learned on the training set and evaluated on the validation set. The best model is selected based on its performance on the validation set, and a final evaluation is performed on the test set.

[0136] Training used the Adam optimizer, with an initial learning rate of 0.001, which was decayed by 0.1 every 30 epochs. To enhance model robustness, a dropout layer (with a dropout rate of 0.3) and L2 regularization (with a regularization coefficient of 0.0001) were added during training.

[0137] Among them: Adam is an adaptive moment estimation optimization algorithm that combines the advantages of AdaGrad and RMSProp; the learning rate is the step size for updating model parameters, and the initial value of 0.001 is suitable for most financial data modeling scenarios; dropout is a regularization technique that randomly discards neurons, and a rate of 0.3 means that 30% of neurons are randomly discarded during each training to prevent overfitting; L2 regularization controls model complexity by adding a weight square sum term to the loss function, and the coefficient 0.0001 represents the regularization strength.

[0138] Model evaluation uses precision, recall, and F1 score as primary metrics. In practical applications for different types of enterprises, the weighting of these metrics can be adjusted based on business needs. For example, risk-sensitive enterprises such as financial institutions may prioritize recall to ensure they capture all possible anomalies (preferring false positives over missed negatives). Meanwhile, enterprises with high levels of normal fluctuation, such as tech startups, may prioritize precision to minimize false positives.

[0139] The calculation formulas for precision, recall and F1 score are as follows:

[0140] Accuracy ,

[0141] Recall ,

[0142] F1 score Accuracy Recall ,

[0143] Among them: TP (True Positive) represents the number of abnormal samples correctly identified; FP (False Positive) represents the number of normal samples incorrectly identified (false positives); FN (False Negative) represents the number of abnormal samples incorrectly identified (false negatives).

[0144] In a real-world application case at a trading company, the model identified 30 anomalies, 25 of which were true anomalies (TP=25) and 5 were false positives (FP=5). Furthermore, 3 true anomalies were not identified (FN=3). Therefore, the precision was 25 / (25+5)=0.833, the recall was 25 / (25+3)=0.893, and the F1 score was 2×0.833×0.893 / (0.833+0.893)=0.862, demonstrating that the model had high accuracy and completeness in detecting financial anomalies at this company.

[0145] The specific implementation method of using time windows to construct financial data series in the present invention is as follows:

[0146] Basic financial data is compiled monthly. During the construction of a basic financial data sequence, a sliding window with a time window of 5 is used to extract the serialized time series data of financial indicators. A depthwise separable time series convolution is used to extract dynamic features of the time series data, and dynamic channel attention is used for feature decoupling. The feature matrix is ​​input into a long short-term memory network to obtain a predicted value for the accounts receivable turnover rate.

[0147] Among them: the sliding window is a data processing technology, and a window size of 5 means that 5 consecutive months of data are considered each time; the long short-term memory network (LSTM) is a special recurrent neural network that can effectively handle long-term dependencies in time series data.

[0148] The sliding window size of 5 means that the model considers five consecutive months of financial data to predict the indicator value for the next month. This window size is designed based on a combination of the seasonality of financial data (usually quarterly or semi-annual) and the company's financial reporting cycle (usually quarterly). In practice, for companies with shorter business cycles (such as fast-moving consumer goods retailers), the window size can be reduced to 3-4; for companies with longer business cycles (such as heavy industry or infrastructure), the window size can be increased to 6-8.

[0149] For example, a retail company processes its financial data using a sliding window with a time window of 5. When predicting anomalies in June 2023, the model considers five consecutive months of data from January to May 2023. In practice, this window size has been found to capture seasonal fluctuations (such as those caused by Spring Festival and summer promotions) without introducing excessive noise, thus balancing the impact of short-term fluctuations with long-term trends.

[0150] During the time series data extraction process, the Z-score standardization method is used to normalize the data:

[0151] ,

[0152] Where: z is the standardized value; x is the original data value; μ is the sample mean, calculated as , where n is the sample size; σ is the sample standard deviation, calculated as .

[0153] This normalization process can mitigate the impact of different scales of different financial indicators and improve the stability of model training. For example, operating revenue may be expressed in billions of yuan, while profit margin is expressed as a percentage. After normalization, both become distributed with a mean of 0 and a standard deviation of 1, allowing the model to handle different indicators fairly.

[0154] The Long Short-Term Memory (LSTM) network, whose unit structure includes an input gate, a forget gate, and an output gate, can effectively capture long-term dependencies in financial data. The LSTM network in this paper is configured as a two-layer LSTM network with 128 neurons per layer. Its input is standardized financial indicators for the previous five months, and its output is the predicted value for the next month. This structure can effectively learn the temporal patterns of financial indicators, such as the seasonal trend of accounts receivable turnover.

[0155] like Figure 6 As shown, the financial data anomaly detection system based on time series feature decoupling provided by the present invention includes a data acquisition module 40, a model construction module 50 and a model training module 60.

[0156] The data collection module 40 is used to collect the company's financial data and construct a sample set of the company's financial data. The model construction module 50 is used to construct a financial data anomaly detection model that decouples time series features. The model training module 60 is used to tune the model in the model construction module 50 and train the model based on the company's financial data sample set.

[0157] The model construction module 50 includes a feature decoupling module 51, a feature separation module 52, an anomaly detection module 53, and an online threshold update module 54. The feature decoupling module 51 is used to construct a time series dataset based on the decoupled data of three types of time series features: periodic fluctuations, trend changes, and sudden noise. The feature separation module 52 is used to separate the three types of time series features in the time series dataset and construct a feature distribution. The anomaly detection module 53 is used to fit the probability distribution of predicted value anomalies through the feature distribution and determine whether the data in the company's financial data sample set is abnormal. The online threshold update module 54 is used to dynamically update the offline training threshold of the anomaly detection model based on the error between the predicted value and the true value, and to update the anomaly detection model online.

[0158] In actual deployment, this system can flexibly adapt to the needs of enterprises of varying sizes. For small and medium-sized enterprises, the system can be deployed on a single server with the recommended configuration of an Intel Xeon E5-2680 v4 or higher CPU, 32GB or higher RAM, and 1TB SSD storage. With this configuration, the system's processing power can support real-time analysis of approximately 100 financial indicators and five years of monthly data, with a response time of less than 3 seconds.

[0159] For large enterprises or financial institutions, a distributed deployment architecture is recommended, using Hadoop or Spark clusters to process large amounts of financial data and GPU-accelerated deep learning model training. For example, a large multinational conglomerate with hundreds of subsidiaries and thousands of financial indicators could complete anomaly detection analysis for the entire group in just 10 minutes using a 10-node Spark cluster, an 80% improvement over traditional methods.

[0160] The system interface supports multiple data input formats, including CSV, Excel, and JSON. It can also be integrated with mainstream ERP systems (such as SAP and Oracle) and financial software, enabling automated data collection and real-time anomaly monitoring. System output includes anomaly detection reports, risk scores, and a visual analysis interface, supporting anomaly analysis and tracing by various dimensions (such as time, department, and business line).

[0161] Preferably, in one embodiment of the present invention, the data acquisition module 40 is used to collect daily data for at least 5 years, including the company's historical financial data and the company's basic financial data; data reconstruction, statistics the basic financial data of each month on a monthly basis, and construct a basic financial data sequence; calculate the company's historical financial data of each month on a monthly basis, including accounts receivable turnover rate, accounts payable turnover rate and inventory turnover rate, and construct a historical financial data sequence.

[0162] During the data collection process, to ensure data quality, the system implemented a multi-level data cleaning and verification mechanism:

[0163] 1. Missing value processing: For missing data with a time interval of less than 3 days, linear interpolation is used to fill in the missing data; for missing data of longer time periods, historical data of the same period combined with trend adjustment are used to fill in the missing data.

[0164] 2. Outlier screening: Use the 3σ principle to initially screen outliers, and then conduct secondary verification based on industry characteristics and company historical data.

[0165] 3. Data consistency check: Ensure the internal consistency of collected data by verifying financial identities (such as assets = liabilities + owner's equity).

[0166] Among them: The linear interpolation formula is y= + (x- ) ( - ) / ( - ), where x is the time point to be interpolated, and y is the interpolation result, ( , )and( , ) are known adjacent data points; the 3σ principle means that data outside the range of μ±3σ are considered as potential anomalies, where μ is the mean and σ is the standard deviation.

[0167] For example, when processing five years of sales data for a manufacturing company, missing data were discovered for February 3 to 5, 2022. Since the missing period was less than three days, the system used linear interpolation to estimate the missing days based on the data for February 2 and 6. Furthermore, the 3σ principle revealed unusually high sales in December 2021, which was verified to be due to year-end promotions. Therefore, this data was retained and specifically marked in the model.

[0168] During the data reconstruction phase, the system uses different aggregation methods based on the characteristics of different financial indicators. For stock indicators (such as assets and liabilities), the end-of-month value is used; for flow indicators (such as revenue and cost), the monthly cumulative value is used; for ratio indicators (such as turnover rate), an appropriate calculation formula is used. For example, the formula for calculating accounts receivable turnover rate is:

[0169] Accounts receivable turnover ,

[0170] Among them: monthly operating income is the total operating income of the month; the accounts receivable balance at the beginning of the month is the accounts receivable balance on the first day of the month; the accounts receivable balance at the end of the month is the accounts receivable balance on the last day of the month; 2 is the adjustment factor to make the calculation result reflect the complete turnover situation.

[0171] This refined data processing ensures the accuracy of subsequent model analysis. In practice, businesses in different industries can adjust the calculation methods for specific indicators based on their business characteristics. For example, businesses with significant seasonal sales characteristics can use a seasonally adjusted turnover rate calculation method to more accurately reflect actual operational efficiency.

[0172] Preferably, in one embodiment of the present invention, after the model training module 60 inputs the financial data sample set into the model, it adopts the offline training threshold of the anomaly detection model based on the three characteristics of periodic fluctuations, trend changes and burst noise; according to the error between the predicted value and the true value, the dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to complete the training of the financial data anomaly detection model.

[0173] Model training adopts a two-stage strategy: the first stage is offline training, which uses historical data to establish a baseline model; the second stage is online fine-tuning, which continuously optimizes model parameters based on new data. This approach combines the advantages of batch learning and incremental learning, ensuring both basic model performance and adaptability to new data patterns.

[0174] Among them: offline training refers to one-time training using historical data before model deployment; online fine-tuning refers to the process of continuously updating model parameters based on new data after model deployment; batch learning refers to training the model at one time using all training data; incremental learning refers to the ability of the model to learn from new data without retraining the entire model.

[0175] Taking a commercial bank as an example, the system first used the bank's historical financial data from 2018 to 2022 for offline training to establish a baseline model. After deployment, the system automatically fine-tuned the model by receiving new financial data monthly. When the bank adjusted its business strategy in early 2023, resulting in structural changes in some financial indicators, the online fine-tuning mechanism enabled the model to quickly adapt to the new pattern, avoiding a large number of false positives and rapidly recovering its detection accuracy from 73% at the beginning of the adjustment to over 92%.

[0176] During offline training, K-fold cross-validation (K=5) was used to evaluate model performance stability. Grid search was used to determine the optimal combination of model hyperparameters. Key hyperparameters included the number of depthwise separable convolution layers (preferably 3-4), convolution kernel size (preferably 3×3 or 5×5), number of attention heads (preferably 4-8), and number of residual blocks (preferably 2-3).

[0177] K-fold cross-validation involves dividing the dataset into K equal subsets, using K-1 subsets to train the model each time, and using the remaining subset for validation. This is repeated K times to average the results. Grid search involves systematically trying all possible hyperparameter combinations and selecting the combination with the best performance.

[0178] Regarding the specific setting of hyperparameters, the present invention found through experiments that for medium-sized enterprises with 20-30 financial indicators, the configuration of 3 layers of depth-wise separable convolution, 3×3 convolution kernels, 6 attention heads, and 2 residual blocks can achieve a good balance between model complexity and computational efficiency; while for large enterprises with more than 50 financial indicators, the number of layers can be increased to 4 layers of convolution, 5×5 convolution kernels, 8 attention heads, and 3 residual blocks to improve the model's expressiveness.

[0179] During the online fine-tuning phase, a sliding window mechanism is used. After receiving new monthly data, the model is fine-tuned using the data from the last 12-24 months. A small learning rate (0.1 times the original learning rate) is used during fine-tuning to avoid overfitting to the latest data and ignoring historical patterns.

[0180] The sliding window mechanism refers to using only the most recent data (e.g., 12-24 months) to update the model. The learning rate refers to the step size of the model parameter update. A smaller learning rate (0.0001) makes the update more stable.

[0181] Model evaluation uses comprehensive indicators, including:

[0182] AUC (area under the curve): evaluates the overall discriminatory ability of the model. The AUC of an excellent model is usually greater than 0.85.

[0183] Precision and recall: Based on the business focus, high-risk areas usually pay more attention to recall.

[0184] Time lead: measures how far in advance the model can detect anomalies. The usual goal is to provide early warnings of at least 1-3 months in advance.

[0185] Wherein: AUC is the area under the ROC curve, and its value range is [0,1]. The closer it is to 1, the better the model performance. Time lead refers to the time difference between the model detecting an anomaly and the actual manifestation of the anomaly. The larger the lead, the earlier the warning.

[0186] In a real-world application case at a retail company, the model achieved an AUC of 0.91, an anomaly detection precision of 88%, and a recall of 85%. It was able to detect potential financial risks an average of 2.5 months in advance. This enabled company management to take early action and avoid multiple potential financial losses.

[0187] Through this systematic training and evaluation process, the reliability and effectiveness of the model in practical applications are ensured.

[0188] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the contents of the present description and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention.

Claims

1. A financial data anomaly detection method based on time series feature decoupling, characterized by: The following steps are involved: Collect the company's financial data and build a sample set of the company's financial data; Build a financial data anomaly detection model that decouples time series features. The model includes: A feature decoupling module based on depthwise separable convolution is used to construct a time series dataset based on the decoupled data of three types of time series features: periodic fluctuations, trend changes, and burst noise; A feature separation module based on an attention mechanism, wherein the feature separation module is used to separate the three types of time series features in the time series dataset and construct a feature distribution; An anomaly detection module based on a dual-stream residual compensation mechanism, the anomaly detection module being used to fit the probability distribution of predicted value anomalies through the feature distribution and determine whether the data in the company's financial data sample set is abnormal; An online threshold updating module is used to dynamically update the offline training threshold of the anomaly detection model based on the error between the predicted value and the true value, and to update the anomaly detection model online.

2. The financial data anomaly detection method based on time series feature decoupling according to claim 1 is characterized in that: Collect the company's financial data and build a sample set of the company's financial data, including: Data collection: Collect daily data for at least 5 years, including the company's historical financial data and basic financial data; Data reconstruction: Statistics are compiled on a monthly basis to build a basic financial data sequence; historical financial data of the company is calculated on a monthly basis, including accounts receivable turnover, accounts payable turnover, and inventory turnover, to build a historical financial data sequence.

3. The financial data anomaly detection method based on time series feature decoupling according to claim 2 is characterized in that: Construct a financial data anomaly detection model that decouples time series features, specifically including: Feature decoupling module: The input dataset of the feature decoupling module is defined as ,in Represents the basic financial data sequence of each month, and m represents the number of basic financial data sequences in the input dataset; the feature decoupling module includes convolutional layers, maximum pooling, batch normalization, and fully connected layers, and uses depthwise separable convolution to extract the two-dimensional feature matrix of the basic financial data sequence of each month. , is the month, and then merged into a multi-scale feature matrix At the same time, we use three types of time series feature extraction modules, namely periodic fluctuation, trend change and burst noise, to reduce the dimension of F, reduce the number of channels of Fi to 1 / 3, and obtain a new data set , Represents the time series characteristics of each month, ∈[1,m]; Feature separation module based on attention mechanism: The input dataset of the feature separation module is defined as , each Contains the time series features of each month extracted by the feature decoupling module; The data is input into a feature separation module consisting of two residual blocks and one fully connected layer. Depthwise separable convolution is used to extract features and generate an N×N feature matrix V. An attention mechanism is then used to extract features related to periodic fluctuations, trend changes, and sudden noise from the feature matrix V corresponding to different months. Using the feature matrix V, a feature separation dataset corresponding to periodic fluctuations, trend changes, and sudden noise is obtained. Anomaly detection module based on dual-stream residual compensation mechanism: using the dataset generated by the feature separation module ; Each data set The input is sent to the feature separation module consisting of 2 residual blocks and 1 fully connected layer. Features are extracted through depthwise separable convolution, and then maximum pooling is performed to generate a new feature matrix. ; Use multi-layer perceptron to transform the feature matrix Projected into the feature space related to the anomaly, the anomaly detection matrix is ​​obtained At the same time, using the data set Construct a data feature matrix related to anomalies ,Will As training data, As training labels, they are input into the anomaly detection module for feature selection, and finally the residual compensation matrix is ​​generated. ; According to the error between the predicted value and the true value, the dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to generate the final anomaly detection threshold.

4. The financial data anomaly detection method based on time series feature decoupling according to claim 3 is characterized in that: Use depth-wise separable convolution to generate a multi-scale feature matrix F, specifically including: The input dataset D is batch normalized and the two-dimensional feature matrix of each month is extracted through a depth-wise separable convolution of size 3×3×n×f. ; Use pooling to reduce the feature matrix The number of channels is 1 / 3, and then the feature matrix of each month is aggregated into a multi-scale feature matrix F in the time dimension; Among them, n represents the feature dimension of each month, and f represents the number of features of each month.

5. The financial data anomaly detection method based on time series feature decoupling according to claim 4 is characterized in that: The attention mechanism is used to extract features related to periodic fluctuations, trend changes, and sudden noise, including: The input features calculated by the attention mechanism are the feature matrix V, and the output features are the feature separation dataset; in, is the feature matrix No. Rank The value of the column, is the attention matrix Middle Rank The value of the column, is the weight vector after average pooling; The feature separation dataset is generated by the feature matrix V and attention matrix E in the feature separation module.

6. The financial data anomaly detection method based on time series feature decoupling according to claim 5 is characterized in that: According to the error between the predicted value and the true value, the dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to generate the final anomaly detection threshold. include: Assume the feature matrix Each point in It represents the jth feature of the kth category anomaly detection for the financial data of the i-th week. The anomaly detection module outputs the anomaly detection matrix , k∈[1,3], l∈[1,m], anomaly detection matrix The element representation The probability of being abnormal; According to the prediction error of the true value, the anomaly detection weight is adjusted to generate the final anomaly detection model and obtain the final anomaly detection result. ; in, represents the j-th feature of the k-th category anomaly detection, Represents the financial data for week i The corresponding true value, is the coefficient of the loss function of the jth feature of the kth category anomaly detection, and is the financial data of the i-th week The probability of being predicted as abnormal, is the residual compensation factor of the jth feature of the kth type of anomaly detection.

7. The financial data anomaly detection method based on time series feature decoupling according to claim 6 is characterized in that: The K-Nearest Neighbors regression algorithm is used to implement anomaly detection, and the nearest k features are calculated to construct the residual compensation matrix. , specifically including: Calculate the distance d between the sample detected for each class of anomaly and other samples in the class, and select the k samples with the smallest distance as the residual compensation matrix ; Among them, KNNregk,l( ) is the K-Nearest Neighbors regression algorithm, used to calculate The distance between the kth class and other samples in the anomaly detection.

8. The financial data anomaly detection method based on time series feature decoupling according to claim 1 is characterized in that: Also includes the model training steps: After the financial data sample set is input into the model, an offline training threshold of an anomaly detection model based on three characteristics: periodic fluctuation, trend change and sudden noise is adopted; According to the error between the predicted value and the true value, a dual-stream residual compensation mechanism is used to dynamically adjust the anomaly detection weight of the prediction model to complete the training of the financial data anomaly detection model.

9. The financial data anomaly detection method based on time series feature decoupling according to claim 2 is characterized in that: The basic financial data of each month is collected on a monthly basis to construct a basic financial data series, including: Extract serialized financial indicator time series data through a sliding window with a time window of 5; Use depth-wise separable temporal convolution to extract dynamic features of temporal data, and perform feature decoupling through dynamic channel attention; The feature matrix is ​​input into the long short-term memory network to obtain the predicted value of accounts receivable turnover rate.

10. A financial data anomaly detection system based on time series feature decoupling is characterized by: include: Data collection module, used to collect the company's financial data and build a sample set of the company's financial data; The model building module is used to build a financial data anomaly detection model that decouples time series features. The model includes: A feature decoupling module based on depthwise separable convolution is used to construct a time series dataset based on the decoupled data of three types of time series features: periodic fluctuations, trend changes, and burst noise; A feature separation module based on an attention mechanism, wherein the feature separation module is used to separate the three types of time series features in the time series dataset and construct a feature distribution; An anomaly detection module based on a dual-stream residual compensation mechanism, the anomaly detection module being used to fit the probability distribution of predicted value anomalies through the feature distribution and determine whether the data in the company's financial data sample set is abnormal; An online threshold updating module, configured to dynamically update an offline training threshold of the anomaly detection model based on an error between a predicted value and a true value, and to update the anomaly detection model online; A model training module is used to tune the model in the model building module and train the model based on the company's financial data sample set.