Mass financial and accounting data fusion auditing feature extraction method, medium and system

By dynamically adjusting the time window boundary and multi-dimensional feature fusion, the problem of inability to adapt to internal changes when extracting the feature of massive accounting data is solved, and a higher precision and comprehensive feature extraction is achieved. It is suitable for scenarios such as enterprise financial management, accounting firm audit, financial institution risk control and government financial supervision.

CN120372255AActive Publication Date: 2025-07-25BEIHAI FORECASTING CENT OF STATE OCEANIC ADMINISTRATION ((QINGDAO MARINE FORECASTING STATION OF STATE OCEANIC ADMINISTRATION) (QINGDAO MARINE ENVIRONMENT MONITORING CENT OF STATE OCEANIC ADMINISTRATION))

Patent Information

Application Number
CN202510854950.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing technology cannot dynamically adapt to the internal changes of data when extracting feature features of massive accounting data, resulting in insufficient accuracy of audit feature extraction, affecting the accuracy and reliability of subsequent analysis.

Method used

The first time window is determined by using the time series clustering algorithm, the independent time window function is constructed to dynamically adjust the window boundary, and the time dimension features are extracted in combination with the accounting time series feature recognition model, and spatial dimension features are extracted based on geographical distribution and organizational structure information. The multi-objective optimization weight determination model is used to fusion to form a comprehensive audit feature vector.

Benefits of technology

It significantly improves the accuracy and adaptability of feature extraction, can accurately locate abnormal changes time nodes, enhances the ability to identify complex accounting data, and ensures the comprehensiveness and accuracy of feature expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372255A_ABST
    Figure CN120372255A_ABST
Patent Text Reader

Abstract

The invention provides a massive financial and accounting data fusion auditing feature extraction method, medium and system, and belongs to the technical field of digital data feature extraction. After a basic time window is determined through time sequence clustering, a self-changing time window function is constructed, and the window boundary is dynamically adjusted according to parameters such as a service fluctuation coefficient; extracting time dimension features by using a finance and accounting time sequence feature recognition model, extracting spatial dimension features based on geographical distribution and organizational structure information, extracting service dimension features according to service types and transaction attributes, calculating an outlier degree through statistical analysis to form outlier degree features, and establishing a second time window for verification; and a multi-objective optimization weight determination model is adopted to solve an optimal weight coefficient, and the four types of features are subjected to weighted fusion to form a comprehensive audit feature vector, so that the technical problem of insufficient audit feature extraction precision caused by incapability of dynamically adapting to internal change features of the data during feature extraction of the massive financial and accounting data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital data feature extraction. Specifically, it relates to a method, medium and system for extracting integrated audit features from a large amount of financial and accounting data. Background Art

[0002] The extraction of audit features from financial and accounting data is the core technology of audit informatization. Traditional methods mainly adopt technical means such as fixed-time window analysis, static feature extraction, and single-dimensional analysis. In application scenarios such as enterprise financial management, accounting firm audit services, financial institution risk control, and government department financial supervision, the existing technology usually divides financial and accounting data based on a preset fixed time period, extracts time series features using statistical analysis methods, and identifies abnormal patterns through rule matching. However, the traditional fixed window method has significant defects and cannot adapt to the irregular change patterns of financial and accounting data, resulting in inaccurate positioning of key abnormal time nodes. The static feature extraction method lacks the ability to perceive the internal dynamic changes of data and is difficult to capture the complex time-varying features of financial and accounting data. The single-dimensional analysis ignores the correlation of multi-dimensional features such as time, space, and business, resulting in insufficient feature expression ability. Traditional technologies are difficult to solve the core technical problem that the extraction accuracy of audit features is insufficient due to the inability to dynamically adapt to the internal change features of data when extracting features from a large amount of financial and accounting data. In the face of complex and variable financial and accounting data, the feature extraction results of existing methods are often not accurate enough, affecting the accuracy and reliability of subsequent audit analysis. Summary of the Invention

[0003] In view of this, the present invention provides a method, medium and system for extracting integrated audit features from a large amount of financial and accounting data, which can solve the technical problem that the extraction accuracy of audit features is insufficient due to the inability to dynamically adapt to the internal change features of data when extracting features from a large amount of financial and accounting data in the prior art.

[0004] The present invention is implemented as follows: In the first aspect of the present invention, a method for extracting massive financial and accounting data fusion audit features includes: determining the length of the first time window by using a time series clustering algorithm, and dividing historical financial and accounting data into multiple first time windows based on the periodic characteristics of the financial and accounting data; constructing a self-varying time window function, inputting the financial and accounting data within the first time window, the business fluctuation coefficient, the time granularity parameter, the data change rate threshold, and the window overlap degree, and outputting a self-varying time window sequence for dynamically adjusting the window boundary to adapt to the internal change characteristics of the financial and accounting data; using a pre-trained financial and accounting time series feature recognition model to perform deep learning processing on the self-varying time window sequence, identifying long-term trend patterns and periodic change rules in the financial and accounting data, and outputting time dimension features; extracting spatial dimension features according to the geographical distribution information and organizational structure information of the financial and accounting data in the self-varying time window sequence; extracting business dimension features based on the business types and transaction attributes of the financial and accounting data in the self-varying time window sequence; calculating the outlier degree of the financial and accounting data within the self-varying time window sequence, and identifying the outliers corresponding to abnormal financial and accounting behaviors through statistical analysis methods to form outlier degree features; constructing a multi-objective optimization weight determination model, using the time dimension features, spatial dimension features, business dimension features, and outlier degree features as input variables, and solving the optimal weight coefficients through a multi-objective optimization algorithm; and using the optimal weight coefficients to perform weighted fusion on the time dimension features, spatial dimension features, business dimension features, and outlier degree features to form a comprehensive audit feature vector.

[0005] Among them, the first time window is specifically a fixed time period determined according to the natural periodicity of the financial and accounting data. The length is automatically determined by analyzing the similarity pattern of historical financial and accounting data through a clustering algorithm, and is used to capture the basic periodic characteristics of the financial and accounting data. It usually corresponds to the standard reporting cycle or settlement cycle of the financial and accounting business, providing a stable time benchmark framework for subsequent feature extraction.

[0006] Among them, the self-varying time window sequence is specifically a set of variable-length time periods further subdivided on the basis of the first time window. The window boundary is adaptively adjusted according to the internal change rate of the financial and accounting data, and is used to accurately locate the time nodes of abnormal changes. It can dynamically adapt to the irregular change patterns of the financial and accounting data and improve the accuracy of feature extraction.

[0007] Among them, the self-varying time window function is used to dynamically adjust the boundary position and length size of the time window according to the real-time change characteristics of the financial and accounting data, so as to adapt to the fluctuation patterns and change rules of the financial and accounting data in different business scenarios. The output is a self-varying time window sequence after adaptive adjustment, and each window is marked with the start time point, end time point, and statistical characteristics of the data within the window.

[0008] Among them, the business fluctuation coefficient is a quantitative indicator specifically reflecting the intensity change of financial and accounting business activities. It is calculated by analyzing the amplitude and frequency characteristics of the business volume in historical financial and accounting data and is used to guide the dynamic adjustment process of the self-varying time window function.

[0009] Among them, the specific structure of the financial and accounting time series feature recognition model is a deep neural network designed based on the long short-term memory network architecture, which includes multiple LSTM units for processing the long-term dependencies of time series data. The entire network structure enhances the recognition ability of key time nodes and important financial and accounting events through residual connections and attention mechanisms.

[0010] Among them, the multi-objective optimization weight determination model is used to solve the optimal weight allocation of time dimension features, space dimension features, business dimension features, and outlier degree features in the comprehensive audit feature vector. It uses the multi-objective genetic algorithm as the core optimization engine and takes the audit accuracy rate, recall rate, F1 score, and feature density after feature fusion as the optimization objective functions.

[0011] Among them, after the step of constructing the self-varying time window function, it also includes establishing a second time window for verifying the feature extraction effect. The length of the second time window is different from that of the first time window. The steps of using the financial and accounting time series feature recognition model to extract time dimension features to form outlier degree features are repeatedly executed to obtain a verification feature set.

[0012] Among them, the outlier points are specifically abnormal financial and accounting data points identified through statistical analysis methods. Their numerical characteristics significantly deviate from the normal range, indicating potential audit risks or abnormal financial and accounting behaviors. The abnormality degree and risk level of the data points are determined through multi-dimensional distance calculation and probability distribution analysis.

[0013] Among them, the outlier degree is a numerical indicator specifically quantifying the abnormality degree of each data point. It is obtained by calculating the statistical distance between the data point and the center of the normal distribution. The larger the value, the higher the abnormality degree. The Mahalanobis distance or kernel density estimation method is used for precise quantification to support subsequent risk assessment and audit decision-making.

[0014] Among them, the feature density is a numerical indicator specifically quantifying the density of non-zero element distribution in the comprehensive audit feature vector. It is obtained through the feature density calculation function. The ratio of the number of non-zero elements to the total number of elements in the feature vector is calculated, and weighted calculation is combined with the numerical size and distribution uniformity of the non-zero elements.

[0015] The second aspect of the present invention provides a computer-readable storage medium. Program instructions are stored in the computer-readable storage medium. When the program instructions run on a computer, they are used to execute the above-mentioned method for extracting fusion audit features from a large amount of financial and accounting data.

[0016] The third aspect of the present invention provides a system for extracting massive financial and accounting data fusion audit features, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing program instructions stored in the computer-readable storage medium is set inside the system.

[0017] The present invention adapts to the internal change characteristics of financial and accounting data by dynamically adjusting the time window boundary. This method can adjust the window length and position in real time according to parameters such as the business fluctuation coefficient and the data change rate threshold, and accurately locate the time nodes of abnormal changes. The present invention effectively solves the limitations of the traditional fixed window method. The self-varying time window function can adaptively adjust the boundary to match the irregular change pattern of the data, significantly improving the positioning accuracy of abnormal time nodes. The multi-dimensional feature fusion mechanism overcomes the deficiencies of single-dimensional analysis. Through the comprehensive processing of the time dimension, space dimension, business dimension, and outlier degree features, a more comprehensive feature expression system is constructed. The pre-trained financial and accounting time series feature recognition model enhances the recognition ability of complex time-varying features, and the multi-objective optimization weight determination model realizes the optimal fusion of features in each dimension; it solves the technical problem that the internal change characteristics of data cannot be dynamically adapted during the extraction of massive financial and accounting data features. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0020] As Figure 1 shown, it is a flowchart of a method for extracting massive financial and accounting data fusion audit features provided by the first aspect of the present invention. The method includes the following steps: A method for extracting massive financial and accounting data fusion audit features includes the following steps: S01. Use a time series clustering algorithm to determine the length of the first time window, and divide the historical financial and accounting data into multiple first time windows based on the periodic characteristics of the financial and accounting data; S02. Construct a self-varying time window function, input the financial and accounting data, business fluctuation coefficient, time granularity parameter, data change rate threshold, and window overlap degree within the first time window, and output a self-varying time window sequence for dynamically adjusting the window boundary to adapt to the internal change characteristics of the financial and accounting data; S03. Deeply process the self-varying time window sequence using a pre-trained financial and accounting time series feature recognition model to identify long-term trend patterns and periodic change rules in financial and accounting data, and output time dimension features; S04. Extract spatial dimension features based on the geographical distribution information and organizational structure information of the financial and accounting data in the self-varying time window sequence;

[0021] S05. Extract business dimension features based on the business types and transaction attributes of the financial and accounting data in the self-varying time window sequence; S06. Calculate the outlier degree of the financial and accounting data within the self-varying time window sequence, identify the outliers corresponding to abnormal financial and accounting behaviors through statistical analysis methods, and form outlier degree features; S07. Establish a second time window to verify the feature extraction effect. The length of the second time window is different from that of the first time window. Repeat steps S02 to S06 to obtain a verification feature set; S08. Construct a multi-objective optimization weight determination model, use the time dimension features, spatial dimension features, business dimension features, and outlier degree features as input variables, and solve the optimal weight coefficients through a multi-objective optimization algorithm; S09. Use the optimal weight coefficients to perform weighted fusion on the time dimension features, spatial dimension features, business dimension features, and outlier degree features to form a comprehensive audit feature vector; S10. Output the comprehensive audit feature vector, including time trend features, periodic pattern features, abnormal behavior features, and multi-dimensional correlation features.

[0022] Among them, the first time window is specifically a fixed time period determined according to the natural periodicity of financial and accounting data. The length is automatically determined by analyzing the similarity patterns of historical financial and accounting data through a clustering algorithm, and is used to capture the basic periodic features of financial and accounting data. It usually corresponds to the standard reporting period or settlement period of financial and accounting operations, providing a stable time benchmark framework for subsequent feature extraction.

[0023] Among them, the second time window is specifically a time period for model verification. Its length is obtained by randomly changing or adjusting according to a preset ratio, and forms a contrast verification with the first time window to ensure the stability and generalization ability of the feature extraction method. The effectiveness and consistency of the extracted features at different time scales are verified through a cross-validation mechanism.

[0024] Among them, the self-varying time window sequence is specifically a set of variable-length time periods further subdivided on the basis of the first time window. The window boundaries are adaptively adjusted according to the internal change rate of financial and accounting data, and are used to accurately locate the time nodes of abnormal changes. It can dynamically adapt to the irregular change patterns of financial and accounting data and improve the accuracy of feature extraction.

[0025] Among them, the outlier is specifically an abnormal financial and accounting data point identified through statistical analysis methods. Its numerical characteristics significantly deviate from the normal range, indicating potential audit risks or abnormal financial and accounting behaviors. The degree of abnormality and risk level of the data point are determined through multi-dimensional distance calculation and probability distribution analysis.

[0026] Among them, the outlier degree is specifically a numerical index that quantifies the degree of abnormality of each data point. It is obtained by calculating the statistical distance between the data point and the center of the normal distribution. The larger the value, the higher the degree of abnormality. Methods such as Mahalanobis distance or kernel density estimation are used for precise quantification to support subsequent risk assessment and audit decisions.

[0027] Among them, the business fluctuation coefficient is specifically a quantitative index that reflects the change in the intensity of financial and accounting business activities. It is calculated by analyzing the amplitude and frequency characteristics of the business volume in historical financial and accounting data and is used to guide the dynamic adjustment process of the self-varying time window function.

[0028] Among them, the time granularity parameter is specifically a numerical parameter that controls the time resolution. It determines the minimum time unit for financial and accounting data analysis and is determined according to the actual needs of financial and accounting operations and the data collection frequency, affecting the fineness of the self-varying time window sequence.

[0029] Among them, the data change rate threshold is specifically a critical value used to determine whether there are significant changes in financial and accounting data. It is determined by statistically analyzing the change distribution characteristics of historical financial and accounting data. When the data change exceeds the threshold, the boundary adjustment mechanism of the self-varying time window function is triggered.

[0030] Among them, the window overlap degree is specifically a proportional parameter that determines the overlap degree of adjacent time windows. Its value range is between 0 and 1, which is used to balance data continuity and calculation efficiency and affects the intersection size of adjacent windows in the self-varying time window sequence.

[0031] Among them, the time dimension feature is specifically a feature vector extracted from the time series of financial and accounting data, including trend features, periodic features, and seasonal features. It is identified and extracted from the self-varying time window sequence through a financial and accounting time series feature recognition model.

[0032] Among them, the space dimension feature is specifically a feature vector extracted based on the geographical distribution information and organizational structure information of financial and accounting data, reflecting the differences in financial and accounting activity patterns between different regions and departments. It is extracted from the geographical and organizational attribute data in the self-varying time window sequence.

[0033] Among them, the business dimension feature is specifically a feature vector extracted according to the business types and transaction attributes of financial and accounting data, reflecting the financial and accounting behavior characteristics and patterns of different business categories. It is extracted from the business classification and transaction records in the self-varying time window sequence.

[0034] Among them, the outlier degree feature is specifically a feature vector that quantifies the abnormal degree of financial and accounting data. It is formed by calculating the outlier degree of data points within a self-varying time window sequence and is used to identify potential audit risk points and abnormal financial and accounting behaviors.

[0035] Among them, the optimal weight coefficient is specifically a combination of weight parameters obtained by solving a multi-objective optimization weight determination model, which is used to balance the importance of time dimension features, space dimension features, business dimension features, and outlier degree features in the comprehensive audit feature vector.

[0036] Among them, the comprehensive audit feature vector is specifically the final feature representation formed by weighted fusion of time dimension features, space dimension features, business dimension features, and outlier degree features according to the optimal weight coefficient, including time trend features, periodic pattern features, abnormal behavior features, and multi-dimensional correlation features.

[0037] The self-varying time window function is used to dynamically adjust the boundary position and length of the time window according to the real-time change characteristics of financial and accounting data, so as to adapt to the fluctuation patterns and change rules of financial and accounting data in different business scenarios. The input includes the original financial and accounting data sequence within the first time window, the business fluctuation coefficient reflecting the change of business activity intensity, the time granularity parameter controlling the time resolution, the data change rate threshold for judging significant data changes, and the window overlap degree parameter determining the overlap degree between adjacent windows. The output is a self-varying time window sequence after adaptive adjustment, and each window is marked with the start time point, end time point, and statistical features of the data within the window.

[0038] The specific structure of the financial and accounting time series feature recognition model is a deep neural network designed based on the long short-term memory network architecture, which includes multiple LSTM units for processing the long-term dependencies of time series data. The sequence length parameter of the network is adjusted according to the periodic characteristics of financial and accounting data to match the data lengths of different business cycles. The hidden state dimension parameter is matched with the complexity and feature dimension of financial and accounting data to ensure the expression ability of the model. The temperature coefficient parameter is used to control the confidence distribution of the model output feature vector so that the model can give a reasonable probability assessment for uncertain financial and accounting patterns. The entire network structure enhances the recognition ability of key time nodes and important financial and accounting events through residual connections and attention mechanisms.

[0039] The steps for establishing the training dataset of the financial and accounting time series feature recognition model specifically include collecting a large amount of historical financial and accounting data and classifying and annotating it according to business types and time periods, converting the original financial and accounting data into a numerical format suitable for neural network training through data cleaning and standardization processing, establishing a supervised learning label system including normal mode labels and abnormal mode labels based on known audit results and expert annotations, using the sliding window technique to slice continuous financial and accounting time series data into multiple training samples, each sample containing an input sequence and a corresponding feature label, augmenting the diversity of training samples through data augmentation techniques including time perturbation and numerical transformation to improve the generalization ability of the model, and finally constructing a large-scale training dataset containing hundreds of thousands of annotated samples for model training and optimization.

[0040] The steps for training the financial and accounting time series feature recognition model specifically include first using a large-scale financial and accounting dataset for pre-training to establish the model's understanding ability of the basic patterns of financial and accounting data, adjusting the sequence length parameter so that the model can process financial and accounting time series data of different lengths and optimizing the hidden state dimension parameter to balance model complexity and computational efficiency, using the temperature coefficient parameter to adjust the gradient update strategy during training so that the model can learn the subtle change patterns in financial and accounting data, using the backpropagation algorithm and an adaptive learning rate scheduler to optimize the model parameters and preventing overfitting through regularization techniques, continuously monitoring the performance of the model on the validation set during training and adjusting the training strategy according to the convergence of the loss function, and finally obtaining a model that can accurately identify financial and accounting time series features through multiple rounds of iterative training for subsequent feature extraction tasks.

[0041] The multi-objective optimization weight determination model is used to solve the optimal weight allocation of time dimension features, space dimension features, business dimension features, and outlier degree features in the comprehensive audit feature vector. It uses the multi-objective genetic algorithm as the core optimization engine, takes the audit accuracy, recall rate, F1 score, and feature density after feature fusion as the optimization objective functions, and finds the optimal weight coefficient combination that balances multiple objectives through population evolution and Pareto optimal solution set search. The model input is the standardized time dimension feature, space dimension feature, business dimension feature, and outlier degree feature vectors, and the output is the corresponding optimal weight coefficients. The sum of the weight coefficients is equal to 1 and each weight value is non-negative. The cross-validation mechanism is used to ensure that the solved optimal weight coefficients have good generalization performance and stability.

[0042] Among them, the feature density is a numerical index specifically quantifying the density of non-zero element distribution in the comprehensive audit feature vector, which is obtained through the feature density calculation function. The feature density calculation function takes the comprehensive audit feature vector as the input, calculates the ratio of the number of non-zero elements to the total number of elements in the feature vector, and performs weighted calculation in combination with the numerical size and distribution uniformity of the non-zero elements, and outputs the feature density value. The larger the feature density value, the richer the effective information contained in the feature vector, which is conducive to improving the accuracy and reliability of audit analysis.

[0043] The following describes the specific implementation manners of the above steps in detail.

[0044] The specific implementation manner of step S01 is to first collect a large amount of historical financial and accounting data as the analysis basis, and convert the original financial and accounting data into a standardized time series format through data preprocessing. The dynamic time warping algorithm is used to calculate the similarity distance between the financial and accounting data in different time periods. This algorithm can effectively handle the problems of time offset and length inconsistency existing in the financial and accounting data. Based on the similarity distance matrix, the hierarchical clustering algorithm is used to perform clustering analysis on the historical financial and accounting data. The similarity threshold is set to 0.85 during the clustering process. When the similarity between data exceeds this threshold, they are classified into the same category. By analyzing the time span distribution characteristics of each category in the clustering results, the optimal length of the first time window is determined. Usually, this length is consistent with the natural periodicity of the financial and accounting business, such as the monthly or quarterly reporting cycle. The historical financial and accounting data is divided using the determined length of the first time window to form multiple regular time periods, and each time period contains a relatively stable financial and accounting business model, laying a time-based reference framework for subsequent feature extraction.

[0045] The specific implementation manner of step S02 is to construct a self-varying time window function based on an adaptive adjustment mechanism. This function takes the financial and accounting data within the first time window as the input basic data. The business fluctuation coefficient is obtained by calculating the coefficient of variation of the business volume in the financial and accounting data. The coefficient of variation is calculated by the ratio of the standard deviation to the mean. When the coefficient of variation exceeds 0.3, it indicates that the business fluctuation is large. The time granularity parameter is determined according to the collection frequency of the financial and accounting data. It is set to 24 hours for daily data and 1 hour for hourly data. The data change rate threshold is determined by analyzing the change distribution of the historical financial and accounting data and is set using the quantile method. Usually, the change rate corresponding to the 95% quantile is taken as the threshold, and the reference value is 15%. The window overlap parameter controls the overlap degree of adjacent time windows, and the setting range is between 0.1 and 0.5, and the reference value is 0.25. The function dynamically calculates the window boundary adjustment amount at each time point according to the input parameters. When it detects that the data change rate exceeds the threshold, it triggers the adaptive adjustment of the window boundary and outputs a self-varying time window sequence containing multiple variable-length time periods.

[0046] The specific implementation of step S03 is to perform deep learning processing on the self-varying time window sequence using a pre-trained financial and accounting time series feature recognition model. This model is designed based on the long short-term memory network architecture and can effectively capture long-term dependencies and complex patterns in financial and accounting data. The financial and accounting data in the self-varying time window sequence is preprocessed according to the input format required by the model, including data standardization and sequence padding operations. The model performs feature learning on the input data through a multi-layer neural network structure to identify long-term trend changes, periodic fluctuation laws, and seasonal adjustment patterns in financial and accounting data. The model output is a time dimension feature set containing trend feature vectors, periodic feature vectors, and seasonal feature vectors. The dimension of each feature vector is determined according to the complexity of the financial and accounting data and is usually set to 64 to 256 dimensions.

[0047] The specific implementation of step S04 is to extract spatial dimension features based on the geographical identification and organizational structure identification of the financial and accounting data in the self-varying time window sequence. First, the geographical distribution information in the financial and accounting data is encoded. The geographical coding algorithm is used to convert the geographical location into a numerical coordinate to establish the mapping relationship between the geographical location and the intensity of financial and accounting activities. The tree structure coding method is used for the organizational structure information to convert the hierarchical relationship into a numerical vector representation. The correlation of financial and accounting activities between different geographical regions is calculated, and the Pearson correlation coefficient method is used to measure the degree of correlation. The correlation coefficient threshold is set to 0.6. The differences in financial and accounting behavior patterns between different organizational departments are analyzed, and the clustering analysis is used to identify the department combinations with similar financial and accounting characteristics. The spatial dimension feature vector containing geographical distribution features and organizational association features is output.

[0048] The specific implementation of step S05 is to extract business dimension features according to the business classification identification and transaction attribute identification of the financial and accounting data in the self-varying time window sequence. The one-hot encoding method is used to numerically represent different business types to establish the mapping relationship between the business type and the feature vector. The multi-dimensional coding method is used for the transaction attributes, including attribute dimensions such as transaction amount level, transaction frequency category, and transaction object type. Statistical analysis methods are used to calculate the financial and accounting behavior feature parameters of different business types, including average transaction amount, transaction frequency distribution, and abnormal transaction ratio. The internal connection between different businesses is identified through business relevance analysis, and the association rule mining algorithm is used to discover the potential laws in the business model. The minimum support threshold is set to 0.1, and the minimum confidence threshold is set to 0.7. The business dimension feature vector containing business type features and transaction mode features is output.

[0049] The specific implementation of step S06 is to calculate the outlier degree quantification index for each financial and accounting data point within the self-varying time window sequence. The Isolation Forest algorithm is used to identify outliers in financial and accounting data. This algorithm detects the outlier degree of data points by constructing random partitioning trees, and the outlier score threshold is set to 0.7. The Mahalanobis distance method is used to calculate the statistical distance between a data point and the center of the normal distribution. The larger the distance value, the higher the outlier degree, and the distance threshold is set to 3 standard deviations. The Kernel Density Estimation method is used to analyze the position of data points in the probability distribution. Data points with a probability density lower than the 5th percentile are identified as outliers. The Local Outlier Factor algorithm is used to calculate the outlier degree of each data point relative to its neighborhood. Data points with an outlier factor greater than 1.5 are considered potential outliers. The results of multiple outlier detection methods are comprehensively evaluated to form an outlier degree feature vector reflecting the data outlier degree.

[0050] The specific implementation of step S07 is to establish a second time window with a length different from that of the first time window to verify the stability of the feature extraction method. The length of the second time window is determined by the random perturbation method, and the perturbation amplitude is set between 10% and 30% of the length of the first time window. The second time window is applied to the same historical financial and accounting data, and the feature extraction process of steps S02 to S06 is repeated to obtain the corresponding verification feature set. The feature similarity measurement method is used to compare the consistency degree of the features extracted under the two time windows. The similarity coefficient threshold is set to 0.8. When the similarity coefficient is lower than this threshold, the feature extraction parameters need to be adjusted. The cross-validation mechanism is used to evaluate the generalization ability of the feature extraction method at different time scales to ensure that the extracted features have good stability and reliability.

[0051] The specific implementation of step S08 is to construct a weight optimization model based on the multi-objective genetic algorithm. This model uses time dimension features, space dimension features, business dimension features, and outlier degree features as input variables. The audit accuracy rate, recall rate, F1 score, and feature density are set as the optimization objective functions, and the weight coefficient combination that balances multiple objectives is found through the Pareto optimal solution set search. The population size of the genetic algorithm is set to 100 individuals, the number of evolutionary generations is set to 200 generations, the crossover probability is set to 0.8, and the mutation probability is set to 0.1. Each individual represents a set of weight coefficients, and the sum of the weight coefficients is equal to 1 and all are non-negative numbers. The fitness function is used to evaluate the comprehensive performance of each weight combination, and the fitness function comprehensively considers the weighted average of each objective function. After multiple rounds of evolutionary iteration, the optimal weight coefficient combination is output.

[0052] The specific implementation of step S09 is to perform weighted fusion processing on the features of the four dimensions using the optimal weight coefficients obtained by solving in step S08. First, the time dimension features, space dimension features, business dimension features, and outlier features are standardized to eliminate the dimensional differences between different features. Calculate the weighted values of the features of each dimension according to the optimal weight coefficients, and fuse the weighted feature vectors through a linear combination method. During the fusion process, feature dimensionality reduction technology is used to process high-dimensional feature vectors, and the principal component analysis method is used to retain the main components with a cumulative contribution rate reaching 95%. Post-process the fused feature vectors, including outlier correction and feature smoothing, and finally form a comprehensive audit feature vector containing time trend features, periodic pattern features, abnormal behavior features, and multi-dimensional correlation features.

[0053] The specific implementation of step S10 is to output the comprehensively audited feature vector after optimized fusion, which contains the complete feature information required for audit analysis. The time trend feature reflects the long-term development and change direction of financial and accounting data, the periodic pattern feature reflects the regular change characteristics of financial and accounting operations, the abnormal behavior feature identifies potential audit risk points, and the multi-dimensional correlation feature reveals the internal connections between different dimensions. The feature vector is output in a standardized format for easy calling and processing by subsequent audit analysis systems. The output result also includes the confidence evaluation and quality evaluation indicators of the feature vector, providing reliable data support for audit decisions.

[0054] The financial and accounting time series feature recognition model is designed using a deep neural network architecture based on long short-term memory networks. The network structure includes an input layer, multiple LSTM hidden layers, an attention layer, and an output layer. The input layer receives preprocessed financial and accounting time series data in the form of a numerical vector sequence of a fixed length. The LSTM hidden layer is designed with a bidirectional structure, including a forward LSTM unit and a backward LSTM unit. The hidden state dimension of each LSTM unit is set to 128 dimensions, which can capture the forward and backward dependencies of financial and accounting data simultaneously. The network contains 3 layers of LSTM hidden layers, and a residual connection mechanism is used between layers to prevent the problem of gradient disappearance. The attention layer is designed with a self-attention mechanism to identify key financial and accounting events and important time nodes by calculating the attention weights between different time steps. The output layer uses a fully connected neural network structure to map the hidden layer features to the final feature vector output through an activation function.

[0055] The process of establishing the training dataset first collects historical financial and accounting data covering multiple industries and enterprises of different scales, with a data time span of no less than 5 years to ensure the representativeness and integrity of the data. Quality inspection and cleaning processing are carried out on the original financial and accounting data, including missing value filling, outlier detection, and data consistency verification. The data is classified and labeled according to the characteristics of financial and accounting operations, and a label system including normal business models and abnormal business models is established. The sliding window technique is used to divide the continuous financial and accounting time series into training samples of a fixed length, and each sample contains an input sequence and a corresponding feature label. The diversity of training samples is expanded through data augmentation techniques, including time shift transformation, numerical perturbation transformation, and sequence length adjustment methods. Finally, a large-scale training dataset containing more than 500,000 labeled samples is constructed, and the dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 for model training, tuning, and performance evaluation.

[0056] It should be noted that the first main technical idea of the present invention is the dynamic adjustment mechanism of the self-varying time window sequence, which realizes the adaptive response to the change characteristics of financial and accounting data by constructing a self-varying time window function. The traditional fixed time window method cannot effectively handle the irregular change patterns and sudden abnormal events existing in financial and accounting data, resulting in the mismatch between the time boundary of feature extraction and the internal change law of the data. Through the comprehensive regulation of parameters such as the business fluctuation coefficient, the data change rate threshold, and the window overlap degree, the present invention enables the time window to dynamically adjust the boundary position and length size according to the real-time change characteristics of financial and accounting data, effectively captures the precise time nodes of abnormal changes, significantly improves the time accuracy and adaptability of feature extraction, and avoids the problem that important financial and accounting events are truncated or missed by the time window boundary in the traditional method.

[0057] The second main technical idea of the present invention is a multi-dimensional feature fusion system based on deep learning, which systematically integrates the features of the time dimension, the space dimension, the business dimension, and the outlier degree dimension. The existing financial and accounting audit methods usually adopt a single dimension or a simple combination of feature analysis methods, lacking the ability to deeply mine the complex correlation relationships in financial and accounting data. The present invention uses a long short-term memory network model to identify the long-term dependence relationships in the time series, combines the geographical distribution and organizational structure information to extract spatial correlation features, and at the same time incorporates the business dimension features of business types and transaction attributes, forming a multi-dimensional collaborative feature representation system, which can comprehensively reflect the complex patterns and potential correlations in financial and accounting data, and greatly improves the integrity and expression ability of audit features.

[0058] The third main technical idea of the present invention is a feature density-guided weight optimization strategy, which quantifies the distribution density of effective information in the comprehensive audit feature vector through a feature density calculation function. Traditional feature fusion methods usually adopt empirical weight assignment or simple average weight strategies, lacking a quantitative evaluation mechanism for feature information density. The present invention takes feature density as an important objective function for multi-objective optimization. By calculating the proportion and distribution uniformity of non-zero elements in the feature vector, it ensures that the fused feature vector contains as rich effective information as possible, avoiding feature redundancy and information sparsity problems, and enabling the comprehensive audit feature vector to have higher information-carrying capacity and discrimination effect.

[0059] The synergistic effect of these three main technical ideas has significantly improved the technical effect. The dynamic adjustment of the self-varying time window provides an accurate time benchmark for multi-dimensional feature extraction, ensuring that features in each dimension can be extracted within the optimal time boundary and avoiding the problem of reduced feature quality caused by time window mismatch. The multi-dimensional feature fusion system provides a rich feature source for feature density optimization, enabling the weight optimization process to find the optimal solution in a larger feature space and improving the quality and stability of the optimization results. The feature density-guided weight optimization strategy provides a quantitative evaluation standard for the self-varying time window adjustment and multi-dimensional feature extraction, forming a closed-loop optimization feedback mechanism, enabling the entire feature extraction system to continuously self-improve and optimize, and achieving a comprehensive improvement in feature extraction accuracy, adaptability, and stability compared with traditional methods.

[0060] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned method for extracting fusion audit features of a large amount of financial and accounting data.

[0061] The third aspect of the present invention provides a system for extracting fusion audit features of a large amount of financial and accounting data, including the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0062] Specifically, the principle of the present invention is as follows: The principle of the present invention to solve the core technical problem lies in constructing a dynamic adaptive feature extraction framework, which can automatically adjust the analysis strategy according to the internal change law of financial and accounting data. First, the basic time window length is determined through the time series clustering algorithm to provide a stable time benchmark for subsequent processing. The key innovation lies in the design of the self-varying time window function, which receives multiple control parameters such as the business fluctuation coefficient, time granularity parameter, data change rate threshold, and window overlap degree, and can real-time monitor the change amplitude and frequency characteristics of financial and accounting data. When the data change exceeds the preset threshold, it automatically triggers the window boundary adjustment mechanism, thereby realizing the accurate tracking of the internal change characteristics of the data.

[0063] The financial and accounting time series feature recognition model is based on the long short-term memory network architecture. Through multiple layers of LSTM units, it processes complex time-dependent relationships, and combines residual connections and attention mechanisms to enhance the recognition ability of key time nodes. The model has been pre-trained on a large-scale financial and accounting data set, and has the ability to deeply understand the basic patterns of financial and accounting data, and can accurately identify long-term trend patterns and periodic change laws from the self-varying time window sequence. The multi-dimensional feature extraction mechanism constructs feature vectors from four dimensions: time, space, business, and anomaly at the same time, and determines the optimal weight coefficient of the model solution through multi-objective optimization of weights, realizing the scientific integration of features in each dimension.

[0064] The fundamental reason why the technical solution of the present invention is logical lies in that it establishes a direct mapping relationship between data change perception and feature extraction strategy adjustment. Through the self-varying time window function, it realizes the active adaptation of the feature extraction process to the internal change law of the data, rather than passively using a fixed strategy to process dynamically changing data, thereby fundamentally solving the fundamental problem that traditional methods cannot dynamically adapt to the internal change characteristics of the data from the technical essence.

[0065] The following provides a specific embodiment 1 of the present invention. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0066] In step S01, the calculation process of the dynamic time warping algorithm similarity distance is described in detail as follows: ; In the formula, is the sequence and the sequence the dynamic time warping distance between them; is the first financial and accounting time series data; is the second financial and accounting time series data; is the path sequence; is the path length; is the the Euclidean distance between the th point pairs on the path; are the coordinates of the path points in sequence ; are the coordinates of the path points in sequence ; is the point serial number on the path, and its value range is from to . The calculation of the coefficient of variation is specifically expressed as: ; In the formula, is the coefficient of variation; is the standard deviation of the financial and accounting data; is the mean value of the financial and accounting data.

[0067] In step S02, the calculation process of the business fluctuation coefficient is described in detail as follows: ; In the formula, is the business fluctuation coefficient; is the standard deviation of the financial and accounting business volume; is the mean value of the financial and accounting business volume; is the adjustment factor, and its value range is 0.8 to 1.2. The calculation of the data change rate is specifically expressed as: ; In the formula, is the data change rate at the th moment; is the financial and accounting data value at the th moment; is the financial and accounting data value at the th moment; is the time index, indicating the current time point. The self-varying time window length adjustment function is specifically expressed as: ; In the formula, is the length of the th self-varying time window; is the basic window length; is the window adjustment coefficient, and its value range is 0.1 to 0.5; is the average data change rate of the th window; is the data change rate threshold; is the window serial number.

[0068] The specific implementation manner of step S03 is the same as the foregoing, and will not be elaborated here in detail.

[0069] In step S04, the calculation process of the Pearson correlation coefficient is described in detail as follows: ; In the formula, is the Pearson correlation coefficient; is the financial and accounting activity data of the th geographical region; is the corresponding reference data of the th geographical region; is the mean of the sequence; is the mean of the sequence; is the sample size; is the geographical region serial number, and the value range is to . The calculation process of geographical distribution correlation is described in detail as follows: ; In the formula, is the geographical distribution correlation coefficient; is the coordinate value of the th geographical region; is the mean of the coordinate values of all geographical regions; is the financial and accounting activity intensity of the th geographical region; is the mean of the financial and accounting activity intensities of all geographical regions; is the total number of geographical regions; is the geographical region serial number, and the value range is to . The calculation of organizational association degree is specifically expressed as: ; In the formula, is the association degree between organization and organization ; is the set of financial and accounting operations of organization ; is the set of financial and accounting operations of organization ; is the organizational hierarchy weight factor, and the value range is 0.5 to 2.0; and are the organization serial numbers.

[0070] In step S05, the calculation process of the confidence degree of business association rules is described in detail as follows: ; In the formula, is the confidence degree of business rule to ; is business and business Support of co-occurrence; For the service Support of single occurrence; Is the antecedent service event; Is the consequent service event. The calculation of the service type feature weight is specifically expressed as: ; In the formula, Is the feature weight of the th type of service; Is the occurrence frequency of the th type of service; Is the total number of service types; Is the total amount of financial data; Is the number of documents containing the th type of service; Is the service type serial number, and the value range is To ; Is the summation variable.

[0071] In step S06, the calculation process of kernel density estimation is described in detail as follows: ; In the formula, Is the kernel density estimate value at the data point ; Is the total number of samples; Is the bandwidth parameter, and the value range is 0.1 to 0.5; Is the kernel function, and the Gaussian kernel function is adopted ; Is the th sample point; Is the data point where the density to be estimated is located; Is the sample serial number, and the value range is To ; Is the input variable of the kernel function; Is the natural constant, approximately equal to 2.718; Is the pi, approximately equal to 3.14159. The calculation process of Mahalanobis distance is described in detail as follows: ; In the formula, Is the Mahalanobis distance of the th data point; Is the feature vector of the th data point; Is the mean vector of the feature vectors; Is the covariance matrix of the feature vectors; is the inverse matrix of the covariance matrix; is the serial number of the data point; represents the vector transpose. The calculation of the local outlier factor is specifically expressed as: ; In the formula, is the local outlier factor of the th data point; is the nearest neighbor set of the data point ; is the local reachability density of the data point ; is the local reachability density of the neighbor point ; is the serial number of the neighbor point; is the number of elements in the nearest neighbor set. The calculation of the outlier degree comprehensive score is specifically expressed as: ; In the formula, is the outlier degree comprehensive score of the th data point; is the normalized Mahalanobis distance; is the normalized local outlier factor; is the normalized isolation forest anomaly score; is the weight coefficient, satisfying .

[0072] The specific implementation manner of step S07 is the same as the foregoing, and will not be elaborated in detail here.

[0073] In step S08, the calculation process of the multi-objective optimization fitness function is described in detail as follows: ; In the formula, is the fitness function value of the weight vector ; is the weight vector of the four-dimensional features, satisfying and ; is the weight coefficient of the objective function, and the value range is 0.2 to 0.3; is the audit accuracy rate function; is the recall rate function; is the F1 score function; is the feature density function; are the weight coefficients of the time dimension, space dimension, business dimension and outlier degree dimension respectively.

[0074] In step S09, the calculation process of the feature weighted fusion is described in detail as follows: ; In the formula, is the comprehensive audit feature vector; is the feature vector in the time dimension; is the feature vector in the space dimension; is the feature vector in the business dimension; is the outlier degree feature vector; is the corresponding optimal weight coefficient.

[0075] In step S10, the feature density calculation function is specifically expressed as: ; In the formula, is the feature density value; is the number of non-zero elements in the feature vector; is the total number of elements in the feature vector; is the th non-zero element value; is the th element position distribution uniformity coefficient, and the calculation method is , where is the th element position index, is the average value of all non-zero element positions, is the maximum position index of the vector; is the element serial number and belongs to the non-zero element set .

[0076] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: A research team, aiming at the financial and accounting audit requirements of a large group, uses the method for extracting massive financial and accounting data fusion audit features of the present invention for actual application verification, including various business types such as procurement, sales, expense expenditure, and asset depreciation. The researchers first collected the complete financial and accounting data of the group in the past 24 months, with a total of 36 million records. The data includes dimension information such as transaction time, amount, business type, geographical location, and organizational structure. In step S01, a time series clustering algorithm is used to analyze the periodic characteristics of historical financial and accounting data, and the similarity distance is calculated through a dynamic time warping algorithm. Calculate according to the formula , where sequence X is the monthly income data and sequence Y is the monthly expenditure data. Through cluster analysis, it is found that the enterprise's financial and accounting data shows obvious quarterly periodicity, so the first time window length is determined to be 90 days. The coefficient of variation calculation result is , indicating that the data fluctuation is moderate.

[0077] During the construction of the self-varying time window in step S02, the researchers set the business fluctuation coefficient parameter. Through the formula The calculated business fluctuation coefficient is 1.23, where , , adjustment factor . The time granularity parameter is set to 1 day, the data change rate threshold is set to 8%, and the window overlap is set to 0.3. Through the data change rate calculation formula , 126 significant change time nodes are identified. The calculation result of the self-varying time window length adjustment function shows that when the basic window length is days, the adjusted window length range is 22 to 38 days, effectively adapting to the dynamic change characteristics of the financial and accounting data. As shown in Table 1.

[0078] Table 1 Statistical table of self-varying time window sequence

[0079] In step S03, the researchers use the financial and accounting time series feature recognition model to process the self-varying time window sequence. This model is designed based on the LSTM architecture, contains 3 layers of LSTM units, the hidden state dimension is set to 128, the sequence length parameter is set to 60, and the temperature coefficient is set to 0.8. The model training data set contains 180,000 labeled samples, of which the normal mode samples account for 73% and the abnormal mode samples account for 27%. Through deep learning processing, the time dimension feature vector is successfully extracted, with a dimension of 64, including 32-dimensional trend features, 20-dimensional periodic features, and 12-dimensional seasonal features.

[0080] In the spatial dimension feature extraction in step S04, based on the geographical distribution information and organizational structure relationship of 23 branches, the Pearson correlation coefficient is calculated. Through the formula The calculated correlation coefficient of financial and accounting activities between geographical regions is 0.67. The calculation result of the geographical distribution correlation indicates that there is a strong correlation between financial and accounting activities in different regions. The calculation result of the organizational correlation shows that the average correlation between the headquarters and each branch is 0.58, and the average correlation between branches is 0.34. As shown in Table 2.

[0081] Table 2 Statistical table of spatial dimension features

[0082] In step S05, the business dimension features are extracted according to the business type and transaction attributes of the financial and accounting data. The enterprise's financial and accounting data involves 34% of procurement business, 28% of sales business, 21% of expense expenditure, and 17% of asset management. Through the business association rule confidence calculation formula , it is found that the correlation confidence between procurement and expense expenditure is 0.73, and the correlation confidence between sales and accounts receivable is 0.81. The calculation results of the business type feature weights show that the weight of the sales business is 0.295, the weight of the procurement business is 0.268, the weight of the expense expenditure is 0.237, and the weight of the asset management is 0.200.

[0083] In the outlier feature extraction of step S06, the researchers used three methods, namely kernel density estimation, Mahalanobis distance, and local outlier factor, to identify abnormal data points. The calculation formula of kernel density estimation sets the bandwidth parameter h to 0.2 and the Gaussian kernel function parameter to the standard normal distribution. The calculation results of the Mahalanobis distance show that a total of 2,847 abnormal data points are identified, accounting for 0.08% of the total data volume. In the calculation of the local outlier factor, the k-nearest neighbor parameter is set to 10, and 3,156 abnormal data points are identified. As shown in Table 3.

[0084] Table 3 Statistical table of outlier features

[0085] In step S07, a second time window is established for verification. The window length is set to 120 days, forming a comparison with the first time window. After repeating the feature extraction process, a verification feature set is obtained, including 68-dimensional features in the time dimension, 45-dimensional features in the spatial dimension, 52-dimensional features in the business dimension, and 36-dimensional outlier features. The cross-validation results show that the consistency of the feature extraction method at different time scales reaches 91.3%.

[0086] In the determination of the multi-objective optimization weights in step S08, a multi-objective genetic algorithm is used to solve the optimal weight coefficients. The calculation formula of the fitness function sets the weight coefficients of the objective functions to , , , . After 500 generations of evolution, the optimal weight coefficient combination is obtained: the weight of the time dimension , the weight of the spatial dimension , the weight of the business dimension , the weight of the outlier . As shown in Table 4.

[0087] Table 4 Statistical table of multi-objective optimization results

[0088] In step S09, the optimal weight coefficients are used for feature weighted fusion. According to the formula , a comprehensive audit feature vector is obtained, with a total dimension of 201. The fused feature vector effectively integrates multi-dimensional information such as time trend, spatial distribution, business association, and abnormal behavior.

[0089] In step S10, through the feature density calculation function The final feature density is calculated to be 0.724, indicating that the comprehensive audit feature vector contains rich and effective information. The output comprehensive audit feature vector includes 64-dimensional time trend features, 58-dimensional periodic pattern features, 43-dimensional abnormal behavior features, and 36-dimensional multi-dimensional association features.

[0090] Traditional financial and accounting audit feature extraction mainly uses single-dimensional statistical analysis methods, such as rule-based anomaly detection, simple time series analysis, and independent spatial distribution statistics. These traditional methods have problems such as incomplete feature extraction, insufficient multi-dimensional information fusion ability, and poor adaptability to dynamic changes. By introducing technical means such as self-varying time windows, multi-dimensional feature fusion, and multi-objective optimization weight determination, the present invention has improved the audit accuracy by 8.9%, the recall rate by 9.7%, the F1 score by 9.6%, and the feature density by 12.6% compared with traditional methods. The present invention can more comprehensively capture complex patterns and potential anomalies in financial and accounting data, improve the accuracy and reliability of audit feature extraction, and provide effective technical support for the intelligent audit of massive financial and accounting data.

[0091] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention.

Claims

1. A method for extracting audit features by fusing a large amount of accounting data, characterized in that, Including: Determine the length of the first time window using a time series clustering algorithm, and divide historical financial and accounting data into multiple first time windows based on the periodic characteristics of the financial and accounting data; Construct a self-varying time window function, input the financial and accounting data within the first time window, business fluctuation coefficient, time granularity parameter, data change rate threshold, and window overlap degree, and output a self-varying time window sequence for dynamically adjusting the window boundary to adapt to the internal change characteristics of the financial and accounting data; Use a pre-trained financial and accounting time series feature recognition model to perform deep learning processing on the self-varying time window sequence, identify long-term trend patterns and periodic change rules in the financial and accounting data, and output time dimension features; Extract spatial dimension features according to the geographical distribution information and organizational structure information of the financial and accounting data in the self-varying time window sequence; Extract business dimension features based on the business types and transaction attributes of the financial and accounting data in the self-varying time window sequence; Calculate the outlier degree of the financial and accounting data within the self-varying time window sequence, identify the outlier points corresponding to abnormal financial and accounting behaviors through statistical analysis methods, and form outlier degree features; Construct a multi-objective optimization weight determination model, use the time dimension features, spatial dimension features, business dimension features, and outlier degree features as input variables, and solve the optimal weight coefficients through a multi-objective optimization algorithm; Use the optimal weight coefficients to perform weighted fusion on the time dimension features, spatial dimension features, business dimension features, and outlier degree features to form a comprehensive audit feature vector.

2. The method for extracting the characteristics of the massive financial and accounting data fusion audit according to claim 1, wherein The first time window is specifically a fixed time period determined according to the natural periodicity of the financial and accounting data. Its length is automatically determined by analyzing the similarity patterns of historical financial and accounting data through a clustering algorithm, and is used to capture the basic periodic characteristics of the financial and accounting data. It usually corresponds to the standard reporting period or settlement period of financial and accounting operations, providing a stable time benchmark framework for subsequent feature extraction.

3. The method for extracting massive financial and accounting data fusion audit features according to claim 2, wherein, The self-varying time window sequence is specifically a set of variable-length time periods further subdivided on the basis of the first time window, and the window boundary is adaptively adjusted according to the internal change rate of the financial and accounting data, and is used to accurately locate the time nodes of abnormal changes.

4. The method for extracting massive financial and accounting data fusion audit features according to claim 3, wherein The self-varying time window function is used to dynamically adjust the boundary position and length size of the time window according to the real-time change characteristics of the financial and accounting data to adapt to the fluctuation patterns and change rules of the financial and accounting data in different business scenarios. The output is a self-varying time window sequence after adaptive adjustment, and each window is marked with the start time point, end time point, and statistical characteristics of the data within the window.

5. The method for extracting massive financial and accounting data fusion audit features according to claim 4, wherein, The business fluctuation coefficient is specifically a quantitative indicator reflecting the change in the intensity of financial and accounting business activities, which is calculated by analyzing the change amplitude and frequency characteristics of the business volume in historical financial and accounting data, and is used to guide the dynamic adjustment process of the self-varying time window function.

6. The method for extracting massive financial and accounting data fusion audit features according to claim 5, wherein, The specific structure of the financial and accounting time series feature recognition model is a deep neural network designed based on the long short-term memory network architecture, including multiple LSTM units for processing the long-term dependence relationship of time series data. The entire network structure enhances the recognition ability of key time nodes and important financial and accounting events through residual connections and attention mechanisms.

7. The method for extracting massive financial and accounting data fusion audit features according to claim 6, wherein The multi-objective optimization weight determination model is used to solve the optimal weight allocation of time dimension features, space dimension features, business dimension features, and outlier degree features in the comprehensive audit feature vector. The multi-objective genetic algorithm is adopted as the core optimization engine, and the audit accuracy rate, recall rate, F1 score, and feature density after feature fusion are used as the optimization objective functions.

8. The method for extracting massive financial and accounting data fusion audit features according to claim 7, wherein After the step of constructing the self-varying time window function, it further includes establishing a second time window for verifying the feature extraction effect. The length of the second time window is different from that of the first time window. The steps of using the financial and accounting time series feature recognition model to extract time dimension features to form outlier degree features are repeatedly executed to obtain a verification feature set.

9. A computer-readable storage medium, characterized in that, Program instructions are stored in the computer-readable storage medium. When the program instructions run on a computer, they are used to execute the method for extracting comprehensive audit features from massive financial and accounting data according to any one of claims 1-8.

10. A system for extracting audit features by fusing a large amount of financial and accounting data, characterized in that, It includes the computer-readable storage medium according to claim 9. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

Citation Information

Patent Citations

  • Time sequence prediction method and system based on multi-scale fusion and attention mechanism

    CN115907203A

  • Multi-dimensional time series data spatio-temporal feature extraction method based on attention mechanism

    CN118606682A

  • Financial data anomaly detection method and system based on artificial intelligence

    CN118673430A

  • Power system load prediction method based on multivariate data fusion

    CN119377809A

  • Dynamically centered setup-time and hold-time window

    US20030223278A1

Cited By

  • Financial voucher generation method, medium and system based on ocean monitoring platform

    CN120672499A