Financial data anomaly analysis and intelligent processing method based on deep learning
By aligning fields and projecting differences in multi-source financial data, and combining cross-path comparison and joint verification with financial business paths, the problem of insufficient cross-platform collaborative data alignment in existing technologies is solved, and efficient and accurate detection and location of financial data anomalies are achieved.
Patent Information
- Application Number
- CN202511321172.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for detecting anomalies in financial data lack cross-platform and cross-path collaborative data consistency alignment and joint modeling, resulting in insufficient anomaly localization accuracy and a lack of effective characterization of the direction and magnitude characteristics of differences between multi-source financial data, making it difficult to achieve efficient anomaly identification and interpretability.
By aligning fields of multi-source data, a collaboratively aligned multi-source data sequence is generated, forming a collaborative difference projection result. The data is then segmented along the time series dimension to generate trend envelope boundaries, merge abnormal segments, and perform cross-path comparison and joint verification in conjunction with financial business paths, outputting anomaly location results.
It significantly improves the comprehensiveness and accuracy of anomaly detection, reduces the false alarm rate, enhances the stability and interpretability of detection results, and facilitates subsequent manual review and decision-making.
Smart Images

Figure CN120823069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a method for analyzing and intelligently processing financial data anomaly based on deep learning. Background Art
[0002] With the development of the digital economy, financial information systems are increasingly becoming multi-source, real-time, and intelligent. Traditional financial data processing relies on manual review and rule-setting, and is suitable for scenarios with a high degree of structure and relatively small data volumes. However, with the increasing complexity of corporate financial activities and the convergence of multi-source data across systems and institutions, problems such as time series inconsistencies, numerical discrepancies, and path anomalies are inevitable during the recording, transmission, and aggregation of financial data. To address these challenges, researchers are using machine learning and deep learning methods to assist in anomaly detection and automated analysis of financial data. Deep learning technology demonstrates significant potential in high-dimensional feature extraction, temporal regularity modeling, and anomaly pattern recognition, potentially improving the efficiency of financial risk control and intelligent auditing.
[0003] However, existing technologies still have many shortcomings. First, most financial anomaly detection methods only target a single data source and lack the ability to consistently align and jointly model collaborative data across platforms and paths, resulting in insufficient anomaly location accuracy. Second, existing methods often rely on static thresholds or simple statistical models when modeling time series, making it difficult to effectively characterize the direction and amplitude characteristics of differences between multi-source financial data, and are prone to false positives or omissions. Third, while some deep learning methods can detect abnormal patterns in large-scale data, they lack a mechanism for linking detection results with actual financial business paths (such as accounting paths, bank confirmation paths, and report summary paths), thereby reducing the interpretability and practicality of anomaly identification. Summary of the Invention
[0004] In view of the problems existing in the existing financial data anomaly analysis technology, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is how to improve the accuracy and positioning capability of anomaly detection.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a method for anomaly analysis and intelligent processing of financial data based on deep learning, which includes: aligning the fields of collected multi-source data according to a unified field mapping relationship to generate a collaboratively aligned multi-source data sequence; encoding the amplitude and direction of the differences between the sources in the multi-source data sequence to form a collaborative difference projection result; segmenting the collaborative difference projection result in the time series dimension according to a preset window length and step size to generate a trend envelope boundary, and merging the continuous windows that break through the trend envelope boundary to obtain an abnormal fragment set; extracting corresponding values from the abnormal fragment set in the original accounting path, bank confirmation path and report summary path to form a cross-path comparison set; performing direction matching on the numerical difference sets of the three paths in the cross-path comparison set, and jointly verifying them with the collaborative difference projection results, and incorporating abnormal fragments with consistent directions into a candidate synthetic set; aggregating the candidate synthetic set according to multiple link relationships, and performing a joint decision based on the collaborative difference projection results, the numerical difference sets and the link relationships to output the abnormal positioning results.
[0007] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the generating of a collaboratively aligned multi-source data sequence includes: aligning the enterprise resource management system data, bank flow data, and the data collected by structured bill image processing according to a unified field mapping relationship to generate an aligned multi-source data sequence; forming a collaborative difference projection result includes: marking missing or misplaced fields in the multi-source data sequence with placeholders, and encoding them with anomaly identifiers to form an alignment sequence with identifiers; in the alignment sequence with identifiers, generating a difference vector for the difference between the three source data in chronological order, and converting them into discrete symbol codes according to the difference direction and size to form a difference trajectory; and combining the difference trajectories into a collaborative difference projection result in chronological order and field order.
[0008] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the formation of the difference trajectory includes: calculating the amplitude of the difference between each field in the multi-source data sequence at adjacent time points, and judging the direction of change; dynamically dividing the amplitude into different levels according to the standard deviation of historical data, and combining the direction and amplitude level to form a preliminary symbol pair; encoding and mapping the preliminary symbol pairs to generate a discrete symbol sequence for each field; combining the discrete symbol sequences in field order to form a cross-field difference trajectory.
[0009] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the formation of the abnormal segment set includes: segmenting the collaborative difference projection results into time series according to the preset window length and step size to obtain a difference projection segment sequence; calculating the fluctuation range within the difference projection segment, and generating the corresponding trend envelope boundary based on the fluctuation range; forming an overlapping area between the trend envelope boundaries of adjacent difference projection segments, tracking the time windows with consistent and continuous directions in the difference trajectories that continuously cross the overlapping area, and merging them to form an abnormal candidate sequence of continuous windows; merging the windows with consistent and continuous directions in the abnormal candidate sequence to output the abnormal segment set.
[0010] As a preferred solution of the method for abnormal analysis and intelligent processing of financial data based on deep learning described in the present invention, wherein: at both ends of the trend envelope boundary of adjacent difference projection segments, the extension length of the overlapping area is determined according to the following rules: the number of high-amplitude symbols in the difference projection segment is counted, and the proportion of the number of high-amplitude symbols to the total number of symbols is calculated. ; The ratio of the longest time window in which high-amplitude symbols appear continuously in the statistical difference projection segment to the total length of the segment ;Extended length The calculation includes: ; in, is the basic length; and is the empirical coefficient; is the time length of the difference projection segment; the overlapping area is the time interval after the trend envelope boundary is extended.
[0011] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the formation of the candidate synthetic set includes: in the abnormal fragment set, according to the time interval and field position, the corresponding numerical values are extracted in the original accounting path, the bank confirmation path and the report summary path respectively to obtain a cross-path comparison set; the numerical values of the three paths in the cross-path comparison set are calculated to form a numerical difference set, and a direction is added to the difference result; the difference result is matched with the collaborative difference projection result according to the direction identifier, and only when the direction remains consistent, the corresponding abnormal fragment is marked as a candidate fragment; the candidate fragments are aggregated according to time continuity and path consistency, and the candidate synthetic set is output.
[0012] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the aggregating candidate segments according to time continuity and path consistency includes: arranging each candidate segment in chronological order and calculating the time interval between adjacent candidate segments; if the time interval between adjacent candidate segments is lower than the preset continuous time threshold, then merging them into continuous aggregated segments in the time dimension; during the merging process, simultaneously checking the direction consistency and amplitude level consistency of each candidate segment in the original accounting path, bank confirmation path and report summary path, and only when the direction and amplitude level are consistent, the candidate segments are merged into continuous aggregated segments; establishing an aggregation index for the aggregated continuous aggregated segments, and recording the fields and path information involved.
[0013] As a preferred solution of the deep learning-based financial data anomaly analysis and intelligent processing method described in the present invention, the determination of the anomaly location result includes: grouping the candidate synthetic set according to the link relationship of matters, subjects, bills and transactions to obtain a link aggregation set; jointly comparing the collaborative difference projection results and the numerical difference sets of each group in the link aggregation set to generate a link comparison result; the joint comparison includes: comparing the direction and amplitude level of the collaborative difference projection results and the direction and amplitude level of the numerical difference sets; in the link comparison results, judging the candidate segments that simultaneously meet the link consistency conditions and marking them as anomaly segment location results; outputting the anomaly segment location results according to the link indexes of matters, subjects, bills and transactions to obtain anomaly location results.
[0014] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the deep learning-based financial data anomaly analysis and intelligent processing method as described in the first aspect of the present invention are implemented.
[0015] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the deep learning-based financial data anomaly analysis and intelligent processing method as described in the first aspect of the present invention are implemented.
[0016] The beneficial effects of the present invention are as follows: the present invention breaks through the limitations of traditional methods that rely on a single data source or static rules, and can uniformly process and comprehensively analyze multi-path data from internal enterprise systems, bank flows, reports, etc., thereby significantly improving the comprehensiveness and accuracy of anomaly detection; when identifying anomalies, it focuses on the dynamic relationship and trend changes between data, which can effectively reduce false alarms caused by short-term fluctuations or isolated differences, and improve the stability and reliability of detection results; and in the output of abnormal results, it fully considers the financial business logic and link relationship, so that the detection results are not only technically effective, but also have strong business interpretability and traceability, which is convenient for subsequent manual review and decision-making.
[0017] Overall, the present invention improves the intelligence, automation, and credibility of financial data anomaly analysis, providing more efficient and scientific technical support for enterprise risk management, audit supervision, and financial compliance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 This is a flowchart of the financial data anomaly analysis and intelligent processing method based on deep learning.
[0020] Figure 2 This is a flowchart of step S2 in the financial data anomaly analysis and intelligent processing method based on deep learning. DETAILED DESCRIPTION
[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0022] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0023] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0024] Figure 1 Flowchart of the method for analyzing and intelligently processing financial data anomalies based on deep learning according to an embodiment of the present invention. Figure 1 As shown in the figure, the financial data anomaly analysis and intelligent processing method based on deep learning includes: S1: Align the collected multi-source data according to the unified field mapping relationship to generate a collaboratively aligned multi-source data sequence; encode the amplitude and direction of the differences between the sources in the multi-source data sequence to form a collaborative difference projection result.
[0025] S1.1: Align the data collected from the enterprise resource management system, bank transaction data, and bill image structured processing according to a unified field mapping relationship to generate an aligned multi-source data sequence.
[0026] First, all fields in each data source are standardized, including unifying field names, normalizing data types, unifying timestamp formats, and adjusting numerical precision.
[0027] For the accounting fields in the enterprise resource management system, the transaction records in the bank flow, and the structured fields obtained by bill image processing, a unified mapping table is established according to the field type and business semantics.
[0028] During the operation, a unique identifier is defined for each field, which is used for index matching during the alignment process to ensure that multi-source data can be directly compared on the same business node.
[0029] The data alignment process uses a field mapping index scanning method to match the corresponding fields of each record in different sources one by one. If a field is missing, a placeholder mark is used in subsequent steps.
[0030] After the mapping is completed, the data from each source is arranged in chronological order to generate a multi-source alignment sequence to ensure that the fields of each record in each source are completely corresponding. This operation can achieve a unified view of data across systems, allowing the records of the same transaction in different sources to be accurately compared, forming a basic collaborative analysis data structure.
[0031] Optionally, during the alignment process, we must also consider potential time misalignments between data sources. For example, bank transactions may lag behind accounting system recordings, and document image recognition may experience scanning delays. During this process, we correct and standardize the time fields, mapping data at different times to a unified time window and annotating the corresponding time segment indexes in the sequence. This time correction method ensures that the aligned multi-source data sequences are completely consistent in the temporal dimension.
[0032] S1.2: Placeholders are used to identify missing or misplaced fields in multi-source data sequences and encode them with exception identifiers to form alignment sequences with identifiers.
[0033] First, missing fields or records with misplaced field order are identified in the aligned multi-source data sequences. Placeholders are inserted into the sequences for missing fields and marked as anomaly identifiers to distinguish between normal and missing values. Placeholders are encoded using a unified symbol, such as "N / A" or "-1," and this is maintained consistently in subsequent differential trajectory calculations.
[0034] The placeholder identifier is used to ensure that the alignment sequence field is complete and to prevent gaps or mismatches in subsequent symbol encoding.
[0035] Secondly, the data is reordered according to the time index and field mapping index, the misplaced fields are adjusted to the correct time series position, and the anomaly identifier is used to distinguish the original misplaced data.
[0036] The above operations not only ensure the integrity of the data sequence but also provide a closed-loop data foundation for generating difference vectors and discrete symbol encoding. By processing placeholders and exception identifiers, an alignment sequence with identifiers is formed, allowing subsequent symbol encoding to accurately represent the magnitude and direction of the difference, while ensuring consistency across multi-source data in both time and field dimensions.
[0037] S1.3: In the aligned sequence with identifiers, the differences between the three source data are used to generate difference vectors in chronological order, and are converted into discretized symbol codes according to the difference direction and magnitude to form difference trajectories.
[0038] The formation of the difference trajectory includes the following steps: First, the amplitude of the difference between each field in the multi-source data sequence at adjacent time points is calculated, and the direction of change is determined; the amplitude is dynamically divided into different levels according to the standard deviation of historical data, and the direction and amplitude level are combined to form a preliminary symbol pair.
[0039] Specifically, the difference between the values of each field at consecutive time points is taken to calculate the amplitude and direction of change. Amplitude is calculated using absolute or percentage changes, and is dynamically categorized as high, medium, or low based on the standard deviation of historical data. Direction is determined by increasing or decreasing the value to determine whether it is positive, negative, or stable. For example, in an aligned time series, a field takes on the following values at five consecutive time points: 1000, 1100, 1250, 1240, and 1245. The difference between these adjacent time points yields +100, +150, -10, and +5. The difference amplitudes are expressed in absolute values, resulting in {100, 150, 10, 5}. Assuming the standard deviation of this field in the past year's historical data is 50, the dynamic classification rule states: high amplitude can be set to be greater than twice the standard deviation; medium amplitude can be set to be less than twice the standard deviation and greater than or equal to the standard deviation; and low amplitude can be set to be less than the standard deviation. Based on this classification, the above differences can be classified as high amplitude, high amplitude, low amplitude, and low amplitude. The difference directions are: positive, positive, negative, and positive. The amplitude level and direction are combined to form a symbol code.
[0040] Each difference vector generated in the operation contains the magnitude level, direction information and time index, forming a preliminary symbol pair.
[0041] Secondly, the preliminary symbol pairs are encoded and mapped. During the operation, different amplitude levels and directions are combined into discrete symbols. For example, positive increasing high amplitude is mapped to P3, positive increasing medium amplitude is mapped to P2, negative decreasing low amplitude is mapped to N1, and stable and unchanged mapping is Z0, generating a discrete symbol sequence for each field.
[0042] S1.4: Combine the difference traces in time order and field order into the collaborative difference projection results.
[0043] S2: Segment the collaborative difference projection results in the time series dimension according to the preset window length and step size to generate the trend envelope boundary, and merge the continuous windows that break through the trend envelope boundary to obtain a set of abnormal segments.
[0044] like Figure 2 As shown, the specific steps of S2 are as follows: S2.1: Segment the collaborative difference projection results into time series according to the preset window length and step size to obtain a difference projection segment sequence.
[0045] In the collaborative difference projection results, the present invention divides the data sequence according to a preset window length, and controls the offset of adjacent segments with a step size to achieve sliding coverage of the time series.
[0046] Specifically, each segment independently computes local amplitude and direction statistics while preserving the time index mapping between consecutive segments.
[0047] Each segment contains a fixed number of time points. The window length can be determined based on the data collection frequency and financial transaction cycle. For example, it can be set to a data segment of one day or one hour. The step size can be set to 50% to 70% of the window length (this is only an example and needs to be set according to actual conditions) to achieve partial overlap between segments to capture cross-window anomalies.
[0048] Through the above operations, a complete difference projection segment sequence is formed. Each segment in the sequence contains a local difference trajectory. The present invention decomposes the continuous time series into quantifiable and computable segment units, making the abnormal features measurable within the local time period while maintaining cross-segment correlation.
[0049] During the segmentation process, the start and end time indexes of each time slice's boundaries must be recorded to facilitate subsequent envelope boundary generation and precise location of overlapping regions. The difference trajectory within each segment contains the aforementioned difference symbol sequence and amplitude level information. By closed-loop association of the time indexes, cross-segment anomaly window tracking can be implemented in subsequent steps.
[0050] S2.2: Calculate the fluctuation range within the difference projection segment, and generate the corresponding trend envelope boundary based on the fluctuation range.
[0051] Among them, the fluctuation range includes indicators such as maximum amplitude, minimum amplitude and sign change density.
[0052] Specifically, the discrete symbol sequence at each time point in the segment is used to count the number of high-amplitude symbols, low-amplitude symbols and direction changes, and the segment local fluctuation characteristics are defined.
[0053] The fluctuation range is used to generate the trend envelope boundary. The boundary is defined as the upper and lower limits of the difference amplitude within the segment. In actual operation, a certain margin can be added to the boundary, for example, adding 5% to 10% based on the historical average fluctuation amplitude to ensure that the envelope boundary can accommodate normal fluctuations without misjudging abnormalities.
[0054] The trend envelope boundary is the start and end time index of the corresponding segment in the time dimension, and the amplitude dimension is the difference amplitude range within the segment, plus the upper and lower bounds after the margin to ensure the clarity and operability of the boundary.
[0055] During the trend envelope boundary generation process, the number of high-amplitude symbols and the distribution of amplitude levels within each segment must be recorded for subsequent adaptive extension calculations of overlapping regions. By quantitatively analyzing the fluctuation range of the differential trajectory within a segment, it is possible to identify time segments near the anomaly threshold, providing a computational basis for tracking.
[0056] S2.3: An overlapping region is formed between the trend envelope boundaries of adjacent difference projection segments, and time windows with consistent directions and continuous in the difference trajectories that continuously cross the overlapping region are tracked and merged to form an abnormal candidate sequence of continuous windows.
[0057] Preferably, at both ends of the trend envelope boundary of adjacent difference projection segments, the extension length of the overlapping area is determined according to the following rules: Count the number of high-amplitude symbols in the difference projection segment and calculate the ratio of the number of high-amplitude symbols to the total number of symbols ; The ratio of the longest time window in which high-amplitude symbols appear continuously in the statistical difference projection segment to the total length of the segment ; Extended length The calculation includes: ; in, As the basic length, it can be 10% of the segment time length; and It is an empirical coefficient, which can be taken as 0.2~0.5; is the time length of the difference projection segment; the obtained extended length is extended at both ends of the trend envelope boundary to form an adaptive overlapping area. The time range of the overlapping area is the start and end indexes of the time interval after the trend envelope boundary is extended.
[0058] Furthermore, in the overlapping area, tracking the difference trajectories that continuously cross the area includes the following steps: First, scan the difference trajectory in chronological order, find the time windows with consistent sign and direction and continuous sequence, and record the start and end time indexes and amplitude levels; For continuous sequences whose symbol directions are consistent within the window and whose amplitudes reach high or medium levels, they are determined to be abnormal candidate windows; During the tracking process, it is also necessary to determine the continuity across segments. If a continuous window spans multiple segments, the time indexes are concatenated in sequence to form an abnormal candidate sequence for the continuous time window.
[0059] Through the above method, the anomaly window can be guaranteed to be coherent in the time dimension, while avoiding the boundary break error caused by segmentation.
[0060] S2.4: Merge the windows with consistent directions and continuity in the abnormal candidate sequence and output the abnormal segment set.
[0061] Specifically, all abnormal candidate windows are scanned in time index order. If the time interval between adjacent abnormal candidate windows is lower than the preset continuous time threshold, they are merged into a single abnormal segment in the time dimension. During the merging process, the amplitude level, direction information and field index of each window are retained to form an abnormal segment data structure.
[0062] After the merge, the abnormal fragments need to be indexed and sorted, and a time index and field index mapping table needs to be established to ensure that each abnormal fragment can be traced back to the original multi-source data sequence and difference trajectory.
[0063] S3: For the set of abnormal fragments, the corresponding values are extracted from the original accounting path, bank confirmation path and report summary path respectively to form a cross-path comparison set; the direction of the numerical difference sets of the three paths in the cross-path comparison set is matched, and jointly verified with the collaborative difference projection results, and the abnormal fragments with consistent directions are included in the candidate synthetic set.
[0064] S3.1: In the abnormal fragment set, according to the time interval and field position, extract the corresponding values in the original accounting path, bank confirmation path and report summary path respectively to obtain a cross-path comparison set.
[0065] Specifically, each abnormal segment is scanned in chronological order, and the start and end time indexes are mapped to the corresponding time points in the original accounting path, bank confirmation path, and report summary path, and the actual values of the corresponding fields are extracted from each path.
[0066] The extracted actual values are arranged in the order of the anomaly segments to form a cross-path comparison set. Each anomaly segment corresponds to a row in the set, and the values for each path correspond to a column vector. This ensures that the multipath data for each anomaly segment is stored in a closed loop, while also preserving the field index, time index, and amplitude level information.
[0067] Furthermore, the extraction process must consider the case where the abnormal fragment spans multiple time windows. If the abnormal fragment covers multiple segments or time windows, it should be extracted step by step in the order of consecutive time indexes to ensure the continuity and integrity of the cross-segment data. During the extraction process, missing fields must be marked and filled to maintain a closed-loop matrix structure, ensuring that each abnormal fragment forms a complete data row in the cross-path comparison set.
[0068] S3.2: Perform a difference calculation on the values of the three paths in the cross-path comparison set to form a numerical difference set, and add a direction to the difference result.
[0069] In this embodiment of the present invention, using the original accounting path as a reference, the numerical differences between the corresponding fields in the bank confirmation path and the report summary path are calculated, generating a difference matrix. The difference result records the difference between each field in each path for each abnormal segment, along with a difference direction. The difference direction is defined based on the sign of the difference: if the reference path value is smaller than the comparison path, the difference direction is increasing; otherwise, it is decreasing. A zero difference indicates stability. This operation ensures that the difference result for each abnormal segment not only quantifies the numerical difference between the paths but also preserves the directional information.
[0070] Furthermore, the absolute values of the differences are classified into high, medium, and low amplitude categories based on the standard deviation of the historical data. The difference direction and amplitude level are combined to generate a difference symbol. Through this operation, each abnormal segment contains complete difference direction and amplitude level information in the difference matrix, achieving comparability with the collaborative difference projection symbol encoding.
[0071] S3.3: Match the difference set results with the collaborative difference projection results according to the direction identification. Only when the direction remains consistent, mark the corresponding abnormal segment as a candidate segment; aggregate the candidate segments according to time continuity and path consistency, and output the candidate synthesis set.
[0072] First, the difference set direction and the collaborative difference projection direction are compared for each abnormal segment one by one. If the direction remains consistent and the amplitude level matches, the corresponding abnormal segment is marked as a candidate segment, and the abnormal segments with inconsistent directions or mismatched amplitude levels are eliminated.
[0073] Furthermore, the candidate segments are aggregated according to time continuity and path consistency, including: Arrange each candidate segment in chronological order and calculate the time interval between adjacent candidate segments; if the time interval between adjacent candidate segments is lower than the preset continuous time threshold, they are merged into a continuous aggregate segment in the time dimension; during the merging process, check the direction consistency and amplitude level consistency of each candidate segment on the original accounting path, bank confirmation path and report summary path at the same time. Only when the direction and amplitude level are consistent, the candidate segments are merged into a continuous aggregate segment; establish an aggregation index for the aggregated continuous aggregate segment, and record the fields and path information involved.
[0074] S4: Aggregate the candidate synthesis set according to the multiple link relationships, and make a joint decision based on the collaborative difference projection results, numerical difference sets and link relationships to output the anomaly positioning results.
[0075] S4.1: Group the candidate synthesis sets according to the link relationships among matters, subjects, bills, and flows to obtain a link aggregation set.
[0076] During the operation, the link attributes of the fields of each candidate segment in the candidate synthesis set are first calibrated.
[0077] The link attributes include a matter identifier (such as a contract or project number), a subject identifier (such as an accounting subject code), a bill identifier (such as a bill number), and a transaction identifier (such as a bank transaction number).
[0078] After the calibration is completed, a group index table is established, and candidate segments with the same matter, the same subject, the same bill or the same flow attribute are classified into corresponding groups to form a link aggregation set.
[0079] It should be noted that this invention expands the candidate composite set from the time and path dimensions to the business semantic dimension, allowing the comparison of abnormal fragments not only to remain at the numerical and directional level, but also to be organized in the context of the actual financial link. By grouping the link aggregation set, the fragments within each group have a high degree of correlation in terms of link attributes, facilitating subsequent joint comparison and consistency determination. If a candidate fragment lacks a certain type of link attribute, a placeholder mark is added in the corresponding group to ensure the integrity of the link aggregation set.
[0080] S4.2: Perform a joint comparison on the collaborative difference projection results and the numerical difference sets of each group in the link aggregation set to generate a link comparison result.
[0081] In this embodiment of the present invention, the joint comparison includes two dimensions: one is the direction sign and amplitude level of the collaborative difference projection result; the other is the direction sign and amplitude level of the numerical difference set. The specific operation is as follows: First, for each candidate segment in each group, the corresponding collaborative difference projection sequence is extracted and compared with the numerical difference sequence within the group, field by field and time point by time point. If the two are consistent in direction and the amplitude difference does not exceed the preset level threshold, the corresponding field comparison result is marked as consistent; otherwise, it is marked as inconsistent. The comparison results are stored in matrix form, with each row corresponding to a candidate segment and each column corresponding to the comparison identifier of a field symbol.
[0082] This comparison process can realize the mutual verification of collaborative difference projection results and numerical difference sets. The collaborative difference projection results can reflect the collaborative anomaly characteristics of the three-source data, while the numerical difference sets directly reflect the numerical deviations between different paths. Through joint comparison, the deviations caused by single-dimensional judgment can be avoided and the accuracy of abnormal judgment of candidate fragments can be improved.
[0083] S4.3: In the link comparison results, the candidate segments that meet the link consistency conditions are judged and marked as abnormal segment positioning results.
[0084] The link consistency conditions include the following rules: First, candidate segments within the same group are marked as consistent in the directional comparison of the collaborative difference projection results and the numerical difference set. Second, the segments have the same directional signs and amplitude levels in the event link, account link, bill link, and flow link. Third, the time interval between adjacent segments in terms of temporal index continuity does not exceed the continuity threshold. Segments that meet these three conditions are marked as abnormal segment positioning results.
[0085] As can be seen, this invention achieves precise screening of anomalous segments from candidate status to location status. Only segments that simultaneously meet consistency conditions at the numerical, symbolic, and link levels are identified as anomalous, ensuring the reliability and interpretability of the final location results. This adjudication process effectively reduces false positives, prevents inconsistent anomalous segments from entering the final result, and thus improves the accuracy and reliability of anomaly location.
[0086] S4.4: Output the abnormal fragment location results according to the link index of the matter, subject, bill and flow to obtain the abnormal location results.
[0087] The output format is a link index table, where each row corresponds to an abnormal fragment positioning result, and each column corresponds to link attributes and symbol information. The output results can be directly connected to the enterprise financial management system to achieve automated abnormality labeling and traceability.
[0088] It should be noted that through the link-indexed output format, each abnormal result can clearly point to a specific matter, subject, bill and flow. Financial personnel can quickly locate the link where the abnormality occurs and realize intelligent financial risk warning and management.
[0089] This embodiment also provides a computer device suitable for the case of anomaly analysis and intelligent processing method of financial data based on deep learning, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the anomaly analysis and intelligent processing method of financial data based on deep learning proposed in the above embodiment.
[0090] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.
[0091] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for analyzing and intelligently processing financial data anomalies based on deep learning as proposed in the above embodiment is implemented.
[0092] In summary, the present invention breaks through the limitations of traditional methods that rely on a single data source or static rules, and can uniformly process and comprehensively analyze multi-path data from internal enterprise systems, bank statements, reports, etc., thereby significantly improving the comprehensiveness and accuracy of anomaly detection; the present invention focuses on the dynamic relationship and trend changes between data when identifying anomalies, which can effectively reduce false alarms caused by short-term fluctuations or isolated differences, and improve the stability and reliability of detection results; the present invention fully considers the financial business logic and link relationship in the output of abnormal results, so that the detection results are not only technically effective, but also have strong business interpretability and traceability, which is convenient for subsequent manual review and decision-making.
[0093] Overall, the present invention improves the intelligence, automation, and credibility of financial data anomaly analysis, providing more efficient and scientific technical support for enterprise risk management, audit supervision, and financial compliance.
[0094] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for analyzing and intelligently processing financial data anomalies based on deep learning, characterized by: include: Align the collected multi-source data according to the unified field mapping relationship to generate a collaboratively aligned multi-source data sequence; Encoding the differences between the sources in the multi-source data sequence in terms of amplitude and direction to form a collaborative difference projection result; In the time series dimension, the collaborative difference projection results are segmented according to the preset window length and step size to generate the trend envelope boundary. The continuous windows that break through the trend envelope boundary are merged to obtain the abnormal segment set. For the set of abnormal fragments, the corresponding values are extracted from the original accounting path, the bank confirmation path, and the report summary path to form a cross-path comparison set. The direction of the numerical difference sets of the three paths in the cross-path comparison set is matched and jointly verified with the collaborative difference projection results. The abnormal fragments with consistent directions are included in the candidate synthesis set. The candidate synthesis set is aggregated according to multiple link relationships, and a joint decision is made based on the collaborative difference projection results, numerical difference sets and link relationships to output the anomaly positioning results.
2. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 1, characterized in that: Generating the collaboratively aligned multi-source data sequences comprises: Align the data collected from the enterprise resource management system, bank transaction data, and bill image structured processing according to a unified field mapping relationship to generate an aligned multi-source data sequence; The forming of the collaborative difference projection result includes: Marking missing or misplaced fields in the multi-source data sequence as placeholders and encoding them with exception identifiers to form an alignment sequence with identifiers; In the alignment sequence with identifiers, the difference values of multi-source data are generated into difference vectors in time sequence, and converted into discretized symbol codes according to the difference direction and magnitude to form difference trajectories; The difference traces are combined into co-difference projection results in time order and field order.
3. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 2, characterized in that: The formation of the difference trajectory includes: Calculate the amplitude of the difference between adjacent time points for each field in the multi-source data series and determine the direction of change; dynamically divide the amplitude into different levels according to the standard deviation of historical data, and combine the direction and amplitude level to form a preliminary symbol pair; Perform encoding mapping on the preliminary symbol pairs to generate a discrete symbol sequence for each field; Discrete symbol sequences are combined in field order to form difference traces across fields.
4. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 3, characterized in that: The formation of the abnormal fragment set includes: The collaborative difference projection results are segmented into time series according to the preset window length and step size to obtain a difference projection segment sequence; Calculate the fluctuation range within the difference projection segment and generate the corresponding trend envelope boundary based on the fluctuation range; An overlapping region is formed between the trend envelope boundaries of adjacent difference projection segments, and the time windows with consistent directions and continuous in the difference trajectories that continuously cross the overlapping region are tracked and merged to form an abnormal candidate sequence of the continuous windows; The windows with consistent directions and continuity in the abnormal candidate sequence are merged to output a set of abnormal fragments.
5. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 4, characterized in that: At both ends of the trend envelope boundary of adjacent difference projection segments, the extension length of the overlapping area is determined according to the following rules: Count the number of high-amplitude symbols in the difference projection segment and calculate the ratio of the number of high-amplitude symbols to the total number of symbols ; The ratio of the longest time window in which high-amplitude symbols appear continuously in the statistical difference projection segment to the total length of the segment ; Extended length The calculation includes: ; in, is the basic length; and is the empirical coefficient; Segment time length for difference projection; The overlapping area is the time interval after the trend envelope boundary is extended.
6. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 5, characterized in that: The forming of the candidate synthesis set includes: In the set of abnormal fragments, the corresponding values are extracted from the original accounting path, bank confirmation path, and report summary path according to the time interval and field position to obtain a cross-path comparison set; Perform a difference calculation on the values of the three paths in the cross-path comparison set to form a numerical difference set, and add a direction to the difference result; Match the difference set result with the collaborative difference projection result according to the direction identification. Only when the direction is consistent, mark the corresponding abnormal segment as a candidate segment; The candidate segments are aggregated according to time continuity and path consistency, and a candidate synthesis set is output.
7. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 6, characterized in that: Aggregating the candidate segments according to time continuity and path consistency includes: Arrange the candidate segments in chronological order and calculate the time interval between adjacent candidate segments; If the time interval between adjacent candidate segments is lower than the preset continuous time threshold, they are merged into a continuous aggregate segment in the time dimension; During the merging process, the direction consistency and amplitude level consistency of each candidate segment on the original accounting path, bank confirmation path, and report summary path are simultaneously checked. Only when the direction and amplitude level are consistent, the candidate segments are merged into a continuous aggregate segment; An aggregation index is created for the continuous aggregation segments after aggregation, and the fields and path information involved are recorded.
8. The method for analyzing and intelligently processing financial data anomalies based on deep learning according to claim 7, characterized in that: Determining the abnormality positioning result includes: The candidate synthesis set is grouped according to the link relationship of matters, subjects, bills and flows to obtain a link aggregation set; Performing a joint comparison of the collaborative difference projection results and the numerical difference sets of each group in the link aggregation set to generate a link comparison result; the joint comparison includes: comparing the direction and amplitude level of the collaborative difference projection results and the direction and amplitude level of the numerical difference sets; In the link comparison results, the candidate segments that meet the link consistency conditions are judged and marked as abnormal segment positioning results; The abnormal segment locating result is output according to the link index of the matter, subject, bill and flow to obtain the abnormal locating result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the deep learning-based financial data anomaly analysis and intelligent processing method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based financial data anomaly analysis and intelligent processing method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Financial risk dynamic prediction method based on deep learning
CN120278527A
Multi-source financial data real-time fusion and anomaly monitoring method and system
CN120493194A