A Multimodal Data-Driven Method and System for Financial and Accounting Supervision and Early Warning

CN122736798APending Publication Date: 2026-09-11BEIJING ZHENGCHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610904189.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0007]本发明实施例提供了基于多模态数据驱动的财会监督预警方法及系统,解决了财会监督预警的准确性、实时性不足,难以满足财会业务系统稳定运行与风险及时感知的实际需求的技术问题

Benefits of technology

1、通过进行财会数据多尺度时序感知,并基于感知结果针对性采取财会数据弱滤波保序处理措施、财会数据自适应折中处理措施或财会数据平滑降噪处理措施,有助于在高精度识别阶梯式缓变异常、随机噪声及过渡区间的基础上,实现差异化时序数据处理,既保留真实微弱递进式异常特征不被过度平滑,又有效剔除随机抖动干扰,提升财会时序监测数据的稳定性与可信度,在财会数据多尺度时序感知结束后,进行财会多模态异构数据表征聚合,有助于充分挖掘并增强时序数值、文本语义、拓扑结构三类异构数据中的异常关联信息与递进式累积特征,实现多维度异常表征互补,财会多模态异构数据表征聚合结束后,基于聚合结果进行财会监督风险预测,有助于实现对财会潜在风险的早期识别、高精度判定与动态预警,提升财会监督的灵敏度与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736798A_ABST
    Figure CN122736798A_ABST
Patent Text Reader

Abstract

This application discloses a method and system for financial and accounting supervision and early warning based on multimodal data, relating to the fields of financial and accounting data processing and supervision and early warning technology. The specific implementation scheme is as follows: First, multi-scale time-series perception of financial and accounting data is performed, and based on the perception results, targeted measures are taken for weak filtering and order preservation, adaptive compromise processing, or smoothing and noise reduction of financial and accounting data. Next, multimodal heterogeneous data representation and aggregation are performed. Finally, financial and accounting supervision risk prediction is performed based on the aggregation results. The technology of this application solves the problem of insufficient real-time performance in financial and accounting supervision, which makes it difficult to meet the actual needs of stable operation and timely risk perception of financial and accounting business systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of accounting data processing and monitoring and early warning technology, and in particular to an accounting monitoring and early warning method and system based on multimodal data-driven approaches. Background Technology

[0002] In the current digital economy and complex business environment, corporate financial and accounting risks are becoming increasingly hidden and interconnected. Traditional supervision methods that rely on structured financial statements and rule engines are no longer sufficient to cope with new risks due to their limited data dimensions, outdated information, and inability to effectively utilize unstructured data such as text, public opinion, and relationship graphs. Therefore, conducting research on financial and accounting supervision and early warning has become an essential requirement for improving the effectiveness of financial and accounting supervision and the accuracy of risk identification.

[0003] The main implementation process of existing financial and accounting supervision and early warning methods is as follows: First, multimodal heterogeneous financial and accounting data is collected from multiple channels such as corporate financial reports, business registration information, and supply chain knowledge graphs. This includes structured data such as data processing latency and CPU usage, semi-structured data such as system log text and alarm text, and image data such as service dependencies. Next, relevant features of the multimodal financial and accounting data are extracted using feature extraction methods such as FinBERT (Financial BERT, a pre-trained language model in the financial field) and GCN (Graph Convolutional Network). Then, the relevant features of the multimodal financial and accounting data are fused using fusion algorithms such as cross-modal Transformer to obtain fused multimodal financial and accounting features. Finally, the fused multimodal financial and accounting features are input into financial and accounting risk prediction models such as random forests for risk prediction. Finally, financial and accounting supervision and early warning are triggered based on risk prediction. For example, when a continuous stepwise increase in data processing latency or a continuous accumulation of cross-module synchronization errors exceeding the early warning threshold is detected, an abnormal system operation warning and risk alert are triggered.

[0004] The above-mentioned technology has at least the following technical problems: In the process of financial and accounting supervision and early warning, the financial and accounting time-series monitoring data of the platform operation, such as data processing latency, interface call response time, service node CPU utilization, log reporting frequency, and cross-module data synchronization errors, is affected by factors such as the gradual drift of hardware sampling accuracy (e.g., CPU processor, router), the gradual accumulation of system load, the slight accumulation of network transmission jitter, and the slow expansion of clock synchronization deviation between modules. These factors typically present weak anomalies that deviate from the normal range in a continuous, small, step-like manner over multiple periods. Current technologies generally use median filtering, moving averages, or fixed threshold truncation to smooth and reduce noise in the financial and accounting time-series monitoring data, which may lead to… These early, subtle anomalies that accumulate over time are treated as random noise and removed. The true anomaly trend is forcibly pulled back to the normal range. Because the multimodal accounting data, which includes time-series monitoring data, has already shown a distribution shift at the time-series level, the overall multimodal accounting data is distorted. Consequently, when extracting features from the distorted multimodal accounting data, the extracted multimodal accounting features lack subtle progressive anomaly representations, and the time-series trend features are distorted and blurred. When fusing multimodal accounting features that lack progressive anomaly representations and have distorted and blurred time-series trend features, the key anomaly features of the fused multimodal accounting features are covered by redundant information, and the risk correlation representation between modalities is broken.

[0005] Meanwhile, when making financial and accounting risk predictions based on financial and accounting risk prediction models such as XGBoost and Random Forest, existing technologies usually make one-time probability judgments based on static features. This may lead to a lag in the model's identification of gradual and hidden operational anomalies. When the time-series anomalies are smoothed and weakened and the feature fusion weights are unbalanced, the strength of effective financial and accounting risk signals is further weakened, ultimately leading to a lag in the response of financial and accounting supervision and early warning.

[0006] In summary, existing financial and accounting supervision and early warning technologies, which employ conventional smoothing and noise reduction methods in the preprocessing stage of multimodal financial and accounting data, struggle to distinguish between random noise and continuous, small-amplitude, gradually changing anomalies. This can easily lead to the removal of weak, progressively changing anomaly signals in the financial and accounting time-series monitoring data, causing a shift in the time-series distribution and resulting in overall distortion of the multimodal financial and accounting data. Furthermore, the feature extraction stage, based on already distorted multimodal data, results in the extraction of features lacking key progressive anomaly representations and time-series trend characteristics. The feature fusion stage fails to compensate for and enhance time-series distorted features, causing key anomaly features to be covered by redundant information and disrupting the risk correlation representation between different modalities. The risk prediction stage relies solely on static features for one-time probability determination, failing to capture the evolutionary patterns of gradual and hidden time-series anomalies, and is prone to recognition lag. These problems accumulate and amplify further in scenarios with increased data sources and more complex system structures, leading to a continuous decline in the accuracy and response efficiency of the entire early warning process. Ultimately, this results in insufficient accuracy and real-time performance of financial and accounting supervision and early warning, failing to meet the actual needs of stable operation and timely risk perception in financial and accounting business systems. Summary of the Invention

[0007] This invention provides a multimodal data-driven method and system for financial and accounting supervision and early warning, which solves the technical problems of insufficient accuracy and real-time performance in financial and accounting supervision and early warning, making it difficult to meet the actual needs of stable operation and timely risk perception in financial and accounting business systems. The technical solution provided by this application is as follows: According to the first aspect of this application, a financial and accounting supervision early warning method based on multimodal data is provided. The method includes: performing multi-scale time-series perception of financial and accounting data, and taking targeted measures such as weak filtering and order preservation processing, adaptive compromise processing, or smoothing and noise reduction processing based on the perception results; after the multi-scale time-series perception of financial and accounting data is completed, performing data aggregation for multimodal heterogeneous data representation; after the aggregation of multimodal heterogeneous data representation, performing financial and accounting supervision risk prediction for continuous state tracking and hierarchical early warning based on the aggregation results.

[0008] According to another aspect of this application, a financial and accounting supervision and early warning system based on multimodal data is provided, comprising: a multi-scale time-series perception module for financial and accounting data, a multimodal heterogeneous data representation and aggregation module for financial and accounting data, and a financial and accounting supervision risk prediction module. The multi-scale time-series perception module is used to: perform multi-scale time-series perception of financial and accounting data, and based on the perception results, take targeted measures such as weak filtering and order preservation processing, adaptive compromise processing, or smoothing and noise reduction processing of financial and accounting data. The multimodal heterogeneous data representation and aggregation module is used to: perform multimodal heterogeneous data representation and aggregation after the multi-scale time-series perception of financial and accounting data is completed. The financial and accounting supervision risk prediction module is used to: perform financial and accounting supervision risk prediction based on the aggregation results after the multimodal heterogeneous data representation and aggregation is completed.

[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. By conducting multi-scale time-series perception of accounting data and taking targeted measures such as weak filtering and order preservation, adaptive compromise processing, or smoothing and noise reduction based on the perception results, it is possible to achieve differentiated time-series data processing while accurately identifying step-like slowly changing anomalies, random noise, and transition intervals. This approach preserves the true, subtle, progressive anomaly characteristics without excessive smoothing, while effectively eliminating random jitter interference, thus improving the stability and reliability of accounting time-series monitoring data. After multi-scale time-series perception of accounting data, multi-modal heterogeneous data representation aggregation helps to fully explore and enhance the anomaly correlation information and progressive cumulative features in the three types of heterogeneous data: time-series numerical data, textual semantic data, and topological data. This achieves complementary multi-dimensional anomaly representations. After the aggregation of multi-modal heterogeneous data representations, accounting supervision risk prediction is performed based on the aggregation results. This helps to achieve early identification, high-precision judgment, and dynamic early warning of potential accounting risks, thereby improving the sensitivity and reliability of accounting supervision.

[0010] 2. By inputting the time-series monitoring sequences of accounting data into a multi-scale sliding observation window, and traversing the time-series monitoring sequences within the multi-scale sliding observation window with a preset step size and progressively increasing window width, the time-series monitoring sequences of accounting data are segmented into a multi-scale local accounting observation fragment set. Based on the multi-scale local accounting observation fragment set, the cumulative amount of the same-direction shift and the variance of the same-direction shift amplitude of accounting data are obtained. This helps to capture the continuous, weak, and gradual same-direction shift characteristics and amplitude stability characteristics in the time-series monitoring sequences of accounting data from multiple time scales. It overcomes the shortcomings of existing technologies that use only a single scale analysis, which are prone to missing gradual trends, are insensitive to small anomalies, or easily misjudge normal fluctuations as anomalies. It can more accurately quantify the accumulation degree and stability of step-like gradual anomalies. The division of accounting time-series intervals based on global trend confidence helps to adaptively divide the time-series data into step-like gradual anomaly intervals, random noise intervals, and noise transition intervals, achieving accurate differentiation of time-series segments with different characteristics and reducing the loss of real anomalies or incomplete noise suppression caused by uniform filtering.

[0011] 3. When there are significant differences in the intensity of abnormal accumulation at nodes during the operation of accounting and finance business, for example, some service modules experience a continuous and progressive increase in CPU utilization due to hardware sampling accuracy drift, while other modules operate normally, or cross-module data synchronization errors accumulate periodically on a certain call chain, but adjacent chains are not abnormal. In this case, the cumulative deviation of each node from the memory value varies greatly. If a graph convolution method based on static initial weights is directly used, the edge weights of nodes with strong abnormal accumulation are the same as those of normal nodes. This causes the abnormal signal to be diluted by normal neighboring nodes during graph convolution aggregation, making it impossible to highlight the progressive accumulation characteristics of abnormal nodes in the feature vector. Therefore, an alternative method for graph structure feature extraction is needed, based on accounting and finance... The business operation topology graph is updated by recursively updating the edge weights in a time sequence to obtain the updated financial and accounting business operation topology graph. The updated financial and accounting business operation topology graph is then input into a preset graph convolution model for feature extraction to obtain an enhanced graph feature vector. This helps to dynamically adjust the topology edge weights according to the actual anomaly accumulation intensity of each node, strengthen the information transmission intensity between nodes with anomaly accumulation, suppress the dilution effect of normal nodes on anomaly signals, and highlight the weak anomaly features of local, gradual, and chain-like propagation in the financial and accounting business system. This makes the features extracted by graph convolution more focused on the anomaly accumulation area and anomaly propagation path, effectively making up for the shortcomings of traditional static weight graph convolution, which cannot adapt to the differences in node anomaly intensity and is prone to drowning out weak anomaly information.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of this application. Wherein: Figure 1 This is a flowchart of the financial and accounting supervision and early warning method based on multimodal data-driven provided in an embodiment of the present invention; Figure 2 This is a flowchart summarizing the financial and accounting supervision and early warning method based on multimodal data-driven methods provided in this embodiment of the invention. Figure 3 This is a logic diagram of the weak filtering and order preservation processing measures for accounting data in the accounting supervision and early warning method based on multimodal data provided in this embodiment of the invention. Figure 4 This is a schematic diagram of the structure of the financial and accounting supervision and early warning system based on multimodal data-driven provided in an embodiment of the present invention; Figure 5 This is a bar chart comparing the performance indicators of different preset attention fusion models of the financial supervision and early warning method based on multimodal data driven by the present invention. Detailed Implementation

[0014] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0015] Example 1, as Figure 1 The flowchart shown is for a multimodal data-driven accounting supervision and early warning method. The processing flow of this method may include the following steps: During the accounting supervision and early warning process, multi-scale time-series perception of accounting data is performed to identify step-like slowly varying anomalies and random noise with progressive unidirectional shifts across multiple sampling periods. Based on the perception results, targeted measures are taken for accounting data: weak filtering and order preservation, adaptive compromise processing, or smoothing and noise reduction. The weak filtering and order preservation measures are used to suppress minor random fluctuations without altering the trend characteristics of the step-like slowly varying anomalies. Adaptive trade-off processing measures are used to preserve weak, unidirectional trends and reduce anomaly loss or noise residue caused by overprocessing. Financial and accounting data smoothing and denoising measures are used to remove interference signals without actual anomaly significance and reduce the interference of random noise on subsequent feature extraction and fusion. After multi-scale time-series perception of financial and accounting data is completed, multi-modal heterogeneous data representation aggregation is performed to enhance the representation of weak, progressive anomalies and fuse multi-source financial and accounting data. After the aggregation of multi-modal heterogeneous data representation, financial and accounting supervision risk prediction is performed based on the aggregation results for continuous state tracking and hierarchical early warning.

[0016] It should be noted that the accounting supervision and early warning method based on multimodal data-driven approach provided in this application relies on a pre-built multimodal accounting supervision and early warning knowledge base. This knowledge base includes core judgment parameters verified by accounting supervision experts, such as preset first-scale window width, preset second-scale window width, and preset third-scale window width. Furthermore, the knowledge base also includes standardized pre-processed historical operational data of the accounting business system, covering accounting time-series monitoring data, text log data, call chain topology data, and their corresponding risk warning results under different business scenarios and anomaly types. The records also include model iteration and optimization data, such as multi-scale time-series trend perception effect evaluation. This multimodal accounting supervision and early warning knowledge base adopts a domain-based heterogeneous storage architecture. It uses relational databases to store structured indicator data, such as the time-series sequence of service node CPU utilization, and uses non-relational databases to store unstructured data, such as system log text. Operation and maintenance personnel can conduct periodic verification and dynamic iterative adjustment of various preset parameters stored in the knowledge base based on the real-time collected data on the operation monitoring of the accounting business system and the feedback on early warning effectiveness. This ensures that the knowledge base can continuously adapt to the complex and ever-changing operating conditions of the accounting business system, such as peak periods of financial settlement.

[0017] As described above, multi-scale time-series perception of accounting data, aggregation of multi-modal heterogeneous accounting data representations, and accounting supervision risk prediction help to build a full-process accounting supervision system from data preprocessing and feature enhancement to risk assessment. Specifically, multi-scale time-series perception of accounting data provides a high-quality, order-preserving time-series data foundation for subsequent feature extraction; aggregation of multi-modal heterogeneous accounting data representations provides highly discriminative joint features for risk prediction; and accounting supervision risk prediction enables early identification and dynamic tracking of potential risks. These three aspects progress sequentially and support each other, improving the comprehensiveness of accounting anomaly monitoring and the accuracy of risk warning.

[0018] like Figure 2The flowchart illustrating the overall overview of the multimodal data-driven accounting supervision and early warning method shows that: Multi-scale time-series perception of accounting data is performed, and the global trend confidence level is obtained. Time-series intervals with a global trend confidence level not lower than a preset first threshold are marked as stepped, gradually changing anomaly intervals. Based on these stepped, gradually changing anomaly intervals, weak filtering and order-preserving processing measures are implemented for the accounting data. Time-series intervals with a global trend confidence level not higher than a preset second threshold are marked as random noise intervals. Based on these random noise intervals, accounting data smoothing and noise reduction processing measures are implemented for the accounting data. Time-series intervals with a global trend confidence level between the preset first and second thresholds are marked as noise transition intervals. Based on these noise transition intervals, adaptive compromise processing of the accounting data is implemented. After the multi-scale time-series perception of financial and accounting data is completed, multi-modal heterogeneous financial and accounting data representation and aggregation are performed to obtain enhanced multi-modal feature vectors. Based on the enhanced multi-modal feature vectors, financial and accounting supervision risks are predicted, and financial and accounting risk probability values ​​are obtained. The financial and accounting risk probability value at the current moment is compared with the preset early warning threshold. If the financial and accounting risk probability value at the current moment is greater than the preset first-level early warning threshold, a first-level financial and accounting supervision prompt is triggered. If the financial and accounting risk probability value at the current moment is not greater than the preset first-level early warning threshold and not less than the preset second-level early warning threshold, a second-level financial and accounting supervision prompt is triggered. If the financial and accounting risk probability value at the current moment is less than the preset second-level early warning threshold, the risk probability value is recorded in the log for subsequent audit analysis.

[0019] As a preferred technical means, the specific process of multi-scale time-series perception of accounting data is as follows: The accounting data time-series monitoring sequence, such as any one of the following sequences (e.g., data processing delay time-series sequence, service node CPU utilization time-series sequence, cross-module data synchronization error time-series sequence), is input into a multi-scale sliding observation window. The accounting data time-series monitoring sequence is traversed within the multi-scale sliding observation window with a preset step size, and the window width increases sequentially. This divides the accounting data time-series monitoring sequence into a multi-scale local accounting observation segment set. For example, the data processing delay time-series sequence can be represented as [s_1, s_2, s_3, ..., s_n], where s represents the overall set of data processing delay time-series sequences, s_i (i = 1, 2, ..., n), i represents the data processing delay value of the i-th sampling point, and n represents the total number of sampling points in the time-series sequence. Each s_i corresponds to the data processing delay monitoring value of a sampling point. The preset step size is set in advance by a preset person. The multi-scale sliding observation window includes a preset first... The system includes a first-scale window, a second-scale window, and a third-scale window. The window widths of the first-scale window, second-scale window, and third-scale window increase sequentially and are all set by designated personnel, meaning the first-scale window is smaller than the second-scale window, which is smaller than the third-scale window. The preset step size is no larger than the first-scale window, ensuring that the time-series data of each sampling point is covered by at least one window, avoiding data omissions, and ensuring continuous capture of time-series trend changes during window sliding. This ensures the continuity and integrity of local observation segments and improves the comprehensiveness and accuracy of trend perception in accounting data time-series monitoring. The multi-scale local accounting observation segment set includes a first-scale accounting observation segment set, a second-scale accounting observation segment set, and a third-scale accounting observation segment set. The first-scale accounting observation segment set contains first-scale accounting observation segments; the second-scale accounting observation segment set contains second-scale accounting observation segments; and the third-scale accounting observation segment set contains third-scale accounting observation segments.

[0020] Based on a multi-scale local accounting observation segment set, the cumulative amount of accounting data offset in the same direction and the variance of the magnitude of accounting data offset in the same direction are obtained. Taking the first-scale accounting observation segment as an example, the specific process of obtaining the cumulative amount of accounting data offset in the same direction and the variance of the magnitude of accounting data offset in the same direction is as follows: Differential operations are performed on adjacent time series points in the first-scale accounting observation segment to obtain the change in accounting time series monitoring data of adjacent sampling points, i.e., the first-order difference sequence of accounting data, which is used to quantify the change magnitude of the accounting data time series monitoring sequence. For example, if the CPU utilization segment is [45%, 46%, 47%], then the corresponding first-order difference sequence of accounting data is [1%, 1%]. The cumulative amount of accounting data offset in the same direction and the variance of the magnitude ... The number of changes in the same direction indicates the number of times adjacent changes in the first-order difference sequence of accounting maintain the same sign. Maintaining the same sign means that both adjacent changes are either positive or both are negative. For example, in the first-order difference sequence of accounting [1%, 1%], the two adjacent changes (1%, 1%) are both positive and have the same sign, so the number of changes in the same direction is 1. If the first-order difference sequence of accounting is [2%, 2%, -3%, -4%, 5%], the four adjacent changes (2%, 2%) are both positive, (2%, -3%) are one positive and one negative, (-3%, -4%) are both negative, and (-4%, 5%) are one positive and one negative, so the number of changes in the same direction is 2. All changes in the same direction within the first-scale accounting observation segment correspond to... The summation of the changes in the same direction is used as the cumulative deviation of accounting data in the same direction, which is used to quantify the strength of the cumulative deviation. For example, if the first-order difference sequence of accounting data is [2%, 2%, -3%, -4%, 5%], where the changes in the same direction are 2%, 2%, -3%, and -4%, then the cumulative deviation of accounting data in the same direction is 2% + 2% + (-3%) + (-4%) = -3%. The variance of all changes in the same direction within the first-scale accounting observation segment is used as the variance of the magnitude of the same direction deviation in accounting data, which reflects the stability of the magnitude of the same direction deviation within the first-scale accounting observation segment. The smaller the variance, the stronger the stability of the deviation. The more stable the magnitude of the same-direction change, the better. For example, if the first-order difference sequence of accounting observation segment corresponding to a certain first-scale accounting observation segment is [-0.8%, 1%, 0.9%, 1%], and the change in the same direction is [1%, 0.9%, 1%], the variance calculation of this set of changes is close to 0, indicating that the magnitude of the same-direction change within this observation segment is relatively stable overall. Using the same method as above, the corresponding calculations are performed on each second-scale accounting observation segment in the second-scale accounting observation segment set and each third-scale accounting observation segment in the third-scale accounting observation segment set to obtain the number of same-direction changes, the cumulative amount of same-direction change in accounting data offset, and the variance of the magnitude of same-direction change in accounting data for each scale.

[0021] Based on the number of co-directional changes at each scale, the cumulative amount of co-directional changes in accounting data offsets, and the variance of the magnitude of co-directional changes in accounting data, a single-scale accounting trend confidence score is obtained to quantify the probability of a step-like gradual change trend in observation segments at each scale. The single-scale accounting trend confidence score includes a first trend confidence score, a second trend confidence score, and a third trend confidence score. The result of calculating the ratio of the number of co-directional changes within a first-scale accounting observation segment to the total number of adjacent time series points within that first-scale accounting observation segment is used as the accounting co-directional change ratio. For example, in a certain first-scale accounting observation... The segment contains 4 time series points, corresponding to 3 pairs of adjacent time series points, for a total of 3 pairs of adjacent time series points. The number of co-directional changes in the accounting first-difference sequence corresponding to this observation segment is 2, therefore the proportion of co-directional changes in accounting data for this observation segment is 2 / 3. The variance of the co-directional changes in accounting data is summed with a preset minimum value, and then 1 is divided by the summation result to obtain the amplitude consistency coefficient. The preset minimum value is pre-set by designated personnel to avoid invalid calculations due to a denominator of 0. The corresponding proportion of co-directional changes in accounting data is then compared with the amplitude... The product of the consistency coefficients is used as the first trend confidence score. The probability of a step-like gradual change trend in the local observation segment is quantified by multiplying the proportion of accounting and financial changes in the same direction with the amplitude consistency coefficient. The higher the proportion of accounting and financial changes in the same direction and the more stable the amplitude of the change, the higher the single-scale trend confidence score. It should be noted that the acquisition process for the second and third trend confidence scores is the same as that for the first trend confidence score. Cross-scale weighted fusion is performed on the single-scale accounting and financial trend confidence scores at each scale to obtain the results used to determine anomalies in each interval of the accounting and financial data time-series monitoring sequence. The type of financial and accounting global trend confidence: cross-scale weighted fusion means that the financial and accounting global trend confidence is the result of weighted summation of the single-scale financial and accounting trend confidence and the corresponding single-scale financial and accounting confidence parameters; the single-scale financial and accounting confidence parameters include the first confidence factor used to quantify the influence of the first trend confidence on the financial and accounting global trend confidence, the second confidence factor used to quantify the influence of the second trend confidence on the financial and accounting global trend confidence, and the third confidence factor used to quantify the influence of the third trend confidence on the financial and accounting global trend confidence.

[0022] The accounting time series intervals are divided based on global trend confidence. The specific division process is as follows: Time series intervals with a global trend confidence not lower than a preset first threshold are marked as step-like slowly changing anomaly intervals. Based on the step-like slowly changing anomaly intervals, weak filtering and order preservation measures are adopted for accounting data. Only a small amount of random jitter within the interval is suppressed, without changing the trend characteristics of the step-like slowly changing anomaly, and the continuous multi-period small-amplitude same-direction shift characteristics of the accounting data time series monitoring sequence within the interval are fully preserved. The time series interval represents the time series interval formed by the common overlapping sampling points of each scale window. For example, the time interval covered by the first-scale accounting observation segment is [1, 2, 3], and the time interval covered by the second-scale accounting observation segment is [1, 2, 3]. The time interval is [1, 2, 3, 4], and the time interval covered by the third-scale accounting observation segment is [1, 2, 3, 4, 5], so the corresponding time series interval is [1, 2, 3]. The accounting data time series monitoring sequence corresponding to the stepped slowly changing anomaly interval has the characteristics of continuous multi-period small same-direction shift, consistent shift direction, and stable amplitude, which is consistent with the weak anomalies caused by factors such as the slow drift of hardware sampling accuracy in the accounting business system. The time series interval with a global trend confidence level not higher than the preset second threshold is marked as a random noise interval. Based on the random noise interval, accounting data smoothing and noise reduction measures are taken to remove interference signals without actual anomaly significance and ensure that the accounting data within this interval is within the range. The system monitors the stability of time-series monitoring data to avoid random noise interfering with subsequent feature extraction and fusion. The preset second threshold is represented by the average global trend confidence level over a historical time period. The absence of actual anomalies indicates that such fluctuations are caused by non-business factors such as transient hardware interference or short-term network jitter, manifesting only as irregular and unsustainable random jumps. These fluctuations do not reflect changes in the operating load of the accounting system, data processing anomalies, or risk accumulation, and are not related to gradual, step-like anomalies. The accounting data time-series monitoring sequence corresponding to the noise interval exhibits characteristics of fluctuations without a fixed direction or continuous unidirectional trend, with random fluctuation amplitudes, and is considered to have no actual anomaly. Random jitter in the conventional sense corresponds to normal random interference during the operation of the accounting business system. The time series interval where the global trend confidence level is between the preset first threshold and the preset second threshold is marked as the noise transition interval. Based on the noise transition interval, an adaptive compromise processing measure for accounting data is adopted to retain the gradual unidirectional change trend within the interval, reduce abnormal loss or noise residue caused by over-processing, and ensure the continuity and integrity of the accounting time series interval division. The preset first threshold is greater than the preset second threshold. The accounting data time series monitoring sequence corresponding to the noise transition interval has both a weak unidirectional change trend and some random fluctuations, without clear abnormal or noise attributes.The accounting data smoothing and denoising processing measures involve applying preset smoothing filtering algorithms, such as median filtering and moving average, to the time-series monitoring sequence of accounting data within the noise range. This results in a smoothed and denoised time-series monitoring sequence of accounting data. Furthermore, based on this smoothed and denoised time-series monitoring sequence, a weakly progressive anomaly representation enhancement method is used to aggregate multimodal heterogeneous accounting data representations.

[0023] It is important to note that the first, second, and third confidence level influence factors involved in this embodiment were pre-constructed by experts in the field of accounting supervision, combining massive amounts of historical operational scenario data from accounting business systems. These factors are stored in a multimodal accounting supervision and early warning knowledge base, providing a core basis for the weighted fusion of trend confidence at different scales. This directly determines the accuracy of the overall accounting trend confidence in distinguishing between step-like, slowly changing anomalies and random noise. Specifically, the construction process of this single-scale accounting confidence parameter first requires collecting a large amount of historical time-series monitoring data from accounting business systems, covering different business scenarios, such as financial settlement. The data includes global trend confidence samples during peak periods, and corresponding actual interval type labels: stepped, gradually changing abnormal intervals, random noise intervals, or noise transition intervals. The peak period for financial settlement refers to the time when the accounting system processes a large volume of critical business operations such as reconciliation and report generation within a specific timeframe. During this period, the system load is significantly higher than normal, and time-series monitoring indicators such as data processing latency and CPU utilization are more prone to cumulative shifts. Identification of this type of scenario involves collecting business volume statistics from the system's historical operation logs. In one specific embodiment, daily, monthly, or quarterly business processing volumes exceeding [a certain threshold] can be selected. The period exceeding a preset business volume threshold is designated as the peak period for financial settlement. This threshold is set by the designated personnel based on statistical analysis of the system's historical average processing capacity and peak business volume distribution. Specifically, the personnel first extract business volume data for at least one complete year, such as the past 12 months, from the historical operation logs of the accounting system. This data includes the number of settlement transactions, reconciliation tasks, and report generation frequency. Then, the business volume data is preprocessed to remove extreme outliers caused by abnormal factors such as system failures, network anomalies, or human error, resulting in a standardized historical business volume dataset. Next, based on statistical analysis... The method calculates the baseline value of the normal processing capacity of the standardized historical business volume dataset. The baseline value of the normal processing capacity is preferably the 75th percentile of the dataset. This baseline value of the normal processing capacity is used as the preset business volume threshold. At the same time, the correspondence between different confidence parameter combinations and the global trend confidence division results is sorted out. Each parameter combination is assigned a weight quantification value based on its influence on the accuracy of interval division. For example, when the proportion of the same direction change captured by the second scale window covering 4 monitoring periods in a certain business scenario is higher than that of the first scale and the third scale, a higher second confidence influence factor is matched to strengthen its proportion in the weighted fusion.The system synchronously records the actual effective values ​​of the first, second, and third confidence level influencing factors under various historical scenarios. Then, through correlation analysis, such as Spearman's rank correlation coefficient, it eliminates factors caused by momentary equipment anomalies, such as sensor sampling drift and network jitter, retaining statistically stable parameter correspondences. Finally, it integrates all effective data to form a complete single-scale accounting confidence parameter. When the system conducts multi-scale time-series trend sensing in accounting, it can quickly retrieve influencing factors at various scales that match the current monitoring data's operating scenario from this system.

[0024] It should also be noted that the preset first threshold in this embodiment is pre-set by preset personnel. Specifically, the preset personnel first collect global trend confidence samples of the accounting business system over multiple historical time windows to form a normal baseline dataset; then, they obtain the mean and standard deviation of the normal baseline dataset, multiply the standard deviation of the normal baseline dataset by a preset multiple, and sum the result of the product operation with the mean as the initial candidate threshold; next, they utilize abnormal global trend confidence samples labeled with real abnormal events in history. This sample of global trend confidence scores, which is manually verified to correspond to time periods of stepwise, slowly changing anomalies, comes from different time windows and does not overlap with the normal baseline dataset. It is used to verify the threshold's ability to identify real anomalies. Based on the receiver operating characteristic (ROC) curve analysis method, the recall and false alarm rates of the global trend confidence scores of the anomalies are comprehensively evaluated using different initial candidate thresholds. The candidate threshold that makes the recall rate not lower than the preset recall requirement and the false alarm rate less than the preset false alarm threshold is selected as the preset first threshold. The preset multiple is set in advance by preset personnel. In an optional embodiment, it can be selected as 2 times or 3 times.

[0025] As described above, multi-scale time-series perception of accounting data helps to capture the distribution characteristics of step-like slowly changing anomalies and random noise in the time-series monitoring sequence of accounting data with high precision across multiple time scales. This enables adaptive division of different types of time-series intervals, reduces noise interference and the probability of misjudgment of anomalies, reduces the loss of weak anomalies or trend distortion caused by single-scale analysis, improves the pertinence and effectiveness of accounting monitoring data preprocessing, and achieves refined, order-preserving, and noise-reducing optimization of the original time-series data.

[0026] like Figure 3The logic diagram of the weak filtering and order preservation measures for accounting data in the multimodal data-driven accounting supervision and early warning method shown illustrates the following: Weak filtering and order preservation measures are implemented for accounting data, along with adaptive local noise point detection and acquisition of local deviation. The local deviation is then assessed to determine if it exceeds a preset noise deviation threshold. If not, the original value of the current sampling point is retained without further processing; otherwise, order-preserving noise replacement is performed based on random jitter noise points to obtain candidate accounting replacement values. Simultaneously, an accounting time-series direction discrimination index is acquired. If the index is greater than 0, the overall trend direction of the stepped, gradually changing anomaly interval is determined to be positive, and order preservation constraint checks are performed based on this overall trend direction. If the index is less than 0, the overall trend direction of the stepped, gradually changing anomaly interval is determined to be negative, and order preservation constraint checks are performed based on this overall trend direction. If the slope is 0, slope-assisted trend judgment is performed. If the slope is zero, the overall trend direction is considered to be without offset, and multimodal heterogeneous data representation and aggregation of accounting data is performed based on the accounting time series monitoring sequence. If the slope is positive, the overall trend direction is positively offset, and order preservation constraint check is performed based on the overall trend direction. If the slope is negative, it is negatively offset, and order preservation constraint check is performed based on the overall trend direction. After the accounting data weak filtering order preservation processing measures are completed, trend preservation verification is performed on the processed accounting time series monitoring sequence. It is determined whether the overall trend direction of the re-acquired accounting time series monitoring sequence is consistent with the overall trend direction before the execution of the accounting data weak filtering order preservation processing measures, and whether the cumulative offset change rate of the accounting time series monitoring sequence after the accounting data weak filtering order preservation processing is less than the preset tolerance threshold. If not, a weak filtering order preservation processing failure prompt is sent; otherwise, multimodal heterogeneous data representation and aggregation of accounting data is performed.

[0027] As a preferred technical approach, the specific process of the weak filtering and order preservation measures for accounting data is as follows: The difference between the value of the last sampling point and the value of the first sampling point in the accounting time-series monitoring sequence within the stepped, gradually changing anomaly interval is used as the accounting time-series direction discrimination index. If the accounting time-series direction discrimination index is greater than 0, the overall trend direction of the stepped, gradually changing anomaly interval is determined to be positive, that is, the monitoring values ​​in the accounting time-series monitoring sequence show an increasing trend, and order preservation constraint checks are performed based on the overall trend direction. If the accounting time-series direction discrimination index is less than 0, the overall trend direction of the stepped, gradually changing anomaly interval is determined to be negative, that is, the monitoring values ​​in the accounting time-series monitoring sequence show a decreasing trend, and order preservation constraint checks are performed based on the overall trend direction. If the accounting time series direction discrimination index is equal to 0, then the slope of the accounting time series monitoring sequence is further obtained based on the linear regression method, and the slope is used to assist in trend discrimination. The process is as follows: if the slope is positive, the overall trend direction is positive; if the slope is negative, it is negative; if the slope is zero, the overall trend direction is considered to be without deviation. Then, based on the accounting time series monitoring sequence, a weak progressive anomaly representation enhancement is performed to aggregate the accounting multimodal heterogeneous data representation. The accounting time series monitoring sequence is obtained based on the linear regression method, which means that the sampling point number, i.e., the time position, is used as the independent variable and the sampling point value is used as the dependent variable. Based on the univariate linear regression fitting algorithm, a univariate linear regression fitting is performed on all sampling points in the stepped slowly changing anomaly interval to obtain the slope of the regression line.

[0028] Adaptive local noise point detection based on the accounting time-series monitoring sequence is performed as follows: For each internal sampling point in the accounting time-series monitoring sequence, i.e., all sampling points except the first and last sampling points, the following operations are performed: Taking the current sampling point as the center, take its previous adjacent sampling point, the current sampling point, and the next adjacent sampling point to form a local window containing three sampling points; the median of the values ​​of the three sampling points in this local window is recorded as the local median value; the absolute difference between the current sampling point value and the local median value is recorded as the accounting sampling point deviation; the standard deviation of the values ​​of the three sampling points in this local window is used as the accounting window standard deviation; the accounting window standard deviation is compared with a preset extreme value. The smaller values ​​are summed, and the result of the division operation between the accounting sampling point deviation and the summation operation is recorded as the local deviation of the current sampling point. The local deviation represents the degree of numerical dispersion of the current sampling point relative to its two adjacent sampling points. The larger the value, the greater the deviation of the current sampling point from the local median value, and the higher the probability that it is a random jitter noise point. It is determined whether the local deviation is greater than the preset noise deviation threshold. If so, the current sampling point is determined to be a random jitter noise point, and order-preserving noise replacement is performed based on the random jitter noise point. Otherwise, the original value of the current sampling point is retained without processing. The preset noise deviation threshold is represented by the average value of the local deviation over a historical time period.

[0029] The specific process of order-preserving noise replacement is as follows: Take the values ​​of the previous and next adjacent sampling points of the current random jitter noise point, and use the median of the two as the accounting candidate replacement value for the current random jitter noise point; Based on the overall trend direction, perform an order-preserving constraint check on the accounting candidate replacement value, specifically as follows: If the overall trend direction is positive, the accounting candidate replacement value is limited to be no less than the value of the previous adjacent sampling point and no greater than the value of the next adjacent sampling point. This is to ensure that the replaced accounting time series monitoring sequence still maintains the monotonicity consistent with the original overall trend direction, reducing the introduction of reverse fluctuations or destruction of the original step-like gradual change characteristics due to the replacement operation. If this constraint is not followed, it may mask the true weak abnormal trend and mislead subsequent feature extraction and risk prediction; If the overall trend direction is negative, the accounting candidate replacement value is limited to be no greater than the value of the previous adjacent sampling point and no less than the value of the next adjacent sampling point; If the accounting candidate replacement value satisfies the above order-preserving constraint, then the value of the current random jitter noise point is... Update to the candidate replacement value for accounting; if the order preservation constraint is not met, the following adjustments are made: when the overall trend direction is positively offset, take the larger of the values ​​of the previous and next adjacent sampling points as the final replacement value for accounting, in order to retain more positive change amplitude and reduce over-compression of the true trend; when the overall trend direction is negatively offset, take the smaller of the two as the final replacement value for accounting, in order to retain more negative change amplitude and reduce over-compression of the true trend; update the value of the current random jitter noise point to the final replacement value for accounting; perform accounting edge sampling point retention; accounting edge sampling point retention means that the original values ​​of the first and last sampling points in the accounting time series monitoring sequence are directly retained. This is because the edge sampling points of the accounting time series monitoring sequence play a key role in determining the overall trend direction, and it is impossible to obtain sampling points in the bilateral neighborhoods for value replacement. Retaining the original values ​​of the first and last sampling points helps to maintain the integrity of the overall trend and the stability of the direction determination.

[0030] After the weak filtering and order preservation measures for accounting data are completed, the trend preservation verification of the processed accounting time-series monitoring sequence is performed. The verification includes the following two conditions: First, the overall trend direction of the accounting time-series monitoring sequence is reacquired, and the overall trend direction is consistent with the overall trend direction before the weak filtering and order preservation measures were implemented. Second, the cumulative offset change rate of the accounting time-series monitoring sequence after the weak filtering and order preservation measures is less than the preset tolerance threshold. The cumulative offset change rate is used to quantify the total change amplitude of the accounting time-series monitoring sequence before and after the processing. The preset tolerance threshold is represented by the average of the cumulative offset change rates over historical time periods. If both of the above conditions are met, then a weak progressive anomaly representation enhancement and aggregation of multimodal heterogeneous accounting data representation are performed based on the accounting time-series monitoring sequence. If either condition is not met, a weak filtering and order preservation failure prompt is sent, and the accounting time-series monitoring sequence data before and after the weak filtering and order preservation measures are uploaded to the accounting supervision data center for manual review or triggering of the backup processing mechanism.

[0031] Specifically, the formula for calculating the cumulative offset change rate is as follows:

[0032] Where R represents the cumulative offset change rate, t=1, 2, 3, ..., N, t represents the index of the sampling point, N represents the total number of sampling points in the accounting time series monitoring sequence, and x t This represents the value of the t-th sampling point in the original accounting time-series monitoring sequence. For example, if the accounting time-series monitoring sequence is a data processing delay sequence, then x... t Let x be the time delay value corresponding to the t-th sampling point, where the denominator is the sum of the absolute values ​​of the differences between all adjacent sampling points in the original accounting time series monitoring sequence, and the numerator is the sum of the absolute values ​​of the differences between all adjacent sampling points in the processed accounting time series monitoring sequence. 1 t x represents the value of the t-th sampling point in the accounting time series monitoring sequence after weak filtering and order preservation processing of accounting data. 1 t-1 x represents the value of the (t-1)th sampling point in the accounting time series monitoring sequence after weak filtering and order preservation processing of accounting data. t-1 This represents the value of the (t-1)th sampling point in the original accounting time-series monitoring sequence.

[0033] As described above, the weak filtering and order preservation measures for accounting data help to suppress fluctuation interference while retaining the true trend characteristics of the step-like gradual anomalies, reduce the risk of feature loss caused by excessive smoothing of abnormal trends, reduce the destruction of the original progressive offset pattern by the filtering process, improve the fidelity and trend integrity of accounting time series data, and achieve accurate order preservation and appropriate smoothing of the step-like gradual anomaly intervals.

[0034] As a preferred technical approach, the specific process of the adaptive compromise processing measure for accounting data is as follows: For the accounting time-series monitoring sequence within the noise transition interval, two processing paths are executed simultaneously: The first path performs strong smoothing and noise reduction processing on the accounting time-series monitoring sequence within the noise transition interval based on the center-weighted moving average filtering algorithm, resulting in a strongly smoothed sequence. This path mainly suppresses random noise. The second path only performs order-preserving noise replacement on sampling points where the local deviation exceeds a preset noise threshold by a preset multiple, while keeping the remaining sampling points unchanged, resulting in a weakly smoothed sequence. This path mainly preserves the slight trend. The strongly smoothed sequence obtained from the first path and the weakly smoothed sequence obtained from the second path are weighted and fused according to the accounting filtering strength factor. The fusion process is as follows: The confidence score of the global accounting trend corresponding to the noise transition interval is normalized and mapped to obtain the result. The accounting filter strength factor has a value between 0 and 1. The normalization mapping process means that the result of the difference between the global maximum and global minimum values ​​of the global trend confidence of the current noise transition interval is used as the confidence normalization index, the result of the difference between the current global trend confidence and global minimum value is used as the confidence difference, and the confidence difference is divided by the confidence normalization index to obtain the normalized value. For each sampling point, the value of that point in the weak smoothing sequence is multiplied by the accounting filter strength factor to obtain the first accounting median value. The difference between 1 and the accounting filter strength factor is calculated, and the result of the difference is multiplied by the value of that point in the strong smoothing sequence to obtain the second accounting median value. The sum of the first accounting median value and the second accounting median value is used as the fused value of that sampling point.

[0035] The sequence preservation check is performed on the merged accounting time-series monitoring sequence. The specific process is as follows: Calculate the overall trend direction of the accounting time-series monitoring sequence. If it is consistent with the overall trend direction of the original accounting time-series monitoring sequence, then perform multimodal heterogeneous data representation aggregation based on the merged accounting time-series monitoring sequence. If it is inconsistent with the overall trend direction of the original accounting time-series monitoring sequence, then send a noise transition interval adaptive compromise processing sequence preservation failure prompt, and upload the original accounting time-series monitoring sequence and the merged accounting time-series monitoring sequence together to the accounting supervision data center for manual review or triggering the backup processing mechanism.

[0036] As described above, adaptive compromise processing of accounting data helps to dynamically balance random noise suppression and gradual trend preservation within the noise transition range. It also allows for adaptive adjustment of the filtering intensity based on the global trend confidence level, reducing the risk of abnormal feature distortion due to excessive smoothing or noise residue caused by insufficient filtering. This approach also reduces the difficulty of adapting a single processing method to the complex data characteristics in the transition range, achieving a high-precision trade-off between noise and trend.

[0037] As a preferred technical means, the multimodal heterogeneous data representation and aggregation of accounting and finance is carried out in the following specific process: Step 1, the accounting and finance time series monitoring sequence is input into a preset time series feature extraction network, and the output is an enhanced time series feature vector for each time point. The enhanced time series feature vector contains the accounting and finance time series monitoring value at the current time point, the cumulative deviation memory value, and the trend confidence. For example, suppose a certain accounting and finance time series monitoring sequence is [x1, x2, x3], which represents the CPU utilization rate of three consecutive sampling points. After inputting into the preset time series feature extraction network, the enhanced time series feature vector output at time t=3 is [x3, m3, c3], where t=1, 2, 3, and t represents the current sampling point number. The enhanced time series feature vector can be represented as [x3, m3, c3], where x3 is the CPU utilization rate value at the current time point, m3 is the value of the cumulative deviation memory unit, which represents the cumulative strength of the positive deviation of two consecutive periods, and c3 is the trend confidence, which represents the probability that the segment has a step-like gradual change trend.

[0038] It should be noted that the preset temporal feature extraction network used in this embodiment is specifically a gated recurrent unit network with a cumulative deviation memory unit, used to capture the cumulative effect of small-amplitude same-direction shifts over multiple consecutive periods. In addition, long short-term memory networks, Transformer temporal encoders, and other time-series models with memory capabilities can also be used. The specific training process is as follows: First, collect financial and accounting time-series monitoring sequence sample data, including normal time-series samples, step-like slowly changing abnormal time-series samples, and noisy time-series samples. Preprocess the sample data, labeling the true cumulative deviation intensity and trend confidence level of each sample as a label; then, based on dividing the preprocessed sample data into training, validation, and test sets, initialize the pre-... The parameters of the time-series feature extraction network are set, and the error between the enhanced time-series feature vector output by the network and the sample label is used as the optimization objective. The network is iteratively trained using the training set. Then, after each round of training, the training effect of the network is verified using the validation set. Based on the validation results, the network parameters, learning rate, and other hyperparameters are adjusted to avoid overfitting or underfitting. Finally, when the error of the network on the validation set reaches the preset error threshold, training is stopped, and the performance of the trained network is tested using the test set to ensure that the network can accurately extract the current financial and accounting time-series monitoring values, cumulative deviation memory values, and trend confidence, thus meeting the requirements for financial and accounting time-series feature extraction. The preset error threshold is represented by the average error of historical time periods.

[0039] Step two involves acquiring accounting text data, such as system logs, error reports, maintenance work orders, regulatory notices, and external public opinion news. After preprocessing the accounting text data, including word segmentation and stop word removal, the data is input into a pre-trained language model to obtain enhanced text feature vectors at each time step, highlighting semantic information related to subtle accounting anomalies and achieving semantic enhancement.

[0040] It is important to note that the pre-trained language model used in this embodiment is specifically the BERT-based model, which is used to mine semantic information related to weak anomalies in the accounting system from text data, thereby achieving semantic enhancement and anomaly association of text features. In addition, pre-trained language models such as RoBERTa and ALBERT can also be used. The specific training process is as follows: First, collect accounting text data, covering various text materials related to accounting anomalies, such as system logs, error reports, maintenance work orders, and regulatory notices. This includes text containing weak anomaly semantics and normal text. Preprocess this text data by removing invalid characters and standardizing the format, then label it with semantic tags related to weak accounting anomalies, such as increased latency and data synchronization anomalies, as training tags. Next, the labeled text data... The dataset is divided into training, validation, and test sets. Parameters of pre-trained language models such as BERT-base are initialized. Preprocessed text data is input into the model for iterative training. During training, the error between the semantic features output by the model and the labeled data is used as the optimization objective, and the model parameters are continuously adjusted. After each training round, the model performance is verified using the validation set. Hyperparameters such as the learning rate and number of iterations are adjusted based on the verification results to avoid overfitting or underfitting. Finally, when the semantic feature extraction accuracy of the model on the validation set reaches a preset accuracy threshold, training is stopped, resulting in a trained model adapted to accounting scenarios and capable of accurately mining weak anomaly semantic information. This model is then used for subsequent semantic enhancement and anomaly association of text features. The preset accuracy threshold is represented by the average semantic feature extraction accuracy.

[0041] Step 3: Extract graph structure features. The specific process is as follows: Based on the call chain tracing algorithm, the financial and accounting operation monitoring data and call logs are used to construct a graph topology, generating a financial and accounting business operation topology graph. Nodes in the graph represent service modules, databases, middleware, or external interfaces, etc., and edges represent call relationships, data dependencies, or synchronization relationships, etc. Each node is associated with the time-series monitoring value of that module at the corresponding time, such as CPU utilization. The financial and accounting business operation topology graph is input into a preset graph convolution model for feature extraction to obtain an enhanced graph feature vector. Step 4: The obtained enhanced time-series feature vectors and enhanced graph features at the same timestamp are compared and contrasted. The text feature vector and the enhanced graph feature vector are aligned along the time dimension. The aligned three feature vectors are then concatenated along the channel dimension to obtain the joint accounting feature vector at each time step. The channel dimension represents the independent components of each feature vector, i.e., each element in the vector is a channel. For example, if the enhanced time-series feature vector is 7-dimensional, the enhanced text feature vector is 768-dimensional, and the enhanced graph feature vector is 256-dimensional, then the concatenated joint accounting feature vector is 7 + 768 + 256 = 1031-dimensional. Time dimension alignment means aligning the enhanced time-series feature vector, the enhanced text feature vector, and the enhanced graph feature vector along the channel dimension. Synchronization is performed according to a preset sampling time benchmark. For example, the sampling period is used as the benchmark. If the sampling frequency of text or graph data is low, it is filled using forward padding or interpolation methods before time dimension alignment. The preset sampling time benchmark is set in advance by preset personnel. Step 5: Input the joint financial and accounting feature vector into the preset attention fusion model and output the normalized weights of the three modalities. The normalized weights include the normalized weights of the temporal modality, the text modality, and the graph structure modality. The enhanced temporal feature vector, enhanced text feature vector, enhanced graph feature vector, and corresponding normalized weights are then processed. Enhanced multimodal feature fusion yields enhanced multimodal feature vectors. This involves multiplying the enhanced temporal feature vector, enhanced text feature vector, and enhanced graph feature vector with their corresponding temporal modality normalized weights, text modality normalized weights, and graph structure modality normalized weights, respectively, and then summing the results. The enhanced multimodal feature vectors at each time step are arranged in a preset time order to form an enhanced multimodal feature sequence. Financial and accounting supervision risk prediction is then performed based on this enhanced multimodal feature sequence, where the preset time order is pre-set by preset personnel.

[0042] It should be noted that the preset graph convolutional model used in this embodiment is specifically a graph attention network model. Graph convolutional networks, graph isomorphic networks, etc., can also be used. The specific training process is as follows: First, the topology graph and node labels corresponding to the historical financial and accounting operation monitoring data, such as normal and abnormal module status, are divided into a training set and a validation set. Then, the cross-entropy loss function is used to calculate the error between the node features output by the model and the true labels, and the backpropagation algorithm and Adam optimizer are used to iteratively update the model parameters until the loss converges or the performance of the validation set no longer improves.

[0043] It should also be noted that the preset attention fusion model used in this embodiment is specifically the multimodal model described in this paper, which is used to adaptively allocate the weights of each modality and enhance the contribution of weakly anomaly-related features. Logistic regression models, support vector machine models, random forest models, weighted summation fusion models, etc., can also be used. The specific training process is as follows: First, collect joint accounting feature vector templates. Each sample contains enhanced temporal feature vectors, enhanced text feature vectors, enhanced graph feature vectors, and corresponding labeled true weights for each modality and the final accounting risk label. Next, divide the collected multimodal sample data into training, validation, and test sets. The model is iteratively trained using the training set, with the dual optimization objectives being the error between the normalized weights of each modality output by the model and the true weights labeled in the samples, and the error between the risk prediction results corresponding to the fused features and the true risk labels. After each training round, the model's weight allocation accuracy and fusion effect are verified using a validation set. Based on the validation results, hyperparameters such as the number of attention heads, learning rate, and number of iterations are adjusted. Finally, when the model's validation parameters on the validation set reach preset thresholds, such as the weight allocation error reaching a preset allocation error threshold, the accuracy reaching a preset accuracy threshold, the recall reaching a preset recall threshold, and the F1 score reaching a preset F1 score, training is stopped, and the performance of the trained model is tested using a test set. The preset allocation error threshold is represented by the average weight allocation error over a historical time period, the preset prediction threshold by the average risk prediction error over a historical time period, the preset accuracy threshold by the average accuracy over a historical time period, the preset recall threshold by the average recall over a historical time period, and the preset F1 score by the average F1 score over a historical time period.

[0044] As described above, the aggregation of multimodal heterogeneous data in accounting helps to fully explore the complementary information of the three modalities: time-series numerical data, textual semantics, and topological structure. It focuses on and strengthens the weak, gradual, and chain-like abnormal accumulation features in multimodal heterogeneous data in accounting, reduces the probability of risk omissions and misjudgments caused by one-sided information from a single modality, reduces the interference caused by information redundancy and feature conflicts between modalities, and achieves adaptive weighted fusion of multidimensional heterogeneous data.

[0045] As a preferred technical means, the specific process of financial and accounting supervision risk prediction is as follows: Obtain a preset deep state space model, such as the improved S4 (Structured State Space for Sequence Modeling) model. The hidden state variables within the preset deep state space model are the cumulative deviation state variables corresponding to the financial and accounting data. The cumulative deviation state variables represent the accumulated financial and accounting risk intensity up to the current time, used to dynamically track the continuous evolution of financial and accounting system risks over time. For example, the cumulative deviation state variables of monitored values ​​such as CPU utilization and data processing latency up to the current time. Perform recursive updates of the cumulative deviation state variables, specifically as follows: For each time step t, the preset deep state space model updates the cumulative deviation state variables according to the following recursive rules:

[0046] Among them, s t s represents the specific value of the cumulative deviation from the state variable at the current moment. t-1 Let t = 0, 1, 2, ..., n, where t represents the current sampling time number and n represents the total number of sampling points. The initial time s0 = 0. α is a learnable decay coefficient, pre-set by a pre-defined person, ranging from 0 to 1, used to control the influence of historical cumulative deviations on the current cumulative deviation. g(⋅) is a fully connected mapping function that maps the current enhanced multimodal feature vector to a single-step risk contribution value. Et represents the enhanced multimodal feature vector.

[0047] The enhanced multimodal feature sequence is input into a preset deep state space model to obtain the accounting risk probability value at the current moment. The accounting risk probability value at the current moment is compared with a preset warning threshold. The specific process is as follows: If the accounting risk probability value at the current moment is greater than the preset first-level warning threshold, a first-level accounting supervision alert is immediately triggered. The trigger time, the corresponding accounting risk probability value, the accumulative state variables, and the associated enhanced multimodal feature fragments are uploaded to the accounting supervision data center. Simultaneously, accounting supervisors are notified via a visual interface or message push. The preset first-level warning threshold is preset by designated personnel in advance. The system is configured such that if the current accounting risk probability value is not greater than the preset level 1 warning threshold and not less than the preset level 2 warning threshold, a level 2 accounting supervision alert is triggered. The alert signal, the corresponding risk probability value, the cumulative state variable value, and the associated enhanced multimodal feature fragments are uploaded to the accounting supervision data center for continuous monitoring. The preset level 2 warning threshold is represented by the average of the accounting risk probability values ​​over a historical period. If the current accounting risk probability value is less than the preset level 2 warning threshold, the risk probability value is recorded in the log for subsequent audit analysis. The preset level 1 warning threshold is greater than the preset level 2 warning threshold.

[0048] It should be noted that the preset deep state space model used in this embodiment is specifically an improved S4 model. Alternatively, models with temporal memory and cumulative update capabilities, such as improved long short-term memory networks, gated recurrent units, and hybrid models of state space models, can also be used. The specific training process is as follows: First, collect multimodal time-series sample data of accounting and finance. Each sample contains an enhanced multimodal feature sequence, the corresponding cumulative deviation state variable's true recursive value, and the final accounting and finance risk probability label, such as the probability values ​​corresponding to normal, low risk, medium risk, and high risk. Next, divide the preprocessed sample data into training, validation, and test sets, and initialize model parameters, including decay. The initial values ​​of coefficient α, the parameters of the fully connected mapping function g(⋅), and the hidden layer parameters of the model are used as dual optimization objectives. The error between the cumulative deviation recursive value of the model output state variable and the true recursive value of the sample, and the error between the output financial risk probability value and the sample label are used as dual optimization objectives. The model is iteratively trained using the training set. After each round of training, the recursive accuracy and risk prediction accuracy of the model are verified using the validation set. Based on the verification results, hyperparameters such as the decay coefficient α, learning rate, and number of iterations are adjusted. Finally, when the recursive error and risk prediction error of the model on the validation set both reach the preset threshold, training is stopped, and the performance of the trained model is tested using the test set.

[0049] As described above, risk prediction through financial and accounting supervision helps to accurately capture potential financial and accounting risks, identify weak abnormal signals in advance, reduce business losses caused by untimely risk detection, reduce subsequent problems caused by hidden risks, achieve dynamic tracking and precise prevention and control of risks, provide a guarantee for the compliant and orderly operation of financial and accounting business, and further improve the financial and accounting supervision system.

[0050] like Figure 4 The diagram shown illustrates the structure of a multimodal data-driven accounting supervision and early warning system. This system includes: a multi-scale time-series perception module for accounting data, a multimodal heterogeneous data representation and aggregation module for accounting data, and an accounting supervision risk prediction module. The multi-scale time-series perception module is used to: perform multi-scale time-series perception of accounting data, and based on the perception results, implement targeted measures such as weak filtering and order preservation processing, adaptive compromise processing, or smoothing and noise reduction processing of accounting data. Through this module, refined processing of different types of data intervals can be achieved, providing high-quality, order-preserving basic data for subsequent feature extraction, reducing noise interference and abnormal feature distortion risks from the source, and strengthening the trend integrity and reliability of accounting time-series data.

[0051] The accounting multimodal heterogeneous data representation aggregation module is used to: perform multimodal heterogeneous data representation aggregation for accounting data after multi-scale time-series perception of accounting data is completed; through the accounting multimodal heterogeneous data representation aggregation module, it helps to fully explore the complementary value of three types of heterogeneous data: time-series numerical data, textual semantic data, and graph structure data, strengthen the feature expression related to weak accounting anomalies, solve the problems of one-sided information and insufficient feature discrimination power of single-modal data, and generate enhanced multimodal features that are both comprehensive and targeted.

[0052] The accounting supervision risk prediction module is used to: predict accounting supervision risks based on the aggregation results after the multimodal heterogeneous accounting data representation is completed. The accounting supervision risk prediction module helps to dynamically track the cumulative evolution of accounting risks, accurately identify potential risks and minor anomalies, improve the timeliness and accuracy of accounting risk warnings, and achieve early prediction and precise prevention and control of risks.

[0053] As described above, the multi-scale time-series perception module for accounting data, the multi-modal heterogeneous data representation and aggregation module for accounting data, and the accounting supervision and risk prediction module help to build a full-process, intelligent accounting supervision system, from data preprocessing and feature enhancement to risk prediction. Specifically, the multi-scale time-series perception module for accounting data provides qualified data input for subsequent modules through refined data processing. The multi-modal heterogeneous data representation and aggregation module for accounting data takes the data processed above and strengthens relevant features through multi-dimensional feature fusion, providing high-value feature input for the accounting supervision and risk prediction module. Based on the fused relevant features, the accounting supervision and risk prediction module achieves high-precision risk prediction and dynamic tracking. The three modules work in a progressive manner.

[0054] like Figure 5 The bar chart comparing the performance metrics of different preset attention fusion models for the multimodal data-driven financial supervision and early warning method shows that the multimodal model in this paper performs excellently in terms of accuracy, recall, and F1 score. The accuracy reaches 0.877, on par with the random forest model and higher than the logistic regression and support vector machine models. The recall is 0.617, on par with the random forest model, slightly lower than the logistic regression model, but better than the support vector machine model. The F1 score is 0.763, also on par with the random forest model and higher than the logistic regression and support vector machine models. Overall, the multimodal model in this paper achieves optimal or tied-optimal performance in all three metrics, verifying its stronger comprehensive classification ability and generalization performance in financial supervision and early warning tasks. It can improve the recall and comprehensive evaluation level of identifying weak financial anomalies while maintaining high accuracy, demonstrating the superiority and reliability of this solution in financial supervision scenarios.

[0055] Example 2, based on Example 1, as an alternative, addresses situations where the accounting system experiences significant differences in the intensity of abnormal node accumulation. For instance, some service modules may experience a continuous increase in CPU utilization over multiple periods due to hardware sampling accuracy drift, while other modules operate normally; or cross-module data synchronization errors may accumulate periodically along a certain call chain, but adjacent chains show no abnormalities. In such cases, the accumulated deviation from the memory value of each node differs greatly. If the graph convolution method based on static initial weights in Example 1 is directly used, the edge weights of nodes with strong abnormal accumulation will be the same as those of normal nodes, leading to… During graph convolution aggregation, anomalous signals are diluted by normal neighbor nodes, failing to highlight the progressively accumulating features of anomalous nodes in the feature vector. Therefore, an alternative method for graph structure feature extraction is required. The specific process is as follows: Based on the accounting business operation topology graph, edge weights are updated sequentially to obtain the updated accounting business operation topology graph. The updated accounting business operation topology graph is then input into a preset graph convolution model for feature extraction to obtain an enhanced graph feature vector. The specific process of the sequential update of edge weights is as follows: For each edge (u, v) of the accounting business operation topology graph, its initial weight w... uv The result is represented by the division of the total number of times node u calls node v over a historical time period by the total number of calls to all nodes; at each time step t, the cumulative deviation memory value m of nodes u and v is used. t (u) and m t (v), dynamically adjust the edge weights, and the specific adjustment formula is as follows:

[0056] Where w uv (t) represents the edge weight of edge (u,v) after dynamic adjustment at time step t, w uv The initial weights of the edges in the accounting and business operation topology graph are represented by γ, which is a preset scaling factor. mt(u) represents the cumulative deviation memory value of node u in the accounting and business operation topology graph. t (v) represents the cumulative deviation memory value of node v in the accounting business operation topology graph. This operation makes the information propagation stronger during graph convolution in areas where weak anomalies accumulate more, thereby highlighting progressive anomaly signals in node feature aggregation.

[0057] As described above, graph structure feature extraction helps adapt to scenarios where the intensity of node anomaly accumulation varies significantly in accounting and finance business systems. It can accurately capture the progressive anomaly accumulation and chain propagation features across modules and nodes, dynamically adjust the edge weights of the accounting and finance business operation topology graph, reduce the dilution of weak accumulation signals from abnormal nodes by normal nodes, and reduce the problems of weakened anomaly features and missed risk detection caused by static edge weights. This enables accurate capture and highlighting of weak progressive anomalies in the accounting and finance business topology.

[0058] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A financial and accounting supervision and early warning method based on multimodal data-driven approach, characterized in that: The method includes: Perform multi-scale time-series perception of financial and accounting data, and based on the perception results, take targeted measures such as weak filtering and order preservation processing, adaptive compromise processing, or smoothing and noise reduction processing of financial and accounting data. After the multi-scale time-series perception of financial and accounting data is completed, the data is aggregated for the representation of multimodal heterogeneous financial and accounting data. After the aggregation of multimodal heterogeneous financial and accounting data is completed, financial and accounting supervision risk prediction is carried out based on the aggregation results for continuous status tracking and hierarchical early warning.

2. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 1, characterized in that, The specific process of multi-scale time-series perception of accounting data is as follows: The time series monitoring sequence of financial and accounting data is input into a multi-scale sliding observation window. The time series monitoring sequence of financial and accounting data is traversed within the multi-scale sliding observation window with a preset step size and an increasing window width, and the time series monitoring sequence of financial and accounting data is divided into a multi-scale local financial and accounting observation segment set. The multi-scale local accounting observation segment set includes a first-scale accounting observation segment set, a second-scale accounting observation segment set, and a third-scale accounting observation segment set. The cumulative amount of same-direction shift in accounting data and the variance of the magnitude of same-direction shift in accounting data are obtained based on a multi-scale local accounting observation fragment set. The specific process for obtaining the cumulative amount of same-direction shift in accounting data and the variance of the magnitude of same-direction shift in accounting data is as follows: By performing a difference operation on adjacent time series points in the first-scale accounting observation segment, the change in accounting time series monitoring data of adjacent sampling points is obtained; The number of times adjacent changes in the same direction occur in a first-difference sequence of accounting data; The summation of all changes in the same direction within the first-scale accounting observation segment is taken as the cumulative amount of accounting data offset in the same direction. The variance of all changes in the same direction within the first-scale accounting observation segment is used as the variance of the magnitude of changes in the same direction of accounting data. Based on the number of same-direction changes at each scale, the cumulative amount of same-direction changes in accounting data, and the variance of the magnitude of same-direction changes in accounting data, the confidence level of accounting trends at a single scale is obtained. The single-scale accounting trend confidence level includes a first trend confidence level, a second trend confidence level, and a third trend confidence level. The ratio of the number of co-directional changes within the first-scale accounting observation segment to the total number of adjacent time series point pairs within the first-scale accounting observation segment is used as the proportion of co-directional changes in accounting. The variance of the same direction change in financial and accounting data is summed with the preset minimum value, and then 1 is divided by the summation result, which is used as the amplitude consistency coefficient. The product of the corresponding proportion of accounting changes in the same direction and the magnitude consistency coefficient is used as the confidence level of the first trend. By performing cross-scale weighted fusion of the single-scale accounting trend confidence scores at each scale, the overall accounting trend confidence score is obtained: The cross-scale weighted fusion means that the result of weighted summation of the single-scale accounting trend confidence score and the corresponding single-scale accounting confidence score is used as the global accounting trend confidence score. The accounting time series intervals were divided based on the global trend confidence level.

3. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 2, characterized in that, The process of dividing accounting time series intervals based on global trend confidence is as follows: The time series intervals with a global trend confidence level not lower than the preset first threshold are marked as step-type slowly changing anomaly intervals. Based on the step-type slowly changing anomaly intervals, weak filtering and order preservation measures for financial and accounting data are adopted to suppress only a small amount of random jitter within the interval, without changing the trend characteristics of the step-type slowly changing anomaly, and fully preserving the continuous multi-period small same-direction shift characteristics of the financial and accounting data time series monitoring sequence within the interval. The time interval refers to the time interval formed by the common overlapping sampling points of each scale window; The time series intervals with global trend confidence not exceeding a preset second threshold are marked as random noise intervals. Based on the random noise intervals, financial data smoothing and noise reduction measures are taken to eliminate interference signals. For time series intervals where the global trend confidence level is between the preset first threshold and the preset second threshold, the noise transition interval is marked. Based on the noise transition interval, an adaptive compromise processing measure for accounting data is adopted to retain the gradual unidirectional change trend within the interval. The aforementioned accounting data smoothing and noise reduction measures involve performing full-range noise suppression and smoothing on the accounting data time-series monitoring sequence within the noise range based on a preset smoothing filtering algorithm, to obtain the accounting data time-series monitoring sequence after accounting data smoothing and noise reduction, and then performing multimodal heterogeneous data representation aggregation based on the accounting data time-series monitoring sequence after accounting data smoothing and noise reduction.

4. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 3, characterized in that, The specific process of the weak filtering and order preservation processing measures for accounting data is as follows: The difference between the value of the last sampling point and the value of the first sampling point in the financial and accounting time series monitoring sequence within the stepped slowly changing anomaly interval is used as the financial and accounting time series direction discrimination index. If the accounting time series direction discrimination index is greater than 0, the overall trend direction of the stepped gradual change abnormal interval is determined to be positive, and the order preservation constraint check is performed based on the overall trend direction. If the accounting time series direction discrimination index is less than 0, the overall trend direction of the stepped gradual change abnormal interval is determined to be negative, and the order preservation constraint check is performed based on the overall trend direction. If the accounting time series direction discrimination index is equal to 0, then the slope of the accounting time series monitoring sequence is obtained based on the linear regression method. Adaptive local noise point detection based on accounting time-series monitoring sequences is performed as follows: For each internal sampling point in the aforementioned accounting time-series monitoring sequence, perform the following operations: Centered on the current sampling point, take its previous adjacent sampling point, the current sampling point, and the next adjacent sampling point to form a local window containing three sampling points; The median of the three sampled values ​​within the local window is recorded as the local median value. The absolute difference between the current sampling point value and the local median value is denoted as the accounting sampling point deviation. The standard deviation of the three sampling points within this local window is taken as the standard deviation of the accounting window. The local deviation of the current sampling point is recorded as the sum of the standard deviation of the accounting window and the preset minimum value, and then the result of the division operation between the accounting sampling point deviation and the sum is recorded as the local deviation of the current sampling point. Determine whether the local deviation is greater than the preset noise deviation threshold. If so, perform order-preserving noise replacement based on random jitter noise points. Otherwise, retain the original value of the current sampling point without processing.

5. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 4, characterized in that, The specific process of the order-preserving noise replacement is as follows: Take the value of the previous and next adjacent sampling points of the current random jitter noise point, and use the median of the two as the accounting candidate replacement value of the current random jitter noise point. Based on the overall trend, an order-preserving constraint check is performed on the candidate alternative values ​​for accounting and finance. The specific process is as follows: If the overall trend is positive, the candidate substitution value for accounting is limited to be no less than the value of the previous adjacent sampling point and no greater than the value of the next adjacent sampling point. If the overall trend is negative, the candidate substitution value for accounting is limited to not being greater than the value of the previous adjacent sampling point and not being less than the value of the next adjacent sampling point. If the accounting candidate replacement value satisfies the above order preservation constraint, then the value of the current random jitter noise point is updated to the accounting candidate replacement value. If the order preservation constraint is not satisfied, the following adjustments will be made: When the overall trend is positive, the larger of the values ​​of the previous and next adjacent sampling points is taken as the final accounting substitution value. When the overall trend is negatively offset, the smaller of the two values ​​is taken as the final accounting substitution value to retain more of the negative change range; Update the current value of the random jitter noise point to the final replacement value for this accounting document; Preservation of edge sampling points in accounting and finance; The retention of the accounting edge sampling points means that the original values ​​of the first and last sampling points in the accounting time series monitoring sequence are directly retained. After the weak filtering and order preservation measures for accounting data are completed, the trend preservation of the processed accounting time series monitoring sequence is verified. The verification includes the following two conditions: The first condition is to re-acquire the overall trend direction of the accounting time-series monitoring sequence; The second condition is that the cumulative offset change rate of the accounting time series monitoring sequence after weak filtering and order preservation processing of accounting data is less than the preset tolerance threshold. The cumulative offset change rate is used to quantify the total change in the accounting time series monitoring sequence before and after processing. If both of the above conditions are met, then multimodal heterogeneous data representation and aggregation of accounting and finance will be performed based on the accounting and finance time series monitoring sequence. If any condition is not met, a failure message for weak filtering and order preservation processing will be sent, and the accounting time series monitoring sequence data before and after the weak filtering and order preservation processing measures will be uploaded to the accounting supervision data center.

6. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 3, characterized in that, The specific process of the adaptive compromise processing measure for accounting data is as follows: For the accounting time series monitoring sequence within the noise transition interval, two processing paths are executed simultaneously: The first approach involves using a center-weighted moving average filtering algorithm to perform strong smoothing and noise reduction on the accounting time series monitoring sequence within the noise transition interval, resulting in a strongly smoothed sequence. The second path only performs order-preserving noise replacement on sampling points where the local deviation exceeds a preset noise threshold by a preset multiple; The strongly smoothed sequence obtained from the first path and the weakly smoothed sequence obtained from the second path are weighted and fused according to the accounting filter intensity factor. The fusion process is as follows: For each sampling point, the value of that point in the weakly smoothed sequence is multiplied by the accounting filter intensity factor to obtain the first accounting intermediate value; The difference between 1 and the accounting filter strength factor is calculated, and then the value of that point in the strongly smoothed sequence is multiplied by the result of the difference calculation to obtain the second accounting intermediate value. The summation of the median value of the first financial accounting report and the median value of the second financial accounting report is taken as the fused value of the sampling point. The sequence preservation check is performed on the merged accounting time-series monitoring sequence. The specific process is as follows: Calculate the overall trend direction of the accounting time series monitoring sequence. If it is consistent with the overall trend direction of the original accounting time series monitoring sequence, then perform multimodal heterogeneous data representation and aggregation based on the fused accounting time series monitoring sequence. If the overall trend direction is inconsistent with the original accounting time-series monitoring sequence, an adaptive compromise processing error message for maintaining order in the noise transition interval will be sent, and the original accounting time-series monitoring sequence and the fused accounting time-series monitoring sequence will be uploaded to the accounting supervision data center together.

7. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 6, characterized in that, The specific process of representing and aggregating multimodal heterogeneous data in accounting is as follows: The accounting time-series monitoring sequence is input into a preset time-series feature extraction network, and the output is an enhanced time-series feature vector for each time point; Obtain accounting and financial text data; After preprocessing the accounting text data, it is input into the pre-trained language model to obtain the enhanced text feature vector at each time step; The graph structure feature extraction process is as follows: Based on the call chain tracing algorithm, the financial and accounting operation monitoring data and graph data are used to construct a topology and generate a financial and accounting business operation topology graph. The accounting business operation topology graph is input into a preset graph convolution model for feature extraction to obtain an enhanced graph feature vector; Align the enhanced temporal feature vector, enhanced text feature vector, and enhanced graph feature vector obtained under the same timestamp according to the time dimension; The three aligned feature vectors are concatenated along the channel dimension to obtain the joint financial and accounting feature vector at each time step. Input the joint feature vector of accounting and finance into the preset attention fusion model, and output the normalized weights of the three modalities; Enhanced temporal feature vectors, enhanced text feature vectors, enhanced graph feature vectors, and corresponding normalized weights are fused to obtain enhanced multimodal feature vectors. The enhanced multimodal feature vectors at each time step are arranged in chronological order to form an enhanced multimodal feature sequence. Financial and accounting supervision risk prediction is then performed based on the enhanced multimodal feature sequence.

8. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 7, characterized in that, The specific process for predicting financial and accounting supervision risks is as follows: Obtain a preset depth state space model, wherein the hidden state variables of the preset depth state space model are the cumulative deviation state variables corresponding to the accounting data; The cumulative deviation state variable represents the intensity of financial and accounting risk accumulated up to the current moment, and is used to dynamically track the continuous evolution of financial and accounting system risk over time. The recursive update of the cumulative deviation state variable is performed as follows: The preset deep state space model updates the cumulative deviation state variables according to a recursive rule; The enhanced multimodal feature sequence is input into a preset deep state space model to obtain the financial and accounting risk probability value at the current moment. The process of comparing the current financial and accounting risk probability value with the preset warning threshold is as follows: If the current financial and accounting risk probability value is greater than the preset first-level warning threshold, a first-level financial and accounting supervision prompt will be triggered immediately, and the time of triggering the warning, the corresponding financial and accounting risk probability value, the accumulative state variables, and the associated enhanced multimodal feature fragments will be uploaded to the financial and accounting supervision data center. If the current financial and accounting risk probability value is not greater than the preset first-level warning threshold and not less than the preset second-level warning threshold, a second-level financial and accounting supervision prompt will be triggered, and the warning signal, the corresponding risk probability value, the cumulative state variable value, and the associated enhanced multimodal feature fragments will be uploaded to the financial and accounting supervision data center for continuous monitoring. If the current financial risk probability value is less than the preset level 2 warning threshold, the risk probability value is recorded in the log for subsequent audit analysis.

9. The financial supervision and early warning method based on multimodal data-driven approach as described in claim 7, characterized in that, The graph structure feature extraction also includes: Based on the accounting business operation topology graph, the edge weights are updated in a time-series recursive manner to obtain the accounting business operation topology graph with updated edge weights. The updated financial and accounting business operation topology graph with edge weights is input into a preset graph convolution model for feature extraction to obtain an enhanced graph feature vector.

10. A multimodal data-driven accounting supervision and early warning system, applied to the multimodal data-driven accounting supervision and early warning method as described in any one of claims 1-9, wherein the multimodal data-driven accounting supervision and early warning system comprises: The modules include: a multi-scale time-series perception module for accounting data, a multi-modal heterogeneous data representation and aggregation module for accounting data, and a risk prediction module for accounting supervision. The multi-scale time-series perception module for accounting data is used to: perform multi-scale time-series perception of accounting data, and take targeted measures such as weak filtering and order preservation processing, adaptive compromise processing, or smoothing and noise reduction processing based on the perception results. The accounting multimodal heterogeneous data representation and aggregation module is used to: perform accounting multimodal heterogeneous data representation and aggregation after the completion of multi-scale time-series perception of accounting data; The accounting supervision risk prediction module is used to predict accounting supervision risks based on the aggregation results after the aggregation of multimodal heterogeneous accounting data is completed.