Dynamic financial prediction and budget control method based on big data
Patent Information
- Application Number
- CN202610942632.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而实际高并发业务中存在资金与资源流转的物理惯性,如集中结算、连锁资源调度等会在数据流中产生连续高频波动
本申请通过提取边界截断点邻域内局部变化度,以局部二阶导数积分量化边缘溢出量,精准识别长期被忽视的离散化谬误,消除周期交界处预测精度周期性崩塌。通过将边缘溢出量非线性映射为重叠时序长度,并对窗口边界数据进行加权融合,生成二阶导数连续的重构特征序列,在不增加模型参数和计算复杂度前提下显著降低跨周期预测误差。通过引入反馈优化参数与学习率系数,利用实际业务值与预测向量的偏差动态更新参考基准,自适应数据分布漂移,保障长期运行下预算调控精度的稳定性与鲁棒性。
Smart Images

Figure CN122779982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of business intelligence management technology, specifically a method for dynamic financial forecasting and budget control based on big data. Background Technology
[0002] With the deep integration of big data and cloud computing, the concurrency and data throughput of distributed business systems are growing exponentially. Accurately extracting time-series features from massive real-time business flows to drive financial forecasting engines has become a core challenge for intelligent budget control. Existing technologies typically use time windows of fixed length and fixed sliding step size to discretize continuous data streams. Their core assumption is that the data distribution within the window is uniform and stable, and that the windows are independent of each other.
[0003] However, in real-world high-concurrency business operations, there is a physical inertia in the flow of funds and resources. For example, centralized settlement and chain resource scheduling can generate continuous high-frequency fluctuations in the data stream. When these fluctuations cross fixed time window boundaries, they are abruptly cut off, causing spectral aliasing and energy loss of momentum information that should have been transmitted across the window. Existing technologies fail to recognize this discretization fallacy and merely perform surface optimization on the fragmented slices by increasing model depth or introducing attention mechanisms, which cannot fundamentally repair the information loss caused by the underlying truncation.
[0004] Because the prediction engine cannot detect and compensate for energy overflow at the boundaries of time windows, a periodic collapse in prediction accuracy inevitably occurs at the intersection of business cycles. This manifests as a normalized increase in error volatility, leading to severe fluctuations in downstream control instructions and resource misallocation at the beginning of the cycle. This problem has become a long-standing common technical bottleneck in large-scale distributed systems. Therefore, how to quantify the degree of energy overflow at the edges of time windows and dynamically adjust the overlap length of adjacent windows and the smooth fusion strategy accordingly, fundamentally breaking the rigid assumption of fixed time-distance truncation and eliminating periodic accuracy collapse, is a key technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a dynamic financial forecasting and budget control method based on big data, aiming to eliminate the periodic collapse of forecast accuracy at the boundary of cycles and significantly reduce cross-cycle forecasting errors without increasing model parameters and computational complexity.
[0006] The objective of this application can be achieved through the following technical solution: Firstly, a dynamic financial forecasting and budget control method based on big data, comprising the following steps: Acquire continuously input business flow data, and perform temporal truncation on the business flow data based on a preset sliding window to generate a first feature tensor and a second feature tensor, as well as the boundary truncation point of the two. Obtain a multidimensional feature vector within a preset neighborhood of the boundary cutoff point, and map the multidimensional feature vector into a feature intensity sequence characterizing the intensity of business fluctuations. Extract the local variability of the feature intensity sequence within the preset neighborhood, and obtain the edge overflow feature of the boundary cutoff point based on the local variability; When the edge overflow feature is greater than a preset reference benchmark, the edge overflow feature is converted into the overlap time length of the first feature tensor and the second feature tensor, and the preset sliding window is updated. Based on the overlap time sequence length, tail data in the first feature tensor and head data in the second feature tensor are obtained. Weighted fusion of the tail data and the head data is performed based on preset weight parameters to obtain a reconstructed feature tensor sequence. Perform temporal state deduction on the reconstructed feature tensor sequence to obtain the business state prediction vector, and generate the corresponding budget resource control instructions; After executing the budget resource control instruction, feedback optimization parameters characterizing the budget control effect are obtained, and the preset reference benchmark is updated based on these parameters.
[0007] Secondly, the dynamic financial forecasting and budget control system based on big data includes the following modules: The data acquisition module is used to acquire continuously input business flow data, and to perform time-series truncation on the business flow data based on a preset sliding window to generate a first feature tensor and a second feature tensor, as well as the boundary truncation point of the two. The vector mapping module is used to obtain a multi-dimensional feature vector in a preset neighborhood of the boundary cutoff point and map the multi-dimensional feature vector into a feature intensity sequence that characterizes the intensity of business fluctuations. The feature extraction module is used to extract the local variability of the feature intensity sequence within the preset neighborhood, and to obtain the edge overflow feature of the boundary cutoff point based on the local variability. The overlap evaluation module is used to convert the edge overflow feature into the overlap time sequence length of the first feature tensor and the second feature tensor when the edge overflow feature is greater than the preset reference benchmark, and update the preset sliding window. The sequence reconstruction module is used to obtain the tail data in the first feature tensor and the head data in the second feature tensor based on the overlap time length, and to perform weighted fusion on the tail data and the head data based on preset weight parameters to obtain a reconstructed feature tensor sequence. The feedback optimization module is used to perform temporal state deduction on the reconstructed feature tensor sequence to obtain a business state prediction vector and generate a corresponding budget resource control instruction; after executing the budget resource control instruction, it obtains feedback optimization parameters characterizing the budget control effect and updates the preset reference benchmark based on them.
[0008] Thirdly, a computer storage medium stores computer-executable instructions, which, when executed, implement the dynamic financial forecasting and budget control method based on big data described in the first aspect.
[0009] Compared with the prior art, the beneficial effects of this application are: This application extracts the local variability within the neighborhood of the boundary cutoff point and quantifies the edge spillover using the local second derivative integral, accurately identifying the long-neglected discretization fallacy and eliminating the periodic collapse of prediction accuracy at the cycle boundary. By nonlinearly mapping the edge spillover to the overlap time series length and weighting and fusing the window boundary data, a reconstructed feature sequence with continuous second derivatives is generated, significantly reducing cross-cycle prediction errors without increasing model parameters or computational complexity. By introducing feedback optimization parameters and learning rate coefficients, the reference benchmark is dynamically updated using the deviation between actual business values and prediction vectors, adaptively adapting to data distribution drift and ensuring the stability and robustness of budget control accuracy under long-term operation. Attached Figure Description
[0010] Figure 1 This is a schematic diagram illustrating the steps of the big data-based dynamic financial forecasting and budget control method of this application; Figure 2 This is a schematic diagram of the modules of the big data-based dynamic financial forecasting and budget control system of this application. Detailed Implementation
[0011] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only to illustrate selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item has been defined in one figure, it does not need to be further defined and explained in subsequent figures. The terms first, second, etc. are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0012] With the deep integration of big data technology and cloud computing platforms, the number of concurrent nodes and data throughput in distributed business systems are growing exponentially. How to accurately extract time-series features from massive real-time business data and efficiently drive the financial forecasting engine has become one of the core challenges in the field of intelligent budget control.
[0013] Mainstream prediction engines in current technologies (such as Long Short-Term Memory networks LSTM or Transformer) typically discretize continuous multidimensional data streams using time windows of fixed length and with fixed sliding steps. The core rigid assumption of this processing mode is that the data distribution within the time window is uniform and stationary, and that the states between each time window are independent of each other.
[0014] However, under normal high-concurrency business conditions, the flow of funds and resources has an inherent physical inertia. For example, large-scale fund transfers during centralized settlement and continuous resource scheduling triggered by chain transactions will generate continuous high-frequency fluctuations in the data stream. When such continuous high-frequency fluctuations cross the boundaries of artificially set fixed time windows, they will be abruptly cut off, causing the momentum information that should have been transmitted across the window to undergo spectral aliasing and energy loss at the boundary.
[0015] Existing technologies have failed to recognize this systematic discretization fallacy caused by fixed-step-size discretization: rigid truncation destroys the local high-frequency continuity of data, and the symptomatic solutions adopted by existing technologies, such as increasing model depth and introducing attention mechanisms, are all surface optimizations on the already contaminated and fragmented discrete slices, which cannot fundamentally repair the information loss caused by the underlying truncation.
[0016] Because the prediction engine cannot sense and compensate for energy overflow at the time window boundary, under normal operating conditions, the prediction accuracy of the prediction engine will inevitably collapse periodically at the intersection of each business cycle, which manifests as a normal increase in error volatility and causes severe fluctuations and resource misallocation in downstream control instructions at the beginning of the cycle. This is a common technical bottleneck in the industry that has long existed in large-scale distributed systems.
[0017] Therefore, how to accurately quantify the degree of energy overflow at the edge of the time window, and thereby dynamically adjust the overlap length of adjacent sampling windows and the smooth fusion strategy, fundamentally breaking the rigid assumption of fixed time distance truncation and eliminating the periodic collapse of prediction accuracy at the periodic boundary, has become a key technical problem that urgently needs to be solved in this field.
[0018] Therefore, such as Figure 1 As shown, this application provides a method for dynamic financial forecasting and budget control based on big data, including the following steps: Acquire continuously input business flow data, and perform temporal truncation on the business flow data based on a preset sliding window to generate a first feature tensor and a second feature tensor, as well as the boundary truncation point of the two. Obtain a multidimensional feature vector within a preset neighborhood of the boundary cutoff point, and map the multidimensional feature vector into a feature intensity sequence characterizing the intensity of business fluctuations. Extract the local variability of the feature intensity sequence within the preset neighborhood, and obtain the edge overflow feature of the boundary cutoff point based on the local variability; When the edge overflow feature is greater than a preset reference benchmark, the edge overflow feature is converted into the overlap time length of the first feature tensor and the second feature tensor, and the preset sliding window is updated. Based on the overlap time sequence length, tail data in the first feature tensor and head data in the second feature tensor are obtained. Weighted fusion of the tail data and the head data is performed based on preset weight parameters to obtain a reconstructed feature tensor sequence. Perform temporal state deduction on the reconstructed feature tensor sequence to obtain the business state prediction vector, and generate the corresponding budget resource control instructions; After executing the budget resource control instruction, feedback optimization parameters characterizing the budget control effect are obtained, and the preset reference benchmark is updated based on these parameters.
[0019] Specifically, business transaction data refers to concurrent business logs continuously written to a distributed database. This includes multi-dimensional business characteristic fields such as timestamps, node identifiers, transaction amounts, and resource call volumes, and is continuously written to the system's data acquisition layer in a streaming manner. The preset sliding window is a pre-defined time length parameter used to divide the continuous business transaction data into several adjacent sampling intervals. The first feature tensor and the second feature tensor correspond to the data aggregation results within two temporally adjacent sampling intervals, respectively, and their common connection point is the boundary cutoff point.
[0020] The preset neighborhood refers to a local time segment centered on the boundary cutoff point and extending forward and backward by a certain time length; the multidimensional feature vector refers to the numerical vector containing all business feature dimensions extracted at each time sampling point within this time segment; the feature intensity sequence is a one-dimensional time series obtained by mapping the multidimensional feature vectors at each time sampling point to scalars through norm operations and arranging them in chronological order, used to characterize the comprehensive business fluctuation intensity within the neighborhood of the boundary cutoff point.
[0021] Local variability is the second-order difference value of the feature intensity sequence in the time dimension, representing the acceleration of business fluctuations at that point in time, i.e. the degree of change in the fluctuation trend. The edge overflow feature is a scalar obtained by integrating (discrete summing) the local variability of all time sampling points in the preset neighborhood, representing the total inertia of the business that is abruptly cut off by a fixed time window at the boundary cutoff point. The larger this value is, the more serious the damage to the data continuity caused by the cutoff at the boundary.
[0022] The preset reference benchmark is a threshold preset by the system for filtering regular background noise. The dynamic overlap repair mechanism is only activated when the edge overflow feature exceeds this benchmark. The overlap time length is the number of sampling points that two adjacent feature tensors need to share on the time axis, calculated by logarithmic nonlinear mapping based on the size of the edge overflow feature. Updating the preset sliding window means shortening the actual sliding step of the next sliding window based on the overlap time length, so that there is a substantial overlap area between adjacent windows on the time axis.
[0023] The tail data is a subset of vectors in the first feature tensor that are located within the segment corresponding to the overlapping time sequence length; the head data is a subset of vectors in the second feature tensor that are located within the segment corresponding to the overlapping time sequence length; weighted fusion refers to using an asymmetric smooth weighting function to perform element-wise weighted superposition of the time-aligned tail data and head data according to the correspondence of sampling points, generating smooth transition node data, replacing the fault data at the original boundary, and finally forming a reconstructed feature tensor sequence.
[0024] Temporal state projection involves inputting a reconstructed feature tensor sequence into a pre-defined deep learning prediction model, which outputs a prediction of the business state for future periods, i.e., a business state prediction vector. Budgetary resource control instructions are execution instructions that contain specific resource scheduling parameters, generated based on the difference between the business state prediction vector and a pre-defined resource quota threshold. Feedback optimization parameters are numerical deviations obtained by comparing the actual business execution values with the predicted values after the control instructions are executed. These deviations are used to form a closed-loop feedback and update the pre-defined reference benchmark, thereby achieving long-term adaptive optimization of the system.
[0025] The core innovation of this application lies in: accurately extracting edge overflow features through local second derivative integration, and using this feature to drive logarithmic nonlinear mapping to dynamically adjust the overlap time series length and improve Hamming window weighted fusion. This introduces the anti-aliasing and overlap-add concepts from the field of digital signal processing into time series data stream management, completely eliminating the accuracy collapse defect of the prediction engine at the business cycle boundary from the data preprocessing level, and achieving a deterministic technological advancement with high returns and low overhead.
[0026] This application further proposes a process for performing time-series truncation on the business flow data, specifically including: storing the business flow data sequentially into a delay buffer queue of a preset length to obtain a continuous data stream; extracting multidimensional feature vectors within the first and second adjacent sampling intervals in the continuous data stream based on a preset sliding window; performing matrix encapsulation to generate a first feature tensor and a second feature tensor accordingly; and using the connection time between the first sampling interval and the second sampling interval as the boundary truncation point.
[0027] The delayed buffer queue is a queue set up in the front end of the data stream processing engine with a preset time length. A first-in, first-out (FIFO) data buffer. Its core function is to provide the system with look-ahead capability: when the system processes the current boundary cutoff point... When accessing data at a specific time, the system can legally access the data at that time through a delayed buffer queue. This allows for the extraction of data from the backward neighborhood of the boundary cutoff point without violating causality. For example, when... When set to 180 seconds (3 minutes), the system processes the current time. When the data is obtained, the business flow information within the next 3 minutes can be predicted, so that the symmetrical neighborhood data centered on the boundary cutoff point can be legally obtained, providing complete neighborhood data support for the accurate calculation of the subsequent edge overflow characteristics.
[0028] A continuous data stream refers to a time-series continuous data sequence formed by arranging discretely written business transaction records according to timestamps within a delayed buffer queue. The first sampling interval corresponds to a complete window length on the time axis before the boundary cutoff point, and the second sampling interval corresponds to a complete window length after the boundary cutoff point. Multidimensional business feature data within these two sampling intervals are matrix-encapsulated, that is, the data at each time sampling point... 3D feature vectors are stacked in chronological order into a shape of The two-dimensional matrix is used to generate the first feature tensor and the second feature tensor, respectively. For the number of sampling points, This represents the number of feature dimensions. This standardized matrix representation lays a unified data structure foundation for subsequent feature strength calculations and weighted fusion operations, while also facilitating efficient parallel computation within streaming processing frameworks.
[0029] This application further proposes a process for obtaining a feature intensity sequence, specifically including: obtaining a multi-dimensional feature vector corresponding to each time sampling point within the preset neighborhood; sequentially performing a squaring operation on the feature components of each dimension in the multi-dimensional feature vector to obtain a sum of the squaring results of the feature components of each dimension; and performing a square root operation on the sum to obtain the comprehensive feature intensity of the corresponding time sampling point. The comprehensive feature intensities corresponding to each time sampling point are spliced together in chronological order to obtain the feature intensity sequence.
[0030] Comprehensive feature strength The physical meaning is: for any time within the predefined neighborhood... Its corresponding multidimensional feature vector contains all the features at that moment. The numerical values of each business characteristic dimension (such as transaction amount, resource consumption, concurrent requests, response latency, etc.) are used. Directly using high-dimensional vectors for boundary energy analysis is computationally complex and difficult to form a single criterion. Instead, the L2 norm (Euclidean norm) of the vector is calculated, which represents the feature components of each dimension in the multi-dimensional feature vector. Perform the squaring operation separately, and then perform the squaring operation on all of them. Summing the squares of all values and then taking the square root of that sum compresses high-dimensional multimodal information into a scalar with a clear physical meaning, representing the comprehensive fluctuation intensity of the business state at that moment. Geometrically, this is equivalent to the eigenvectors in... Length in 3D space; ; In the formula, For the first The combined feature intensity of each time sampling point The number of dimensions of the multidimensional feature vector. For the first The th time sampling point Each dimension of feature components. Preset neighborhood. All time sampling points within The values are concatenated in chronological order to obtain a one-dimensional feature intensity sequence. This dimensionality reduction operation preserves the overall fluctuation information of multi-dimensional business features while unifying the input of subsequent boundary energy analysis into a one-dimensional time-series signal, significantly reducing computational complexity. At the same time, it preserves the rotation invariance inherent in the L2 norm, ensuring that the comprehensive feature intensity is independent of the direction of the business feature vector, making it universally applicable to data of different business types.
[0031] This application further proposes a process for obtaining edge overflow features, specifically including: extracting the combined feature intensity of the current time sampling point, its previous time sampling point, and its next time sampling point based on the feature intensity sequence. , , Obtain the local variability of the current time sampling point. Based on the local variability of each time sampling point within the preset neighborhood, a numerical value representing the total inertia of the truncated service is obtained. And use it as a feature of edge overflow.
[0032] Local variation The calculation employs the central difference numerical differentiation method. The first-order difference represents the rate of change of the characteristic intensity sequence (i.e., the speed of business fluctuations), while the second-order difference represents the change in the rate of change itself, i.e., the acceleration of business fluctuations, corresponding to the inertial force of data flow. When data smoothly crosses the boundary, the first-order difference (slope) may not be zero, but the second-order difference (acceleration) approaches zero; however, when there are drastic high-frequency fluctuations or sudden trend changes at the boundary truncation point, the amplitude of the second-order difference will surge significantly, accurately capturing the business inertia of abruptly cut-off points. Local variability at the current time sampling point. By extracting the current time sampling point The sampling point of the previous time. and subsequent sampling points The comprehensive characteristic intensity value is calculated according to the central difference formula; ; In the formula, The time interval between adjacent sampling points is the minimum sampling time interval set at the system's underlying layer. The above formula is known as the second-order central difference approximation in discrete signal processing. It is a standard numerical approximation method for the second derivative of continuous functions, possessing second-order accuracy and achieving an optimal balance between numerical error and computational cost.
[0033] Furthermore, for the preset neighborhood The edge overflow feature is obtained by summing and integrating the local variability of all time sampling points within the time frame. ; ; In the formula, This indicates the corresponding time of the boundary cutoff point. This represents the time length during which the preset neighborhood expands outwards from the boundary cutoff point. Edge overflow characteristic. By discretizing and summing the local variability of all time sampling points within a preset neighborhood, the total inertia of the business momentum within the entire neighborhood is comprehensively measured. The larger this characteristic value, the more severe the actual business fluctuations at the truncated time window boundary, the more severe the information spillover caused by forced truncation, and the higher the degree of distortion of the boundary state of the prediction model. This method of quantifying boundary energy spillover using local second derivative integrals introduces the spectral aliasing sensing theory in the field of digital signal processing into time-series financial data processing in the form of lightweight differential operations, realizing high-precision aliasing risk quantification with O(1) level time complexity, which constitutes a strict causal driving relationship with the subsequent dynamic overlap mapping link.
[0034] This application further proposes a process for updating the preset sliding window, specifically including: adjusting the edge overflow feature. With preset sensitivity coefficient After performing the multiplication operation, the natural logarithm mapping function and the preset scaling factor are used. Obtain the basic overlap length value ; Obtain the overlap time length ,in, The maximum overlap time sequence length is preset. The updated preset sliding window is obtained by subtracting the overlap time sequence length from the current preset sliding window's corresponding time length.
[0035] Overlapping timing length The following formula is used to determine: First, the edge overflow feature is... With preset sensitivity coefficient After performing the multiplication operation, add 1, then take the natural logarithm of the result to obtain the logarithmically transformed overflow strength measure; then, compare this result with a preset scaling factor. Multiply by the base overlap length to obtain the base overlap length value; finally, take the base overlap length value and the preset maximum overlap timing length. The smaller value is used as the actual overlap timing length.
[0036] Preset sensitivity coefficient control Trigger sensitivity for overlap length, preset scaling factor Map the output value of the logarithmic function to the order of magnitude of the actual number of sampling points; The maximum overlap timing length preset for the system is limited by constraints such as memory capacity and communication latency. and All of these are engineering coefficients calibrated based on historical peak data of the system. They can be initially configured through offline statistical analysis and corrected through feedback optimization parameters during long-term system operation.
[0037] The core reason for using the natural logarithmic function instead of a nonlinear mapping is that business fluctuations often exhibit exponential growth. When extremely high concurrency triggers severe overflows, a linear mapping could cause the overlap length to expand explosively, leading to a collapse of system computing resources. The inherent saturated decay damping characteristic of the logarithmic function ensures that even under extreme conditions, the overlap length grows at a sublinear rate, eventually being... Hard constraint truncation ensures that the entire mapping mechanism converges absolutely under both extreme and normal conditions, thus exhibiting engineering robustness.
[0038] When the edge overflow feature When the preset reference threshold is not exceeded, the system maintains the original fixed sliding step size and does not activate the dynamic overlap mechanism, thereby avoiding unnecessary computational overhead; only when... The dynamic overlap calculation process described above is only triggered when the preset reference baseline is exceeded, thus adjusting the overlap time series length. The updated actual sliding step size is obtained by subtracting the preset sliding window time length, so that the next sampling window and the current window are separated by a length of [value missing] on the time axis. The overlapping areas are used to achieve dynamic compensation for boundary truncation.
[0039] This application further proposes a process for obtaining the reconstructed feature tensor sequence, specifically including: obtaining the number of time sampling points corresponding to the overlap temporal length. In the first feature tensor, the reciprocal... The vector subset from the first time sampling point to the last time sampling point is taken as its tail data. The tail data is then multiplied element-wise with the descent weight parameter of the corresponding time sampling point to obtain the first weighted data. In the second feature tensor, the vector subset from the first time sampling point to the last time sampling point is... A subset of vectors from each time sampling point is used as its header data. The header data is multiplied element-wise with the rising weight parameter of the corresponding time sampling point to obtain the second weighted data. Using the boundary cutoff point as a reference, the tail data and the head data are time-aligned to obtain multiple sampling point pairs with consistent timestamps within the corresponding overlapping time length. The first weighted data and the second weighted data corresponding to each sampling point pair are vector-added to obtain smooth transition node data. The tail data and head data of the first feature tensor and the second feature tensor at the boundary cutoff point are replaced according to the smooth transition node data to generate the corresponding reconstructed feature tensor sequence.
[0040] The meaning of timing alignment is: using boundary cutoff points Based on this, the first feature tensor tail data of the first feature tensor The sampling point (counting from the tail forward) and the first sampling point in the second feature tensor head data Each sampling point (counted from head to tail) has the same physical time mapping within the overlapping region; that is, the two points form a sampling point pair, representing a mixed estimate of the business state described by two different sampling directions at the same time. For each sampling point pair, the first weighted data and the second weighted data are vector-added (i.e., all data corresponding to that sampling point pair are vector-added). (Adding each feature dimension sequentially) yields the smooth transition node data at that moment. The smooth transition node data of all sampling point pairs within the overlapping region replaces the fault data of the original feature tensor at the boundary, forming a reconstructed feature tensor sequence.
[0041] The reconstructed sequence achieves a smooth transition from the first feature tensor to the second feature tensor within the overlapping region, eliminating the data jumps caused by the original fixed time window truncation. This ensures that the feature sequence received by the prediction engine locally satisfies the second derivative continuity constraint, guaranteeing that the hidden states of models such as LSTM can be smoothly transmitted at the boundary of the business cycle, thus eliminating the periodic collapse of prediction accuracy from the root.
[0042] This application further proposes a process for obtaining preset weight parameters, specifically including: the preset weight parameters include a decreasing weight parameter and an increasing weight parameter. Wherein, the decreasing weight parameter... Smooth decreasing trend is achieved using a cosine function, with increasing weight parameters. It is obtained by subtracting the descent weight parameter from 1. The calculation formulas for both are as follows: ; ; In the formula, Indicates the first The descent weight parameter for each sampling point relative to its corresponding time sampling point. Indicates the first The ascending weight parameter for each sampling point relative to its corresponding time sampling point. This represents the index value of the sampling point pair within the corresponding overlap time series length. This represents the number of time sampling points corresponding to the overlap time sequence length.
[0043] The above weighting function is an improved asymmetric generalization of the standard Hamming window, when hour, , This indicates that the tail data of the first feature tensor is used completely at the beginning of the overlapping region; when hour, , This indicates that at the end of the overlapping region, the dominant weights have switched to the header data of the second feature tensor. The entire overlapping process presents a smooth transition curve from the first feature tensor to the second feature tensor, completely eliminating the data jumps at the original truncation point.
[0044] This weighting scheme has the following key characteristics: First, it supports arbitrary indexes. Satisfying That is, the energy conservation constraint, which ensures that the weighted fusion operation will not amplify or reduce the absolute amplitude of business features in terms of numerical value, maintain the numerical conservation of business data in the overlapping area, and do not introduce systematic bias. Second, the smooth transition characteristics of the cosine function ensure that the reconstructed feature tensor sequence has continuous first derivatives at both ends of the overlapping region, which approximates the continuity of local second derivatives in an engineering sense, satisfying the requirement for smoothness of the input data of the prediction engine. Third, the Hamming window coefficients (0.54 and 0.46) correspond to better sidelobe suppression effects than rectangular windows in the frequency domain, which can more effectively suppress the high-frequency aliasing components introduced by truncation, providing the theoretically optimal window function selection for the overlap-addition operation from the perspective of signal processing.
[0045] This application further proposes a process for performing time-series state deduction on the reconstructed feature tensor sequence and updating the preset reference benchmark, specifically including: inputting the reconstructed feature tensor sequence into a preset long short-term memory network model, extracting the future time period business feature prediction value output by the long short-term memory network model as a business state prediction vector, comparing the business state prediction vector with the preset resource quota threshold of the target execution node, and generating a budget resource control instruction containing resource scheduling parameters based on the difference between the two. After the target execution node executes the budget resource control instruction, its actual business execution value is recorded, and the numerical deviation between the actual business execution value and the business status prediction vector is used as a feedback optimization parameter. ; Obtain the feedback optimization parameters and the preset tolerance threshold. The difference, combined with the preset learning rate coefficient For the current preset reference benchmark Perform an update to obtain the updated preset reference baseline. .
[0046] ; In the formula, and The minimum and maximum preset reference bases are used to limit the update range of the preset reference bases and prevent them from drifting into invalid intervals.
[0047] The Long Short-Term Memory (LSTM) network model is the implementation vehicle of the prediction engine in this scheme. Its gating mechanism (forget gate, input gate, and output gate) can effectively capture long-range dependencies in time-series data and realize information transfer across time steps through cell states. After receiving the reconstructed feature tensor sequence, the hidden state of the LSTM can achieve seamless cross-period state transfer at the smoothed boundary, avoiding the hidden state jumps caused by the original truncation faults, thus outputting a more stable and accurate business state prediction vector, that is, the prediction result of the resource demand of each business node in the future period.
[0048] The generation logic of the budget resource control instruction is as follows: the business status prediction vector is compared with the preset resource quota threshold of each target execution node one by one. For nodes whose predicted value exceeds the quota threshold, the required expansion capacity is calculated. For nodes whose predicted value is lower than the quota threshold, the shrinkable capacity is calculated. A budget resource control instruction containing parameters such as resource type, scheduling direction (expansion / shrinkage), scheduling amount and execution time is generated and sent to the distributed resource scheduling system for execution.
[0049] Feedback optimization parameters The overall error level of this prediction was quantified. A preset tolerance threshold was set. This represents the system's tolerance range for normal prediction errors. Only when the feedback optimization parameters exceed the tolerance threshold will an effective preset reference benchmark adjustment be triggered. This constitutes the update increment of the reference benchmark, where the preset learning rate coefficients are... The magnitude of each adjustment was controlled. When When this happens, it means the current preset reference benchmark is too high, causing some edge overflow features that should be identified to be missed. The system then increases the recognition sensitivity by decreasing the preset reference benchmark. At the same time, the preset reference benchmark is appropriately raised to reduce unnecessary dynamic overlap triggering and reduce computational overhead.
[0050] Dual constraints ensure that the updated preset reference benchmark is always located at Within the effective range, this prevents parameter runaway caused by extreme feedback values. Through the aforementioned closed-loop feedback mechanism, the system's preset reference benchmark can continuously and adaptively adjust according to the distribution characteristics of business flow data and the historical statistical patterns of prediction errors. This ensures the stability of prediction accuracy when facing data distribution drift during long-term operation, while also eliminating the need for manual periodic parameter calibration, significantly reducing system operation and maintenance costs.
[0051] Through the above steps, the present invention can achieve the following technical effects: 1) By extracting the local change degree of the multidimensional feature vector in the preset neighborhood of the boundary cutoff point, the edge overflow feature is quantified in the form of local second derivative integral, accurately sensing the cross-window energy overflow caused by fixed time window cutoff, fundamentally identifying and quantifying the discretization error that has been ignored in the prior art, providing a driving quantity with clear physical meaning for subsequent dynamic overlap repair, and completely eliminating the periodic collapse of prediction accuracy at the periodic boundary.
[0052] 2) By nonlinearly mapping the edge overflow feature to the overlapping time series length through a logarithmic constraint function, the tail data and head data of adjacent windows are weighted and fused using an asymmetric improved Hamming window to generate a reconstructed feature tensor sequence with continuous local second derivatives. Without increasing the number of parameters and computational complexity of the prediction model, the prediction mean square error at the intersection of time periods is significantly reduced. High-yield, low-cost deterministic technological progress is achieved with an additional computational cost of O(1).
[0053] 3) By introducing feedback optimization parameters and preset learning rate coefficients, the numerical deviation between the actual business execution value and the business status prediction vector is used as a closed-loop optimization signal to dynamically update the preset reference benchmark, enabling the system to continuously adapt to the distribution drift of business flow data and ensure the stability and robustness of budget control accuracy under long-term operation conditions.
[0054] In another implementation, such as Figure 2 As shown, this application also provides a dynamic financial forecasting and budget control system based on big data, including the following modules: The data acquisition module is used to acquire continuously input business transaction log data. Based on a preset sliding window, it performs time-series truncation on the business transaction log data to generate a first feature tensor, a second feature tensor, and their boundary truncation points. The data acquisition module refers to a system component that reads concurrent business transaction logs from a distributed database in real time, manages the delay buffer queue, and encapsulates the feature tensor matrix. Specifically, it can be implemented using a streaming data processing framework combined with a first-in-first-out delay buffer queue. Its function is to provide standardized tensor input for subsequent feature extraction and edge analysis, while simultaneously granting the system compliant forward-looking data access capabilities through the delay buffer queue.
[0055] The vector mapping module is used to obtain multi-dimensional feature vectors within a preset neighborhood of the boundary cutoff point, and map these multi-dimensional feature vectors into a feature intensity sequence representing the intensity of business fluctuations. The vector mapping module is a system component that performs L2 norm dimensionality reduction calculations on the multi-dimensional feature vectors within the neighborhood of the boundary cutoff point. Its core function is to uniformly map high-dimensional multimodal business features to a one-dimensional scalar time-series signal, providing a single comparable input for subsequent local variability calculations, while reducing the dimensionality complexity of subsequent calculations.
[0056] The feature extraction module is used to extract the local variability of the feature intensity sequence within the preset neighborhood and obtain the edge overflow feature of the boundary truncation point based on the local variability. The feature extraction module is a system component that performs central difference second-order differential operations and neighborhood integration operations. It is achieved by performing discrete second-order central difference on the one-dimensional feature intensity sequence and summing all difference values within the neighborhood. It is the core computational node of the entire system for perceiving the danger of time window boundary truncation, and the output edge overflow feature directly drives subsequent dynamic overlap decisions.
[0057] The overlap evaluation module is used to convert the edge overflow feature into the overlap time-series length of the first feature tensor and the second feature tensor when the edge overflow feature is greater than a preset reference benchmark, and to update the preset sliding window. The overlap evaluation module is a system component that performs threshold judgment and logarithmic nonlinear mapping calculation. Its function is to transform abstract edge energy indicators into specific executable sampling window adjustment parameters, driving the subsequent extraction and fusion of overlap data; a logarithmic constraint mapping function is used to ensure that the overlap length does not expand indefinitely under extreme conditions, guaranteeing the system's engineering robustness.
[0058] The sequence reconstruction module is used to obtain the tail data in the first feature tensor and the head data in the second feature tensor based on the overlap time series length. It then performs weighted fusion on the tail data and the head data based on preset weight parameters to obtain a reconstructed feature tensor sequence. The sequence reconstruction module is a system component that performs asymmetric Hamming window weight generation, element-wise weighted multiplication, temporal alignment vector addition, and smooth transition node replacement operations. It is the core functional node of this invention for achieving cross-window data continuity restoration. Its output reconstructed feature tensor sequence satisfies the local second derivative continuity constraint within the overlapping region, providing high-quality smoothing input for the downstream prediction engine.
[0059] The feedback optimization module performs temporal state deduction on the reconstructed feature tensor sequence to obtain a business state prediction vector and generates corresponding budget resource control instructions. After executing the budget resource control instructions, it obtains feedback optimization parameters characterizing the budget control effect and updates the preset reference benchmark based on these parameters. The feedback optimization module integrates three functions: a long short-term memory network prediction engine, a budget resource control instruction generator, and a preset reference benchmark adaptive updater. By constructing a complete closed loop from prediction output to parameter adjustment, it achieves the system's adaptive optimization capability during long-term operation, automatically adapting to the long-term evolution of business pipeline data distribution without manual intervention.
[0060] In another embodiment, this application also provides a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned big data-based dynamic financial forecasting and budget control method. The computer storage medium can be a non-volatile storage medium in various forms such as flash memory, hard disk drive, solid-state drive (SSD), and optical disc (CD / DVD), or a volatile storage medium such as dynamic random access memory (DRAM).
[0061] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.
Claims
1. A dynamic financial forecasting and budget control method based on big data, characterized in that, Includes the following steps: Acquire continuously input business flow data, and perform temporal truncation on the business flow data based on a preset sliding window to generate a first feature tensor and a second feature tensor, as well as the boundary truncation point of the two. Obtain a multidimensional feature vector within a preset neighborhood of the boundary cutoff point, and map the multidimensional feature vector into a feature intensity sequence characterizing the intensity of business fluctuations. Extract the local variability of the feature intensity sequence within the preset neighborhood, and obtain the edge overflow feature of the boundary cutoff point based on the local variability; When the edge overflow feature is greater than a preset reference benchmark, the edge overflow feature is converted into the overlap time length of the first feature tensor and the second feature tensor, and the preset sliding window is updated. Based on the overlap time sequence length, tail data in the first feature tensor and head data in the second feature tensor are obtained. Weighted fusion of the tail data and the head data is performed based on preset weight parameters to obtain a reconstructed feature tensor sequence. Perform temporal state deduction on the reconstructed feature tensor sequence to obtain the business state prediction vector, and generate the corresponding budget resource control instructions; After executing the budget resource control instruction, feedback optimization parameters characterizing the budget control effect are obtained, and the preset reference benchmark is updated based on these parameters.
2. The dynamic financial forecasting and budget control method based on big data according to claim 1, characterized in that, The process of performing time-series truncation on the aforementioned business flow data includes: The business flow data is sequentially stored into a delay buffer queue of a preset length to obtain a continuous data stream. Based on a preset sliding window, multidimensional feature vectors in the first and second sampling intervals that are adjacent to each other are extracted from the continuous data stream. Matrix encapsulation is then performed to generate a first feature tensor and a second feature tensor. The connection time between the first sampling interval and the second sampling interval is taken as the boundary cutoff point.
3. The dynamic financial forecasting and budget control method based on big data according to claim 1, characterized in that, The process of obtaining the feature intensity sequence includes: Obtain the multidimensional feature vector corresponding to each time sampling point within the preset neighborhood. Perform a squaring operation on each feature component of the multidimensional feature vector in turn to obtain the sum of the squaring results of each feature component. Perform a square root operation on the sum to obtain the comprehensive feature intensity of the corresponding time sampling point. ; ; in, For the first The combined feature intensity of each time sampling point The number of dimensions of the multidimensional feature vector. For the first The th time sampling point The feature components of each dimension are concatenated according to the time sequence of the comprehensive feature intensities corresponding to each time sampling point to obtain the feature intensity sequence.
4. The dynamic financial forecasting and budget control method based on big data according to claim 3, characterized in that, The process of obtaining edge overflow characteristics includes: Based on the feature intensity sequence, extract the combined feature intensity of the current time sampling point, its previous time sampling point, and its next time sampling point. , , Obtain the local variability of the current time sampling point. ; ; in, Given the time interval between adjacent time sampling points, a value representing the total inertia of the truncated service is obtained based on the local variability of each time sampling point within the preset neighborhood. And use it as a feature of edge overflow; ; in, This indicates the corresponding time of the boundary cutoff point. This represents the time length during which the preset neighborhood expands outwards from the boundary cutoff point.
5. The dynamic financial forecasting and budget control method based on big data according to claim 1, characterized in that, The process of updating the preset sliding window includes: The edge overflow feature With preset sensitivity coefficient After performing the multiplication operation, the natural logarithm mapping function and the preset scaling factor are used. Obtain the basic overlap length value ; Obtain the overlap time length ,in, The maximum overlap time sequence length is preset. The updated preset sliding window is obtained by subtracting the overlap time sequence length from the current preset sliding window's corresponding time length.
6. The dynamic financial forecasting and budget control method based on big data according to claim 1, characterized in that, The process of obtaining the reconstructed feature tensor sequence includes: Obtain the number of time sampling points corresponding to the overlap time series length. In the first feature tensor, the reciprocal... The vector subset from the first time sampling point to the last time sampling point is taken as its tail data. The tail data is multiplied element-wise with the descent weight parameter of the corresponding time sampling point to obtain the first weighted data. The first time sampling point is transferred to the second feature tensor. A subset of vectors from each time sampling point is used as its header data. The header data is multiplied element-wise with the rising weight parameter of the corresponding time sampling point to obtain the second weighted data. Based on the boundary cutoff point, the tail data and the head data are time-aligned to obtain multiple sampling point pairs with consistent timestamps within the corresponding overlapping time length. The first weighted data and the second weighted data corresponding to each sampling point pair are vector-added to obtain smooth transition node data. Replace the tail and head data of the first and second feature tensors at the boundary truncation point with the smooth transition node data to generate the corresponding reconstructed feature tensor sequence.
7. The dynamic financial forecasting and budget control method based on big data according to claim 6, characterized in that, The process of obtaining the preset weight parameters includes: The preset weight parameters include the decreasing weight parameters and the increasing weight parameters. The decreasing weight parameters... The rising weight parameter ; in, Indicates the first The descent weight parameter for each sampling point relative to its corresponding time sampling point. Indicates the first The ascending weight parameter for each sampling point relative to its corresponding time sampling point. This represents the index value of the sampling point pair within the corresponding overlapping time series length.
8. The dynamic financial forecasting and budget control method based on big data according to claim 1, characterized in that, The process of updating the preset reference benchmark includes: The reconstructed feature tensor sequence is input into a preset long short-term memory network model. The predicted value of future time period business features output by the long short-term memory network model is extracted as a business state prediction vector. The business state prediction vector is compared with the preset resource quota threshold of the target execution node, and a budget resource control instruction containing resource scheduling parameters is generated based on the difference between the two. After the target execution node executes the budget resource control instruction, its actual business execution value is recorded, and the numerical deviation between the actual business execution value and the business status prediction vector is used as a feedback optimization parameter. ; Obtain the feedback optimization parameters and the preset tolerance threshold. The difference, combined with the preset learning rate coefficient For the current preset reference benchmark Perform an update to obtain the updated preset reference baseline. , and These are the preset minimum and maximum preset reference bases.
9. A dynamic financial forecasting and budget control system based on big data, characterized in that: Includes the following modules: The data acquisition module is used to acquire continuously input business flow data, and to perform time-series truncation on the business flow data based on a preset sliding window to generate a first feature tensor and a second feature tensor, as well as the boundary truncation point of the two. The vector mapping module is used to obtain a multi-dimensional feature vector in a preset neighborhood of the boundary cutoff point and map the multi-dimensional feature vector into a feature intensity sequence that characterizes the intensity of business fluctuations. The feature extraction module is used to extract the local variability of the feature intensity sequence within the preset neighborhood, and to obtain the edge overflow feature of the boundary cutoff point based on the local variability. The overlap evaluation module is used to convert the edge overflow feature into the overlap time sequence length of the first feature tensor and the second feature tensor when the edge overflow feature is greater than the preset reference benchmark, and update the preset sliding window. The sequence reconstruction module is used to obtain the tail data in the first feature tensor and the head data in the second feature tensor based on the overlap time length, and to perform weighted fusion on the tail data and the head data based on preset weight parameters to obtain a reconstructed feature tensor sequence. The feedback optimization module is used to perform temporal state deduction on the reconstructed feature tensor sequence to obtain a business state prediction vector and generate a corresponding budget resource control instruction; after executing the budget resource control instruction, it obtains feedback optimization parameters characterizing the budget control effect and updates the preset reference benchmark based on them.
10. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement the dynamic financial forecasting and budget control method based on big data as described in any one of claims 1-8.