Dynamic weighted feature scoring method
Through the dynamic feature weighting method based on the ADF test, using three-point median smoothing and dynamic quantile threshold, combined with QACF and HSIC weighting, the problems of high misjudgment rate and data drift in financial risk control and industrial forecasting are solved, achieving higher data analysis accuracy and adaptability.
Patent Information
- Application Number
- CN202510804192.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies in financial risk control and industrial forecasting have problems such as high misjudgment rate, information loss, and fixed thresholds that cannot adapt to data drift.
A dynamic feature weighting method based on the ADF test is adopted. Through three-point median smoothing and dynamic quantile threshold, combined with QACF and HSIC weighting, composite conditional judgment is realized to ensure that the trend strength exceeds the historical 75% quantile level. The general statistical significance standard p_t<0.05 is used to activate QACF+HSIC weighting or only HSIC weighting.
It significantly reduces the misjudgment rate, improves the AUC value and the accuracy of data analysis, optimizes the weight calculation process, adapts to data drift, and improves the performance of financial risk control and industrial forecasting.
Smart Images

Figure CN120705525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence time series algorithms, and specifically provides a dynamic feature weighting method based on ADF statistics, which is suitable for high-dimensional time series data analysis scenarios such as financial risk control and industrial forecasting. Background Art
[0002] There are three major defects in the existing technology:
[0003] 1. High false positive rate: Single ADF test is sensitive to noise (see: Dickey DA, Fuller WA. Distribution of the Estimators for Autoregressive Time Series with a UnitRoot.
[0004] Journal of the American Statistical Association,1979,74(366):427-431,Section 4);
[0005] 2. Hard switching defect: the choice between autocorrelation and kernel metric leads to information loss (see: Sugiyama M et al. Direct Importance Estimation for Covariate Shift Adaptation.
[0006] Annals of the Institute of Statistical Mathematics, 2008, 2(3):999-1010, Section 5.2); 3. Threshold rigidity: fixed thresholds cannot adapt to data drift (see: Gama J et al. Learning with drift detection.
[0007] Brazilian Symposium on ArtificialIntelligence,2004:286-295,Section 4) Summary of the Invention
[0008] 1. Technical Solution
[0009] Core Process:
[0010] graph TD
[0011] A[Input time series data]-->B[ADF test]
[0012] B-->C[three-point median smoothing]
[0013] C-->D[dynamic quantile threshold]
[0014] D-->E{compound condition judgment}
[0015] E-->| And p_t<0.05F[activate QACF+HSIC weighting]
[0016] E-->|otherwise|G[HSIC weighting only]
[0017] Compound condition design principle:
[0018] Ensure that the trend strength exceeds the historical 75% percentile level, excluding minor perturbations (based on Tukey IQR criterion)
[0019] p_t<0.05 was used as the general statistical significance standard to ensure non-random fluctuations (see Table 4 in [2]).
[0020] 2. Pseudocode
[0021] Dynamic weight calculation (corresponding to step d of weight 1)
[0022]
[0023] BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a flow chart of a time series feature weighting method based on ADF test and dynamic weight generation. DETAILED DESCRIPTION
[0024] 3.3 Experimental Verification
[0025] Experimental design and verification methods:
[0026] Dataset: S&P 500 high-frequency trading data (2014-2024, N=50,000 samples), containing 12 features (such as transaction amount, timestamp, merchant category).
[0027] Experimental setup:
[0028] Sample division: 10-fold cross-validation, repeated 30 times and averaged.
[0029] Class balance: Fraud samples account for 0.1% (processed by SMOTE oversampling).
[0030] Hardware environment: Intel Xeon Gold 6348 CPU, 128GB RAM, Python 3.9 implementation.
[0031] Evaluation indicators: AUC (area under the curve), false positive rate (False Positive Rate).
[0032] Compare to baseline
[0033] Comparison Method False positive rate (↓) AUC(↑) Calculation time (s) Fixed Weight Method [1] 0.28±0.03 0.72 8.2 ADF threshold method [2] 0.19±0.02 0.81 12.7 The present invention 0.09±0.01 0.91 15.3
[0034] Significance test (α=0.01):
[0035] The present invention vs. fixed weight method: t = 6.32, p = 3.2 × 10 -7
[0036] The present invention vs. ADF threshold method: t = 4.15, p = 1.8 × 10 -5
[0037] 3.4 Parameter optimization analysis
[0038] 1. Sensitivity of window length (W) and lag order (L):
[0039] AUC values for different parameter combinations tested on the S&P 500 dataset:
[0040] W\L 3 5 7 10 50 0.82 0.85 0.87 0.84 80 0.86 0.89 **0.91** 0.88 100 0.87 0.90 0.90 0.87 200 0.83 0.85 0.86 0.84
[0041] Note 1: When W>150, the performance is reduced due to the decay of historical data timeliness, but it is still usable (such as economic cycle analysis).
[0042] Note 2: L = 10 is applicable to long lag scenarios (such as power load forecasting)
[0043] Conclusion: When W∈[80,100] and L=7, the performance is optimal (AUC≥0.90), which is consistent with the parameter range of claim 1.
[0044] 2. Quantile threshold verification:
[0045] Based on Tukey's IQR criterion (reference [3]), the 0.75 quantile balances sensitivity and robustness (Q3 + 1.5 × IQR):
[0046] Quantile False positive rate AUC 0.60 0.15 0.82 0.70 0.12 0.86 **0.75** **0.09** **0.91** 0.90 0.23 0.75
[0047] 3. Zero threshold ε stability:
[0048] According to the IEEE 754 double precision standard (reference [4]), ε∈[10 -8 ,10 -6] when the weight volatility is < 0.3%:
[0049] ε Weighted Volatility AUC fluctuation range <![CDATA[10 -8 ]]> ±0.3% 0.90±0.002 <![CDATA[**10 -6 **]]> **±0.1%** **0.91±0.001**
[0050] References
[0051] [1] Wind Financial Terminal User Manual v5.2, 2023
[0052] [2]Dickey DA et al.(1979)*JASA*74(366):427Table 4
[0053] [3]Tukey JW.*Exploratory Data Analysis*.1977:43
[0054] [4]IEEE Std 754-2019.
Claims
1. Sovereign Regenerator, a feature dynamic weighting method based on ADF statistics, characterized by, The following steps are involved: a) Perform an ADF test on the input sequence at time t to obtain the ADF statistic And the corresponding p-value p_t; b) Calculate smoothing statistics: c) Based on the past W moments Calculate the dynamic threshold: d) When satisfied And p_t<0.05: β t =1-α t Otherwise: α t =0,β t =1; e) Output weighted score of feature Xk: S t (k) = α t ·QACF(X k )+β t ·HSIC(X k ,Y), where ε∈[10 -8 ,10 -6 ] is used to avoid division by zero. The dynamic weighted scoring can be achieved through a sliding window averaging method.
2. A delayed correlator, as claimed in claim 1, wherein QACF(X k ) is calculated as: And the lag order L∈[3,10].
3. Gaussian kernel processor, the method as claimed in claim 1, wherein HSIC (X k ,Y) is calculated using the independence criterion based on Gaussian kernel.
4. A window optimizer, as claimed in the method of claim 1, wherein the dynamic threshold calculation window W∈[50,200].
5. A cross-domain adapter, the method as described in claim 1, is applicable to financial risk control, wind power forecasting, and industrial sensor network scenarios. Instruction manual experimental comparison Parameter sensitivity test: >Experiments show that the best effect is achieved when W=100, L=5.