Commodity abnormal behavior early warning method and system based on correlation analysis

By designing a dynamic threshold determination method and filtering formula, combined with K-means clustering and principal component analysis, multiple optimal time lag values ​​are identified, and multi-dimensional correlation analysis of products, platforms, regions, etc. is achieved. This solves the problems of single perspective and noise interference in correlation analysis in existing technologies, and provides a more comprehensive technical application.

CN120672380APending Publication Date: 2025-09-19SHANDONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510810761.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing time series data association analysis methods cannot effectively shield the common trends between subjects, are subject to noise interference, have a single analysis perspective, cannot identify multiple groups of time-lagged association patterns, have a limited scope of application, and are difficult to meet the needs of cross-subject, cross-platform, and cross-regional linkage warnings.

Method used

The implementation plan includes: using new equipment, materials, processes or combinations, etc., which reflects the innovative methods adopted by the applicant.

Benefits of technology

It realizes multi-angle and multi-dimensional correlation analysis of different subjects such as commodities, platforms, and regions, identifies multi-optimal lag value correlation patterns, and provides data support and decision support for cross-subject, cross-platform, and cross-region linkage early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672380A_ABST
    Figure CN120672380A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of transaction monitoring, and provides a commodity abnormal behavior early warning method and system based on correlation analysis, and the method comprises the steps: obtaining the transaction data of a plurality of commodities, and processing the transaction data into a real time sequence: obtaining a prediction time sequence through a time sequence prediction model based on the real time sequence, calculating the difference between the real time sequence and the predicted time sequence to obtain an original fluctuation sequence of each commodity; for a real time sequence, calculating a mean value and a standard deviation in the sliding window, taking a weighted value of the mean value and the standard deviation as a dynamic threshold value, and filtering an original fluctuation sequence by adopting a filtering function to obtain a fluctuation sequence; based on the fluctuation sequences of the two commodities, calculating a synchronous association strength sequence and a lagging association strength sequence of the two commodities; and when a certain commodity is abnormal, carrying out early warning on other commodities according to the synchronous association strength sequence and the lagged association strength sequence. And cross-subject, cross-platform and cross-region linkage early warning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of transaction monitoring, and in particular relates to a method and system for early warning of abnormal commodity behavior based on association analysis. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of online transactions, a large number of transactions have occurred on online trading platforms, and the accompanying transaction risks have also intensified. The current market lacks efficient and intelligent online transaction risk early warning methods, which cannot effectively meet the needs of coordinated early warning across entities, platforms, and regions. Online transaction risk supervision faces severe challenges.

[0004] To address this issue, there is an urgent need for a technology that can deeply analyze the temporal correlations between different entities in historical transaction data, supporting coordinated risk warnings for multiple dimensions such as commodities, platforms, and regions. For example, if correlation analysis reveals a strong contemporaneous correlation between commodity A and commodity B, and real-time data detects a risk for commodity A, a coordinated risk warning for commodity B should be issued simultaneously. If correlation analysis reveals optimal time lags of 5 and 21 days between e-commerce platforms α and β, then when a risk is detected on platform α, the correlation analysis can be used to promptly issue warnings for platform β 5 and 21 days later. Similarly, if a strong contemporaneous correlation is determined between Province A and provinces B, F, H, and K, then when a risk is detected in Province A, the system can automatically issue coordinated risk warnings for provinces B, F, H, and K.

[0005] Existing time series data association analysis methods have the following problems: First, existing association analysis algorithms are generally unable to effectively screen out common trends between entities. Year-round abnormal behavior time series data for various commodities are influenced by many common factors, exhibiting certain common fluctuation trends, such as seasonal fluctuations and fluctuations caused by holidays or specific events (such as "Double Eleven" and "May 18th"). In reality, these common fluctuations are meaningless for identifying essential connections between commodities. We need an algorithm that can screen out common trends and focus solely on the unique fluctuations of each commodity. In this case, only when two commodities share unique fluctuations can we consider them to have a dependent or influential relationship.

[0006] Second, there is noise interference after the fluctuation series is extracted. Sometimes the deviation between the predicted value and the true value is not necessarily due to actual fluctuations, but rather noise generated during the prediction process. An algorithm is needed to clarify the boundary between noise and true fluctuations, masking the noise while emphasizing the true fluctuations. However, current algorithms do not use a dynamic approach to selecting the threshold for noise masking, making them inadequate for non-stationary time series affected by complex factors. For example, during a certain period of time, when the sales market is affected by seasonal factors, holidays, special events, and other factors, abnormal product behavior is more likely to occur in large numbers and exhibit unstable trends, making it more susceptible to noise. Therefore, the threshold should be set higher, while the opposite should be true.

[0007] Third, existing algorithms for analyzing associations in time series data often suffer from limited applicability and a single analytical perspective. Most algorithms lack both holistic and targeted analytical tools, making it difficult to simultaneously support global association pattern mining (such as cluster analysis) and refined analysis between specific entities. Furthermore, most existing algorithms only support analysis of contemporaneous associations, failing to fully uncover time-lagged association patterns between data entities. Overall, existing algorithms are insufficiently comprehensive in terms of the level, perspective, and temporal coverage of association analysis, failing to meet the diverse needs of association patterns and, consequently, struggling to adapt to the early warning needs of complex online transaction environments.

[0008] Fourth, when using the CCF (Cross-Correlation Function) for cross-correlation analysis, existing algorithms often overlook the presence of multiple time-lagged associations. For example, product A may have two time-lagged leading effects on product B: one with a 7-day lag, meaning product A leads product B by approximately one week; the other with a 58-day lag, meaning product A leads product B by approximately two months. Current association analysis methods lack the ability to define and identify multiple optimal time-lagged values, making it difficult to fully exploit the existence of multiple time-lagged association patterns. Summary of the Invention

[0009] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a method and system for warning of abnormal commodity behavior based on association analysis, which can effectively eliminate the common trends between different subjects and focus on the essential associations, and use dynamic threshold determination methods and filtering formulas to shield noise and emphasize real fluctuations on non-stationary data. At the same time, it supports multi-dimensional and multi-level association analysis, and identifies multi-optimal lag value association patterns. It can realize multi-angle and multi-dimensional association analysis of historical transaction data of different subjects such as commodities, platforms, and regions, and provide data support and decision support for cross-subject, cross-platform, and cross-regional linkage warnings.

[0010] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides a method for early warning of abnormal commodity behavior based on association analysis, comprising: Get the transaction data of several commodities and process them into real time series: Based on the real time series, the predicted time series is obtained through the time series prediction model. The difference between the real time series and the predicted time series is calculated to obtain the original fluctuation series of each commodity. For real time series, the mean and standard deviation within the sliding window are calculated, and the weighted value of the mean and standard deviation is used as the dynamic threshold. Based on the dynamic threshold, the filter function is used to filter the original fluctuation sequence to obtain the fluctuation sequence; Based on the fluctuation series of the two commodities, the concurrent correlation strength series and the lagged correlation strength series of the two commodities are calculated; When an abnormality occurs in a certain product, an early warning will be issued for other products based on the concurrent correlation strength sequence and the lagged correlation strength sequence.

[0011] Furthermore, the filtering function is: ; in, x is a fluctuation value in the original fluctuation sequence, β is the dynamic threshold, α is the amplification factor.

[0012] Furthermore, the contemporaneous correlation strength in the contemporaneous correlation strength sequence is: ; in, is the fluctuation value at each time slice in the fluctuation sequence of category A commodities, is the fluctuation value at each time slice in the fluctuation series of category b commodities, and are the means of the fluctuation values ​​at all time slices in the fluctuation series of category A and category B commodities respectively.

[0013] Furthermore, it also includes: performing K-means clustering on all commodities based on the fluctuation sequence, and based on the clustering results, using the principal component analysis method to reduce the dimension of the fluctuation sequence and then performing visual representation.

[0014] Furthermore, the method further includes: finding all local peaks in the lag correlation strength sequence and arranging them in descending order, and saving the lag values ​​of the local peaks. and the corresponding lagged correlation strength As a candidate set; set similarity threshold and correlation threshold , traverse the candidate set if and , then the hysteresis value and Join the set of optimal time lag values; if or , the screening ends.

[0015] A second aspect of the present invention provides a commodity abnormal behavior early warning system based on association analysis, comprising: The file management module is configured to: obtain transaction data of several commodities and process it into real time series: based on the real time series, obtain a predicted time series through a time series prediction model; calculate the difference between the real time series and the predicted time series to obtain the original fluctuation series of each commodity; The module for targeted analysis of concurrent correlation is configured to: for a real time series, calculate the mean and standard deviation within a sliding window, use the weighted value of the mean and standard deviation as a dynamic threshold, and filter the original fluctuation sequence using a filter function based on the dynamic threshold to obtain a fluctuation sequence; A cross-correlation analysis module is configured to: calculate a concurrent correlation strength series and a lagged correlation strength series of the two commodities based on the fluctuation series of the two commodities; The early warning module is configured to: when an abnormality occurs in a certain product, issue an early warning to other products based on the concurrent correlation strength sequence and the lagged correlation strength sequence.

[0016] Furthermore, the filtering function is: ; in, x is a fluctuation value in the original fluctuation sequence, β is the dynamic threshold, α is the amplification factor.

[0017] Furthermore, it also includes a concurrent correlation holistic analysis module, which is configured as follows: Based on the fluctuation sequence, all commodities are clustered using K-means, and based on the clustering results, the fluctuation sequence is reduced in dimension using the principal component analysis method and then visualized.

[0018] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for warning of abnormal commodity behavior based on association analysis.

[0019] The fourth aspect of the present invention provides a computer device, comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein when the processor executes the program, the steps of the method for warning of abnormal commodity behavior based on association analysis as described above are implemented.

[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention achieves the shielding of common fluctuations and the extraction of unique fluctuations in association analysis: an overall prediction model for all commodities is trained to allow the model to learn the common fluctuation trends of abnormal behaviors of commodities, and then the model is used to make predictions for each commodity separately, and the deviation between the predicted value and the true value is used as the unique fluctuation of each commodity. This achieves the shielding of common fluctuations in association mining and focuses on unique fluctuations, thereby mining the essential associations.

[0021] The present invention realizes the fluctuation sequence filtering processing on non-stationary time series data: a dynamic threshold determination algorithm is designed, and the statistical characteristics of different sliding window data are used to calculate the boundary value between the noise and the real fluctuation belonging to the sliding window, thereby avoiding the problem of interference of non-stationary time series data on filtering; at the same time, a filtering function is designed to achieve the goal of amplifying the real fluctuation while shielding the noise, providing more accurate and clear fluctuation sequence data for subsequent association mining.

[0022] This invention addresses the issues of single-level perspective and incomplete time series coverage in time series data association analysis. By combining holistic cluster analysis with targeted analysis, it achieves both a macroscopic understanding of the overall commodity landscape and a series of targeted analyses of risky commodities. By combining concurrent association mining with cross-correlation mining, it accurately locates high-risk entities when risks occur and identifies the timing and likelihood of future risk occurrences. This provides multi-level, cross-temporal and spatial data support for cross-subject coordinated early warning across commodities, platforms, and regions.

[0023] This invention solves the problem of identifying cross-correlations at multiple time scales. By employing a set of optimal time lag value definition and identification algorithms, the proposed method ensures that the correlation strengths corresponding to the optimal time lag value groups found form relatively independent peak regions. This algorithm improves the stability and significance of correlation analysis, ensuring that the selected lag value combinations have real correlation significance and more effectively characterizing multi-scale time lag correlation patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0025] Figure 1 This is a flow chart of a method for early warning of abnormal commodity behavior based on correlation analysis according to the first embodiment of the present invention; Figure 2 This is a flow chart of a fluctuation sequence extraction algorithm according to the first embodiment of the present invention; Figure 3 Schematic diagram of the filtering algorithm flow in the first embodiment of the present invention; Figure 4 Schematic diagram of the cross-correlation analysis algorithm flow in the first embodiment of the present invention; Figure 5 This is an architecture diagram of a product abnormal behavior early warning system based on correlation analysis according to the second embodiment of the present invention; Figure 6 It is a structural diagram of a computer device according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0027] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0028] Example 1 This embodiment provides a method for early warning of abnormal commodity behavior based on association analysis.

[0029] This embodiment provides a method for warning of abnormal commodity behavior based on association analysis. First, a prediction model that has learned the overall fluctuation trend is used to predict each commodity. The deviation between the predicted value and the true value is used as the original fluctuation sequence. The fluctuation sequence of non-stationary time series data is filtered through a dynamic threshold setting algorithm and a filtering formula. Then, a comprehensive and multi-dimensional correlation analysis is performed on each subject by calculating the Pearson correlation coefficient, K-means clustering, CCF function calculation, and a multi-optimal time lag value definition and identification algorithm. This can provide data support for cross-subject risk linkage warnings across commodities, platforms, regions, and other areas.

[0030] This embodiment provides a method for warning of abnormal commodity behavior based on association analysis, such as Figure 1 As shown, the following steps are included: Step 1: Preprocess the raw transaction data uploaded and selected by the user in the file management module: select days as time slices, count the number of abnormal transactions on each date as the value assigned to each time slice; if there is no transaction data for a certain product on a certain date, assign a value of -1 to that time slice for that product; standardize or normalize the data to eliminate scale differences between different products, thereby obtaining a time series of the number of abnormal behaviors for each product throughout the year. The series length of each product is equal, equal to the total number of days in the year.

[0031] For example, the time series of the number of abnormal behaviors of a certain product throughout the year is: 0, 0, 0, 0, 0, 0, 4, 0, 0,0, 0, 7, 0, 0, 0, 0, 0, 0, 0, 0, 3, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 8, 0,0, 0, 1, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…, where each number represents the number of abnormal behaviors of the product on that day, and the length of the sequence is the total number of days in the year.

[0032] The raw data in this example covers transaction data for 100 commodities in 2023. Each data entry contains fields such as commodity ID, transaction time, and whether it is abnormal.

[0033] The data is labeled, and any anomalies are anomalies. The cause of the anomaly may be suspected violations in the transaction, etc.

[0034] In this example, the linkage warning subject selected by the user is the product, that is, the user needs to perform correlation analysis on different products.

[0035] The number of abnormal transactions of each commodity on each date is counted as the value of each time slice, thereby generating an abnormal behavior sequence for each commodity.

[0036] Step 2: If Figure 2 As shown in the figure, a time series prediction model is trained using the time series of all commodities. The trained time series prediction model is then used to predict the time series of the number of abnormal behaviors for each commodity. The difference between the predicted value and the true value of each time slice is taken as the fluctuation, thereby deriving the original fluctuation series of each commodity.

[0037] Step 201: Select a time series prediction model. In this example, the LSTM model is selected and trained based on the time series data of all products. The time series data of all products are integrated into a training set to train the LSTM model, ensuring that the LSTM model learns the overall patterns and common trends of all products. Step 202: Evaluate the prediction performance of the time series prediction model by using cross-validation or other methods to ensure that the time series prediction model understands the overall time series pattern.

[0038] Step 203: For each time slice of the 100 commodities, the difference between the actual value and the predicted value is calculated, and the difference is regarded as the fluctuation value; these differences are regarded as the fluctuation sequence of each commodity and recorded in the result data set.

[0039] Step 3: If Figure 3 As shown in the figure, the original fluctuation sequences of 100 commodities are processed to remove noise while amplifying the real fluctuations, thereby obtaining the final fluctuation sequence of each commodity for subsequent analysis.

[0040] Step 301: Determine the noise-fluctuation threshold β by a dynamic threshold determination method.

[0041] Step 30101. Select an appropriate sliding window size for the dynamic threshold setting method. Select the sliding window size N, which is the number of historical data points referenced each time the dynamic threshold is calculated. The selection of the window size N should be consistent with the fluctuation characteristics of the data. It should not be too small (sensitive to short-term fluctuations) nor too large (unable to keep up with changes). In this example, 30 days is selected.

[0042] Step 30102: In each sliding window, calculate the mean μ and standard deviation σ of the N data points in the current window; the mean μ reflects the central tendency of the data in the window; the standard deviation σ represents the fluctuation range of the data in the window, which is used to evaluate the discreteness of the data.

[0043] The calculation formulas for the mean and standard deviation are as follows: Mean: ; Standard Deviation: ; in, is each data point in the sliding window. In this example, it is the number of abnormal behaviors of a certain product in each time slice within the time window. N is the window size, which is 30 in this example.

[0044] The explanation of the above formula is: The mean calculation process of the current sliding window data point is: the sum of all data points in the current sliding window The number of data points in the current sliding window; the standard deviation of the data points in the current sliding window is calculated as follows: the sum of the squares of the differences between all data points in the current sliding window and the mean The number of data points in the current sliding window.

[0045] Step 30103: Set the dynamic threshold β at the current time point based on the calculated mean μ and standard deviation σ: ; Here, k is the adjustment factor, typically 1.5 or 2, but in this example, 2 is used. The value of k determines the sensitivity of the threshold. A larger k value results in a higher threshold, making it more sensitive to larger fluctuations. A smaller k value results in a lower threshold, making it more sensitive to smaller fluctuations.

[0046] The explanation of the above formula is: The calculation process of the dynamic threshold at the current time point is: the mean of the current sliding window data points Adjustment Factor The standard deviation of the data points in the current sliding window.

[0047] Step 30104: As new data points arrive, the sliding window moves forward, updates the mean and standard deviation of the latest N data points, and recalculates the dynamic threshold β accordingly; removes the earliest data point in the window and incorporates the latest data point into the calculation to keep the window size unchanged; recalculates the mean and standard deviation so that the threshold is always based on the latest data changes.

[0048] Step 302: For each commodity, the fluctuation value of each time slice is determined according to the dynamically determined threshold and filter function. Shielding noise and emphasizing real fluctuations.

[0049] in, As shown below: ; Among them, x is the fluctuation value, β is the threshold value determined dynamically above, and α is the amplification factor, which is 1.5 in this example. When the absolute value of the fluctuation When the absolute value of the fluctuation is less than or equal to the threshold, the output is directly set to 0, indicating that fluctuations less than or equal to the threshold are ignored; when ... When , that is, to amplify the part exceeding the threshold and maintain the direction of the original fluctuation; here express The sign function ensures the consistency of positive and negative signs.

[0050] The explanation of the above formula is: the calculation process of the filtered fluctuation value of a certain commodity in a certain time slice is: if the absolute value of the original fluctuation of the time slice is The dynamic threshold at this moment, then the filtered fluctuation value is 0; if the absolute value of the original fluctuation of the time slice is The dynamic threshold at this moment, then the absolute value of the filtered fluctuation value is the amplification factor The difference between the absolute value of the original fluctuation and the dynamic threshold has the same sign as the original fluctuation value of the time slice.

[0051] Step 4: Calculate the Pearson correlation coefficient of the fluctuation series of two commodities to conduct contemporaneous correlation mining for commodity pairs, achieving targeted analysis of contemporaneous correlations. By calculating and ranking the Pearson correlation coefficients of each commodity's fluctuation series, information on the strongest correlations for a given commodity and a comparison chart of abnormal behavior sequences are provided, or the contemporaneous correlation strength and abnormal behavior sequence comparison chart for a given commodity pair are provided, providing targeted data support for linkage warnings.

[0052] For example, if a user specifies a product, information about the products with the strongest associations with it and a detailed comparison chart of abnormal behavior sequences will be displayed. If a user specifies a pair of products, the association between the two products and a comparison chart of abnormal behavior sequences will be displayed.

[0053] In this example, we select Category A and Category B products to calculate the correlation strength during the same period. The Pearson correlation coefficient is calculated as follows: ; in, and is each pair of values ​​in the data, and They are and In this example, is the fluctuation value of category A commodities in each time slice, is the fluctuation value of category b goods in each time slice; and are the means of fluctuation values ​​of category A and category B commodities over all time slices respectively.

[0054] Step 5: K-means clustering is performed on the 100 products based on their fluctuation sequences. Based on the clustering results, principal component analysis is used to reduce the dimensionality of the product fluctuation sequences. This allows for visualization of the clustering results and a holistic analysis of contemporaneous correlations. This identifies clusters of strongly correlated products (i.e., products that fall into the same category), providing comprehensive data support for coordinated alerts. The Pearson correlation coefficients for all pairs of products are also displayed, allowing for detailed information and a comparison chart of abnormal behavior sequences.

[0055] Because clustering is based on the fluctuation sequence of each product, which consists of many values, visualizing the clustering results in a vector space suitable for display on a two-dimensional computer screen requires dimensionality reduction before visualization. In other words, clustering is performed using the fluctuation sequence, but the clustering results are visualized using the reduced dimensionality vectors.

[0056] In this example, the K-means classification is set to 10 categories, and the output dimension of the principal component analysis is set to 2 dimensions.

[0057] In this example, 100 products are clustered into 10 categories through K-means clustering, and the clustering results are visualized using vectors reduced in dimensionality using principal component analysis, allowing users to understand the overall situation of product associations. In addition, a sequential list of the association strengths of all product pairs is provided. Clicking to view the list displays the details of the association, and a sequence comparison diagram of abnormal behavior of two products is provided.

[0058] Step 6: Figure 4 As shown in the figure, the CCF function between two products is calculated. Then, an optimal set of time lag values ​​is selected, ensuring that their corresponding correlation strengths are sufficiently close and significantly different from subsequent lag values, and that the correlation strength exceeds a set threshold. This optimal set of time lag values ​​and their corresponding correlation strengths are obtained, thus enabling cross-correlation analysis (correlation analysis with time lags) for the product pairs. The CCF function graph is displayed, and a cross-correlation analysis report is generated, indicating whether the two products have a strong correlation at a certain time lag value. If so, the optimal time lag value and the corresponding lagged correlation strength are provided.

[0059] For example, if the user specifies leading product a and led product b, the CCF function graph will be presented, and a cross-correlation analysis report will be given. The specific content includes: whether product a has a leading effect on product b in terms of a certain time lag value; if so, all leading time lag values ​​and corresponding correlation strengths will be given.

[0060] Step 601: Calculate the lag values ​​between the two sequences, from lag 1 to max_lag, and their corresponding lag correlation strength sequence, ccf_values. Specifically, one of the sequences is shifted backward or forward by different lag steps, and the correlation of the overlapping portions of the two sequences is calculated at each lag position, using the Pearson correlation coefficient as the metric. This yields the correlation strength corresponding to each lag step, ultimately forming a mapping between lags and correlations.

[0061] max_lag is set as follows: ; in, is the sequence length, i.e. the number of days in a year.

[0062] In this example, the year is 2023, and the total number of days is 365, so the sequence length is 365, so max_lag is selected as 182, that is, .

[0063] Step 602: Find all local peaks in ccf_values ​​and arrange them in descending order, and save the lag values ​​lag and corresponding correlation strengths ccf_value of these local peaks as a candidate set candidate_lags.

[0064] Step 603: Set a similarity threshold , used to determine whether the correlation strength of the lagged values ​​is similar; set a correlation threshold , used to determine whether there is a strong association. In this example, Take 0.1, The above parameters are set based on business experience and expert domain knowledge, making the algorithm more flexible and controllable, suitable for specific data, and ensuring that the results conform to business logic. , for each local peak , and compare it to the strength of association at the next lagged value. If ,and , then the hysteresis value is considered and The correlation strength is similar and has a strong correlation, so add them .if or , the screening ends. The final optimal time lag value set is , The time lag values ​​in represent the time intervals at which the two products exhibit the strongest temporal correlation. This means that changes in one series may affect the other series after these lags. This set can be used in feature construction, causal analysis, or time series forecasting. For example, using lagged features as model input can improve the ability to predict future trends or events in the target series.

[0065] Step 7: When an abnormality occurs in a certain product, an early warning is issued for other products based on the concurrent correlation strength and the lagged correlation strength.

[0066] Example 2 This embodiment provides a commodity abnormal behavior early warning system based on association analysis, such as Figure 5 As shown, it specifically includes: The file management module is configured to upload, download, delete, process, select, and query user files; obtain transaction data for several commodities and process it into real time series: based on the real time series, a predicted time series is obtained through a time series prediction model; the difference between the real time series and the predicted time series is calculated to obtain the original fluctuation series for each commodity; the obtained fluctuation series for all commodities are stored in the form of files in a server-side folder; the selection function assigns client variables, and all association analyses are performed around the selected data; The module for targeted concurrent correlation analysis is configured to: calculate the mean and standard deviation of a real time series within a sliding window, use the weighted value of the mean and standard deviation as a dynamic threshold, and filter the original fluctuation sequence using a filtering function based on the dynamic threshold to obtain a fluctuation sequence; display information and sequence comparison charts of the strongest correlations between specified products, and display information and sequence comparison charts of the correlation strength of specified product pairs; The module for analyzing the overall contemporaneous correlation is configured to: perform K-means clustering on all products based on the fluctuation sequence. Based on the clustering results, the fluctuation sequence is reduced in dimension using principal component analysis and then visualized. The module also displays a product clustering result graph and a list of the correlation strength order of all product pairs, allowing for the review of abnormal behavior sequence comparison graphs and detailed information. A cross-correlation analysis module is configured to: calculate the concurrent correlation strength series and the lagged correlation strength series of the two commodities based on the fluctuation series of the two commodities; display the CCF function graph of a specified leading-leading commodity pair and inform the cross-correlation analysis conclusion of the two commodities; The early warning module is configured to: when an abnormality occurs in a certain product, issue an early warning to other products based on the concurrent correlation strength sequence and the lagged correlation strength sequence.

[0067] The filtering function is: ; Among them, x is a fluctuation value in the original fluctuation sequence, β is the dynamic threshold, and α is the amplification factor.

[0068] The contemporaneous correlation strength in the contemporaneous correlation strength sequence is: ; in, is the fluctuation value at each time slice in the fluctuation sequence of category A commodities, is the fluctuation value at each time slice in the fluctuation series of category b commodities, and are the means of the fluctuation values ​​at all time slices in the fluctuation series of category A and category B commodities respectively.

[0069] The cross-correlation analysis module is further configured to: find all local peaks in the lagged correlation strength sequence and sort them in descending order, and save the lagged values ​​of the local peaks and the corresponding lagged correlation strength As a candidate set; set similarity threshold and correlation threshold , traverse the candidate set if and , then the hysteresis value and Join the set of optimal time lag values; if or , the screening ends.

[0070] This embodiment provides a commodity abnormal behavior early warning system based on association analysis, which can implement file name management, status management, and storage management of multi-user files.

[0071] This embodiment provides a product abnormal behavior warning system based on association analysis, which realizes multi-angle and comprehensive association analysis functions, including cluster analysis, strong association query for a certain product, association ranking table and detailed analysis of all product pairs, association analysis for specified product pairs, cross-correlation analysis for specified product pairs, etc.

[0072] This embodiment provides a commodity abnormal behavior early warning system based on association analysis, which can realize manual setting of parameters such as thresholds, improves flexibility and controllability, can be applied to specific data, and can integrate expert domain knowledge into the analysis process to make the results more consistent with business logic.

[0073] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.

[0074] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for early warning of abnormal commodity behavior based on association analysis as described in the first embodiment above are implemented.

[0075] Example 4 This embodiment provides a computer device, such as Figure 6As shown, the system includes a display device, an input device, a computer-readable storage medium (volatile memory and non-volatile storage medium), a processor, a communication interface (i.e., a network interface), and a computer program stored on the computer-readable storage medium and executable by the processor. The processor, the communication interface, and the computer-readable storage medium may be connected via a bus or other means. The communication interface is used to receive and transmit data, and when the processor executes the program, it implements the steps of the method for early warning of abnormal commodity behavior based on association analysis as described in the first embodiment above.

[0076] Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0080] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for early warning of abnormal commodity behavior based on association analysis, characterized in that: include: Get the transaction data of several commodities and process them into real time series: Based on the real time series, the predicted time series is obtained through the time series prediction model. The difference between the real time series and the predicted time series is calculated to obtain the original fluctuation series of each commodity. For real time series, the mean and standard deviation within the sliding window are calculated, and the weighted value of the mean and standard deviation is used as the dynamic threshold. Based on the dynamic threshold, the filter function is used to filter the original fluctuation sequence to obtain the fluctuation sequence; Based on the fluctuation series of the two commodities, the concurrent correlation strength series and the lagged correlation strength series of the two commodities are calculated; When an abnormality occurs in a certain product, an early warning will be issued for other products based on the concurrent correlation strength sequence and the lagged correlation strength sequence.

2. The method for early warning of abnormal commodity behavior based on association analysis according to claim 1, characterized in that: The filtering function is: ; in, x is a fluctuation value in the original fluctuation sequence, β is the dynamic threshold, α is the amplification factor.

3. The method for early warning of abnormal commodity behavior based on association analysis according to claim 1, characterized in that: The contemporaneous correlation strength in the contemporaneous correlation strength sequence is: ; in, is the fluctuation value at each time slice in the fluctuation sequence of category A commodities, is the fluctuation value at each time slice in the fluctuation series of category b commodities, and are the means of the fluctuation values ​​at all time slices in the fluctuation series of category A and category B commodities respectively.

4. The method for early warning of abnormal commodity behavior based on association analysis according to claim 1, characterized in that: Also includes: Based on the fluctuation sequence, K-means clustering is performed on all commodities. Based on the clustering results, the fluctuation sequence is reduced in dimension using principal component analysis and then visualized.

5. The method for early warning of abnormal commodity behavior based on association analysis according to claim 1, characterized in that: Also includes: Find all local peaks in the lagged correlation strength sequence and sort them in descending order, saving the lagged values ​​of the local peaks and the corresponding lagged correlation strength As a candidate set; set similarity threshold and correlation threshold , traverse the candidate set if and , then the hysteresis value and Join the set of optimal time lag values; if or , the screening ends.

6. A commodity abnormal behavior early warning system based on association analysis, characterized in that: include: The file management module is configured to: obtain transaction data of several commodities and process it into real time series: based on the real time series, obtain a predicted time series through a time series prediction model; calculate the difference between the real time series and the predicted time series to obtain the original fluctuation series of each commodity; The module for targeted analysis of concurrent correlation is configured to: for a real time series, calculate the mean and standard deviation within a sliding window, use the weighted value of the mean and standard deviation as a dynamic threshold, and filter the original fluctuation sequence using a filter function based on the dynamic threshold to obtain a fluctuation sequence; A cross-correlation analysis module is configured to: calculate a concurrent correlation strength series and a lagged correlation strength series of the two commodities based on the fluctuation series of the two commodities; The early warning module is configured to: when an abnormality occurs in a certain product, issue an early warning to other products based on the concurrent correlation strength sequence and the lagged correlation strength sequence.

7. The commodity abnormal behavior early warning system based on association analysis according to claim 6, characterized in that: The filtering function is: ; in, x is a fluctuation value in the original fluctuation sequence, β is the dynamic threshold, α is the amplification factor.

8. The commodity abnormal behavior early warning system based on association analysis according to claim 6, characterized in that: It also includes a concurrent correlation holistic analysis module, which is configured as follows: Based on the fluctuation sequence, all commodities are clustered using K-means. Based on the clustering results, the fluctuation sequence is reduced in dimension using principal component analysis and then visualized.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for early warning of abnormal commodity behavior based on association analysis according to any one of claims 1 to 5 are implemented.

10. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein: When the processor executes the program, the steps of the method for early warning of abnormal commodity behavior based on association analysis according to any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Early warning method and system for ventilation equipment of smart laboratory

    CN121140142A