Data relevance mining method of big data analysis platform

By collecting and analyzing the fluctuation trend of data on the big data analysis platform, calculating strain slope data and using regression algorithms, the problem of low accuracy of traditional correlation analysis is solved, and a higher precision data correlation analysis is achieved.

CN120197682APending Publication Date: 2025-06-24BEIJING FABO HONGYE TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304015.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditionally, the correlation results of data are analyzed through the similarity of data fluctuations and trends. The correlation results obtained are too general and the correlation analysis accuracy is low, making it easy to obtain two unrelated data whose analysis results are related.

Method used

The data correlation mining method of the big data analysis platform is adopted to collect data of the correlation term to be analyzed for a fixed time period, establish the change trend coordinate system of the first historical correlation data and the second historical correlation data, calculate the strain slope data, and calculate the correlation coefficient ratio and residual coefficient ratio based on the regression algorithm to determine the correlation coefficient ratio.

Benefits of technology

By more accurately reflecting the change correspondence between parent sequence data and sub sequence data, the correlation of data fluctuations can be measured with higher accuracy, avoid interference from unrelated data, and improve the accuracy of data correlation analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197682A_ABST
    Figure CN120197682A_ABST
Patent Text Reader

Abstract

The invention discloses a data relevance mining method for a big data analysis platform, which comprises the following steps of: acquiring to-be-analyzed relevance item data with fixed time duration to obtain historical first relevance data and historical second relevance data, and determining a relevance task of the to-be-analyzed relevance item data; establishing a change trend coordinate system of the historical first associated data and the historical second associated data, and calculating first strain slope data and second strain slope data; calculating first strain regression slope data of the first strain slope data and the second strain slope data based on a regression algorithm; changing item associated data of the historical second associated data, and calculating second strain slope data corresponding to the historical first associated data; and the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data and the second strain slope data are calculated to judge the correlation of the data, so that the analysis precision of the data correlation can be greatly improved, and the anti-interference performance of data correlation analysis is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and particularly relates to a method for mining data relevance in a big data analysis platform. Background Art

[0002] In recent years, with the rapid development of big data, it is very necessary to mine the value and relevance between data. Data association analysis is an important part of data mining. It can help us discover the association relationships between data, so as to discover the knowledge hidden in the data. In the business field, data association analysis can help enterprises understand customers' purchase behaviors, thereby increasing sales. In the medical field, data association analysis can help doctors understand patients' medical histories, thereby improving the diagnostic accuracy. In the financial field, it is used to evaluate the price associations between different assets, so as to better manage the risks of investment portfolios, and for the association analysis between the production link and sales to optimize the supply chain process. The analysis of relevance can provide valuable information for decision-making, help us better understand and predict various phenomena and trends, and thus make more informed decisions in many fields such as business, finance, and social sciences.

[0003] Traditional association analysis methods often analyze the relevance of data by calculating the support data and confidence data of the data. However, if the support threshold is too high, many potentially meaningful patterns are deleted because they contain items with small support. If the support threshold is too low, the calculation cost is very high and a large number of association patterns are generated. The confidence measure ignores the support of the item set in the rule consequent. Analyzing the relevance of data through the similarity of the fluctuation trends of the data results in overly general association results, with relatively low association analysis accuracy, and it is easy to obtain two unrelated data whose analysis results are associated. Summary of the Invention

[0004] The technical problem solved by the present invention is that: when analyzing the relevance of data through the similarity of the fluctuation trends of the data in the traditional way, the obtained association results are too general, the association analysis accuracy is relatively low, and it is easy to obtain two unrelated data whose analysis results are associated.

[0005] To solve the above technical problem, the present invention provides the following technical solution: A method for mining data relevance in a big data analysis platform, including the following steps: Step S100, collect the data of the associated items to be analyzed for a fixed time duration, obtain the historical first association data and historical second association data of the data of the associated items to be analyzed, and determine the association task of the data of the associated items to be analyzed; Step S200, establish a change trend coordinate system for the historical first association data and historical second association data, and calculate the first strain slope data of the historical first association data and the second strain slope data of the second association data; Step S300, calculate the first strain regression slope data of the first strain slope data and the second strain slope data based on a regression algorithm; Step S400, change the item association data of the historical second association data, and calculate the second strain slope data corresponding to the historical first association data at this time based on the first strain regression slope data; Step S500, calculate the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data and the second strain slope data to determine the relevance of the data.

[0006] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, wherein: Step S100 specifically includes: Step S101, retrieve the database storage platform to obtain the data of the item to be analyzed for association, and the data of the item to be analyzed for association includes historical first association data and historical second association data; Step S102, intercept the historical first association data and the historical second association data with a fixed time duration; Step S103, use the historical first association data as the mother sequence data, and use the historical second association data as the subsequence data; Step S104, determine the association tasks of the data of the item to be analyzed for association, and the association tasks include a first association task and a second association task.

[0007] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, wherein: when the number of item association data of the subsequence data is less than 1, determine the association task as the first association task; When the number of item association data of the subsequence data is greater than 1, determine the association task as the second association task.

[0008] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, wherein: Step S200 specifically includes: Step S201, establish a change trend coordinate system for the historical first association data and the historical second association data, use the change amounts of the mother sequence data and the subsequence data as the ordinate, and use the fixed time duration as the abscissa to establish the change trend coordinate system for the historical first association data and the historical second association data, and the length of the change trend coordinate system is equal to the fixed time duration; Step S202, divide the fixed time duration into N duration sub-regions, and the mother sequence data and the subsequence data are divided into corresponding N data sub-regions; Step S203: Calculate the strain slope data of the mother region of the mother sequence data and the strain slope data of the mother region of the subsequence data for the N data sub-regions respectively. The mathematical expression is: Strain slope data of the mother region = (End slope value data of the sub-region of the mother sequence data - Start slope value data of the sub-region of the mother sequence data) / (End time point of the duration sub-region - Start time point of the end time point of the duration sub-region); Strain slope data of the sub-region = (End slope value data of the sub-region of the subsequence data - Start slope value data of the sub-region of the subsequence data) / (End time point of the duration sub-region - Start time point of the duration sub-region); Step S204: Integrate the N strain slope data of the mother regions into a mother vector data set (Y1, Y2, Y3... YN) to obtain the first strain slope data of the historical first correlation data; Integrate the N strain slope data of the sub-regions to obtain a mother vector data set (X1, X2, X3... XN) to obtain the second strain slope data of the historical second correlation data.

[0009] As a preferred solution of the data correlation mining method of the big data analysis platform described in the present invention, wherein: Step S300 specifically includes: Step S301: Input the first strain slope data (Y1, Y2, Y3... YN) and the second strain slope data (X1, X2, X3... XN) into N linear regression equations. The mathematical expression of the linear regression equation is: Y1 = aX1 + b, where a represents the correlation coefficient and b represents the residual coefficient, to obtain N correlation coefficients a1, a2, a3... an and N residual coefficients b1, b2, b3... bn; Step S302: Add Y1 and Y2 in the first strain slope data (Y1, Y2, Y3... YN) and X1 and X2 in the second strain slope data (X1, X2, X3... XN), and input them into the linear regression equation to obtain N / 2 correlation coefficients a1, a2, a3... an / 2 and N / 2 residual coefficients b1, b2, b3... bn / 2; Add Y1, Y2, and Y3 in the first strain slope data (Y1, Y2, Y3... YN) and X1, X2, and X3 in the second strain slope data (X1, X2, X3... XN), and input them into the linear regression equation to obtain N / 3 correlation coefficients a1, a2, a3... an / 3 and N / 3 residual coefficients b1, b2, b3... bn / 3. By gradually changing the size of N, N = 1, 2, 3... N, calculate the corresponding correlation coefficients and residual coefficients; Step S303: Integrate the calculated correlation coefficient and residual coefficient to obtain an N×N-dimensional matrix. The first row of the matrix has the first row data as a1, a2, a3... an, b1, b2, b3... bn. The second row of the matrix is a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0..., and the last row of the matrix is a1, 0, 0,... b1, 0, 0,....

[0010] Step S304: Multiply the preset N-dimensional weight matrix by the N×N-dimensional matrix to obtain the first strain regression slope data of the first strain slope data and the second strain slope data.

[0011] As a preferred solution of the data correlation mining method of a big data analysis platform according to the present invention, wherein: Step S400 specifically includes: Step S401: When the association task is the second association task, change the item association data of the historical second association data. The number of X1 in the second strain slope data (X1, X2, X3... XN) is equal to the number of item association data K; Step S402: Input the first strain slope data (Y1, Y2, Y3... YN) and the second strain slope data (X1, X2, X3... XN) into N linear regression equations to obtain N correlation coefficients a1, a2, a3... an and N residual coefficients b1, b2, b3... bn. The number of the correlation coefficient a1 is equal to the number of item association data K, and the number of the residual coefficient b1 is equal to the number of item association data K; Step S403: Integrate the calculated correlation coefficient and residual coefficient to obtain an N×N×K-dimensional matrix. The first row of the matrix is a1, a2, a3... an, b1, b2, b3... bn. The second row of the matrix is a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0..., and the last row of the matrix is a1, 0, 0,... b1, 0, 0,... 0. The number of the correlation coefficient a1 is equal to the number of item association data K, and the number of the residual coefficient b1 is equal to the number of item association data K; Step S404: Multiply the preset N-dimensional weight matrix by the N×N×K-dimensional matrix to obtain the second strain regression slope data.

[0012] As a preferred solution of the data correlation mining method of a big data analysis platform according to the present invention, wherein: Step S500 specifically includes: Step S501: Successively combine the data of each row after multiplying the preset N - dimensional weight matrix by the N×N - dimensional matrix. The combination rules include a first combination rule and a second combination rule; Step S502: Calculate the correlation coefficient ratio of the correlation coefficient a1 and the correlation coefficient ratio of the residual coefficient b1 of the combined matrix; Step S503: According to different association tasks, use different association comparison methods to judge the relevance of data. The association comparison methods include a first association comparison method and a second association comparison method.

[0013] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, among them: Compare a1 and a2 in the first - row matrix a1, a2, a3... an, b1, b2, b3... bn with a1 in the second - row matrix a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0... 0; When a2 - a1>0 and a1>0, combine according to the first combination rule; When a2 - a1<0 and a1<0, combine according to the first combination rule; When a2 - a1<0 and a1>0, combine according to the second combination rule; When a2 - a1>0 and a1<0, combine according to the second combination rule.

[0014] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, among them: The specific steps of combining according to the first combination rule include: Add a1 and a2 in the first - row matrix a1, a2, a3... an, b1, b2, b3... bn to a1 in the second - row matrix a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0... 0; The specific steps of combining according to the second combination rule include: Remove a1 and a2 in the first - row matrix a1, a2, a3... an, b1, b2, b3... bn.

[0015] As a preferred solution of the data relevance mining method of a big data analysis platform according to the present invention, among them: When the association task is the first association task, select the first association comparison method. The first association comparison method specifically includes: Calculate the correlation coefficient ratio and the correlation coefficient ratio; Add all the combined correlation coefficients a1 and residual coefficients b1 to obtain the total a1 and total b1, and add all the correlation coefficients a1 and residual coefficients b1 before combination to obtain the total a2 and total b2; The correlation coefficient ratio = a1 total / a2 total, and the residual coefficient ratio = b1 total / b2 total; the larger the correlation coefficient ratio and the smaller the residual coefficient ratio, the greater the data correlation. When the association task is the second association task, select the second association comparison method. The first association comparison method specifically includes: calculating the correlation coefficient ratio and the residual coefficient ratio between the first strain regression slope data and the second strain regression slope data. The smaller the difference between the correlation coefficient ratio and the residual coefficient ratio, the smaller the data correlation. The larger the difference between the correlation coefficient ratio and the residual coefficient ratio, the greater the data correlation.

[0016] Advantages of the present invention: By dividing the data of the association item to be analyzed with a fixed time duration into a mother sequence data and a sub-sequence data, using the change amounts of the mother sequence data and the sub-sequence data as the ordinate and the fixed time duration as the abscissa to establish a change trend coordinate system, the change trend directions of the mother sequence data and the sub-sequence data can be intuitively observed. Divide the mother sequence data and the sub-sequence data into corresponding N data sub-regions, calculate the mother region strain slope data and the sub-region strain slope data, and input the mother region strain slope data and the sub-region strain slope data into a linear regression equation to obtain the correlation coefficient and the residual coefficient of the mother sequence data and the sub-sequence data in each data sub-region, which can more accurately reflect the change correspondence relationship between the mother sequence data and the sub-sequence data. Integrate all the correlation coefficients and residual coefficients to obtain an N×N-dimensional matrix, multiply the preset N-dimensional weight matrix by the N×N-dimensional matrix, obtain the first strain regression slope data and merge them row by row according to the merging rule, calculate the corresponding correlation coefficient ratio and residual coefficient ratio, which can well measure the correlation of the local data fluctuations of the mother sequence data and the sub-sequence data, and can accurately know the magnitude of the correlation of the change trends of the mother sequence data and the sub-sequence data. When the number of item association data of the mother sequence data is greater than 1, compare the difference between the correlation coefficient ratio and the residual coefficient ratio, avoiding the interference of uncorrelated data on the data correlation analysis, and making the result accuracy of the data correlation analysis higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic flow chart of a method for mining data correlation of a big data analysis platform provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.

[0019] Embodiment, refer to Figure 1, as an embodiment of the present invention, provides a method for mining data correlation in a big data analysis platform, including the following steps: Step S100, collect the data of the associated items to be analyzed for a fixed time duration, obtain the historical first associated data and historical second associated data of the data of the associated items to be analyzed, and determine the association task of the data of the associated items to be analyzed; Step S200, establish a change trend coordinate system for the historical first associated data and historical second associated data, and calculate the first strain slope data of the historical first associated data and the second strain slope data of the second associated data; Step S300, calculate the first strain regression slope data of the first strain slope data and the second strain slope data based on the regression algorithm; Step S400, change the item associated data of the historical second associated data, and calculate the second strain slope data of the corresponding historical first associated data at this time based on the first strain regression slope data; Step S500, calculate the correlation coefficient ratio and residual coefficient ratio of the first strain regression slope data and the second strain slope data to determine the data correlation.

[0020] Step S100 specifically includes: Step S101, retrieve the database storage platform to obtain the data of the associated items to be analyzed, and the data of the associated items to be analyzed includes historical first associated data and historical second associated data; Step S102, intercept the historical first associated data and historical second associated data for a fixed time duration; Step S103, use the historical first associated data as the mother sequence data and the historical second associated data as the sub-sequence data; Step S104, determine the association task of the data of the associated items to be analyzed, and the association task includes the first association task and the second association task.

[0021] In this embodiment, the data of the associated items to be analyzed for a fixed time duration in the database storage platform is retrieved. The fixed time duration is one quarter. The data of the associated items to be analyzed includes mother sequence data and sub-sequence data. The change of the sub-sequence data will cause the change of the mother sequence data. The association task of the data of the associated items to be analyzed is determined as the first association task or the second association task.

[0022] When the number of item associated data of the sub-sequence data is less than 1, determine the association task as the first association task; When the number of item associated data of the sub-sequence data is greater than 1, determine the association task as the second association task.

[0023] In this embodiment, the association task of the associated item data to be analyzed is determined according to the number of item association data of the subsequence data. The first association task is one-to-one data association analysis, and the second association task is many-to-one data association analysis.

[0024] Step S200 specifically includes: Step S201: Establish a change trend coordinate system for the historical first association data and the historical second association data. Using the change amounts of the parent sequence data and the subsequence data as the ordinate and the fixed time duration as the abscissa, establish the change trend coordinate system for the historical first association data and the historical second association data. The length of the change trend coordinate system is equal to the fixed time duration. Step S202: Divide the fixed time duration into N duration sub-regions, and the parent sequence data and the subsequence data are divided into corresponding N data sub-regions. Step S203: Calculate the parent region strain slope data of the parent sequence data and the parent region strain slope data of the subsequence data for the N data sub-regions respectively. The mathematical expressions are: Parent region strain slope data = (End slope value data of the sub-region of the parent sequence data - Start slope value data of the sub-region of the parent sequence data) / (End time point of the duration sub-region - Start time point of the end time point of the duration sub-region); Sub-region strain slope data = (End slope value data of the sub-region of the subsequence data - Start slope value data of the sub-region of the subsequence data) / (End time point of the duration sub-region - Start time point of the duration sub-region). Step S204: Integrate the N parent region strain slope data into a parent vector data set (Y1, Y2, Y3... YN) to obtain the first strain slope data of the historical first association data. Integrate the N sub-region strain slope data to obtain a parent vector data set (X1, X2, X3... XN) to obtain the second strain slope data of the historical second association data.

[0025] In this embodiment, the associated item data to be analyzed with a fixed time duration is divided into parent sequence data and subsequence data. Using the change amounts of the parent sequence data and the subsequence data as the ordinate and the fixed time duration as the abscissa to establish a change trend coordinate system can visually observe the change trend directions of the parent sequence data and the subsequence data. Dividing the parent sequence data and the subsequence data into corresponding N data sub-regions and calculating the parent region strain slope data and the sub-region strain slope data can take into account the influence effect of the change of the sub-region strain slope data within each data sub-region on the parent region strain slope data, making the data association analysis more coherent and comprehensive.

[0026] Step S300 specifically includes: Step S301: Input the first strain slope data (Y1, Y2, Y3... YN) and the second strain slope data (X1, X2, X3... XN) into N linear regression equations. The mathematical expression of the linear regression equation is: Y1 = aX1 + b, where a represents the correlation coefficient and b represents the residual coefficient, to obtain N correlation coefficients a1, a2, a3... an and N residual coefficients b1, b2, b3... bn; Step S302: Add Y1 and Y2 in the first strain slope data (Y1, Y2, Y3... YN) and X1 and X2 in the second strain slope data (X1, X2, X3... XN), and input them into the linear regression equation to obtain N / 2 correlation coefficients a1, a2, a3... an / 2 and N / 2 residual coefficients b1, b2, b3... bn / 2; Add Y1, Y2, and Y3 in the first strain slope data (Y1, Y2, Y3... YN) and X1, X2, and X3 in the second strain slope data (X1, X2, X3... XN), and input them into the linear regression equation to obtain N / 3 correlation coefficients a1, a2, a3... an / 3 and N / 3 residual coefficients b1, b2, b3... bn / 3. And so on, gradually change the value of N, where N = 1, 2, 3... N, and calculate the corresponding correlation coefficients and residual coefficients; Step S303: Integrate the calculated correlation coefficients and residual coefficients to obtain an N×N-dimensional matrix. The first row matrix has the first row data as a1, a2, a3... an, b1, b2, b3... bn, the second row matrix is a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0..., and the last row matrix is a1, 0, 0,... b1, 0, 0,...

[0027] Step S304: Multiply the preset N-dimensional weight matrix by the N×N-dimensional matrix to obtain the first strain regression slope data of the first strain slope data and the second strain slope data.

[0028] In this embodiment, inputting the first strain slope data and the second strain slope data into the linear regression equation can obtain the correlation coefficient and the residual coefficient of the parent sequence data and the subsequence data of each data sub-region, which can more accurately reflect the change correspondence relationship between the parent sequence data and the subsequence data. At this time, the number of the correlation coefficient and the residual coefficient = N / N. Gradually change the size of the denominator N, N = 1, 2, 3... N, and integrate all the correlation coefficients and the residual coefficients to obtain an N×N-dimensional matrix. The N-dimensional weight matrix follows a standard normal distribution, and the first strain regression slope data of the parent sequence data and the subsequence data is obtained. The first strain regression slope data can reflect the change of the data correlation of the sub-region data of different sizes.

[0029] Step S400 specifically includes: Step S401, when the association task is the second association task, change the item association data of the historical second association data. The number of X1 in the second strain slope data (X1, X2, X3... XN) is equal to the number of the item association data K; Step S402, input the first strain slope data (Y1, Y2, Y3... YN) and the second strain slope data (X1, X2, X3... XN) into N linear regression equations to obtain N correlation coefficients a1, a2, a3... an and N residual coefficients b1, b2, b3... bn. The number of the correlation coefficient a1 is equal to the number of the item association data K, and the number of the residual coefficient b1 is equal to the number of the item association data K; Step S403, integrate the calculated correlation coefficients and residual coefficients to obtain an N×N×K-dimensional matrix. The first row matrix is a1, a2, a3... an, b1, b2, b3... bn, the second row matrix is a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0..., and the last row matrix is a1, 0, 0,... b1, 0, 0,... 0. The number of the correlation coefficient a1 is equal to the number of the item association data K, and the number of the residual coefficient b1 is equal to the number of the item association data K; Step S404, multiply the preset N-dimensional weight matrix by the N×N×K-dimensional matrix to obtain the second strain regression slope data.

[0030] In this embodiment, when the associated task is the second associated task, the item associated data of the historical second associated data is changed. The number of each item of data corresponding to the second strain slope data is equal to the number of item associated data of the subsequence data. The number of correlation coefficients and residual coefficients included in each item of data of the N×N×K-dimensional matrix is equal to the number of item associated data of the subsequence data. Multiply the N-dimensional weight matrix by the N×N×K-dimensional matrix to obtain the second strain regression slope data, and the second strain regression slope data can represent the combined influence of multiple item associated data on the mother sequence data.

[0031] Step S500 specifically includes: Step S501, successively combine the data of each row after multiplying the preset N-dimensional weight matrix by the N×N-dimensional matrix. The combination rules include the first combination rule and the second combination rule; Step S502, calculate the correlation coefficient ratio of the correlation coefficient a1 and the correlation coefficient ratio of the residual coefficient b1 of the combined matrix; Step S503, according to the different associated tasks, use different association comparison methods to judge the relevance of the data. The association comparison methods include the first association comparison method and the second association comparison method.

[0032] Compare a1 and a2 in the first row matrix a1, a2, a3... an, b1, b2, b3... bn with a1 in the second row matrix a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0... 0; When a2 - a1 > 0 and a1 > 0, combine according to the first combination rule; When a2 - a1 < 0 and a1 < 0, combine according to the first combination rule; When a2 - a1 < 0 and a1 > 0, combine according to the second combination rule; When a2 - a1 > 0 and a1 < 0, combine according to the second combination rule.

[0033] In this embodiment, by comparing the magnitude relationship of the correlation coefficients of each row of the first strain regression slope data and the second strain regression slope data, the corresponding combination rule is obtained. According to different combination rules, as the sub-region data increases from small to large, eliminate the data of the corresponding items that do not conform, and leave the data of the corresponding items that conform until the last row is combined.

[0034] The specific steps of the first combination rule combination include: add a1 and a2 in the first row matrix a1, a2, a3... an, b1, b2, b3... bn to a1 in the second row matrix a1, a2, a3... an / 2, 0, 0,... b1, b2, b3... bn / 2, 0, 0... 0; The specific steps of the second merging rule include: removing a1 and a2 from the first row matrices a1, a2, a3... an, b1, b2, b3... bn.

[0035] When the associated task is the first associated task, select the first associated comparison method. The first associated comparison method specifically includes: calculating the correlation coefficient ratio and the correlation coefficient ratio; adding all the merged correlation coefficients a1 and the residual coefficients b1 to obtain the total of a1 and b1, and adding all the correlation coefficients a1 and the residual coefficients b1 before merging to obtain the total of a2 and b2. The correlation coefficient ratio = a1 total / a2 total, and the residual coefficient ratio = b1 total / b2 total; the larger the correlation coefficient ratio and the smaller the residual coefficient ratio, the greater the data correlation. When the associated task is the second associated task, select the second associated comparison method. The first associated comparison method specifically includes: calculating the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data and the second strain regression slope data. The smaller the difference between the correlation coefficient ratio and the residual coefficient ratio, the smaller the data correlation. The larger the difference between the correlation coefficient ratio and the residual coefficient ratio, the greater the data correlation.

[0036] In this embodiment, the larger the correlation coefficient ratio and the smaller the residual coefficient ratio, the more synchronous the changes of the subsequence data and the parent sequence data. The smaller the correlation coefficient ratio and the larger the residual coefficient ratio, the more asynchronous the changes of the subsequence data and the parent sequence data. When the associated task is the second associated task, select the second associated comparison method, calculate the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data and the second strain regression slope data. The smaller the difference between the correlation coefficient ratio and the residual coefficient ratio, the smaller the change of the item-associated data on the parent sequence data, and the smaller the data correlation. The larger the difference between the correlation coefficient ratio and the residual coefficient ratio, the greater the change of the item-associated data on the parent sequence data, and the greater the data correlation.

[0037] By dividing the associated item data to be analyzed with a fixed time duration into parent sequence data and child sequence data, using the change amounts of the parent sequence data and the child sequence data as the ordinate and the fixed time duration as the abscissa to establish a change trend coordinate system, the change trend directions of the parent sequence data and the child sequence data can be visually observed. The parent sequence data and the child sequence data are divided into corresponding N data sub-regions, the strain slope data of the parent region and the strain slope data of the sub-regions are calculated, and inputting the strain slope data of the parent region and the strain slope data of the sub-regions into a linear regression equation can obtain the correlation coefficient and the residual coefficient of the parent sequence data and the child sequence data in each data sub-region, which can more accurately reflect the change correspondence relationship between the parent sequence data and the child sequence data. Integrating all the correlation coefficients and residual coefficients to obtain an N×N-dimensional matrix, multiplying the preset N-dimensional weight matrix by the N×N-dimensional matrix to obtain the first strain regression slope data and merging them row by row according to the merging rule, calculating the corresponding correlation coefficient ratio and residual coefficient ratio, when it can well measure the correlation of the local data fluctuations of the parent sequence data and the child sequence data, the magnitude of the correlation of the change trends of the parent sequence data and the child sequence data can be accurately known. When the number of item correlation data of the parent sequence data is greater than 1, comparing the difference between the correlation coefficient ratio and the residual coefficient ratio avoids the interference of uncorrelated data on the data correlation analysis, making the result accuracy of analyzing the data correlation higher.

[0038] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk, or optical disk. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 functions specified in one box or multiple boxes.

[0039] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A data correlation mining method for a big data analysis platform, characterized in that: include: Step S100, collecting the to-be-analyzed associated item data of a fixed time duration, obtaining the historical first associated data and the historical second associated data of the to-be-analyzed associated item data, and determining the associated tasks of the to-be-analyzed associated item data; Step S200, establishing a change trend coordinate system of the historical first associated data and the historical second associated data, and calculating first strain slope data of the historical first associated data and second strain slope data of the second associated data; Step S300, calculating first strain regression slope data of the first strain slope data and the second strain slope data based on a regression algorithm; Step S400, changing the item association data of the historical second association data, and calculating the second strain slope data corresponding to the historical first association data based on the first strain regression slope data; Step S500, calculating the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data and the second strain slope data to determine the correlation of the data.

2. The data correlation mining method of a big data analysis platform according to claim 1, characterized in that: Step S100 specifically includes: Step S101, calling a database storage platform to obtain the to-be-analyzed associated item data, wherein the to-be-analyzed associated item data includes historical first associated data and historical second associated data; Step S102, intercepting the historical first associated data and the historical second associated data of a fixed time duration; Step S103, taking the historical first associated data as parent sequence data, and taking the historical second associated data as child sequence data; Step S104: determining the associated tasks of the associated item data to be analyzed, wherein the associated tasks include a first associated task and a second associated task.

3. The data correlation mining method of a big data analysis platform according to claim 2, characterized in that: When the number of item-related data of the subsequence data is less than 1, determining the related task to be the first related task; When the number of item-associated data of the subsequence data is greater than 1, the associated task is determined to be a second associated task.

4. The data correlation mining method of a big data analysis platform according to claim 3, characterized in that: Step S200 specifically includes: Step S201, establishing a change trend coordinate system of the historical first associated data and the historical second associated data, using the change amount of the parent sequence data and the child sequence data as the ordinate and the fixed time duration as the abscissa, to establish a change trend coordinate system of the historical first associated data and the historical second associated data, wherein the length of the change trend coordinate system is equal to the fixed time duration; Step S202, dividing the fixed time duration into N duration sub-regions, and the parent sequence data and the child sequence data are divided into corresponding N data sub-regions; Step S203, respectively calculating the parent region strain slope data of the parent sequence data and the parent region strain slope data of the sub-sequence data of the N data sub-regions, the mathematical expression is: parent region strain slope data = end slope value data of the sub-region of the parent sequence data - start slope value data of the sub-region of the parent sequence data / end time point of the duration sub-region - start time point of the end time point of the duration sub-region; Sub-region strain slope data = end slope value data of sub-region of sub-sequence data - start slope value data of sub-region of sub-sequence data / end time point of duration sub-region - start time point of duration sub-region; Step S204, integrating the strain slope data of N parent regions into a parent vector data set (Y1, Y2, Y3 . . . YN) to obtain the first strain slope data of the historical first associated data; The strain slope data of the N sub-regions are integrated to obtain a mother vector data set (X1, X2, X3 . . . XN), and the second strain slope data of the historical second correlation data are obtained.

5. The data correlation mining method of a big data analysis platform according to claim 4, characterized in that: Step S300 specifically includes: Step S301, inputting the first strain slope data (Y1, Y2, Y3...YN) and the second strain slope data (X1, X2, X3...XN) into N linear regression equations, the mathematical expression of the linear regression equation is: Y1=aX1+b, wherein a represents the correlation coefficient, b represents the residual coefficient, and obtains N correlation coefficients a1, a2, a3...an, and N residual coefficients b1, b2, b3...bn; Step S302, adding Y1, Y2 in the first strain slope data (Y1, Y2, Y3 ... YN) and X1, X2 in the second strain slope data (X1, X2, X3 ... XN), inputting the linear regression equation, and obtaining N / 2 correlation coefficients a1, a2, a3 ... an / 2, and N / 2 residual coefficients b1, b2, b3 ... bn / 2; Add Y1, Y2, Y3 in the first strain slope data (Y1, Y2, Y3...YN) and X1, X2, X3 in the second strain slope data (X1, X2, X3...XN), input the linear regression equation, obtain N / 3 correlation coefficients a1, a2, a3...an / 3, N / 3 residual coefficients b1, b2, b3...bn / 3, and gradually change the size of N, N=1, 2, 3...N, and calculate the corresponding correlation coefficients and residual coefficients; Step S303, integrating the calculated correlation coefficients and residual coefficients to obtain an N×N dimensional matrix, wherein the first row of the matrix is ​​a1, a2, a3 ... an, b1, b2, b3 ... bn, the second row of the matrix is ​​a1, a2, a3 ... an / 2, 0, 0, ... b1, b2, b3 ... bn / 2, 0, 0 ..., and the last row of the matrix is ​​a1, 0, 0, ... b1, 0, 0, ...; Step S304 : multiplying a preset N-dimensional weight matrix by the N×N-dimensional matrix to obtain first strain regression slope data of the first strain slope data and the second strain slope data.

6. The data correlation mining method of a big data analysis platform according to claim 5, characterized in that: Step S400 specifically includes: Step S401, when the associated task is a second associated task, changing the item associated data of the historical second associated data, the number of X1 of the second strain slope data (X1, X2, X3 . . . XN) is equal to the number K of the item associated data; Step S402, inputting the first strain slope data (Y1, Y2, Y3 ... YN) and the second strain slope data (X1, X2, X3 ... XN) into N linear regression equations to obtain N correlation coefficients a1, a2, a3 ... an, and N residual coefficients b1, b2, b3 ... bn, wherein the number of the correlation coefficients a1 is equal to the number K of the item-related data, and the number of the residual coefficients b1 is equal to the number K of the item-related data; Step S403, integrating the calculated correlation coefficients and residual coefficients to obtain an N×N×K dimensional matrix, the first row of the matrix is ​​a1, a2, a3...an, b1, b2, b3...bn, the second row of the matrix is ​​a1, a2, a3...an / 2, 0, 0, ...b1, b2, b3...bn / 2, 0, 0..., the last row of the matrix is ​​a1, 0, 0, ...b1, 0, 0, ...0, the number of the correlation coefficients a1 is equal to the number K of the item-related data, and the number of the residual coefficients b1 is equal to the number K of the item-related data; Step S404: multiplying the preset N-dimensional weight matrix by the N×N×K-dimensional matrix to obtain the second strain regression slope data.

7. The data correlation mining method of a big data analysis platform according to claim 6, characterized in that: Step S500 specifically includes: Step S501, successively merging each row of data after multiplying a preset N-dimensional weight matrix with the N×N-dimensional matrix, wherein a merging rule includes a first merging rule and a second merging rule; Step S502, calculating the correlation coefficient ratio of the correlation coefficient a1 and the correlation coefficient ratio of the residual coefficient b1 of the combined matrix; Step S503: according to different association tasks, different association comparison methods are used to judge the association of data, wherein the association comparison methods include a first association comparison method and a second association comparison method.

8. The data correlation mining method of a big data analysis platform according to claim 7, characterized in that: Compare a1, a2 in the first row matrix a1, a2, a3 ... an, b1, b2, b3 ... bn with a1 in the second row matrix a1, a2, a3 ... an / 2, 0, 0, ... b1, b2, b3 ... bn / 2, 0, 0 ... 0; When a2-a1 is greater than 0 and a1 is greater than 0, merge according to the first merge rule; When a2-a1 is less than 0 and a1 is less than 0, merge according to the first merge rule; When a2-a1 is less than 0 and a1 is greater than 0, merge according to the second merge rule; When a2-a1 is greater than 0 and a1 is less than 0, merge according to the second merge rule.

9. The data correlation mining method of a big data analysis platform according to claim 8, characterized in that: The first merging rule merging specific steps include: adding a1, a2 in the first row matrix a1, a2, a3 ... an, b1, b2, b3 ... bn and a1 in the second row matrix a1, a2, a3 ... an / 2, 0, 0, ... b1, b2, b3 ... bn / 2, 0, 0 ... 0; The specific steps of merging by the second merging rule include: removing a1 and a2 from the first row matrices a1, a2, a3 ... an, b1, b2, b3 ... bn.

10. The data association mining method of a big data analysis platform according to claim 9, characterized in that: When the associated task is the first associated task, the first associated comparison method is selected, and the first associated comparison method specifically includes: calculating the correlation coefficient ratio and the correlation coefficient ratio; adding all the combined correlation coefficients a1 and the residual coefficient b1 to obtain the sum of a1 and the sum of b1, and adding all the correlation coefficients a1 and the residual coefficient b1 before the combination to obtain the sum of a2 and the sum of b2; The correlation coefficient ratio = a1 total / a2 total, the residual coefficient ratio = b1 total / b2 total; the larger the correlation coefficient ratio, the smaller the residual coefficient ratio, indicating that the data correlation is greater; When the associated task is the second associated task, the second associated comparison method is selected, and the first associated comparison method specifically includes: calculating the correlation coefficient ratio and the residual coefficient ratio of the first strain regression slope data to the second strain regression slope data; The smaller the difference between the correlation coefficient ratio and the residual coefficient ratio, the smaller the data correlation is; The larger the difference between the correlation coefficient ratio and the residual coefficient ratio, the greater the data correlation.