E-commerce data monitoring method and platform based on big data
By constructing a comparison data group and explicit and implicit correlation data mapping table, combining gradient values and distribution coefficients, the problem of insufficient implicit correlation identification in e-commerce data monitoring is solved, real-time and accurate monitoring and risk warning of e-commerce data is realized, and the intelligence and security of data management are improved.
Patent Information
- Application Number
- CN202510757878.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-11
AI Technical Summary
Existing e-commerce data monitoring technologies cannot accurately identify the abnormal fluctuations hidden in complex business scenarios and data risks under multi-platform collaboration, and lack monitoring of the implicit correlation and cross-effects between different data dimensions, resulting in sales losses or inventory backlogs.
By constructing a comparison data group of each sales platform, obtaining platform distribution coefficients, labeling explicit and implicit correlation data, using gradient values and distribution coefficients to make real-time abnormal judgments, establishing explicit and implicit correlation data mapping tables, and realizing cross-platform and multi-time data comparison analysis.
Real-time, accurate and in-depth monitoring of e-commerce data, quickly identify abnormal platforms and data types, reduce the risk of sales losses or inventory backlogs, and improve the intelligence and security of data management.
Smart Images

Figure CN120296607A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of e-commerce data monitoring, and specifically relates to a method and platform for e-commerce data monitoring based on big data. Background Art
[0002] With the rapid development of e-commerce, the sales data on various e-commerce platforms are becoming increasingly large and complex. For e-commerce enterprises, timely and effectively monitoring and analyzing these massive sales data is the key to maintaining market competitiveness and optimizing operation strategies.
[0003] However, there are many deficiencies in the existing e-commerce data monitoring technologies. Traditional methods mostly rely on the abnormal threshold triggering of a single data indicator, single-dimensional abnormal data identification, or spatio-temporal network analysis, and cannot accurately identify the abnormal fluctuations hidden in complex business scenarios and the data risks under multi-platform collaboration. They cannot monitor the low-significance abnormal data hidden under the normal business appearance according to the implicit associations and cross-influences between different data dimensions; that is, monitor the abnormal data types that seem scattered and isolated but are actually highly implicitly associated, and it is difficult to reveal the explicit and implicit association relationships between different data types while lacking accurate judgment of abnormal data types; resulting in sales losses or inventory backlogs. Based on this, a method and platform for e-commerce data monitoring based on big data are proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and platform for e-commerce data monitoring based on big data, which solves the technical problem that traditional methods cannot monitor the low-significance abnormal data hidden under the normal business appearance according to the implicit associations and cross-influences between different data dimensions.
[0005] A method for e-commerce data monitoring based on big data includes the following steps: Step 1: According to the historical data of the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products in multiple preset time periods on each sales platform, construct a control data group corresponding to each sales platform respectively; Step 2: Obtain the platform distribution coefficients corresponding to various types of data according to the historical data of various types of data; Step 3: Mark the explicit association data and implicit association data corresponding to different types of data respectively; Step 4: Generate a real-time data group corresponding to each sales platform according to the real-time sales data corresponding to various types of data of the sold products on each sales platform at the same time, and obtain the gradient value of each sales platform, and determine and mark the abnormal platform according to the gradient value; Step 5: When the abnormal platform is unique, based on the real-time sales data of various types of data, obtain the real-time distribution coefficients corresponding to various types of data respectively. According to the comparison results with the corresponding platform distribution coefficients, judge and mark the abnormal data types, obtain the explicit associated data and implicit associated data corresponding to the abnormal data types, and output the abnormal data types corresponding to the abnormal platform and their corresponding explicit associated data and implicit associated data together.
[0006] As a further solution of the present invention: The specific way to construct the control data groups corresponding to each sales platform is as follows: Obtain the mean value of the sum of the maximum and minimum values corresponding to various types of data in the sales order volume, sales traffic, product inventory, and sales unit price respectively within multiple preset time periods of the same sales platform for various types of data of the sold products, and establish a control data group Ba (C1a, C2a, C3a, C4a) of various types of data within the corresponding sales platform, where a represents different sales platforms, a = 1, 2,..., q; q represents the number of sales platforms, and q is a positive integer greater than or equal to 2.
[0007] As a further solution of the present invention: The specific way to obtain the platform distribution coefficients corresponding to various types of data is as follows: Randomly select a type of data from various types of data as the target type of data; mark the historical data Ab corresponding to the target type of data within multiple preset time periods of all sales platforms respectively, and use the discrete values of different historical data Ab corresponding to the target type of data as the platform distribution coefficient F1 corresponding to the target type of data. b represents different historical data corresponding to the target type of data, b = 1, 2,..., e, and e represents the total number of different historical data corresponding to the target type of data, b = q × n; use the same analysis method to analyze the historical data corresponding to various types of data within multiple preset time periods of each sales platform respectively, and then the platform distribution coefficients Fk corresponding to various types of data can be obtained. k represents different types of data, k = 1, 2, 3, 4.
[0008] As a further solution of the present invention: The specific way to respectively mark the explicit associated data and implicit associated data corresponding to different types of data is as follows: S1: Randomly select a type of data from various types of data as the analysis type of data; S2: Randomly select one from multiple preset time periods as the analysis time period; Take the mean value of the historical data E11 corresponding to the analysis type of data within the analysis time period of each sales platform as the standard value D11 corresponding to the analysis type of data at the analysis time period; S3: Repeat step S2, and the standard values D1t corresponding to the analysis data at each preset time period can be obtained; take the time period label corresponding to each data acquisition period as the abscissa t, and take the standard values D1t corresponding to each preset time period t as the ordinate, and obtain the data points D1t(t, D1t) corresponding to each preset time period t. Mark each data point in the two-dimensional coordinate system in sequence and connect them in sequence, and then obtain the change curve Q1 corresponding to the analysis data, where t is different preset time periods, and at the same time t is the time period label corresponding to each preset time period respectively, t = 1, 2,..., n, which is a time period of every ten days starting from the current time of data acquisition, n is the number of preset time periods, the time interval between different preset time periods is 10 days, n is a positive integer and greater than or equal to 3; S4: Repeat steps S1 - S3, and the change curves corresponding to the remaining other types of data can be obtained, and then the change curves Qk corresponding to each type of data can be obtained, where k represents different types of data, k = 1, 2, 3, 4; S41: According to each data point Dkt(t, Dkt) corresponding to the change curve Qk of each type of data, according to the correlation coefficient Xi between the analysis data and the other types of data, i = 1, 2, 3, the specific method is as follows; Randomly select one from the other types of data except the analysis data as the comparison data: According to the standard values D1t and D2t corresponding to the analysis data and the comparison data at each preset time period, calculate the correlation coefficient X1 between the change curves of the analysis data and the comparison data. When the absolute value of the correlation coefficient X1 is greater than the preset value Y1, then mark the comparison data as the dominant associated data of the analysis data. Otherwise, obtain the mutual coefficient M1 between the analysis data and the comparison data. If the mutual coefficient is higher than the threshold Y2, then mark the comparison data as the recessive associated data of the analysis data. Otherwise, do not perform any processing, Y1 = 0.7 and Y2 > 0.6; S42: Repeat step S41, and the dominant associated data and recessive associated data corresponding to different types of data can be obtained.
[0009] As a further solution of the present invention: The specific method for obtaining the mutual coefficient M1 between the analysis data and the comparison data is as follows: Obtain the covariance R(D1, D2) between the standard values D1t and D2t of the analysis data and the comparison data at each preset time period. At the same time, obtain the variance V1 of the standard value D1t of the analysis data at each preset time period and the variance V2 of the standard value D2t of the comparison data. Through the formula: M1 = R(D1, D2) / (V1×V2); calculate the mutual coefficient M1 between the analysis data and the comparison data.
[0010] As a further solution of the present invention, the specific method for obtaining the gradient values corresponding to each sales platform is as follows: According to the real-time sales data corresponding to various types of data of the sold products on each sales platform within the same time, generate real-time data groups Ga (G1a, C2a, C3a, C4a) corresponding to each sales platform respectively. Through Ua = |G1a - C1a| × θ1 + |G2a - C2a| × θ2 + |G3a - C3a| × θ3 + |G4a - C4a| × θ4, obtain the gradient values Ua corresponding to each sales platform respectively, where θ1, θ2, θ3, and θ4 are preset weight coefficients corresponding to various types of data respectively, satisfying 1 = θ1 + θ2 + θ3 + θ4. The specific method for judging and marking abnormal platforms is as follows: Mark the sales platform with the gradient value Ua greater than the preset threshold Y3 as an abnormal platform, and do not perform any processing otherwise.
[0011] As a further solution of the present invention, the specific method for judging and marking abnormal data types is as follows: When the abnormal platform is unique, according to the real-time data G1a, C2a, C3a, and C4a corresponding to various types of data on each sales platform respectively, obtain the real-time distribution coefficients Jk corresponding to various types of data respectively. Compare and analyze the real-time distribution coefficients Jk corresponding to various types of data with the corresponding platform distribution coefficients Fk. Mark the corresponding data types with the absolute value of the difference between the real-time distribution coefficient Jk and the platform distribution coefficient Fk greater than the threshold β1 as abnormal data types, otherwise, do not perform any processing.
[0012] When the abnormal platforms are not unique, directly generate a data anomaly signal, and output the data anomaly signal and the corresponding abnormal platforms simultaneously.
[0013] An e-commerce data monitoring platform based on big data, which is implemented by a method for monitoring e-commerce data based on big data, including: Data acquisition end: Acquire the historical data corresponding to various types of data of the sold products on each sales platform within multiple preset time periods. Control data group construction end: According to the historical data of the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products on each sales platform within multiple preset time periods, construct control data groups corresponding to each sales platform respectively. Platform distribution coefficient acquisition end: Obtain the platform distribution coefficients corresponding to various types of data respectively according to the historical data of various types of data. Explicit and implicit association data marking end: Mark the explicit association data and implicit association data corresponding to different types of data respectively. Abnormal platform marking end: Generate real-time data groups corresponding to each sales platform according to the real-time sales data of various types of products on each sales platform within the same time, obtain the gradient value of each sales platform, and determine and mark the abnormal platform according to the gradient value; Abnormal data type determination end: When the abnormal platform is unique, obtain the real-time distribution coefficient corresponding to each type of data according to the real-time sales data of various types of data, and judge and mark the abnormal data type according to the comparison result with the corresponding platform distribution coefficient; Output end: Obtain the explicit associated data and implicit associated data corresponding to the abnormal data type, and output the abnormal data type corresponding to the abnormal platform and its corresponding explicit associated data and implicit associated data together.
[0014] Compared with the prior art, the beneficial effects of the present invention are: (1) In the present invention, through a dual analysis mechanism, by performing correlation analysis on all data types in pairs, a comprehensive mapping table is finally established, complex explicit and implicit correlation relationships between various types of data are mined, and through the establishment of a correlation relationship mapping table, it is clearly shown which data types are explicitly associated with each data type and which data types are implicitly associated with, solving the pain points of the traditional solution lacking the ability of in-depth correlation analysis and being unable to effectively identify low-significance and implicitly associated abnormal data types; (2) In the present invention, by obtaining data in real time and calculating the gradient value, abnormal platforms are quickly identified. When there are multiple abnormal platforms, the system will directly generate a data anomaly signal, and at the same time output the signal and all corresponding abnormal platforms, and further locate the specific abnormal data type and its associated data when there is a unique anomaly, realizing real-time, accurate and in-depth monitoring of e-commerce data, helping operation personnel quickly lock the root cause of the anomaly, enabling operation personnel to quickly formulate targeted countermeasures, and effectively reducing the potential sales loss or inventory backlog risk; (3) In the present invention, by performing comparative analysis and distribution monitoring on e-commerce data across multiple platforms and multiple time periods, the sensitivity and recognition accuracy of abnormal data types are effectively improved; through the explicit / implicit correlation mapping table, the correlation traceability of data anomalies is realized, and low-significance anomalies hidden under the normal business appearance and seemingly scattered and isolated but actually highly implicitly associated abnormal data types can be discovered. When a data anomaly occurs, it can quickly locate other data that may be associated behind it, providing an important direction for subsequent problem troubleshooting, effectively preventing the spread of anomalies and risk aggregation between multiple platforms, ensuring the operation safety of e-commerce, facilitating enterprises to accurately identify risk points, and effectively improving the intelligence and security of e-commerce data management. Brief Description of the Drawings
[0015] Figure 1Schematic diagram of the method framework structure of the present invention; Figure 2 Schematic diagram of the platform framework structure of the present invention. Detailed implementation manners
[0016] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0017] Embodiment 1: Please refer to Figure 1 , the present application provides an e-commerce data monitoring method based on big data, including the following steps: Step 1: Obtain the historical data corresponding to various types of data of the sold products in multiple preset time periods on each sales platform, and construct a control data group corresponding to the sold products on each sales platform respectively; Various types of data of the sold products include the sales order volume, sales traffic, product inventory, and sales unit price corresponding to each sales platform respectively; The specific method for constructing the control data group corresponding to the sold products on each sales respectively is; Take the mean value of the sum of the maximum and minimum values of each type of data in the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products in multiple preset time periods on the same sales platform as the control data group of each type of data within the corresponding sales platform, and then obtain the control data groups Ba (C1a, C2a, C3a, C4a) of various types of data within different sales platforms, where a represents different sales platforms, a = 1, 2,..., q; q represents the number of sales platforms, q is a positive integer and greater than or equal to 2, where C1a, C2a, C3a, and C4a are respectively the mean values of the sum of the maximum and minimum values of each type of data in the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products in multiple preset time periods on the same sales platform; It should be noted that multiple preset time periods refer to different data acquisition time periods t, which is set here as the time period of every ten days starting from the current time of data acquisition, and n times are obtained in sequence. t is a different data acquisition time period, and at the same time t is the time period label corresponding to each data acquisition time period, that is, t = 1, 2,..., n; n is the number of preset time periods, the time interval between different preset time periods is 10 days, and n is a positive integer and greater than or equal to 3; For example, if today is May 20th and n = 3, then the different data acquisition time periods are respectively: The data acquisition period 1 is from May 10th to May 20th; the data acquisition period 2 is from April 30th to May 10th; the data acquisition period 3 is from April 20th to April 30th; Traditional methods are difficult to integrate data from different sales platforms, resulting in one-sided analysis. By collecting historical sales data on different sales platforms and within different time periods, integrating multi-platform historical data through unified calculation rules, establishing a cross-platform data benchmark, eliminating data fragmentation problems, providing a historical reference standard for subsequent real-time data comparison, enabling enterprises to comprehensively understand the performance of products on each platform, and avoiding misjudging the market situation due to data deviation on a single platform.
[0018] Step 2: Analyze the historical data corresponding to various types of sales product data on each sales platform within multiple preset time periods, and obtain the platform distribution coefficients corresponding to each type of data according to the analysis results. The specific method is as follows: Randomly select one type of data from various types of data as the target type of data; mark the historical data corresponding to the target type of data on all sales platforms within multiple preset time periods as Ab, where b represents different historical data corresponding to the target type of data, b = 1, 2,..., e, and e represents the total number of different historical data corresponding to the target type of data, b = q × n; Take the discrete value of the different historical data Ab corresponding to the target type of data as the platform distribution coefficient F1 corresponding to the target type of data; That is, through the formula, , calculate the platform distribution coefficient F1 corresponding to the target type of data, where Ap is the mean value of Ab; Use the same analysis method to analyze the historical data corresponding to each type of data on each sales platform within multiple preset time periods, and the platform distribution coefficients Fk corresponding to each type of data can be obtained, where k represents different types of data, k = 1, 2, 3, 4; Traditional methods lack quantitative analysis of data discrete characteristics and are difficult to judge the distribution law of data on each platform. By quantifying the discrete value, revealing the distribution differences of data on different platforms, providing a basis for data distribution characteristics for anomaly detection, helping enterprises identify whether data fluctuations exceed the normal distribution range, and timely discovering potential anomalies. By judging real-time data anomalies through the platform distribution coefficient, it can reflect whether the cross-platform distribution of a certain type of data at the current moment deviates from its historical normal distribution.
[0019] Step 3: According to the historical data corresponding to various types of sales product data on each sales platform within multiple preset time periods, respectively mark the explicit associated data and implicit associated data corresponding to different types of data, and establish an explicit and implicit data association relationship mapping table. The specific method is as follows: S1: Randomly select one type of data from various types of data as the analysis type of data; S2: Randomly select one from multiple preset time periods as the analysis time period; Obtain the historical data E11 corresponding to the analysis type data in each sales platform during the analysis time period; take the mean value of E11 as the standard value D11 corresponding to the analysis type data at the analysis time period; S3: Repeat step S2 to obtain the standard values D1t corresponding to the analysis type data at each preset time period; use the time period label corresponding to each data acquisition time period as the abscissa t, and use the standard values D1t corresponding to each preset time period t as the ordinate to obtain the data points D1t(t, D1t) corresponding to each preset time period t, mark each data point D1t(t, D1t) in the two-dimensional coordinate system in sequence, and connect them in sequence to obtain the change curve Q1 corresponding to the analysis type data; S4: Repeat steps S1 - S3 to obtain the change curves corresponding to the remaining other types of data respectively, and then obtain the change curves Qk corresponding to each type of data, where k represents different types of data, k = 1, 2, 3, 4; S41: According to each data point Dkt(t, Dkt) corresponding to the change curve Qk of each type of data, obtain the correlation coefficient Xi between the analysis type data and the other types of data, i = 1, 2, 3. The specific method is as follows: Randomly select one from the other types of data excluding the analysis type data as the comparison data; Mark each data point on the change curve of the comparison data as D2t(t, D2t); where D2t is the standard value corresponding to the comparison data at each preset time period; Calculate the correlation coefficient X1 between the change curves of the analysis type data and the comparison data through the Pearson similarity calculation formula, that is: ; where D1p and D2p are the mean values of D1t and D2t respectively; When the absolute value of the correlation coefficient X1 is greater than the preset value Y1, mark the comparison data as the dominant associated data of the analysis type data. The specific value of the preset value Y1 is determined by relevant personnel according to actual needs, Y1 = 0.7; Otherwise, the covariance R(D1, D2) between the standard values D1t and D2t of the analysis data and the comparison data at each preset time period is obtained, and the variance V1 of the standard value D1t of the analysis data at each preset time period and the variance V2 of the standard value D2t of the comparison data are obtained at the same time, through the formula: M1=R(D1, D2) / (V1×V2); the mutual coefficient M1 between the analysis data and the comparison data is calculated, if the mutual coefficient M1 is higher than the threshold value Y2, the comparison data is marked as implicitly associated data of the analysis data, otherwise, no processing is performed, the specific value of the threshold value Y2 is formulated by relevant personnel according to actual needs, and meets, Y2>0.6; pass ; Calculate the covariance between D1t and D2t by and , calculate the variance V1 of the standard value D1t and the variance V2 of the comparison data standard value D2t; Among them, the role of covariance R (D1, D2) is to reflect the synergy of the deviation of the two variables of the analysis data and the comparison data from the mean. If the covariance is positive, it means that the two variables fluctuate in the same direction. If the covariance is negative, it means that the two variables fluctuate in opposite directions. Then, by dividing by V1×V2 to eliminate the dimension effect, the value range of the mutual coefficient M1 is [−1, 1], which is convenient for horizontal comparison of the correlation strength between the analysis data and the comparison data. The same analysis method is used to analyze the remaining data of other categories, thereby obtaining explicit related data and implicit related data corresponding to the analyzed data; First, the Pearson similarity calculation formula is used to find those explicit correlation data with very strong linear relationships regardless of volatility. Then, for those linear relationships that are not so strong, a secondary screening is performed using the mutual coefficient M1. The mutual coefficient will screen out those data whose linear relationships are not top-notch but whose relationships are very stable and have low volatility, and define them as implicit correlation data.
[0020] S42: Repeat step S41 to obtain explicit associated data and implicit associated data corresponding to different types of data, and establish an explicit and implicit data association relationship mapping table; By analyzing the associations between all data types, a comprehensive mapping table is finally established, which clearly shows which data types are explicitly associated with each data type and which data types are implicitly associated with each data type. Traditional methods cannot deeply analyze the explicit and implicit correlations between data, and it is difficult to locate the root cause of the problem. Through the dual analysis mechanism, the correlations of different intensities between data can be excavated to help enterprises understand the intrinsic connections between data. When anomalies occur, the relevant influencing factors can be quickly traced, the efficiency of problem diagnosis can be improved, normal fluctuations can be distinguished from abnormal behaviors, and the comprehensiveness of detection logic can be improved.
[0021] Step 4: Obtain the real-time sales data corresponding to various types of data of the products sold within the same time on each sales platform, thereby obtaining the real-time data groups corresponding to each sales platform respectively, and analyze and compare them with the control data groups corresponding to each sales platform respectively to obtain the gradient values corresponding to each sales platform respectively, and determine and mark the abnormal platforms according to the gradient values. The specific method is as follows: Generate the real-time data groups Ga (G1a, C2a, C3a, C4a) corresponding to each sales platform respectively according to the real-time sales data corresponding to various types of data of the products sold within the same time on each sales platform. Obtain the gradient value Ua corresponding to each sales platform respectively through Ua = |G1a - C1a|×θ1 + |G2a - C2a|×θ2 + |G3a - C3a|×θ3 + |G4a - C4a|×θ4, where θ1, θ2, θ3, and θ4 are the preset weight coefficients corresponding to various types of data respectively, and satisfy 1 = θ1 + θ2 + θ3 + θ4; Mark the sales platforms with the gradient value Ua greater than the preset threshold Y3 as abnormal platforms, and do not perform any processing otherwise: The specific value of the preset threshold Y3 is determined by relevant personnel according to actual needs; When there are multiple abnormal platforms, directly generate a data anomaly signal, and output both the data anomaly signal and the corresponding abnormal platforms; When there are multiple abnormal platforms, the system will directly generate a data anomaly signal, and output this signal and all the corresponding abnormal platforms at the same time, which means that the entire e-commerce system may face large-scale problems and requires a global investigation.
[0022] Traditional methods lack effective detection of abnormal platforms and cannot detect abnormal data fluctuations in a timely manner. By comprehensively calculating multi-dimensional data and weights and accurately locating abnormal platforms, enterprises can quickly lock in abnormal sales platforms, avoid sales losses or inventory backlogs caused by abnormal platform data, and achieve the timeliness and accuracy of anomaly detection. Compared with the traditional method that relies on periodic reports or manual analysis, it can monitor in real time and quickly locate the sales platforms where anomalies occur, thus avoiding the problem of lagging anomaly detection and winning precious response time for enterprises.
[0023] Step 5: Obtain the real-time distribution coefficients corresponding to various types of data respectively according to the real-time sales data corresponding to various types of data of the products sold within the same time on each sales platform. Compare and analyze the real-time distribution coefficients corresponding to various types of data with the corresponding platform distribution coefficients, judge and mark abnormal data types, and output the abnormal data types corresponding to the abnormal platforms and their corresponding explicit associated data and implicit associated data. The specific method is as follows; When the abnormal platform is unique, the real-time distribution coefficient Jk corresponding to each type of data is obtained according to the real-time data G1a, C2a, C3a and C4a corresponding to each type of data on each sales platform. ; Calculate and obtain the real-time distribution coefficient Jk corresponding to each type of data; Compare and analyze the real-time distribution coefficient Jk and the corresponding platform distribution coefficient Fk corresponding to each type of data, and mark the corresponding data type whose absolute value of the difference between the real-time distribution coefficient Jk and the platform distribution coefficient Fk is greater than the threshold β1 as an abnormal data type. Otherwise, no processing is performed, that is, the corresponding data type whose Hk is greater than β1 in Hk=|Jk-Fk| is marked as an abnormal data type; The specific value of the threshold β1 is determined by relevant personnel based on actual needs. By searching the explicit and implicit data association relationship mapping table, the explicit associated data and implicit associated data corresponding to the abnormal data type are obtained, and the abnormal data type corresponding to the abnormal platform and its corresponding explicit associated data and implicit associated data are output together; Traditional methods make it difficult to determine the type of abnormal data and are unable to solve the problem in a targeted manner. By comparing distribution coefficients and clarifying the specific dimensions of abnormal data, enterprises can accurately identify abnormal data types, formulate targeted strategies, and improve the accuracy of operational decisions.
[0024] By being able to conduct comparative analysis and distribution monitoring of e-commerce data across multiple platforms and multiple time periods, the sensitivity and recognition accuracy of data anomalies can be effectively improved; through explicit / implicit association mapping tables, the association tracing of data anomalies can be achieved to improve the efficiency of anomaly root cause analysis; real-time gradient and distribution coefficient determination can achieve more intelligent anomaly risk warning and data intervention decisions; effectively prevent anomaly propagation and risk aggregation between multiple platforms to ensure the operational security and data health of e-commerce; it can realize dynamic monitoring of e-commerce data across platforms, multiple time periods, and multiple data types, and flexibly identify multi-source data anomalies and complex associations, so that enterprises can accurately identify risk points, trace the root causes of anomalies, and intervene and adjust, effectively improving the intelligence and security of e-commerce data management, and is suitable for complex and changeable online transaction environments; It solves the problem of the lack of in-depth analysis of explicit and implicit correlations between various types of data in traditional methods, as well as the pain point that it is difficult to reveal deep-seated problems based on changes in a single data indicator. It can discover low-significance anomalies hidden under the appearance of normal business, and abnormal data types that seem scattered and isolated but are actually highly implicitly correlated, and identify implicit correlations and cross-influences between different data dimensions. In this way, when a piece of data is abnormal, other data that may be associated with it can be quickly located, providing an important direction for subsequent troubleshooting.
[0025] Example 2: As the second example of the present invention, please refer to Figure 2 , and a big data-based e-commerce data monitoring platform is provided. This platform is used to implement the previously disclosed big data-based e-commerce data monitoring method, specifically including: Data acquisition end: Acquire the historical data corresponding to various types of data of the sold products in multiple preset time periods on each sales platform respectively. After obtaining the corresponding permissions and licenses, acquiring the historical data corresponding to various types of data of the sold products in multiple preset time periods on each sales platform from the e-commerce platform or the e-commerce data center is prior art, so no further elaboration will be made here. Control data group construction end: Construct the control data groups corresponding to each sales platform respectively according to the historical data of sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products in multiple preset time periods on each sales platform. Platform distribution coefficient acquisition end: Obtain the platform distribution coefficients corresponding to various types of data respectively according to the historical data of various types of data. Explicit and implicit association data marking end: Mark the explicit association data and implicit association data corresponding to different types of data respectively. An innovative method combining Pearson similarity and mutual coefficient is introduced to deeply explore the complex explicit and implicit association relationships between various types of data, and an association relationship mapping table is established. It solves the pain points of the traditional solution lacking the ability to analyze deep associations, and being unable to effectively identify low-significance, implicit association abnormal data types and predict the risk of scattered abnormal aggregation. It can reveal the more complex causal logic behind the data, rather than just staying at the fluctuations of surface indicators.
[0026] Abnormal platform marking end: Generate the real-time data groups corresponding to each sales platform respectively according to the real-time sales data of various types of data of the sold products on each sales platform at the same time, and obtain the gradient values of each sales platform. Determine and mark the abnormal platforms according to the gradient values. Abnormal data type determination end: When the abnormal platform is unique, obtain the real-time distribution coefficients corresponding to various types of data according to the real-time sales data of various types of data, and judge and mark the abnormal data types according to the comparison results with the corresponding platform distribution coefficients. Output end: Acquire the explicit association data and implicit association data corresponding to the abnormal data types, and output the abnormal data types corresponding to the abnormal platforms and their corresponding explicit association data and implicit association data together. By querying the association relationship mapping table according to the abnormal data type, outputting its explicit and implicit association data, and providing the impact chain of the abnormal data through the output of the association data. By obtaining data in real time and calculating the gradient value to quickly identify the abnormal platform, and further locating the specific abnormal data type and its associated data when there is a unique abnormality, this solution realizes the real-time, accurate, and in-depth monitoring of e-commerce data. This enables enterprises to quickly respond to market changes, transforming the dilemma of lagging abnormal discovery and difficult accurate judgment of the root cause of problems into the immediate discovery of abnormalities and efficient root cause location.
[0027] Greatly improves the problem-solving efficiency and helps operators quickly lock in the root cause of the abnormality. For example, if the order volume decreases while the traffic is an explicit association, then prioritize checking the traffic problem; if the inventory is an implicit association, then it is necessary to combine other information to deeply analyze the inventory risk. This deep auxiliary analysis ability enables operators to quickly formulate targeted countermeasures and effectively reduce the potential risks of sales losses or inventory backlogs.
[0028] Embodiment 3: As Embodiment 3 of the present invention, when the present application is specifically implemented, compared with Embodiment 1 and Embodiment 2, the technical solution of this embodiment lies in combining the solutions of the above-mentioned Embodiment 1 and Embodiment 2 for implementation.
[0029] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.
[0030] As described above, only the specific implementation manners of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An e-commerce data monitoring method based on big data, characterized in that, It includes the following steps: Step 1: According to the historical data of sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of sold products in multiple preset time periods on each sales platform, construct a control data group corresponding to each sales platform respectively; Step 2: Obtain the platform distribution coefficients corresponding to various types of data according to the historical data of various types of data; Step 3: Mark the explicit association data and implicit association data corresponding to different types of data respectively; Step 4: Generate a real-time data group corresponding to each sales platform according to the real-time sales data of various types of data of sold products on each sales platform at the same time, and obtain the gradient value of each sales platform. Determine and mark the abnormal platform according to the gradient value; Step 5: When the abnormal platform is unique, obtain the real-time distribution coefficients corresponding to various types of data according to the real-time sales data of various types of data, judge and mark the abnormal data type according to the comparison result with the corresponding platform distribution coefficient, obtain the explicit association data and implicit association data corresponding to the abnormal data type, and output the abnormal data type corresponding to the abnormal platform and its corresponding explicit association data and implicit association data together.
2. The e-commerce data monitoring method based on big data according to claim 1, wherein, The specific method for constructing a control data group corresponding to each sales platform is: Obtain the mean value of the sum of the maximum and minimum values of various types of data among the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of sold products on the same sales platform in multiple preset time periods, and establish a control data group Ba (C1a, C2a, C3a, C4a) of various types of data within the corresponding sales platform, where a represents different sales platforms, a = 1, 2,..., q; q represents the number of sales platforms, q is a positive integer and greater than or equal to 2.
3. The e-commerce data monitoring method based on big data according to claim 2, characterized in that, The specific method for obtaining the platform distribution coefficients corresponding to various types of data is: Randomly select one type of data from various types of data as the target type of data; mark the historical data Ab corresponding to the target type of data in multiple preset time periods on all sales platforms respectively, and use the discrete value of different historical data Ab corresponding to the target type of data as the platform distribution coefficient F1 corresponding to the target type of data. b represents different historical data corresponding to the target type of data, b = 1, 2,..., e, e represents the total number of different historical data corresponding to the target type of data, b = q×n; use the same analysis method to analyze the historical data corresponding to various types of data in multiple preset time periods on each sales platform respectively, and then the platform distribution coefficients Fk corresponding to various types of data can be obtained, where k represents different types of data, k = 1, 2, 3, 4.
4. The e-commerce data monitoring method based on big data according to claim 3, wherein, The specific method for marking the explicit association data and implicit association data corresponding to different types of data respectively is: S1: Randomly select one type of data from various types of data as the analysis type of data; S2: Randomly select one from multiple preset time periods as the analysis time period; Take the mean value of the historical data E11 corresponding to the analysis type of data on each sales platform within the analysis time period as the standard value D11 of the analysis type of data at the analysis time period; S3: Repeat step S2 to obtain the standard values D1t corresponding to the analysis data at each preset time period. Use the time period labels corresponding to each data acquisition period as the abscissa t, and the standard values D1t corresponding to each preset time period t as the ordinate to obtain the data points D1t(t, D1t) corresponding to each preset time period t. Mark each data point in the two-dimensional coordinate system in sequence and connect them in sequence to obtain the change curve Q1 corresponding to the analysis data, where t is different preset time periods, and at the same time t is the time period label corresponding to each preset time period, t = 1, 2,..., n, which is a time period of every ten days starting from the current time of data acquisition, n is the number of preset time periods, the time interval between different preset time periods is 10 days, n is a positive integer and greater than or equal to 3; S4: Repeat steps S1 - S3 to obtain the change curves corresponding to the remaining other types of data, and then obtain the change curves Qk corresponding to each type of data, where k represents different types of data, k = 1, 2, 3, 4; S41: According to each data point Dkt(t, Dkt) corresponding to the change curve Qk of each type of data, according to the correlation coefficient Xi between the analysis data and the rest of the types of data, i = 1, 2, 3, the specific method is as follows; Randomly select one from the other types of data except the analysis data as the comparison data: According to the standard values D1t and D2t corresponding to the analysis data and the comparison data at each preset time period, calculate the correlation coefficient X1 between the change curves of the analysis data and the comparison data. When the absolute value of the correlation coefficient X1 is greater than the preset value Y1, mark the comparison data as the explicit associated data of the analysis data. Otherwise, obtain the mutual coefficient M1 between the analysis data and the comparison data. If the mutual coefficient is higher than the threshold Y2, mark the comparison data as the implicit associated data of the analysis data. Otherwise, do not perform any processing, Y1 = 0.7 and Y2 > 0.6; S42: Repeat step S41 to obtain the explicit associated data and implicit associated data corresponding to different types of data.
5. The method for monitoring e-commerce data based on big data according to claim 4, characterized in that, The specific method for obtaining the mutual coefficient M1 between the analysis data and the comparison data is as follows: Obtain the covariance R(D1, D2) between the standard values D1t and D2t of the analysis data and the comparison data at each preset time period. At the same time, obtain the variance V1 of the standard value D1t of the analysis data at each preset time period and the variance V2 of the standard value D2t of the comparison data. Through the formula: M1 = R(D1, D2) / (V1×V2); calculate the mutual coefficient M1 between the analysis data and the comparison data.
6. The method for monitoring e-commerce data based on big data according to claim 4, wherein The specific method for obtaining the gradient value corresponding to each sales platform is as follows: Based on the real-time sales data corresponding to various types of data of the sold products on each sales platform within the same time, real-time data groups Ga (G1a, C2a, C3a, C4a) corresponding to each sales platform are generated. Through Ua = |G1a - C1a|×θ1 + |G2a - C2a|×θ2 + |G3a - C3a|×θ3 + |G4a - C4a|×θ4, the gradient values Ua corresponding to each sales platform are obtained, where θ1, θ2, θ3, and θ4 are preset weight coefficients corresponding to various types of data, and 1 = θ1 + θ2 + θ3 + θ4.
7. The e-commerce data monitoring method based on big data according to claim 6, characterized in that, The specific method for determining and marking abnormal platforms is as follows: Mark the sales platforms with gradient value Ua greater than the preset threshold Y3 as abnormal platforms, and do not perform any processing otherwise.
8. A method for monitoring e-commerce data based on big data according to claim 7, characterized in that, The specific method for judging and marking abnormal data types is as follows: When the abnormal platform is unique, based on the real-time data G1a, C2a, C3a, and C4a corresponding to various types of data on each sales platform, the real-time distribution coefficients Jk corresponding to various types of data are obtained. The real-time distribution coefficients Jk corresponding to various types of data are compared and analyzed with the corresponding platform distribution coefficients Fk. The data types corresponding to the absolute value of the difference between the real-time distribution coefficient Jk and the platform distribution coefficient Fk greater than the threshold β1 are marked as abnormal data types, otherwise, no processing is performed.
9. The method for monitoring e-commerce data based on big data according to claim 7, wherein, When the abnormal platforms are not unique, directly generate a data anomaly signal and output the data anomaly signal and the corresponding abnormal platforms simultaneously.
10. A method for monitoring e-commerce data based on big data according to claim 4, characterized in that The specific method for obtaining the correlation coefficient X1 between the change curves of the analysis data and the comparison data is as follows: By the formula: ; Calculate to obtain the correlation coefficient X1 between the change curves of the analysis data and the comparison data, where D2t is the standard value corresponding to the comparison data at each preset time period, and D1p and D2p are the means of D1t and D2t respectively.
11. An e-commerce data monitoring platform based on big data, characterized in that, This platform implements the method for monitoring e-commerce data based on big data described in any one of claims 1 - 10, including: A control data group construction end, which constructs control data groups corresponding to each sales platform according to the historical data of the sales order volume, sales traffic, product inventory, and sales unit price corresponding to various types of data of the sold products on each sales platform within multiple preset time periods; A platform distribution coefficient acquisition end, which obtains the platform distribution coefficients corresponding to various types of data according to the historical data of various types of data; An explicit and implicit association data marking end, which respectively marks the explicit association data and implicit association data corresponding to different types of data; An abnormal platform marking end, which generates real-time data groups corresponding to each sales platform according to the real-time sales data of various types of data of the sold products on each sales platform within the same time, and obtains the gradient values of each sales platform, and determines and marks abnormal platforms according to the gradient values; An abnormal data type determination end, when the abnormal platform is unique, obtains the real-time distribution coefficients corresponding to various types of data according to the real-time sales data of various types of data, and judges and marks abnormal data types according to the comparison results with the corresponding platform distribution coefficients; The output end obtains the explicit associated data and implicit associated data corresponding to the abnormal data type, and outputs the abnormal data type corresponding to the abnormal platform and its corresponding explicit associated data and implicit associated data together.