Electronic commerce management system and method based on big data
By collecting and analyzing data from each operating node on the e-commerce platform, identifying abnormal interactive nodes and calculating the coefficients and risk coefficients of the failed nodes, the problem of inaccurate risk assessment in the existing technology is solved, real-time monitoring and risk prediction of the operation status of the e-commerce platform is realized, and system stability and user experience are improved.
Patent Information
- Application Number
- CN202510346679.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The risk assessment of existing e-commerce platforms relies on traditional monitoring and analysis methods, and is difficult to meet the requirements of real-time and accuracy, and cannot comprehensively analyze the various factors of nodes. It lacks the intelligent risk prevention and control capabilities of big data and advanced algorithms.
By collecting data from each operating node of the e-commerce platform, analyzing the data transmission volume, flow direction and interaction frequency, identifying abnormal interactive nodes, and calculating the coefficients and risk coefficients of the failed nodes, monitoring and evaluating the platform's operating status in real time, and early warning and emergency treatment are carried out.
Real-time monitoring and risk prediction of the operation status of e-commerce platforms is realized, system stability is improved, the risk of operational interruption is reduced, and user experience and platform competitiveness is improved.
Smart Images

Figure CN119990778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of e-commerce dynamic management, and in particular to an e-commerce management system and method based on big data. Background Art
[0002] With the rapid development of e-commerce, platform operations have become increasingly complex, involving multiple interactive nodes, such as user behavior, product display, payment, order management, logistics, and customer feedback. These nodes together constitute the core operating framework of the platform, and the abnormality or failure of each node may lead to the degradation of platform performance, damage to user experience, and even affect the continued operation of the overall business. Therefore, real-time monitoring and timely detection of potential abnormal problems, especially the prediction of node failure and operation risks, have become the key to maintaining efficient operation of e-commerce platforms. In order to improve the stability and responsiveness of the platform, a technical solution is needed that can evaluate the status of each operating node of the platform in real time, effectively predict potential risks, and handle emergencies.
[0003] The prior art has the following deficiencies:
[0004] At present, the risk assessment of e-commerce platforms mostly relies on traditional monitoring and analysis methods, such as rule-based detection, static performance evaluation models, etc. These methods often rely on manual intervention and outdated historical data, and it is difficult to meet the requirements of real-time and accuracy. Many existing systems have problems with low processing efficiency and poor prediction accuracy when processing large-scale, multi-dimensional data, and it is difficult to provide effective early warning before problems occur at the nodes. In addition, existing technologies are usually unable to comprehensively analyze the impact of various factors of nodes (such as transaction volume, interaction frequency, inventory fluctuations) on node operation, and lack intelligent risk prevention and control capabilities based on big data and advanced algorithms. Therefore, existing technologies have failed to effectively solve the technical difficulties of how to accurately monitor node status, predict potential risks and intervene in a timely manner through big data analysis, and there is room for improvement. Summary of the invention
[0005] The purpose of the present invention is to provide an e-commerce management system and method based on big data to solve the above-mentioned problems.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] An e-commerce management method based on big data comprises the following steps:
[0008] S1: Collecting data from each business operation node in the e-commerce platform, including: user behavior data, transaction data, product information, logistics information and customer feedback data;
[0009] S2: Process the data collected from each business operation node, analyze the data transmission volume, flow direction and interaction frequency, monitor the interaction mode between different operation nodes, detect the fluctuation pattern and abnormal behavior of data traffic, and identify potential abnormal interaction nodes;
[0010] The nodes include: user behavior nodes, product display nodes, payment nodes, order management nodes, logistics nodes and customer feedback nodes;
[0011] S3: Based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and operation success rate are integrated and calculated to obtain the failure node coefficient, which is used to determine whether the abnormal interaction node is a failure node;
[0012] The failure nodes include delayed payment response, delayed order processing, untimely inventory updates, and inaccurate logistics distribution;
[0013] S4: Based on the judgment results, each interactive node is divided into failed nodes and non-failed nodes, and early warning processing is performed on failed nodes. For non-failed nodes, a multi-dimensional risk assessment model is constructed by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and the risk coefficient is calculated to predict the operation risk of the node in the future and perform emergency processing based on the prediction results.
[0014] As a further solution of the present invention: the identifying of potential abnormal interaction nodes specifically includes:
[0015] Obtain the transmission volume and flow direction of real-time data, and compare and analyze them with the transmission volume and flow direction of historical data. According to the degree of deviation between historical data and real-time data, calculate the consistency deviation coefficient to evaluate the degree of deviation of the real-time data of the current node;
[0016] Obtain the interaction frequency of real-time data, and calculate the abnormal fluctuation index of the interaction frequency of real-time data according to the fluctuation of the interaction frequency of real-time data, so as to evaluate the abnormal fluctuation degree of the interaction frequency of real-time data;
[0017] The consistency deviation coefficient of real-time data and the abnormal fluctuation index of interaction frequency are fused and normalized to obtain the node abnormality coefficient.
[0018] As a further solution of the present invention: the acquisition logic of the consistency deviation coefficient is:
[0019] Collect real-time and historical data of each business operation node in the e-commerce platform;
[0020] Pre-process real-time data and historical data to eliminate the impact of different dimensions or data differences;
[0021] The dynamic time warping algorithm is used to calculate the alignment cost between real-time data and historical data, which reflects the similarity of the two time series, where:
[0022] The real-time data series is , represents the real-time data at each time point;
[0023] The historical data series is , represents the historical data in the corresponding time period;
[0024] By calculating the local distance and the minimum alignment cost , the final dynamic time warping distance is the minimum alignment path cost;
[0025] According to the final dynamic time warping distance, the consistency deviation coefficient is calculated;
[0026] The calculation expression of the consistency deviation coefficient is:
[0027] ;
[0028] in, represents the final dynamic time warping distance, Represents the time series corresponding to real-time data, Represents the time series corresponding to historical data, Indicates the total amount of real-time data. Indicates the total amount of historical data, Represents real-time data, Represents historical data, represents the consistency deviation coefficient, represents the local distance, represents the minimum cost, Indicates The standardized value of real-time data, Indicates The standardized value of historical data.
[0029] As a further solution of the present invention: the acquisition logic of the abnormal fluctuation index of interaction frequency is:
[0030] Collect interaction frequency data of each operation node in the e-commerce platform;
[0031] The interaction frequency data is preprocessed, and the real-time interaction frequency data is normalized by using a standardized method for subsequent analysis;
[0032] The standardized interaction frequency data is modeled by a variational autoencoder, where the encoder maps the input data to the latent space, the decoder converts the variables in the latent space into reconstructed data, and calculates the mean square error between the interaction frequency of the actual data and the reconstruction frequency of the reconstructed data to obtain the reconstruction error at each time point;
[0033] The calculation expression of the reconstruction error is:
[0034] ;
[0035] In the formula, Represents each time point of the collection, Indicates The standardized interaction frequency at each time point is Indicates The normalized reconstruction frequency at each time point is Indicates The reconstruction error at each time point;
[0036] Calculate the standard deviation of the reconstruction error over a period of time and mean ;
[0037] Calculate the degree of deviation of the reconstruction error at each time point from the mean and standard deviation to obtain the abnormal fluctuation index of the interaction frequency at each time point. Sum the abnormal fluctuation index of the interaction frequency at each time point to obtain the abnormal fluctuation index of the interaction frequency in the corresponding time period. The calculation expression is:
[0038] ;
[0039] In the formula, represents the abnormal fluctuation index of interaction frequency, Indicates the total duration, represents the mean of the reconstruction error, Represents the standard deviation of the reconstruction error.
[0040] As a further solution of the present invention: based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and the operation success rate are integrated and calculated to obtain the failure node coefficient, which specifically includes:
[0041] Collect data from each business operation node in the e-commerce platform, including: operation frequency data and operation success rate data;
[0042] A Bayesian network model is constructed based on the operation frequency data and the operation success rate data, and a conditional dependency relationship between the operation frequency and the operation success rate and the failure node probability is defined, where the failure node probability indicates the probability of whether a node fails;
[0043] Set a prior distribution for each variable, where the operation frequency Normal distribution , operation success rate Normal distribution , the probability of failed nodes The conditional probabilities between the operation frequency and the operation success rate are expressed as and ;
[0044] Among them, the prior distribution of the operation frequency is:
[0045] ;
[0046] Among them, the prior distribution of the operation success rate is:
[0047] ;
[0048] Based on the prior distribution of operation frequency, the prior distribution of operation success rate and the conditional probability table, the joint probability of failed nodes is calculated using Bayesian network reasoning , and according to the values of operation frequency and operation success rate, the probability of node failure is obtained;
[0049] The calculation expression of the conditional probability of the failed node is:
[0050] ;
[0051] ;
[0052] Based on the conditional probability of failed nodes, the joint probability of failed nodes is obtained through Bayesian network reasoning , the calculation expression is:
[0053] ;
[0054] By calculating the failure probability of the node, the failure node coefficient is obtained. The calculation expression is:
[0055] ;
[0056] in, Represents each time point of the collection, Indicates The operation frequency data at each collection time point, Indicates The operation success rate data at each collection time point, represents the mean value of operation frequency, represents the standard deviation of the operation frequency, represents the mean of the operation success rate, represents the standard deviation of the operation success rate, Indicates the operation frequency The average impact on failed nodes, Indicates the operation frequency The standard deviation of the impact on failed nodes, represents the failure coefficient of the node, The prior distribution of operation frequency, represents the prior distribution of the operation success rate, represents the conditional probability of the operation frequency of the failed node, represents the conditional probability of the failed node operation qualification rate, represents the joint probability.
[0057] As a further solution of the present invention: the method of dividing each interactive node into a failed node and a non-failed node according to the judgment result specifically includes:
[0058] It is determined whether the failure coefficient of each interactive node is greater than or equal to a preset threshold. If so, it is recorded as a failed node; if not, it is recorded as a non-failed node.
[0059] As a further solution of the present invention: the construction of a multi-dimensional risk assessment model and calculation of the risk coefficient specifically include:
[0060] Collect relevant data from non-failed nodes, including product sales data, inventory fluctuations, payment channel usage frequency, and logistics delay history;
[0061] Preprocess the collected relevant data of each non-failed node;
[0062] The processed data is converted into a feature matrix, where each row represents the feature of an interaction node and each column represents a feature;
[0063] The random forest algorithm is used to model the feature matrix and train multiple decision trees. Each tree is trained according to the node characteristics, and the prediction probability of whether the node is invalid is obtained through ensemble learning.
[0064] The training data of each decision tree is randomly extracted from the historical data of the interactive node, and each tree is trained using a different feature subset. The result of the ensemble learning is calculated by weighted average to obtain the predicted probability of node failure.
[0065] Based on the trained random forest model, each interactive node is predicted to obtain the predicted probability of node failure, and the failure risk coefficient is calculated based on the predicted probability. The calculation expression is:
[0066] ;
[0067] In the formula, represents the predicted probability of node failure, is the weight of failure risk, Represents the failure risk factor.
[0068] An e-commerce management system based on big data, comprising:
[0069] A data collection module, which is used to collect data from various business operation nodes in the e-commerce platform, including: user behavior data, transaction data, product information, logistics information and customer feedback data;
[0070] A data flow monitoring and anomaly identification module, which processes the data collected from each commercial operation node, analyzes the data transmission volume, flow direction and interaction frequency, monitors the interaction mode between different operation nodes, detects the fluctuation mode and abnormal behavior of data flow, and identifies potential abnormal interaction nodes;
[0071] A failure node evaluation module, which analyzes the operation frequency and operation success rate of each abnormal interaction node based on the abnormal interaction node, performs fusion processing and calculation on the operation frequency and the operation success rate, and obtains a failure node coefficient for determining whether the abnormal interaction node is a failure node;
[0072] A multi-dimensional risk assessment and emergency response module, which divides each interactive node into failed nodes and non-failed nodes according to the judgment results, performs early warning processing on failed nodes, and for non-failed nodes, builds a multi-dimensional risk assessment model by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and calculates the risk coefficient to predict the operation risk of the node in the future period of time, and performs emergency processing based on the prediction results.
[0073] Beneficial effects of the present invention:
[0074] (1) The present invention identifies potential abnormal interaction nodes in the e-commerce platform by combining real-time data collection with advanced data analysis technology, and monitors and evaluates the operation status of the platform in real time by calculating the node abnormality coefficient and the failure node coefficient. Specifically, based on the advanced technologies of time series analysis, dynamic time warping algorithm and variational autoencoder, the system can continuously track the data flow, interaction frequency and operation success rate of each operation node of the platform, and automatically identify data fluctuations, operation anomalies and their possible sources of failure. This anomaly detection mechanism based on data consistency deviation and interaction frequency fluctuation enables the platform to quickly identify potential faulty nodes and issue early warnings in time before abnormal behavior occurs. Through accurate identification and early warning processing of failed nodes, the platform can effectively avoid operational interruptions caused by node performance degradation or operational failures, thereby significantly improving the stability of the system, reducing losses and enhancing user experience. At the same time, the platform can make timely optimization and adjustments based on the failure risk and performance changes of the nodes to ensure continuous improvement of service quality, thereby enhancing the competitiveness and market adaptability of the platform.
[0075] (2) The present invention introduces a multi-dimensional risk assessment model and combines it with an advanced random forest algorithm to accurately predict the risk of non-failed nodes of the e-commerce platform. By comprehensively considering multiple dimensions of data such as commodity sales data, inventory fluctuations, payment channel usage frequency, and logistics delay history, and using deep learning and machine learning techniques, it is possible to extract potential trends and abnormal patterns from historical data and real-time data, thereby predicting the future operating risks of nodes. By constructing a decision tree integration model based on random forests, the system can analyze the risk factors of each node and calculate the failure prediction probability of each node. Based on these predictions, the platform can proactively identify high-risk nodes before potential failures occur, and take emergency measures in a timely manner, such as optimizing inventory management, expanding payment channel bandwidth, and adjusting logistics resources, so as to avoid or mitigate possible operational interruptions. This method can not only effectively reduce the operational risks faced by the platform and reduce the costs and losses caused by failures, but also optimize resource allocation, improve overall operational efficiency and responsiveness, and enable the platform to respond quickly to complex and changing market environments, ensure the sustainable and stable development of the business, and enhance the platform's operational resilience and market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The present invention will be further described below in conjunction with the accompanying drawings.
[0077] Figure 1 It is a flowchart of the specific steps of an e-commerce management method based on big data of the present invention;
[0078] Figure 2 It is a flow chart of an e-commerce management method system based on big data in the present invention. DETAILED DESCRIPTION
[0079] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0080] See also Figure 1 As shown, the present invention is an e-commerce management method based on big data, comprising the following steps:
[0081] S1: Collect data from each business operation node in the e-commerce platform, including: user behavior data (such as browsing, searching, clicking), transaction data (such as order generation, payment behavior, transaction amount), product information (such as inventory, product price, promotion activities), logistics information (such as delivery status, delivery time) and customer feedback data (such as evaluation, return and exchange information);
[0082] S2: Analyze the collected data traffic, including: data transmission volume, flow direction and interaction frequency, monitor the interaction mode between different operation nodes, detect the fluctuation mode and abnormal behavior of data traffic, and identify potential abnormal interaction nodes, including: user behavior nodes, product display nodes, payment nodes, order management nodes, logistics nodes and customer feedback nodes;
[0083] S3: Based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and operation success rate are integrated and calculated to obtain the failure node coefficient, which is used to determine whether the abnormal interaction node is a failure node;
[0084] The failure nodes include delayed payment response, delayed order processing, untimely inventory updates, and inaccurate logistics distribution;
[0085] S4: Based on the judgment results, each interactive node is divided into failed nodes and non-failed nodes, and early warning processing is performed on failed nodes. For non-failed nodes, a multi-dimensional risk assessment model is constructed by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and the risk coefficient is calculated to predict the operation risk of the node in the future and perform emergency processing based on the prediction results.
[0086] In S1: data from each business operation node in the e-commerce platform is collected, including: user behavior data (such as browsing, searching, clicking), transaction data (such as order generation, payment behavior, transaction amount), product information (such as inventory, product price, promotional activities), logistics information (such as delivery status, delivery time) and customer feedback data (such as evaluation, return and exchange information); specifically including:
[0087] In the data collection process, firstly, real-time data collection is performed on each business operation node in the e-commerce platform in time series;
[0088] Specifically, the platform continuously records and collects various types of data through log systems, databases and real-time data streaming platforms, including user behavior data (such as browsing, searching, clicking), transaction data (such as order generation, payment behavior, transaction amount), product information (such as inventory, product prices, promotional activities), logistics information (such as delivery status, delivery time) and customer feedback data (such as evaluation, return and exchange information);
[0089] The collected data is marked with timestamps to form a time series to ensure that subsequent analysis can accurately reflect the trend of data changes over time;
[0090] The data collection module periodically collects data from each operation node at a set time interval (such as every minute) and stores it in a data warehouse or real-time processing system;
[0091] Through time serialization processing, data is organized and indexed in chronological order, allowing the system to trace back to data changes at a specific point in time at any time, and then analyze the behavior patterns, sales trends, inventory changes, etc. of each node in different time periods, providing basic support in the time dimension for subsequent risk assessment and early warning.
[0092] In S2, the collected data traffic is analyzed, including: data transmission volume, flow direction and interaction frequency, monitoring the interaction mode between different operation nodes, detecting the fluctuation mode and abnormal behavior of data traffic, and identifying potential abnormal interaction nodes, which include: user behavior nodes, product display nodes, payment nodes, order management nodes, logistics nodes and customer feedback nodes, specifically including:
[0093] Obtain the transmission volume and flow direction of real-time data, and compare and analyze them with the transmission volume and flow direction of historical data. According to the degree of deviation between historical data and real-time data, calculate the consistency deviation coefficient to evaluate the degree of deviation of the real-time data of the current node;
[0094] Obtain the interaction frequency of real-time data, and calculate the abnormal fluctuation index of the interaction frequency of real-time data according to the fluctuation of the interaction frequency of real-time data, so as to evaluate the abnormal fluctuation degree of the interaction frequency of real-time data;
[0095] The consistency deviation coefficient of real-time data and the abnormal fluctuation index of interaction frequency are fused and normalized to obtain the node abnormality coefficient.
[0096] The logic for obtaining the consistency deviation coefficient is:
[0097] Collect real-time and historical data of each business operation node in the e-commerce platform;
[0098] Preprocess real-time data and historical data, including standardization, to eliminate the impact of different dimensions or data differences;
[0099] The dynamic time warping algorithm is used to calculate the alignment cost between real-time data and historical data, which reflects the similarity of the two time series, where:
[0100] The real-time data series is , represents the real-time data at each time point;
[0101] The historical data series is , represents the historical data in the corresponding time period;
[0102] By calculating the local distance and the minimum alignment cost , the final dynamic time warping distance is the minimum alignment path cost;
[0103] According to the final dynamic time warping distance, the consistency deviation coefficient is calculated;
[0104] The calculation expression of the consistency deviation coefficient is:
[0105] ;
[0106] in, represents the final dynamic time warping distance, Represents the time series corresponding to real-time data, Represents the time series corresponding to historical data, Indicates the total amount of real-time data. Indicates the total amount of historical data, Represents real-time data, Represents historical data, represents the consistency deviation coefficient, represents the local distance, represents the minimum cost, Indicates The standardized value of real-time data, Indicates The standardized value of historical data.
[0107] The acquisition logic of the abnormal fluctuation index of interaction frequency is:
[0108] Collect interaction frequency data of each operation node in the e-commerce platform;
[0109] The interaction frequency data is preprocessed, and the real-time interaction frequency data is normalized by using a standardized method for subsequent analysis;
[0110] The standardized interaction frequency data is modeled by a variational autoencoder, where the encoder maps the input data to the latent space, the decoder converts the variables in the latent space into reconstructed data, and calculates the mean square error between the interaction frequency of the actual data and the reconstruction frequency of the reconstructed data to obtain the reconstruction error at each time point;
[0111] The calculation expression of the reconstruction error is:
[0112] ;
[0113] In the formula, Represents each time point of the collection, Indicates The standardized interaction frequency at each time point is Indicates The normalized reconstruction frequency at each time point is Indicates The reconstruction error at each time point;
[0114] Calculate the standard deviation of the reconstruction error over a period of time and mean ;
[0115] Calculate the degree of deviation of the reconstruction error at each time point from the mean and standard deviation to obtain the abnormal fluctuation index of the interaction frequency at each time point. Sum the abnormal fluctuation index of the interaction frequency at each time point to obtain the abnormal fluctuation index of the interaction frequency in the corresponding time period. The calculation expression is:
[0116] ;
[0117] In the formula, represents the abnormal fluctuation index of interaction frequency, Indicates the total duration, represents the mean of the reconstruction error, Represents the standard deviation of the reconstruction error.
[0118] The calculation expression of the node abnormality coefficient is:
[0119] ;
[0120] In the formula, represents the node abnormality coefficient, represents the consistency deviation coefficient, represents the abnormal fluctuation index of interaction frequency, and is the preset scale factor, and and are greater than 0, Represents the logarithm of natural number coefficients.
[0121] Compare the node anomaly coefficient of each interaction node with a preset threshold;
[0122] If the node abnormality coefficient is greater than or equal to the preset threshold, it means that the corresponding node is abnormal;
[0123] If the node abnormality coefficient is less than the preset threshold, it means that the corresponding node is normal.
[0124] It should be noted that the node anomaly coefficient is a numerical indicator used to measure whether the performance of each operating node in the e-commerce platform is abnormal within a specific time period. By comparing the differences between historical data and real-time data, a coefficient value is calculated. If the behavior or state of the node deviates from the normal mode beyond a certain threshold, the anomaly coefficient is high, indicating that the node has potential failure or abnormal risk. The corresponding coefficient can help the platform promptly identify nodes with poor performance or potential failures, such as payment delays, untimely inventory updates, and delayed order processing, so as to provide early warning and intervention to avoid affecting the overall platform's operational efficiency.
[0125] In S3, based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and operation success rate are integrated and calculated to obtain the failure node coefficient, which is used to determine whether the abnormal interaction node is a failure node, including:
[0126] Collect interaction frequency data of each operation node in the e-commerce platform;
[0127] The interaction frequency data is preprocessed, and the real-time interaction frequency data is normalized by using a standardized method for subsequent analysis;
[0128] The standardized interaction frequency data is modeled by a variational autoencoder, where the encoder maps the input data to the latent space, the decoder converts the variables in the latent space into reconstructed data, and calculates the mean square error between the interaction frequency of the actual data and the reconstruction frequency of the reconstructed data to obtain the reconstruction error at each time point;
[0129] The calculation expression of the reconstruction error is:
[0130] ;
[0131] In the formula, Represents each time point of the collection, Indicates The standardized interaction frequency at each time point is Indicates The normalized reconstruction frequency at each time point is Indicates The reconstruction error at each time point;
[0132] Calculate the standard deviation of the reconstruction error over a period of time and mean ;
[0133] Calculate the degree of deviation of the reconstruction error at each time point from the mean and standard deviation to obtain the abnormal fluctuation index of the interaction frequency at each time point. Sum the abnormal fluctuation index of the interaction frequency at each time point to obtain the abnormal fluctuation index of the interaction frequency in the corresponding time period. The calculation expression is:
[0134] ;
[0135] In the formula, represents the abnormal fluctuation index of interaction frequency, Indicates the total duration, represents the mean of the reconstruction error, Represents the standard deviation of the reconstruction error;
[0136] Collect data from each business operation node in the e-commerce platform, including: operation frequency data and operation success rate data;
[0137] A Bayesian network model is constructed based on the operation frequency data and the operation success rate data, and a conditional dependency relationship between the operation frequency and the operation success rate and the failure node probability is defined, where the failure node probability indicates the probability of whether a node fails;
[0138] Set a prior distribution for each variable, where the operation frequency Normal distribution , operation success rate Normal distribution , the probability of failed nodes The conditional probabilities between the operation frequency and the operation success rate are expressed as and ;
[0139] Among them, the prior distribution of the operation frequency is:
[0140] ;
[0141] Among them, the prior distribution of the operation success rate is:
[0142] ;
[0143] Based on the prior distribution of operation frequency, the prior distribution of operation success rate and the conditional probability table, the joint probability of failed nodes is calculated using Bayesian network reasoning , and according to the values of operation frequency and operation success rate, the probability of node failure is obtained;
[0144] The calculation expression of the conditional probability of the failed node is:
[0145] ;
[0146] ;
[0147] Based on the conditional probability of failed nodes, the joint probability of failed nodes is obtained through Bayesian network reasoning , the calculation expression is:
[0148] ;
[0149] By calculating the failure probability of the node, the failure node coefficient is obtained. The calculation expression is:
[0150] ;
[0151] in, Represents each time point of the collection, Indicates The operation frequency data at each collection time point, Indicates The operation success rate data at each collection time point, represents the mean value of operation frequency, represents the standard deviation of the operation frequency, represents the mean of the operation success rate, represents the standard deviation of the operation success rate, Indicates the operation frequency The average impact on failed nodes, Indicates the operation frequency The standard deviation of the impact on failed nodes, represents the failure coefficient of the node, The prior distribution of operation frequency, represents the prior distribution of the operation success rate, represents the conditional probability of the operation frequency of the failed node, represents the conditional probability of the failed node operation qualification rate, represents the joint probability.
[0152] It is determined whether the failure coefficient of each interactive node is greater than or equal to a preset threshold. If so, it is recorded as a failed node; if not, it is recorded as a non-failed node.
[0153] It should be noted that the failure coefficient of each interactive node reflects whether the corresponding node can continue to be used, and the larger the value of the failure coefficient of the interactive node, the higher the failure degree of the corresponding node.
[0154] In S4, according to the judgment results, each interactive node is divided into failed nodes and non-failed nodes, and early warning processing is performed on failed nodes. For non-failed nodes, a multi-dimensional risk assessment model is constructed by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and the risk coefficient is calculated to predict the operation risk of the node in the future. Emergency processing is performed based on the prediction results, including:
[0155] Collect relevant data from non-failed nodes, including product sales data, inventory fluctuations, payment channel usage frequency, and logistics delay history;
[0156] Preprocess the collected relevant data of each non-failed node;
[0157] The processed data is converted into a feature matrix, where each row represents the feature of an interaction node and each column represents a feature;
[0158] The random forest algorithm is used to model the feature matrix and train multiple decision trees. Each tree is trained according to the node characteristics, and the prediction probability of whether the node is invalid is obtained through ensemble learning.
[0159] The training data of each decision tree is randomly extracted from the historical data of the interactive node, and each tree is trained using a different feature subset. The result of the ensemble learning is calculated by weighted average to obtain the predicted probability of node failure.
[0160] Based on the trained random forest model, each interactive node is predicted to obtain the predicted probability of node failure, and the failure risk coefficient is calculated based on the predicted probability. The calculation expression is:
[0161] ;
[0162] In the formula, represents the predicted probability of node failure, is the weight of failure risk, represents the failure risk factor;
[0163] Based on multiple dimensions such as commodity sales data, inventory fluctuations, payment channel usage frequency, and logistics delays, the risk factor of each node is weighted and calculated to obtain the comprehensive risk factor of each node.
[0164] All nodes are sorted according to the comprehensive risk factor, and high-risk nodes are prioritized. For high-risk nodes, emergency response processes are initiated, such as increasing payment channel bandwidth, optimizing inventory management, and strengthening logistics scheduling to reduce the probability of failure events.
[0165] See also Figure 2 As shown, an e-commerce management system based on big data includes:
[0166] A data collection module, which is used to collect data from various business operation nodes in the e-commerce platform, including: user behavior data, transaction data, product information, logistics information and customer feedback data;
[0167] A data flow monitoring and anomaly identification module, which processes the data collected from each commercial operation node, analyzes the data transmission volume, flow direction and interaction frequency, monitors the interaction mode between different operation nodes, detects the fluctuation mode and abnormal behavior of data flow, and identifies potential abnormal interaction nodes;
[0168] A failure node evaluation module, which analyzes the operation frequency and operation success rate of each abnormal interaction node based on the abnormal interaction node, performs fusion processing and calculation on the operation frequency and the operation success rate, and obtains a failure node coefficient for determining whether the abnormal interaction node is a failure node;
[0169] A multi-dimensional risk assessment and emergency response module, which divides each interactive node into failed nodes and non-failed nodes according to the judgment results, performs early warning processing on failed nodes, and for non-failed nodes, builds a multi-dimensional risk assessment model by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and calculates the risk coefficient to predict the operation risk of the node in the future period of time, and performs emergency processing based on the prediction results.
[0170] The working principle of the present invention is to improve the operational efficiency and risk management ability of the e-commerce platform through real-time data collection and analysis; the method of the present invention includes four main steps: first, the data of each commercial operation node in the e-commerce platform is collected through the log system and the real-time data flow platform, including user behavior, transactions, commodities, logistics and customer feedback information, and the time series processing is used to ensure the time series of the data; secondly, the collected data flow is analyzed, the consistency deviation coefficient and the abnormal fluctuation index of the interaction frequency are calculated, and the abnormal coefficient of the node is calculated by combining the two to identify potential abnormal interaction nodes; then, based on the abnormal coefficient of the node, combined with the operation frequency and success rate, the failure node coefficient is calculated by the Bayesian network model, and each node is judged to be a failure node; finally, for non-failed nodes, a multi-dimensional risk assessment model is constructed by analyzing commodity sales, inventory fluctuations, payment channel usage frequency and logistics delays, and the risk coefficient is calculated by using the random forest algorithm to predict the future operation risk of the node and take emergency measures in advance. The method can not only monitor the abnormal state of each node of the platform in real time, but also optimize operational decisions through intelligent risk prediction and reduce the occurrence of potential risks.
[0171] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0172] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0173] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0174] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0175] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. An e-commerce management method based on big data, characterized in that: The following steps are involved: S1: Collecting data from each business operation node in the e-commerce platform, including: user behavior data, transaction data, product information, logistics information and customer feedback data; S2: Process the data collected from each business operation node, analyze the data transmission volume, flow direction and interaction frequency, monitor the interaction mode between different operation nodes, detect the fluctuation pattern and abnormal behavior of data traffic, and identify potential abnormal interaction nodes; The nodes include: user behavior nodes, product display nodes, payment nodes, order management nodes, logistics nodes and customer feedback nodes; S3: Based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and operation success rate are integrated and calculated to obtain the failure node coefficient, which is used to determine whether the abnormal interaction node is a failure node; The failure nodes include delayed payment response, delayed order processing, untimely inventory updates, and inaccurate logistics distribution; S4: Based on the judgment results, each interactive node is divided into failed nodes and non-failed nodes, and early warning processing is performed on failed nodes. For non-failed nodes, a multi-dimensional risk assessment model is constructed by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and the risk coefficient is calculated to predict the operation risk of the node in the future and perform emergency processing based on the prediction results.
2. The e-commerce management method based on big data according to claim 1, characterized in that: The identifying of potential abnormal interaction nodes specifically includes: Obtain the transmission volume and flow direction of real-time data, and compare and analyze them with the transmission volume and flow direction of historical data. According to the degree of deviation between historical data and real-time data, calculate the consistency deviation coefficient to evaluate the degree of deviation of the real-time data of the current node; Obtain the interaction frequency of real-time data, and calculate the abnormal fluctuation index of the interaction frequency of real-time data according to the fluctuation of the interaction frequency of real-time data, so as to evaluate the abnormal fluctuation degree of the interaction frequency of real-time data; The consistency deviation coefficient of real-time data and the abnormal fluctuation index of interaction frequency are fused and normalized to obtain the node abnormality coefficient.
3. The e-commerce management method based on big data according to claim 2 is characterized in that: The acquisition logic of the consistency deviation coefficient is: Collect real-time and historical data of each business operation node in the e-commerce platform; Pre-process real-time data and historical data to eliminate the impact of different dimensions or data differences; The dynamic time warping algorithm is used to calculate the alignment cost between real-time data and historical data, which reflects the similarity of the two time series, where: The real-time data series is , represents the real-time data at each time point; The historical data series is , represents the historical data in the corresponding time period; By calculating the local distance and the minimum alignment cost , the final dynamic time warping distance is the minimum alignment path cost; According to the final dynamic time warping distance, the consistency deviation coefficient is calculated; The calculation expression of the consistency deviation coefficient is: ; in, represents the final dynamic time warping distance, Represents the time series corresponding to real-time data, Represents the time series corresponding to historical data, Indicates the total amount of real-time data. Indicates the total amount of historical data, Represents real-time data, Represents historical data, represents the consistency deviation coefficient, represents the local distance, represents the minimum cost, Indicates The standardized value of real-time data, Indicates The standardized value of historical data.
4. The e-commerce management method based on big data according to claim 2, characterized in that: The acquisition logic of the abnormal fluctuation index of interaction frequency is: Collect interaction frequency data of each operation node in the e-commerce platform; The interaction frequency data is preprocessed, and the real-time interaction frequency data is normalized by using a standardized method for subsequent analysis; The standardized interaction frequency data is modeled by a variational autoencoder, where the encoder maps the input data to the latent space, the decoder converts the variables in the latent space into reconstructed data, and calculates the mean square error between the interaction frequency of the actual data and the reconstruction frequency of the reconstructed data to obtain the reconstruction error at each time point; The calculation expression of the reconstruction error is: ; In the formula, Represents each time point of the collection, Indicates The standardized interaction frequency at each time point, Indicates The normalized reconstruction frequency at each time point is Indicates The reconstruction error at each time point; Calculate the standard deviation of the reconstruction error over a period of time and mean ; Calculate the degree of deviation of the reconstruction error at each time point from the mean and standard deviation to obtain the abnormal fluctuation index of the interaction frequency at each time point. Sum the abnormal fluctuation index of the interaction frequency at each time point to obtain the abnormal fluctuation index of the interaction frequency in the corresponding time period. The calculation expression is: ; In the formula, represents the abnormal fluctuation index of interaction frequency, Indicates the total duration, represents the mean of the reconstruction error, Represents the standard deviation of the reconstruction error.
5. The e-commerce management method based on big data according to claim 1, characterized in that: Based on the abnormal interaction nodes, the operation frequency and operation success rate of each abnormal interaction node are analyzed, and the operation frequency and the operation success rate are integrated and calculated to obtain the failure node coefficient, which specifically includes: Collect data from each business operation node in the e-commerce platform, including: operation frequency data and operation success rate data; A Bayesian network model is constructed based on the operation frequency data and the operation success rate data, and a conditional dependency relationship between the operation frequency and the operation success rate and the failure node probability is defined, where the failure node probability indicates the probability of whether a node fails; Set a prior distribution for each variable, where the operation frequency Normal distribution , operation success rate Normal distribution , the probability of a failed node The conditional probabilities between the operation frequency and the operation success rate are expressed as and ; Among them, the prior distribution of the operation frequency is: ; Among them, the prior distribution of the operation success rate is: ; Based on the prior distribution of operation frequency, the prior distribution of operation success rate and the conditional probability table, the joint probability of failed nodes is calculated using Bayesian network reasoning , and according to the values of operation frequency and operation success rate, the probability of node failure is obtained; The calculation expression of the conditional probability of the failed node is: ; ; Based on the conditional probability of failed nodes, the joint probability of failed nodes is obtained through Bayesian network reasoning , the calculation expression is: ; By calculating the failure probability of the node, the failure node coefficient is obtained. The calculation expression is: ; in, Represents each time point of the collection, Indicates The operation frequency data at each collection time point, Indicates The operation success rate data at each collection time point, represents the mean value of operation frequency, represents the standard deviation of the operation frequency, represents the mean of the operation success rate, represents the standard deviation of the operation success rate, Indicates the operation frequency The average impact on failed nodes, Indicates the operation frequency The standard deviation of the impact on failed nodes, represents the failure coefficient of the node, The prior distribution of operation frequency, represents the prior distribution of the operation success rate, represents the conditional probability of the operation frequency of the failed node, represents the conditional probability of the failed node operation qualification rate, represents the joint probability.
6. The e-commerce management method based on big data according to claim 1, characterized in that: According to the judgment result, each interactive node is divided into a failed node and a non-failed node, specifically including: It is determined whether the failure coefficient of each interactive node is greater than or equal to a preset threshold. If so, it is recorded as a failed node; if not, it is recorded as a non-failed node.
7. The e-commerce management method based on big data according to claim 1, characterized in that: The multi-dimensional risk assessment model is constructed and the risk coefficient is calculated, specifically including: Collect relevant data from non-failed nodes, including product sales data, inventory fluctuations, payment channel usage frequency, and logistics delay history; Preprocess the collected relevant data of each non-failed node; The processed data is converted into a feature matrix, where each row represents the feature of an interaction node and each column represents a feature; The random forest algorithm is used to model the feature matrix and train multiple decision trees. Each tree is trained according to the node characteristics, and the prediction probability of whether the node is invalid is obtained through ensemble learning. The training data of each decision tree is randomly extracted from the historical data of the interactive node, and each tree is trained using a different feature subset. The result of the ensemble learning is calculated by weighted average to obtain the predicted probability of node failure. Based on the trained random forest model, each interactive node is predicted to obtain the predicted probability of node failure, and the failure risk coefficient is calculated based on the predicted probability. The calculation expression is: ; In the formula, represents the predicted probability of node failure, is the weight of failure risk, Represents the failure risk factor.
8. An e-commerce management system based on big data, characterized in that: An e-commerce management method based on big data as claimed in any one of claims 1 to 7, comprising: A data collection module, which is used to collect data from various business operation nodes in the e-commerce platform, including: user behavior data, transaction data, product information, logistics information and customer feedback data; A data flow monitoring and anomaly identification module, which processes the data collected from each commercial operation node, analyzes the data transmission volume, flow direction and interaction frequency, monitors the interaction mode between different operation nodes, detects the fluctuation mode and abnormal behavior of data flow, and identifies potential abnormal interaction nodes; A failure node evaluation module, which analyzes the operation frequency and operation success rate of each abnormal interaction node based on the abnormal interaction node, performs fusion processing and calculation on the operation frequency and the operation success rate, and obtains a failure node coefficient for determining whether the abnormal interaction node is a failure node; A multi-dimensional risk assessment and emergency response module, which divides each interactive node into failed nodes and non-failed nodes according to the judgment results, performs early warning processing on failed nodes, and for non-failed nodes, builds a multi-dimensional risk assessment model by analyzing commodity sales data, inventory fluctuations, payment channel usage frequency and logistics delay history, and calculates the risk coefficient to predict the operation risk of the node in the future period of time, and performs emergency processing based on the prediction results.
Citation Information
Cited By
Cross-border e-commerce operation information monitoring and analyzing method and device and medium
CN120894023A
Sales behavior monitoring method and system based on data visualization
CN120996852A
Intelligent process configuration engine system supporting globalized business
CN121119658A
E-commerce promotion flow abnormity identification method and system
CN122196854A
An e-commerce promotion traffic anomaly identification method and system
CN122196854B