Cross-border e-commerce product sales prediction system based on big data correlation analysis

By combining the time difference analysis module and the topic analysis module, dynamically adjusting the prediction model is solved, and the cross-border e-commerce product sales forecasting system is insufficient in real-time data fusion and model flexibility, achieving efficient and accurate sales forecasts and rapid market response.

CN120338867AInactive Publication Date: 2025-07-18QUANZHOU MUYUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445651.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cross-border e-commerce product sales forecasting system based on big data correlation analysis is insufficient in real-time data fusion and dynamic modeling, and it is difficult to respond to sudden market changes and emerging consumption trends in a timely manner. The model flexibility is poor, the prediction accuracy is limited, and it is difficult to respond to market changes quickly.

Method used

Through the combination of the time difference analysis module and the topic analysis module, we can dynamically capture the changes in hot spots in domestic and foreign markets, build a prediction model, and adjust the model parameters in real time through the optimization feedback unit to achieve continuous optimization of the prediction model and effective use of data.

Benefits of technology

It improves the timeliness and accuracy of predictions, can accurately predict future sales values, supports enterprises to quickly respond to market changes, and optimize inventory management and marketing strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338867A_ABST
    Figure CN120338867A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cross-border e-commerce, in particular to a cross-border e-commerce product sales prediction system based on big data correlation analysis. The system comprises a data integration unit, a data analysis unit, a prediction modeling unit, a business application unit and an optimization feedback unit. The data integration unit collects domestic and overseas data to construct a unified data warehouse, the data analysis unit sorts and analyzes the data, the time difference analysis module calculates the time difference of the same domestic and overseas topic degree, the topic degree analysis module compares domestic and overseas sales volume, and if the domestic sales volume is larger than or equal to the overseas sales volume, a judgment module marks a topic degree keyword; the prediction modeling unit constructs a prediction model based on the matched keyword sales volume and time difference information to obtain a future sales volume value; and the optimization feedback unit compares the actual sales volume with a prediction result, judges the accuracy rate through a judgment module, feeds back keyword sales volume information with low accuracy rate to the topic degree analysis module, dynamically adjusts model parameters and recalculates prediction data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-border e-commerce, and more specifically, to a cross-border e-commerce product sales volume prediction system based on big data correlation analysis. Background Art

[0002] Cross-border e-commerce refers to an e-commerce model in which goods are sold to consumers in different countries or regions through an Internet platform, and payment, logistics, customs clearance and other links are completed.

[0003] For the existing cross-border e-commerce product sales volume prediction system based on big data correlation analysis, its operation steps are as follows: collect comprehensive business data including historical sales data, user behavior data, product attributes, market trends, competitor information, etc. from multiple business systems of the cross-border e-commerce platform, clean, standardize and integrate the collected data, and build a unified product sales volume prediction data warehouse as the analysis basis. Subsequently, through correlation analysis, feature engineering and machine learning model training, deeply mine and model the data, predict future product sales volume, and output the prediction results, which are presented to users in the form of charts, reports, etc., to help enterprises accurately grasp market demand, optimize inventory management, marketing strategies and product recommendations, and improve the operation efficiency and market competitiveness of cross-border e-commerce.

[0004] However, in the existing cross-border e-commerce product sales volume prediction system based on big data correlation analysis, although comprehensive business data including market trends is collected, the existing system mainly relies on historical data for prediction in actual application, and is insufficient in real-time data fusion and dynamic modeling. For example: the data update frequency is low (such as batch processing by week / month instead of real-time stream processing), the hot spot detection is delayed (lack of real-time NLP event extraction ability), and the model response lags (static feature engineering is not convenient for dynamically incorporating emerging hot spot features), resulting in difficulty in timely responding to sudden market changes and emerging consumption trends;

[0005] Secondly, the flexibility of the existing system model has defects. Even if advanced algorithms such as LSTM and XGBoost are used, the deployed model still has problems such as: static feature space (unable to dynamically incorporate features generated in real time; such as the instantaneous discount rate of competitors), parameter solidification (lack of an online learning mechanism and unable to automatically adjust weights with market changes), and single prediction dimension (no elastic prediction framework is established under multiple scenarios), making it difficult to flexibly combine real-time market trends (such as competitor promotion activities, changes in consumer preferences) to adjust the prediction model, resulting in limited prediction accuracy; moreover, after predicting the sales volume of cross-border e-commerce products, it is difficult for the system to perform real-time verification and recalculation of the prediction data according to the changes in big data hot spots at different times, which is not convenient for quickly responding to market changes and affects the timeliness of decision-making.

[0006] To solve the above problems, there is an urgent need for a cross-border e-commerce product sales prediction system based on big data correlation analysis. Summary of the Invention

[0007] The purpose of the present invention is to provide a cross-border e-commerce product sales prediction system based on big data correlation analysis to solve the problems raised in the above background technology.

[0008] To achieve the above purpose, a cross-border e-commerce product sales prediction system based on big data correlation analysis is provided, including a data integration unit, a data analysis unit, a prediction modeling unit, a business application unit, and an optimization feedback unit;

[0009] The data integration unit collects domestic and foreign sales data and time information to build a data warehouse. The data analysis unit processes and analyzes the sales data and time difference. The prediction modeling unit generates sales prediction data based on the analysis results. The business application unit records the actual sales data after the product is launched in real time, where:

[0010] The data analysis unit includes a time difference analysis module, a topic popularity analysis module, and a determination module. The time difference analysis module is used to calculate the sales peak time difference of the same topic products at home and abroad; the topic popularity analysis module is used to compare the domestic and foreign sales peak data. When the domestic sales volume ≥ the foreign sales volume, the determination module marks the topic keyword to obtain the marked keyword; the product keyword is extracted from the product information to be predicted. The prediction modeling unit obtains the corresponding sales data by matching the product keyword with the marked keyword, constructs a prediction model in combination with the sales peak time difference information, and outputs the future sales prediction value;

[0011] The optimization feedback unit compares the deviation between the actual sales volume and the prediction result. When the accuracy rate is lower than the threshold value, the relevant product data is fed back to the topic popularity analysis module for model parameter optimization and prediction value recalculation.

[0012] As a further improvement of this technical solution, the time difference analysis module establishes a time threshold through statistical analysis. By collecting N samples of the sales peak time difference of domestic and foreign topics with the same topic popularity, the mean μ and standard deviation σ of the N time difference samples are calculated. According to the formula:

[0013]

[0014] where: z is the z value corresponding to the confidence level.

[0015] As a further improvement of this technical solution, when the topic popularity analysis module is used to compare and train the domestic and foreign sales volume under the same topic popularity, it includes the following method steps:

[0016] Input layer: Input domestic sales volume data and foreign sales volume data;

[0017] Alignment layer: Ensure that the domestic sales data and the overseas sales data are aligned according to the same topic popularity. If the data is not aligned, interpolation processing is performed, and then re-matching is carried out;

[0018] Calculation layer: Calculate the difference between the domestic sales volume and the overseas sales volume for each topic popularity Perform calculations;

[0019] Filtering layer: Mark the topic popularity keywords with a difference ≥0, and filter out the topic popularity where the domestic sales volume is greater than or equal to the overseas sales volume;

[0020] Output layer: Output the marked topic popularity list and its corresponding domestic sales volume, overseas sales volume, and sales volume difference for subsequent analysis.

[0021] As a further improvement of this technical solution, the prediction modeling unit is used to calculate the sales volume and time of the product that needs to be predicted for sales volume, including the following method steps:

[0022] S1. Match the product keywords of the predicted product with the marked keywords marked in the topic popularity analysis module, and filter out several topic popularities related to this product;

[0023] S2. Extract the domestic sales volume data and overseas sales volume data of each topic popularity under the matched keywords;

[0024] S3. Combine the sales volume data under the marked keywords with the time threshold calculated by the time difference analysis module, and calculate the sales volume prediction value at the future time point through the time series analysis model;

[0025] S4. Transmit the prediction result to the front-end application in real time through the GraphQL interface, output the sales volume prediction value at the future time point, as well as the fluctuation range of the predicted sales volume, display the prediction result on the data visualization platform, and view the prediction result through the visualization platform.

[0026] In another technical solution, the optimization feedback unit is used to compare the actual result with the sales volume prediction value, and quantify the accuracy of the prediction result and obtain the prediction accuracy rate by calculating the absolute error, relative error, mean absolute error (MAE), and mean absolute percentage error (MAPE).

[0027] As a further improvement of this technical solution, the determination module dynamically updates the storage and entry of topic popularity marks and actual sales data through accuracy classification, ensuring the continuous optimization of the prediction model and the effective utilization of data. If

[0028]

[0029] Store and retain the data with high accuracy (≥90%), unmark the data with low accuracy (<90%), and enter it into the topic analysis module for the next analysis.

[0030] As a further improvement of this technical solution, bring the time when the domestic and overseas topic degrees are generated in the actual sales data into the time difference analysis module, update the data, and bring the domestic and overseas sales data in the actual sales data of into the topic analysis module, recalculate the difference data, re-screen the topic degrees where the domestic sales volume is greater than or equal to the overseas sales volume, and re-match and mark.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] 1. In the cross-border e-commerce product sales prediction system based on big data correlation analysis, through the combination of the time difference analysis module and the topic analysis module, it can dynamically capture the changes in domestic and overseas market hotspots, flexibly adjust the prediction model, solve the problem in the prior art that it is difficult to predict hot-selling products based on the big data hotspots every year, and improve the timeliness and accuracy of prediction; furthermore, through the prediction modeling unit combined with the time threshold of the time difference analysis module and the marked keywords of the topic analysis module, it can accurately predict the sales volume value in a future period of time, and output the confidence interval and relevant analysis reports, providing scientific decision-making support for enterprises and improving the efficiency of inventory management, marketing strategies and product recommendations.

[0033] 2. In the cross-border e-commerce product sales prediction system based on big data correlation analysis, through the synergistic effect of the optimization feedback unit and the determination module, and through the classification processing of the prediction accuracy rate by the determination module, it can compare the actual sales volume with the prediction result in real time, dynamically update the storage and entry of the topic degree mark and the actual sales data, ensure the continuous optimization of the prediction model and the effective utilization of data; and dynamically adjust the data model parameters of the topic analysis module and recalculate the prediction data, solve the problem in the prior art that it is difficult to verify and recalculate the prediction data according to the changes in big data hotspots at different times, and realize the rapid response and continuous optimization to market changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the overall system structure block diagram of the present invention;

[0035] Figure 2 is the module structure block diagram of the data analysis unit assembly components of the present invention.

[0036] The meanings of each label in the figure are:

[0037] 100, Data Integration Unit; 110, Domestic and Overseas Module;

[0038] 200, Data Analysis Unit; 210, Time Difference Analysis Module; 220, Topic Popularity Analysis Module; 230, Judgment Module;

[0039] 300, Prediction Modeling Unit;

[0040] 400, Business Application Unit;

[0041] 500, Optimization Feedback Unit. Detailed Implementation Manner

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Please refer to Figure 1 As shown, a cross-border e-commerce product sales volume prediction system based on big data correlation analysis is provided, including a data integration unit 100, a data analysis unit 200, a prediction modeling unit 300, a business application unit 400, and an optimization feedback unit 500;

[0044] Through the data integration unit 100, domestic and overseas sales data and sales time are collected to construct a unified data warehouse. Based on the data warehouse, the data analysis unit 200 organizes and analyzes the sales data and sales time difference of the data. The prediction modeling unit 300 performs operations on the analyzed sales data and time difference information to obtain predicted sales volume data. After the predicted product is put on the market for actual sales, the business application unit 400 records the actual sales volume, where:

[0045] The data analysis unit 200 includes a time difference analysis module 210, a topic popularity analysis module 220, and a judgment module 230. The time difference analysis module 210 is used to calculate the time difference information between the sales volume peak values under the same topic popularity at home and abroad. The topic popularity analysis module 220 is used to compare the sales volume peak value data at home and abroad under the same topic popularity. If the domestic sales volume ≥ overseas sales volume, the judgment module 230 marks the keywords of this topic popularity to obtain marked keywords, extracts the product information to be predicted to obtain product keywords, and through information matching of the product keywords and the marked keywords, obtains the sales volume data under the corresponding marked keywords. The prediction modeling unit 300 is used to establish a prediction model based on the analysis results, and uses the matched marked keyword sales volume and time difference information as the input training set of the model to predict the sales volume value in the next period of time;

[0046] The optimization feedback unit 500 compares the actual sales volume recorded by the business application unit 400 with the prediction result, judges the accuracy level through the judgment module 230, brings the sales volume information of product keywords with low accuracy into the topic degree analysis module 220, and dynamically adjusts the data model parameters of the topic degree analysis module 220 and recalculates the prediction data.

[0047] In specific use, during the calculation of the sales volume of the product to be predicted, first, the data integration unit 100 collects all-round data including historical sales data, product attributes, market trends, etc. from multiple domestic and overseas data sources (such as e-commerce platforms, social media, market reports, etc.), and cleans, standardizes, and integrates the collected data according to domestic and overseas, and constructs a unified data warehouse;

[0048] Subsequently, the data analysis unit 200 sorts out and analyzes the data in the data warehouse. The time difference analysis module 210 calculates the time difference information under the same topic degree at home and abroad to determine the time offset of the topic degree at home and abroad; the topic degree analysis module 220 compares the sales volumes at home and abroad under the same topic degree. If the domestic sales volume ≥ the overseas sales volume, the judgment module 230 marks the keywords of this topic degree; the judgment module 230 matches according to the keyword information of the product and the marked keyword information to screen out the keywords related to the predicted product; the prediction modeling unit 300 analyzes the sales volume of the matched keywords and the time difference information, constructs a prediction model based on the analysis result, calculates the sales volume value in the next period of time, and outputs the prediction result, including the predicted sales volume value, confidence interval, and relevant analysis reports;

[0049] After the product is put on the market for actual sales, the business application unit 400 extracts the actual sales amount and time information, and the optimization feedback unit 500 compares the actual sales volume of the business application unit 400 with the prediction result to calculate the prediction accuracy. If the accuracy ≥ 90%, the mark of this topic degree is retained, and the actual sales data is stored in the marked topic degree; if the accuracy < 90%, the mark of this topic degree is cancelled, and the actual sales data is input into the topic degree analysis module 220, and the data model parameters are dynamically adjusted and the prediction data is recalculated; during the dynamic update and recalculation process, the time difference analysis module 210 and the topic degree analysis module 220 re-analyze the updated actual sales data: bring the time generated by the domestic and overseas topic degrees in the actual sales data into the time difference analysis module 210 to update the time difference information; bring the domestic and overseas sales data in the actual sales data into the topic degree analysis module 220 to recalculate the difference data, screen out the topic degrees with domestic sales volume ≥ overseas sales volume, and re-match and mark according to the database information.

[0050] Further, the data integration unit 100 includes a domestic and overseas module 110, which is used to match the domestic and overseas topic degrees of the same topic, and classify the sales volumes and the time of generating sales volumes of products under the same topic degree at home and abroad.

[0051] In specific use, during the data collection process, first, the domestic and overseas module 110 collects data from multiple domestic and overseas data sources (such as e-commerce platforms, social media, search engines, market reports, etc.), divides them into two parts: a domestic database and an overseas database, and matches the topic degrees of the contents of the two databases to identify the same topic degrees in domestic and overseas data (such as popular products, star effects, market trends, etc.); classifies the data under the same topic degree to ensure the alignment of domestic and overseas data in terms of topic degree; then classifies the domestic and overseas data under the same matched topic degree by sales volume and time, providing a structured data basis for the subsequent data analysis unit 200.

[0052] Furthermore, the time difference analysis module 210 establishes a time threshold through statistical analysis. By collecting N samples of the time differences between the domestic and overseas topic degree peak times of the same topic degree, the mean μ and the standard deviation σ of the N time difference samples are calculated.

[0053] In specific use, by collecting N samples of the time differences between the domestic and overseas topic degree peak times of the same topic degree from the data warehouse, denoted as: ;

[0054] Among them, the time difference sample ( ) = overseas topic degree peak time ( ) - domestic topic degree peak time ( );

[0055] By calculating the mean μ and the standard deviation σ of the time difference samples;

[0056] Mean:

[0057]

[0058] Among them, represents the mean of the time difference samples, representing the average level of the time difference samples; represents the number of time difference samples; the k-th time difference sample value, .

[0059] Standard deviation:

[0060]

[0061] Among them, Represents the standard deviation of the time difference samples, measuring the degree of dispersion of the time difference sample data. The larger the standard deviation, the more dispersed the data; Represents the number of time difference samples; The k-th time difference sample value, ; Represents the mean of the time difference samples.

[0062] Find the corresponding z value (such as 1.96) according to the preset confidence level (such as 95%), and calculate the confidence interval:

[0063]

[0064] Among them, Represents the mean of the time difference samples; Represents the quantile of the standard normal distribution determined according to the preset confidence level; Represents the standard deviation of the time difference samples; Represents the number of time difference samples.

[0065] For example: Suppose the following time difference samples are collected: (unit: days)

[0066] Calculate the mean:

[0067] Calculate the standard deviation:

[0068]

[0069] Determine the confidence interval (assuming z = 1.96):

[0070]

[0071] Output result: The time threshold range is from 8.24 days to 13.36 days.

[0072] Furthermore, when the topic popularity analysis module 220 is used to perform comparative training on domestic and foreign sales under the same topic popularity, it includes the following method steps:

[0073] Input layer: Input domestic sales data and foreign sales data;

[0074] Domestic sales data: , where Represents the domestic sales volume of the i-th topic popularity;

[0075] Foreign sales data: , where Represents the domestic sales volume of the i-th topic popularity

[0076] Alignment layer: Ensure that the domestic sales data and the overseas sales data are aligned according to the same topic popularity. If the data is not aligned, perform matching or interpolation processing;

[0077] Calculation layer: Calculate the difference between the domestic sales volume and the overseas sales volume for each topic popularity Perform the calculation: ;

[0078] Where is the domestic sales volume, is the overseas sales volume,

[0079] Filter layer: Mark the topic popularity keywords with a difference ≥0, and filter out the topic popularity where the domestic sales volume is greater than or equal to the overseas sales volume: ;

[0080] Among them, represents the set of topic popularity where the domestic sales volume ≥ overseas sales volume;

[0081] Output layer: Output the marked list of topic popularity and its corresponding domestic sales volume, overseas sales volume, and sales volume difference for subsequent analysis;

[0082] For example: Suppose there is the following data:

[0083] Domestic sales data:

[0084] Overseas sales data:

[0085] List of topic popularity: ;

[0086] Input the domestic sales data and the overseas sales data, ensure that the domestic and overseas data are aligned according to the topic popularity (assuming they are already aligned), and calculate the sales volume difference:

[0087]

[0088]

[0089]

[0090]

[0091] At this time, the topic popularity that satisfies is , , ;

[0092] Marked list of topic popularity: ,

[0093] Corresponding domestic sales volume, overseas sales volume, and sales volume difference: ;

[0094] Thus, by comparing domestic and overseas sales volume data under the same topic popularity, select the topic popularity with domestic sales volume ≥ overseas sales volume.

[0095] Furthermore, the prediction modeling unit 300 is used to calculate the sales volume and time for products that need sales volume prediction, including the following method steps:

[0096] S1. Query domestic sales volume data and overseas sales volume data related to the matching topic popularity from the data warehouse: According to the matching topic popularity keywords (such as "Spring Festival promotion"), query domestic and overseas sales volume data in the data warehouse; the query conditions include topic popularity keywords and time range (such as data in the past year); subsequently, extract the time point lists of domestic and overseas data, align the domestic and overseas data by time point to form a unified time series data; among them, if a certain time point is missing in the domestic or overseas data, supplement the missing value through interpolation methods (such as linear interpolation, nearest neighbor interpolation); subsequently, store the aligned domestic sales volume data and overseas sales volume data in a temporary data table for subsequent analysis;

[0097] S2. Extract domestic sales volume data and overseas sales volume data for each topic popularity under the matching keywords;

[0098] S3. Through the time threshold range ([8.24, 13.36] days, for example) calculated by the time difference analysis module 210, select a suitable model according to the data characteristics, such as:

[0099] ARIMA (Autoregressive Integrated Moving Average Model): Suitable for linear time series data;

[0100] Prophet: Suitable for time series data with seasonality and trend;

[0101] LSTM (Long Short-Term Memory Network): Suitable for non-linear time series data;

[0102] Subsequently, organize the topic popularity sales volume data in time series, and perform smoothing processing (such as moving average) or differencing processing (such as removing trends) on the data, train the time series model using historical sales volume data, adjust the model parameters to optimize the model performance; determine the prediction time range according to the time difference threshold; use the trained model to predict the sales volume value at future time points; calculate the confidence interval through the prediction error (such as standard deviation) of the model;

[0103] S4. Output the sales volume prediction results at future time points, including predicted sales volume values, confidence intervals, and relevant analysis reports.

[0104] Furthermore, by optimizing the feedback unit 500 to compare the actual results with the sales forecast values, the steps to calculate the prediction accuracy are as follows:

[0105] Extract the actual sales data And the predicted sales data , where represents the actual sales volume at the i-th time point, represents the predicted sales volume at the i-th time point, ensuring that the actual sales data and the predicted sales data are aligned in terms of time points;

[0106] Subsequently, calculate the absolute error: represents the absolute difference between the predicted value and the actual value at each time point;

[0107] Among them, represents the absolute error at the i-th time point, reflecting the magnitude of the absolute difference between the predicted value and the actual value at this time point; represents the actual value at the i-th time point; The predicted value at the i-th time point.

[0108] Calculate the relative error: represents the relative difference percentage between the predicted value and the actual value at each time point;

[0109] Among them represents the relative error at the i-th time point, indicating the relative difference degree between the predicted value and the actual value at this time point in percentage form; represents the actual value at the i-th time point; The predicted value at the i-th time point.

[0110] Calculate the mean absolute error: represents the average absolute difference between the predicted value and the actual value;

[0111] Among them, represents the mean absolute error, measuring the overall average absolute difference degree between the predicted value and the actual value; represents the total number of time points; represents the actual value at the i-th time point; The predicted value at the i-th time point.

[0112] Calculate the mean absolute percentage error: represents the average relative difference percentage between the predicted value and the actual value;

[0113] Among them, represents the mean absolute percentage error, reflecting the overall average relative difference degree between the predicted value and the actual value in percentage form; represents the total number of time points; Denote the actual value at the i-th time point; Denote the predicted value at the i-th time point.

[0114] Calculate the prediction accuracy rate: Denote the accuracy level of the prediction result. For example, if , then the prediction accuracy rate is 90%;

[0115] In addition, the determination module 230 classifies the accuracy rate, dynamically updates the storage and entry of the topic popularity tags and the actual sales data, and ensures the continuous optimization of the prediction model and the effective utilization of the data. The specific steps are as follows:

[0116] Input data: prediction accuracy rate (calculated through the optimization feedback unit 500), actual sales data ; and the list of topic popularity tags (the marked topic popularity); if:

[0117]

[0118] For case:

[0119] Associate the actual sales data with the corresponding topic popularity tag, and store it in the database; and retain the topic popularity tag for subsequent analysis and prediction;

[0120] For case:

[0121] Cancel the tag of this topic popularity, and remove it from the tag list ; enter the actual sales data into the topic popularity analysis module 220 as new input data for the next analysis operation;

[0122] For example: Suppose there is the following data:

[0123] Prediction accuracy rate: ;

[0124] Actual sales data: ;

[0125] Topic popularity tag: ;

[0126] For case:

[0127] Since , enter the data storage and tag update process; associate the actual sales data with the topic popularity tag and store it in the database, retaining the topic popularity tag ; At this time, the updated topic popularity marking list is , and the stored actual sales data is ;

[0128] For 's situation:

[0129] Since , enter the process of canceling the mark and data entry; cancel the topic popularity mark , update the marking list to , and enter the actual sales data into the topic popularity analysis module 220 for the next analysis and calculation; at this time, the updated topic popularity marking list: is , and the actual sales data entered into the topic popularity analysis module is ;

[0130] Furthermore, bring the time when the domestic and overseas topic popularity in the actual sales data is generated into the time difference analysis module 210, and update the data. At the same time, bring the domestic and overseas sales data in the actual sales data into the topic popularity analysis module 220, recalculate the difference data, and filter out the topic popularity with domestic sales volume greater than or equal to overseas sales volume. The specific operation steps are as follows:

[0131] First, update the time difference analysis module 210:

[0132] S1. Input the time data when the domestic and overseas topic popularity in the actual sales data is generated and the historical time difference data;

[0133] S2. Extract the time points when the domestic and overseas topic popularity is generated from the actual sales data;

[0134] For example:

[0135] Domestic time: ;

[0136] Overseas time: ;

[0137] S3. According to the formula , calculate the domestic and overseas time difference for each topic popularity, and merge the new time difference data with the historical time difference data;

[0138] S4. Calculate the new mean and standard deviation , and update the confidence interval.

[0139] Secondly, update the topic popularity analysis module 220:

[0140] S1. Input the domestic and overseas sales volume data in the actual sales data and the historical sales volume difference data;

[0141] S2. Extract domestic and overseas sales volume data from the actual sales data;

[0142] For example:

[0143] Domestic sales volume: ;

[0144] Overseas sales volume: ;

[0145] S4. According to the formula calculate the difference between the domestic sales volume and the overseas sales volume for each topic popularity;

[0146] S5. Mark the topic popularity of the difference and screen out the topic popularity with domestic sales volume greater than or equal to overseas sales volume according to the formula ;

[0147] S6. Combine the newly screened topic popularity marks with the historical mark list and remove duplicates.

[0148] Finally, through the above steps, the system can dynamically update the data of the time difference analysis module 210 and the topic popularity analysis module 220, re-screen out the topic popularity with domestic sales volume greater than or equal to overseas sales volume, and update the mark list to ensure the continuous optimization of the prediction model and the effective utilization of data.

[0149] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A cross-border e-commerce product sales prediction system based on big data correlation analysis, characterized in that: It includes a data integration unit (100), a data analysis unit (200), a prediction modeling unit (300), a business application unit (400), and an optimization feedback unit (500); The data integration unit (100) collects domestic and overseas sales data and time information to build a data warehouse. The data analysis unit (200) processes and analyzes the sales data and time difference. The prediction modeling unit (300) generates sales prediction data based on the analysis results. The business application unit (400) records the actual sales data of the product after it is launched in real time, where: The data analysis unit (200) includes a time difference analysis module (210), a topic popularity analysis module (220), and a determination module (230). The time difference analysis module (210) is used to calculate the time difference between the sales peak times of domestic and overseas products with the same topic. The topic popularity analysis module (220) is used to compare the domestic and overseas sales peak data. When the domestic sales volume ≥ overseas sales volume, the determination module (230) marks the topic keyword to obtain a marked keyword. The product information to be predicted is extracted to obtain a product keyword. The prediction modeling unit (300) obtains the corresponding sales data by matching the product keyword with the marked keyword, constructs a prediction model in combination with the sales peak time difference information, and outputs the future sales prediction value; The optimization feedback unit (500) compares the deviation between the actual sales volume and the prediction result. When the accuracy rate is lower than the threshold, it feeds back the relevant product data to the topic popularity analysis module (220) for model parameter optimization and recalculation of the prediction value.

2. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 1, characterized in that: The data integration unit (100) includes a domestic and overseas module (110). The domestic and overseas module (110) uses a text similarity algorithm to perform semantic matching on the domestic and overseas topic popularity texts, identifies the same topic popularity, and aligns the domestic and overseas sales data in time series for the matched same topic popularity.

3. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 1, wherein: The time difference analysis module (210) establishes a time threshold through statistical analysis. By collecting N samples of the time differences between the domestic and overseas topic popularity peak times with the same topic popularity, the mean μ and standard deviation of the N time difference samples are calculated.

4. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 1, characterized in that: When the topic popularity analysis module (220) is used to compare and train the domestic and overseas sales volumes under the same topic popularity, it includes the following method steps: Input layer: Input domestic sales volume data and overseas sales volume data; Alignment layer: Ensure that the domestic sales volume data and overseas sales volume data are aligned according to the same topic popularity. If the data is not aligned, interpolation processing is performed and then re-matched; Calculation layer: Calculate the difference between the domestic sales volume and the overseas sales volume for each topic popularity Perform the calculation; Screening layer: For the difference Mark the topical keywords with a value of ≥0, and screen out the topicality with domestic sales greater than or equal to overseas sales; Output layer: Output the marked topic popularity list and its corresponding domestic sales volume, overseas sales volume, and sales volume difference for subsequent analysis.

5. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 1, characterized in that: The prediction modeling unit (300) is used to calculate the sales volume and time of the product that needs to be predicted for sales, including the following method steps: S1. Match the product keyword of the predicted product with the marked keywords marked in the topic popularity analysis module (220), and screen out several topic popularities related to the product; S2. Extract the domestic sales volume data and overseas sales volume data of each topic popularity under the matched marked keywords; S3. Combine the sales volume data under the marked keywords with the time threshold calculated by the time difference analysis module (210), and calculate the sales volume prediction value at a future time point through a time series analysis model; S4. Transmit the prediction result to the front-end application in real time through the GraphQL interface, output the sales volume prediction value at a future time point and the fluctuation range of the predicted sales volume, display the prediction result on the data visualization platform, and view the prediction result through the visualization platform.

6. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 1, wherein: The optimization feedback unit (500) is used to compare the actual result with the sales volume prediction value, quantify the accuracy of the prediction result by calculating the absolute error, relative error, mean absolute error (MAE), and mean absolute percentage error (MAPE), and obtain the prediction accuracy rate.

7. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 6, characterized in that: The determination module (230) dynamically updates the storage and entry of the topic degree mark and the actual sales data through accuracy classification to ensure the continuous optimization of the prediction model and the effective utilization of data. If 8. Store and retain the data with high accuracy (≥90%), cancel the mark for the data with low accuracy (<90%), and enter it into the topic degree analysis module (220) for the next analysis.

9. The cross-border e-commerce product sales prediction system based on big data correlation analysis according to claim 7, characterized in that: Bring the time when the topic popularity inside and outside the country in the actual sales data into the time difference analysis module (210) to update the data, and bring the domestic and overseas sales data in the actual sales data of into the topic popularity analysis module (220) to recalculate the difference data, re-screen the topic popularity where the domestic sales volume is greater than or equal to the overseas sales volume, and re-match and mark.