A Cross-border E-commerce Data Statistics and Management Method for Data Processing and Analysis

By establishing an order data storage database and anomaly database in cross-border e-commerce data statistics management, combined with Pearson's correlation coefficient calculation and correlation judgment, the data processing problem with the same order number but different other data is solved, and more efficient data deduplication and analysis are achieved, reducing the risk of data loss.

CN119576919BActive Publication Date: 2025-06-13中浩昌泽(北京)科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510115539.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the statistical management process of cross-border e-commerce data, it is difficult for the existing technology to effectively handle situations where the order numbers are the same but other data are different, resulting in the loss of important e-commerce data and affecting data analysis and business decisions.

Method used

By establishing an order data storage database and connecting with cross-border e-commerce platforms, obtain and store e-commerce data. Deduplication processing is used to store e-commerce data with the same order number but different other data into an abnormal database. Use Pearson's correlation coefficients to calculate the correlation coefficients of different categories of e-commerce data, and set the correlation coefficient threshold to determine the data correlation. The exception data is processed and deleted according to the correlation coefficient results.

Benefits of technology

It effectively reduces the amount of data, reduces the complexity of data processing, quickly locates problem data, improves the efficiency of data analysis, and improves the accuracy of deduplication by comprehensively considering the various characteristics of e-commerce data, and reduces the risk of data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576919B_ABST
    Figure CN119576919B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing. Specifically, it relates to a cross-border e-commerce data statistics management method for data processing and analysis. First, different e-commerce data except for the same order number are determined, and such e-commerce data are separately stored in an abnormal database. Different order numbers and corresponding other data are stored in an order data storage database. Then, the Pearson correlation coefficients of different categories of e-commerce data in the order data storage database are calculated, and the correlation between different categories of data is determined. According to the calculated Pearson correlation coefficients, the abnormal e-commerce data in the order data storage database are deleted. The separate storage and processing of abnormal data, and the deletion of abnormal data in the order data storage database reduce the data volume and the complexity of data processing. The calculation of the Pearson correlation coefficient and the determination of abnormal data based on this can quickly locate problem data and improve the efficiency of data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and more specifically, to a cross-border e-commerce data statistical management method for data processing and analysis. Background Art

[0002] In the statistical management process of cross-border e-commerce data, duplicate removal is a key precondition. Currently, the common duplicate removal technology mainly compares based on the order number. Once the same order number is found, one of them will be deleted. However, this method has obvious limitations;

[0003] In actual operation, the staff may enter the order number incorrectly for various reasons. At this time, other information in the e-commerce data, such as product details, customer information, transaction time, etc., may be completely different. If one of the data is simply deleted based on the same order number, it is very likely to cause the loss of important e-commerce data, thus having a serious negative impact on subsequent data analysis, business decision-making, and customer service.

[0004] In addition, in the cross-border e-commerce field, order numbers are usually arranged in ascending order. When there are cases where the order numbers are the same but other data are different, if the data with the same order number cannot be re-numbered, it is impossible to accurately distinguish and manage these orders, which will not only cause chaos in data management but also affect the smooth progress of the business process. In view of this, we propose a cross-border e-commerce data statistical management method for data processing and analysis. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem of directly deleting e-commerce data when the order numbers in the e-commerce data are the same but other data are different.

[0006] To achieve the above object, the present invention provides a cross-border e-commerce data statistical management method for data processing and analysis, including the following steps:

[0007] S1. Establish an order data storage database, establish a connection with the cross-border e-commerce platform through the HTTP protocol, obtain e-commerce data and store it;

[0008] S2. Perform duplicate removal processing on the e-commerce data in the order data storage database, and store the e-commerce data with the same order number but different other data in the abnormal database;

[0009] S3. Calculate the correlation coefficients of different categories of all e-commerce data in the order data storage database after deduplication using the Pearson correlation coefficient, and set a correlation coefficient threshold. Based on the correlation coefficient threshold and the correlation coefficient, determine whether there is a correlation between different categories in the e-commerce data. When there is a correlation, feedback the correlation in the e-commerce data to the staff, and determine the abnormal data in the e-commerce data;

[0010] S4. Sort the e-commerce data in the order data storage database in ascending order of the order number, and analyze whether there are discontinuous orders. When there are discontinuous order numbers, supplement the order numbers in descending order, and define the supplemented order numbers as supplementary numbers;

[0011] S5. Compare the supplementary numbers with the order numbers in the abnormal database, compare whether the order numbers are the same, whether the number of the same order numbers and the number of adjacent supplementary numbers are the same. When they are all the same, replace the order numbers in the abnormal database with the supplementary numbers, and analyze whether there is abnormal data again.

[0012] As a further improvement of this technical solution, the working principle of the connection between the order data storage database and the cross-border e-commerce platform is as follows:

[0013] The order data storage database sends a connection request to the cross-border e-commerce platform server address. After the cross-border e-commerce platform server receives the connection request from the order data storage database, it extracts the API key in the request and searches and compares it in the API key list stored in itself;

[0014] If the server compares and finds that the API key in the request is exactly the same as the API key stored in itself, it indicates that the identity verification of the order data storage database is successful, and the e-commerce data in the cross-border e-commerce platform is obtained. Otherwise, the request is resent and verified.

[0015] As a further improvement of this technical solution, the specific steps for deduplicating the e-commerce data in the order data storage database in S2 are as follows:

[0016] Step 1: Sense the query statement for querying the order number of the e-commerce data stored in the order data storage database, and input the query statement into the order data storage database. Traverse the e-commerce data stored in the order data storage database according to the query statement, and retrieve the order number corresponding to the query statement, which is the order number of the e-commerce data in the order data storage database, and define it as the order number set;

[0017] Step 2: Randomly select an order number from the order number set based on the random selection logic, and compare it with the unselected order numbers based on the duplicate determination logic to determine whether there are the same numbers;

[0018] If there is no order number in the order number set that is the same as the selected order number, it is determined that the e-commerce data stored in the order data storage database is not duplicated;

[0019] Step 3: If there is the same order number, retrieve other data with the same order number in the order data storage database. If the other data is still the same, delete the selected order number and the corresponding other data in the order data storage database, and continue to store the e-commerce data corresponding to the non-deleted order numbers in the order data storage database;

[0020] If the other data is different, store the same order number and the e-commerce data corresponding to the order number separately in the exception database;

[0021] Randomly select another order number in the order number set that has not been selected, and compare it with the order numbers in the order number set that have not been compared until all the order numbers in the order number set have been compared.

[0022] As a further improvement of this technical solution, the expression corresponding to the random selection logic is as follows: For the order number set S, there are N order numbers in S. Generate a random order number r through a random number generator, where the range is [0, N - 1]. The selected order number is n = S[r], expressed as n = S[random(0, N - 1)], where random(a, b) represents a function that generates a random number between a and b.

[0023] As a further improvement of this technical solution, the expression corresponding to the duplicate determination logic is as follows: For the selected order number n 1 and another order number n in S 2 , to determine whether they are duplicates is expressed as n 1 = n j . For the comparison of all order numbers in the order data set, perform the above judgment on all n 1 except n j , expressed as for and n j ≠ n 1 , check whether n j = n 1 holds.

[0024] As a further improvement of this technical solution, to determine whether other data is the same, based on the data consistency logic expression as follows: For the other data r 1 and r 2 corresponding to two identical order numbers, the order numbers of the other data r 1 and r 2 are respectively denoted as f 1, f 2 , …, f k , other data consistency judgment is expressed as:

[0025] For judgment r 1 (f i ) = r 2 (f i ) whether it holds, is expressed by the logical expression: (r 1 (f i ) = r 2 (f 1 )) ∧ (r 1 (f 2 ) = r 2 (f 2 )) ∧ … ∧ (r 1 (f k ) = r 2 (f k ));

[0026] Among them, ∧ represents the logical AND operation. When the result of this expression is true, other data is consistent; when the result is false, other data is inconsistent.

[0027] As a further improvement of this technical solution, the working principle of calculating the correlation coefficient by the Pearson correlation coefficient in S3 is as follows:

[0028] Perceive the e-commerce data after deduplication processing in the database of perceived order data, and analyze the same categories of all e-commerce data, and calculate the Pearson correlation coefficient between different categories;

[0029] The calculation formula of the Pearson correlation coefficient is:

[0030] Among them, x i and y i are the i-th numerical values corresponding to the e-commerce data, and are the average values of multiple selected e-commerce data, and n is the number of selected e-commerce data.

[0031] As a further improvement of this technical solution, the working principle of determining whether there is a correlation between multiple e-commerce data in S3 is as follows:

[0032] If the correlation coefficient between e-commerce data > the maximum value of the correlation coefficient threshold, it is determined that there is a correlation between e-commerce order data, which is a positive linear correlation. If the correlation coefficient < the minimum value of the correlation coefficient threshold, it is determined that there is a correlation between e-commerce order data, which is a negative linear correlation. Otherwise, it is determined that there is no linear correlation between e-commerce order data.

[0033] As a further improvement of this technical solution, the working principle of determining whether there is a correlation between different categories in the e-commerce data in S3 is as follows: After analyzing the correlation between e-commerce data, randomly select more than 2 e-commerce data again through random selection logic, calculate the correlation coefficient between different category data in the selected e-commerce data, and then determine whether there is a correlation between different category data through the correlation coefficient threshold. If there is a correlation in the whole e-commerce data analysis process, and there is no correlation at this time, the randomly selected e-commerce data will be defined as an abnormal data set;

[0034] Randomly select e-commerce data again. During the selection process, add any one e-commerce data in the abnormal data set. If it is determined that there is a correlation between different category data through the correlation coefficient threshold at this time, the added e-commerce data will be removed from the abnormal data set. Otherwise, it is determined that the added e-commerce data is abnormal data, and the abnormal data will be deleted from the order data storage database.

[0035] As a further improvement of this technical solution, in S4, sorting is performed in the order of the order number size. The basic idea is to select a reference order number, place the order numbers smaller than the reference order number on the left, and the numbers larger than the reference order number on the right, and then sort the left and right parts separately;

[0036] In S4, traverse and check the sorted order number sequence. By calculating the difference between adjacent order numbers, if the difference between adjacent order numbers is greater than 1, it is determined that the two adjacent numbers are not continuous.

[0037] Compared with the prior art, the beneficial effects of the present invention:

[0038] In this cross-border e-commerce data statistics and management method for data processing and analysis, first determine different e-commerce data except for the same order number, and store such e-commerce data separately in the abnormal database, store different order numbers and corresponding other data in the order data storage database, then calculate the Pearson correlation coefficient of different categories of e-commerce data in the order data storage database, and determine the correlation between different category data, and delete the abnormal e-commerce data in the order data storage database according to the calculated Pearson correlation coefficient;

[0039] The separate storage and processing of abnormal data, and the deletion of abnormal data in the order data storage database reduce the data volume and the complexity of data processing. The calculation of the Pearson correlation coefficient and the determination of abnormal data based on this can quickly locate problem data and improve the efficiency of data analysis;

[0040] Finally, by sorting the e-commerce data according to the order numbers of the e-commerce data, supplementing the missing order numbers during the sorting process, and re-numbering the e-commerce data in the abnormal database, and then calculating whether the re-numbered e-commerce data is abnormal data through the Pearson correlation coefficient, it is urged that when the enterprise performs deduplication processing, it no longer only focuses on the order numbers, but comprehensively considers other e-commerce data, explores a more comprehensive and accurate deduplication algorithm, such as comparing by combining multiple data features, improves the accuracy of deduplication, and reduces the risk of data loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of the overall steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] A cross-border e-commerce data statistical management method for data processing and analysis includes the following steps:

[0044] S1. Establish an order data storage database, establish a connection with the cross-border e-commerce platform through the HTTP protocol, obtain and store the e-commerce data, and the specific working principle is as follows:

[0045] The e-commerce data includes: order data (including order number, order time, customer information, product details, payment amount, etc.), product data (product ID, name, description, price, inventory, etc.), customer evaluation data and other data;

[0046] Before the order data storage database establishes a connection with the cross-border e-commerce platform, the cross-border e-commerce platform provides a server address and an API key. The server address is the target network location for the order data storage database to communicate with the platform, and the API key is the key information for identity authentication. It is unique, just like a key, and only the order data storage database holding the correct key can legally access the platform data;

[0047] The order data storage database sends a connection request to the cross-border e-commerce platform server address (the connection request includes the API key and the e-commerce data to be retrieved). After receiving the connection request from the order data storage database, the cross-border e-commerce platform server extracts the API key in the request. The server searches and compares it in the list of API keys stored in itself (the list of API keys is a database or a similar data storage structure that stores all legal API keys). If the server finds that the API key in the request is exactly the same as a certain API key stored in itself, it indicates that the authentication of the order data storage database is successful, and the e-commerce data in the cross-border e-commerce platform is retrieved. Otherwise, the request is resent and verified.

[0048] S2. Dedup the e-commerce data in the order data storage database. The specific steps are as follows:

[0049] Step 1: Sense the query statement for querying the order number of the e-commerce data stored in the order data storage database (the query statement is an instruction for retrieving, filtering, and operating on the data stored in the e-commerce data of the order data storage database), and input the query statement into the order data storage database. Traverse the e-commerce data stored in the order data storage database according to the query statement, and retrieve the order number corresponding to the query statement, which is the order number of the e-commerce data in the order data storage database, and define it as the order number set. Since the order number is set as the key field that uniquely identifies each order in the entire e-commerce data system, the order number can ensure that each order has a unique identity identifier.

[0050] Step 2: Randomly select an order number from the order number set based on the random selection logic, and compare it with the unselected order numbers based on the duplicate determination logic to determine whether there are identical numbers.

[0051] The expression corresponding to the random selection logic is as follows: For the order number set S with N order numbers in S, a random order number r is generated by a random number generator, where the range is [0, N - 1]. The selected order number is n = S[r], denoted as n = S[random(0, N - 1)], where random(a, b) represents a function that generates a random number between a and b.

[0052] The expression corresponding to the duplicate determination logic is as follows: For the selected order number n 1 and another order number n 2 in S, determine whether they are duplicates, denoted as n 1 = n j . For the comparison of all order numbers in the order data set, perform the above judgment on all n 1 except n j . It is denoted as for and n j ≠n 1 , check n j =n 1 to see if it holds;

[0053] If there is no order number in the order number set that is the same as the selected order number, it is determined that the e-commerce data stored in the order data storage database is not repeated.

[0054] In cross-border e-commerce, the management of e-commerce data is extremely crucial. With the expansion of business, the volume of e-commerce data has increased sharply. Due to manual input errors, system failures, data synchronization deviations, etc., duplicate order numbers often appear in the order data storage database;

[0055] If it is only determined based on this that the e-commerce data is repeated, on the one hand, valuable data may be mistakenly deleted. For example, when different customers place orders at similar times and the same order number is generated due to a system failure, but other data such as products and delivery addresses are different. Simply processing according to the repeated number will result in the loss of business information, affecting inventory, logistics, and customer service, damaging the customer experience, and even causing customer loss;

[0056] On the other hand, it interferes with the analysis results of data. When analyzing sales, customer behavior, and market strategies, inaccurate e-commerce data will cause deviations in the results and mislead decision-making. For example, mistakenly deleting valid order data leads to a low evaluation of regional sales performance, and the enterprise reduces investment, missing market growth opportunities. At the same time, it increases the workload and cost of data processing, requires more manpower and time to check and correct errors, and reduces operational efficiency.

[0057] Step 3: If there are the same order numbers, retrieve other data (other data are all e-commerce data except the order number) of the same order numbers in the order data storage database. If the other data are still the same, delete the selected order number and the corresponding other data in the order data storage database, and continue to store the e-commerce data corresponding to the undelete order numbers in the order data storage database;

[0058] Judging whether the other data are the same is based on the data consistency logical expression as follows: For the other data r 1 and r 2 corresponding to two identical order numbers, the order numbers of the other data r 1 and r 2 are respectively recorded as f 1 , f 2 , …, f k , and the other data consistency judgment is expressed as:

[0059] For judging r 1 (f i ) = r 2 (f iWhether it holds is expressed by the logical expression: (r 1 (f i ) = r 2 (f 1 )) ∧ (r 1 (f 2 ) = r 2 (f 2 )) ∧ … ∧ (r 1 (f k ) = r 2 (f k ));

[0060] Where ∧ represents the logical AND operation. When the result of this expression is true, other data is consistent; when the result is false, other data is inconsistent;

[0061] If other data is different, the same order numbers and the corresponding e-commerce data of the order numbers are stored separately in the exception database;

[0062] Randomly select another unselected order number from the set of order numbers, and compare it with the uncompared order numbers in the set of order numbers until all the order numbers in the set of order numbers are compared.

[0063] S3. Calculate the correlation coefficients of different categories of all e-commerce data in the deduplicated order data storage database using the Pearson correlation coefficient, and set a correlation coefficient threshold (the correlation coefficient threshold refers to the general standard for judging the relevance of different indicators in the industry. In the e-commerce industry, it is usually considered that when the absolute value of the correlation coefficient between the commodity price and the sales volume is greater than 0.6, there is a strong correlation, and this 0.6 is an empirical value obtained based on a large number of practices and studies in the industry). Based on the correlation coefficient threshold and the correlation coefficient, determine whether there is a correlation between multiple e-commerce data, and when there is a correlation, feedback the correlation in the e-commerce data to the staff;

[0064] And delete the e-commerce data that does not meet the correlation in the order data storage database according to the correlation between the e-commerce data;

[0065] The working principle of calculating the correlation coefficient using the Pearson correlation coefficient is as follows:

[0066] Perceive the e-commerce data after deduplication processing in the order data storage database, analyze the same categories of all e-commerce data, and calculate the Pearson correlation coefficient between different categories;

[0067] The calculation formula of the Pearson correlation coefficient is:

[0068] Where x i and y i are the i-th values corresponding to the e-commerce data, and is the average value of multiple e-commerce data, and n is the number of selected e-commerce data;

[0069] The working principle of determining whether there is a correlation between multiple e-commerce data is as follows:

[0070] If the correlation coefficient between e-commerce data > the maximum value of the correlation coefficient threshold, it is determined that there is a correlation between e-commerce single data, which is a positive linear correlation. If the correlation coefficient < the minimum value of the correlation coefficient threshold, it is determined that there is a correlation between e-commerce single data, which is a negative linear correlation. Otherwise, it is determined that there is no linear correlation between e-commerce single data;

[0071] Positive linear correlation: It means that when one data in e-commerce single data increases, another data will also increase accordingly, and the increasing trend shows a linear characteristic. When inventory and price show a positive linear correlation, it may mean that as inventory increases, price will also increase accordingly, indicating that there may be an inherent positive relationship between them, which may be due to factors such as inventory costs.

[0072] Negative linear correlation: When one data in e-commerce single data increases, another data decreases in approximately the same proportion. For example, in some promotional activities, there may be a situation where the price decreases while the sales volume (which can be reflected indirectly from the change in inventory) increases. At this time, the changes in price and inventory may show a negative linear correlation.

[0073] By performing a deduplication operation on e-commerce data, the uniqueness of the obtained e-commerce data is ensured. After this step, duplicate e-commerce data is removed, reducing data redundancy, making subsequent analysis more efficient. And because the Pearson correlation coefficient between different categories is calculated through all e-commerce data, due to the particularity of abnormal data, the relationship between it and other data may not conform to normal business logic or data rules. However, due to its small proportion, in the process of calculating the correlation coefficient, its contribution to the overall calculation result is relatively small, and it will not significantly distort or change the correlation coefficient value calculated based on normal data, and will not cause a large deviation in the correlation coefficient, and thus will not cause substantial interference to the analysis of the correlation between different categories;

[0074] After the analysis of the correlation between e-commerce data is completed, again randomly select more than 2 e-commerce data through random selection logic, calculate the correlation coefficient between different category data in the selected e-commerce data, and again determine whether there is a correlation between different category data through the correlation coefficient threshold. If there is a correlation during the analysis of all e-commerce data, and there is no correlation at this time, then the randomly selected e-commerce data is defined as an abnormal data set;

[0075] Then, randomly select e-commerce data again. During the selection process, add any e-commerce data from the abnormal data set. If it is determined at this time that there is a correlation between different categories of data through the correlation coefficient threshold, remove the added e-commerce data from the abnormal data set; otherwise, determine that the added e-commerce data is abnormal data and delete the abnormal data from the order data storage database.

[0076] Through continuous screening and adjustment of the abnormal data set, this process finally deletes the abnormal data from the order data storage database, which can gradually optimize the data quality in the database. After multiple screenings and verifications, it is ensured that the data remaining in the database is more in line with the true correlation between the data, avoiding abnormal data from misleading subsequent data analysis, decision support, etc., and providing a high-quality data foundation for more accurate business decisions.

[0077] S4. Sort the e-commerce data in the order data storage database in ascending order of the order number.

[0078] The basic idea of sorting in ascending order of the order number is to select a reference order number, place the order numbers smaller than the reference order number on the left, and the numbers larger than the reference order number on the right, and then sort the left and right parts separately.

[0079] For the order number a stored separately i and a j , if a i < a j , then swap the positions of a i and a j , and determine the sorting of the order numbers stored separately through multiple comparisons.

[0080] And analyze whether there are discontinuous orders. When there are discontinuous order numbers, supplement the order numbers in descending order, and define the supplemented order numbers as supplementary numbers.

[0081] Traverse and check the sorted order number sequence. By calculating the difference between adjacent two order numbers, if the difference between adjacent order numbers is greater than 1, it is determined that the two adjacent numbers are not continuous. For the sequence [1003, 1001, 1000], after sorting, it is [1003, 1001, 1000], and the difference between adjacent elements 1003 and 1001 is 2, so it is determined to be discontinuous.

[0082] When discontinuous order numbers are found, generate supplementary numbers according to the difference between adjacent numbers. For 1003 and 1001 in the above example, the supplementary number is 1002. The generation of supplementary numbers is calculated in sequence and according to the difference, filling the discontinuous gaps in descending order.

[0083] S5. Compare the supplementary number with the order number in the exception database. Check whether the order numbers are the same, whether the quantities of the same order numbers are the same, and whether the quantities of adjacent supplementary numbers are the same. If all are the same, then replace the order number in the exception database with the supplementary number and analyze again whether there is any abnormal data.

[0084] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A cross-border e-commerce data statistical management method for data processing and analysis, characterized in that: The following steps are involved: S1. Establish an order data storage database, and establish a connection with the cross-border e-commerce platform through the HTTP protocol to obtain and store e-commerce data; S2. Based on the random selection logic and the duplication determination logic, the e-commerce data in the order data storage database is deduplicated, and the e-commerce data with the same order number but different data other than the order number is stored in the abnormal database; S3. Calculate the correlation coefficients of different categories of all e-commerce data in the order data storage database after deduplication using the Pearson correlation coefficient, set a correlation coefficient threshold, determine whether there is a correlation between different categories in the e-commerce data based on the correlation coefficient threshold and the correlation coefficient, and if there is a correlation, feedback the correlation in the e-commerce data to the staff, and determine abnormal data in the e-commerce data, including after the correlation analysis between the e-commerce data is completed, randomly select more than 2 e-commerce data through random selection logic, calculate the correlation coefficients between different categories of data in the selected e-commerce data, and again determine whether there is a correlation between different categories of data through the correlation coefficient threshold. If there is a correlation in the process of analyzing all e-commerce data, but there is no correlation at this time, the randomly selected e-commerce data is defined as an abnormal data set; E-commerce data is randomly selected again. During the selection process, any e-commerce data in the abnormal data set is added. If the correlation coefficient threshold is used to determine that there is a correlation between data of different categories, the added e-commerce data is removed from the abnormal data set. Otherwise, the added e-commerce data is determined to be abnormal data and deleted from the order data storage database. S4. Select a base order number for sorting, sort the e-commerce data in the order data storage database in the order of the order numbers, and analyze whether there are discontinuous order numbers. If there are discontinuous order numbers, supplement the order numbers in descending order, and define the supplemented order numbers as supplementary numbers; S5. Compare the supplementary number with the order number in the abnormal database to see if they are the same, and whether the number of identical order numbers and the number of adjacent supplementary numbers are the same. If they are the same, replace the order number in the abnormal database with the supplementary number, and analyze again whether there is abnormal data therein.

2. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 1 is characterized by: The working principle of establishing a connection between the order data storage database and the cross-border e-commerce platform is as follows: The order data storage database sends a connection request to the cross-border e-commerce platform server address. After receiving the order data storage database connection request, the cross-border e-commerce platform server extracts the API key in the request, and the server searches and compares it in the API key list stored in itself; If the server finds that the API key in the request is exactly the same as the API key stored in its own storage, it indicates that the authentication of the order data storage database is successful, and the e-commerce data in the cross-border e-commerce platform is obtained. Otherwise, the request is resent and verified.

3. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 2 is characterized by: In S2, the e-commerce data in the order data storage database is deduplicated. The specific steps are as follows: Step 1: sensing a query statement for querying an order number of e-commerce data stored in the order data storage database, and inputting the query statement into the order data storage database, traversing the e-commerce data stored in the order data storage database according to the query statement, calling out the order number corresponding to the query statement, that is, the order number of the e-commerce data in the order data storage database, and defining it as an order number set; Step 2: randomly select an order number from the order number set based on the random selection logic, and compare it with the unselected order numbers based on the duplication determination logic to determine whether there is an identical number; If there is no order number identical to the selected order number in the order number set, it is determined that the e-commerce data stored in the order data storage database is not repeated; Step 3: If the same order number exists, call out other data of the same order number in the order data storage database; if the other data is still the same, delete the selected order number and the corresponding other data in the order data storage database, and continue to store the e-commerce data corresponding to the undeleted order number in the order data storage database; If other data are different, the same order number and the e-commerce data corresponding to the order number are stored separately in the abnormal database; An order number that has not been selected in the order number set is randomly selected again, and is compared with the order number that has not been compared in the order number set, until all order numbers in the order number set are compared.

4. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 3 is characterized by: The expression corresponding to the random selection logic is as follows: an order number set S, there are N order numbers in S, a random order number r is generated by a random number generator, the range is [0, N-1], the selected order number is n=S[r], expressed as n=S[random(0,N-1)], where random(0,N-1) represents a function for generating random numbers between 0 and N-1.

5. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 4 is characterized by: The expression corresponding to the duplication judgment logic is as follows: For the selected order number n1 and another order number n in S j , to determine whether it is repeated, it is represented by n1=n j , for the comparison of all order numbers in the order data set, for all n except n1 j The above judgment is expressed as And n1≠n j Order number, check n1 = n j Is it true? 6. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 5 is characterized by: The logical expression for judging whether other data are the same is based on data consistency as follows: For two identical order numbers r1 and r2, the other data types of r1 and r2 are recorded as f1, f2, …, f k , other data consistency judgments are expressed as: for Determine r1(f i )=r2(f i ) is true, expressed as a logical expression: (r1(f i )=r2(f1))∧(r1(f2)=r2(f2))∧…∧(r1(f k )=r2(f k )); The ^ represents a logical AND operation. When the result of the expression is true, the other data are consistent; when the result is false, the other data are inconsistent.

7. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 3 is characterized by: The working principle of calculating the correlation coefficient of the Pearson correlation coefficient in S3 is as follows: Sense the e-commerce data after deduplication in the order data storage database, analyze the categories that are the same in all e-commerce data, and calculate the Pearson correlation coefficient between different categories; The calculation formula of Pearson correlation coefficient is: Among them, x i and i is the i-th value corresponding to the e-commerce data, and is the average value of multiple e-commerce data, and m is the number of selected e-commerce data.

8. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 7 is characterized by: The working principle of determining whether there is a correlation between multiple e-commerce data in S3 is as follows: If the correlation coefficient between the e-commerce data is greater than the maximum value of the correlation coefficient threshold, it is determined that there is a correlation between the e-commerce data, which is a positive linear correlation. If the correlation coefficient is less than the minimum value of the correlation coefficient threshold, it is determined that there is a correlation between the e-commerce data, which is a negative linear correlation. Otherwise, it is determined that there is no linear correlation between the e-commerce data.

9. The cross-border e-commerce data statistics management method for data processing and analysis according to claim 1 is characterized by: In S4, the order numbers are sorted in order of size. The basic idea is to select a base order number, put the order numbers smaller than the base order number on the left, and the order numbers larger than the base order number on the right, and then sort the left and right parts respectively; In S4, the sorted order number sequence is traversed and checked, and the difference between two adjacent order numbers is calculated. If the difference between the adjacent order numbers is greater than 1, it is determined that the two adjacent numbers are not continuous.

Citation Information

Patent Citations

  • Data cleaning algorithm based on Internet trading information

    CN105045807A

  • Melon and fruit planting water and fertilizer consumption data collection and analysis system and method

    CN119106259A