Electronic commerce data processing method

Through the improved DBSCAN algorithm and the distributed particle swarm feature selection method based on rough sets, the problems of slow data processing speed and low accuracy are solved, and more efficient data classification and accurate information push are achieved.

WO2025123479A1PCT designated stage expired Publication Date: 2025-06-19HEBEI CHEM & PHARMA COLLEGE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/075926
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-02-05
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing e-commerce data processing methods have problems such as slow data processing speed, low efficiency and low data push accuracy.

Method used

The improved DBSCAN algorithm is used to reduce the dimensionality of labelless data, and the distributed particle swarm feature selection method based on rough set is used to process labeled data to improve data processing speed and classification accuracy.

Benefits of technology

It improves the speed of e-commerce data processing and classification accuracy, and ensures the accurate push of relevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024075926_19062025_PF_FP_ABST
    Figure CN2024075926_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electronic commerce, and disclosed is an electronic commerce data processing method. The method comprises the following steps: information acquisition: acquiring data of sub-platforms under an electronic commerce platform, comprising platforms and products browsed by users and data generated when purchasing products; information extraction: verifying and extracting the acquired data to extract labeled data and unlabeled data; information processing: respectively processing the labeled data and the unlabeled data, and performing data dimension reduction and classification; information storage: storing data having undergone information processing; and information pushing: on the basis of a data dimension reduction and classification result, pushing and showing relevant products or data to a target customer. The method improves the speed of electronic commerce data processing and the accuracy of electronic commerce data classification, thereby accurately pushing relevant information to target customers.
Need to check novelty before this filing date? Find Prior Art

Description

E-commerce data processing methods Technical Field

[0001] The present invention relates to the technical field of e-commerce, and in particular to an e-commerce data processing method with strong processing capability. Background Art

[0002] E-commerce is a business activity that uses information network technology as a means and is centered on the exchange of goods. It can also be understood as the activities of conducting transactions and related services in the form of electronic transactions on the Internet, Intranet and Value Added Network (VAN). It is the electronicization, networking and informatization of all aspects of traditional business activities.

[0003] E-commerce generally refers to a new business model that allows buyers and sellers to conduct various transactions face-to-face, across a wide range of commercial and trade activities globally, using the open internet environment and browser / server applications. This allows for online shopping for consumers, online transactions between merchants, online electronic payments, and various other business, trading, financial, and related integrated service activities. Various definitions have been offered by various sectors, depending on their respective positions, perspectives, and levels of participation in e-commerce. E-commerce is categorized as: ABC, B2B, B2C, C2C, B2M, M2C, B2A (i.e., B2G), C2A (i.e., C2G), O2O, and others. A variety of e-commerce data processing systems and methods exist in the prior art, but these methods suffer from shortcomings such as slow data processing speed, low efficiency, and low data push accuracy.

[0004] Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to provide an e-commerce data processing method that can improve the speed of data processing and ensure accurate push of information.

[0006] To solve the above technical problems, the present invention adopts a technical solution: an e-commerce data processing method, comprising the following steps:

[0007] Information collection: Collecting data from various sub-platforms under the e-commerce platform, including the platforms and products browsed by users, as well as data generated when purchasing products;

[0008] Information extraction: review and extract the collected data to extract labeled data and unlabeled data;

[0009] Information processing: process labeled data and unlabeled data separately, perform data dimensionality reduction and classification;

[0010] Information storage: storing the processed data;

[0011] Information push: Based on the results of data dimensionality reduction and classification, relevant products or data are pushed and displayed to target customers.

[0012] A further technical solution is to use an improved DBSCAN algorithm to reduce the dimensionality of unlabeled data: re-divide e-commerce websites according to the data density of business websites, divide websites with different densities into different areas, and then use the improved DBSCAN algorithm to perform clustering in different areas.

[0013] A further technical solution is to use a distributed particle swarm feature selection method based on rough sets to process labeled data: assuming a distributed environment formed by K physical hosts, the data set is divided into multiple sub-datasets according to the resource conditions of the K physical hosts; each physical host area is regarded as the activity range of the particle swarm, and the entire data set is fragmented into more particle swarms.

[0014] The entire DR_PSO algorithm is executed in parallel. Each physical host executes a Map_Reduce process to update the particle swarm position. In the Map phase, the main task is to generate subsets and generate the relevant data of the fitness function in the form of key-value pairs and pass them to the Reduce function. In the Reduce phase, the fitness function is mainly calculated in slices. Finally, the sum factor is called on the reduce result to obtain the optimal value of the current phase, and the global optimal value is updated accordingly.

[0015] A further technical solution is that the DR_PSO algorithm is used to select relevant features of high-dimensional large data sets for attribute reduction. The algorithm starts to calculate the feature kernel of the rough set and uses a greedy strategy based on feature importance to initialize the population to obtain a relatively better initial population, thereby reducing the later search process.

[0016] Secondly, the algorithm begins to execute particle swarm updates in parallel in multiple environments. The algorithm adopts a quantum particle swarm update strategy to increase population diversity by adding multiple random factors to avoid falling into local optimality.

[0017] In the Map phase, the position of the particle swarm is updated and calculated, the subset key-value pairs corresponding to the evaluation function are calculated, and the key-value pairs are formed and passed to the corresponding Reduce nodes. Finally, in the Reduce phase, the evaluation function is calculated and the global optimal value and the local optimal value of each particle swarm are updated.

[0018] The beneficial effect of adopting the above technical solution is that when processing e-commerce data, the method described in this application adopts an improved DBSCAN algorithm to reduce the dimensionality of unlabeled data, and adopts a distributed particle swarm feature selection method based on rough sets to process labeled data, thereby improving the speed of e-commerce data processing and the accuracy of e-commerce data classification, so that relevant information can be accurately pushed to target customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] FIG1 is a flow chart of a method according to an embodiment of the present invention;

[0021] FIG2 is a flow chart of the DBSCAN algorithm according to an embodiment of the present invention;

[0022] FIG3 is a flow chart of the DR_PSO algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0024] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0025] In general, as shown in FIG1 , an embodiment of the present invention discloses an e-commerce data processing method, comprising the following steps:

[0026] Information collection: Collecting data from various sub-platforms under the e-commerce platform, including the platforms and products browsed by users, as well as data generated when purchasing products;

[0027] Information extraction: review and extract the collected data to extract labeled data and unlabeled data;

[0028] Information processing: process labeled data and unlabeled data separately, perform data dimensionality reduction and classification;

[0029] Information storage: storing the processed data;

[0030] Information push: Based on the results of data dimensionality reduction and classification, relevant products or data are pushed and displayed to target customers.

[0031] The following describes the above method in detail:

[0032] In the process of information processing, this application mainly solves the problems of sparse data samples and distance calculation caused by the high dimension and low density of data, and extracts data dimensionality reduction and classification methods for labeled data and unlabeled data respectively.

[0033] Aiming at the problem of high-dimensional data mining in e-commerce data processing, a data mining method based on principal component analysis and improved DBSCAN clustering is proposed. Data dimensionality reduction is achieved through principal component analysis, and a principal component composition model of indicators is constructed. This model is used to design a new distance measurement method, improve the DBSCAN clustering algorithm, and solve the analysis problem of high-dimensional data.

[0034] Given the uneven data density of e-commerce websites, we first need to re-divide these websites based on density, assigning websites with different densities to different regions. We then cluster these regions using the improved DBSCAN algorithm, which uses an enhanced similarity metric. Therefore, the basic steps of the improved DBSCAN algorithm are as follows:

[0035] 1) Use density detection method to classify e-commerce websites. Use density detection method to detect the density variability of sample e-commerce demonstration websites.

[0036] 2) Clustering different densities using enhanced similarity metrics. After dividing the sample e-commerce websites, the DBSCAN algorithm combined with the e-commerce website similarity metric is used to obtain reasonable clustering results.

[0037] Use density detection method to classify e-commerce websites:

[0038] Two DBSCAN global parameters (Eps and MinPts) are used to filter data of varying densities. First, set the MinPts parameter to 4. A random e-commerce website is selected, and the three nearest neighbors are found. The Eps for that website is calculated. Similarly, the Eps for the three nearest neighbors are calculated, and finally, the Eps for each website is calculated. E-commerce websites are then assigned to different clusters based on their Eps.

[0039] As shown in Figure 2, the improved DBSCAN algorithm includes the following steps:

[0040] Step 1: Calculate the distance from point k to its three neighboring points {m, n, p}, Distance(k, y), y∈m,n,p;

[0041] Step 2: Set the Eps value of point k: Eps = max(Distance(k, y)), y∈m,n,p;

[0042] Step 3: Sort the Eps values ​​in descending order;

[0043] Step 4: Determine the category of e-commerce websites based on different density areas. These density areas are merged based on the ratio of two nearby EPS values. If the ratio is large enough, it indicates that data with different densities are captured.

[0044] Step 5: Through the enhanced similarity measure, the DBSCAN algorithm can generate different clusters from each type of e-commerce website.

[0045] Cluster different densities using an enhanced similarity measure:

[0046] In the traditional DBSCAN algorithm, similarity can be measured using the Euclidean distance function. However, Euclidean distance ignores the weighting of three influencing factors. The first commonality factor summarizes the number of visits per million users and the number of website links; the second commonality factor reflects the average number of page views and speed per person; and the third commonality factor represents the size of the website. For e-commerce websites, the number of visits per million users and the number of website links are the most important factors, so the weight of F1 should be much greater than that of F2 and F3.

[0047] Furthermore, a distributed particle swarm feature selection method based on rough sets is used to process labeled data: assuming a distributed environment formed by K physical hosts, the data set is divided into multiple sub-datasets according to the resource conditions of the K physical hosts; each physical host area is regarded as the activity range of the particle swarm, and the entire data set is sharded into more particle swarms; the entire DR_PSO algorithm is executed in parallel, and each physical host executes a Map_Reduce process to update the position of the particle swarm, among which: in the Map stage, the main task is to generate subsets, and generate the relevant data of the fitness function in the form of key-value pairs and pass them to the Reduce function; in the Reduce stage, the fitness function is mainly calculated in shards; finally, the sum factor is called on the reduce result to obtain the optimal value of the current stage, and the global optimal value is updated accordingly.

[0048] The DR_PSO algorithm primarily selects relevant features from large, high-dimensional datasets for attribute reduction. Based on rough set theory, it calculates the optimal attribute subset and transforms the problem into a particle swarm optimization problem. This algorithm executes multiple particle swarm optimization algorithms in parallel within a MapReduce environment, aiming to expand the search space and improve efficiency. The divide-and-conquer approach to feature subset evaluation functions in a distributed environment ensures the global nature of the optimal feature subset.

[0049] The algorithm begins by calculating the kernel of the rough set's features and initializing the population using a greedy strategy based on feature importance. This results in a relatively optimal initial population, reducing the need for subsequent searches. Secondly, the algorithm begins executing particle swarm updates in parallel across multiple environments. The algorithm employs a quantum particle swarm update strategy, increasing population diversity by adding multiple random factors to avoid falling into local optima.

[0050] In the Map phase, the position of the particle swarm is updated and calculated, the subset key-value pairs corresponding to the evaluation function are calculated, and the key-value pairs are formed and passed to the corresponding Reduce nodes. Finally, in the Reduce phase, the evaluation function is calculated and the global optimal value and the local optimal value of each particle swarm are updated.

[0051] In summary, as shown in Figure 3, the DR_PSO algorithm specifically includes the following steps:

[0052] 1) Calculate the feature core CORE;

[0053] Obtain the discriminant matrix and assign the single variable attribute to CORE. Define whether the feature kernel is the optimal reduction according to the rough set dependency. If so, directly return REDU = CORE, otherwise continue execution.

[0054] 2) Initialize the particle swarm;

[0055] Calculate the importance of non-core attributes in turn and initialize the particle swarm;

[0056] 3) Initialize the local optimal solution pbest and the global optimal solution gbest for each particle;

[0057] 4) When the stop condition is not met, each node calls the MapReduce process;

[0058] Map phase: Update the position of each particle and obtain the key-value pairs after the fitness function is decomposed;

[0059] Reduce stage: sum the results and obtain the fitness function of each particle swarm, and use it to update the pbest and gbest of each particle swarm;

[0060] 5) Return the optimal feature subset REDU.

[0061] When processing e-commerce data, the method described in the present application adopts an improved DBSCAN algorithm to perform dimensionality reduction on unlabeled data, and adopts a distributed particle swarm feature selection method based on rough sets to process labeled data, thereby improving the speed of e-commerce data processing and the accuracy of e-commerce data classification, and thus being able to accurately push relevant information to target customers.

Claims

1. An e-commerce data processing method, characterized in that The steps include: Information collection: Collect data from each sub-platform under the e-commerce platform, including the platforms and products browsed by users and the data generated when purchasing products; Information extraction: review and extract the collected data to extract labeled data and unlabeled data; Information processing: process labeled data and unlabeled data separately, perform data dimension reduction and classification; Information storage: storing the processed data; Information push: Based on the results of data dimensionality reduction and classification, relevant products or data are pushed and displayed to target customers.

2. The electronic commerce data processing method according to claim 1, characterized in that: For unlabeled data, the improved DBSCAN algorithm is used to reduce the dimension of the data: The e-commerce websites are re-divided according to the data density of the business websites, and the websites with different densities are divided into different areas. Then the improved DBSCAN algorithm is used to cluster them in different areas.

3. The electronic commerce data processing method according to claim 1, wherein: For labeled data, a distributed particle swarm feature selection method based on rough sets is used for processing: Assuming there is a distributed environment formed by K physical hosts, the data set is divided into multiple sub-data sets according to the resource conditions of the K physical hosts; each physical host area is regarded as the activity range of the particle swarm, and the entire data set is sharded into more particle swarms; The entire DR_PSO algorithm is executed in parallel. Each physical host executes a Map_Reduce process to update the position of the particle swarm. In the Map phase, the main task is to generate subsets, generate the relevant data of the fitness function in the form of key-value pairs, and pass them to the Reduce function. In the Reduce phase, the fitness function is mainly calculated in slices. Finally, the sum factor is called on the reduce result to obtain the optimal value of the current phase, and the global optimal value is updated accordingly.

Citation Information

Patent Citations

  • Electronic commerce data processing system

    CN106447464A

  • Customer classification method and system based on adaptive particle swarm

    CN110909773A

  • Commodity parallel dynamic pushing method in e-commerce platform

    CN110941771A

  • Internet-based e-commerce platform data analysis and decision-making system

    CN113240455A

  • Scene construction and evaluation method for park-level data value-added service

    CN113609177A