Big data-based e-commerce transaction platform and e-commerce transaction information collection method

By collecting and processing big data on an e-commerce trading platform, including mobile interaction behavior and third-party market price data, combining dynamic layered annotation and abnormal mode detection, a time-series transaction knowledge graph is built, and differential privacy processing is implemented, the problems of incomplete data collection, high fraud detection rate and conflicts between encryption technology and privacy compliance are solved, and more efficient user portraits, product recommendations and transaction monitoring are achieved.

CN120198199AInactive Publication Date: 2025-06-24LUBEI TECHNICIAN COLLEGE (BINZHOU AVIATION SECONDARY VOCATIONAL SCHOOL BINZHOU ENTREPRENEURSHIP UNIV BINZHOU ENTREPRENEURSHIP INCUBATION CENT)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510306752.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-15
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120198199A_ABST
    Figure CN120198199A_ABST
Patent Text Reader

Abstract

The invention relates to the field of e-commerce data collection. The e-commerce transaction information collection method based on big data comprises the following steps: collecting user mobile terminal interaction behavior data through an embedded SDK, synchronously calling an e-commerce platform API to obtain structured transaction data, and generating an original transaction data set; performing dynamic hierarchical labeling on the original transaction data through a hierarchical rule engine to obtain a transaction data set with hierarchical labels; dynamically adjusting the sampling weight of the high-risk transaction, and generating an optimized sampling instruction set; analyzing transaction data corresponding to the optimized sampling instruction set, and converting unstructured data into feature vectors; according to the method, differential privacy processing is performed on sensitive fields in a time sequence transaction knowledge graph, an auditing statistical report is generated through a homomorphic encryption algorithm, and a desensitized multi-dimensional analysis result is output, so that the precision of a user portrait and a commodity recommendation model is improved, the transaction fraud omission ratio is reduced, and the problem that a traditional encryption technology is incompatible with privacy compliance requirements is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of e-commerce data collection, and particularly to an e-commerce trading platform based on big data and a method for collecting e-commerce trading information. Background Art

[0002] With the exponential growth of the scale of e-commerce transactions, the transaction data generated by the platform every day has covered multi-source heterogeneous data such as user behavior logs, payment records, product evaluations, and cross-platform price comparison information.

[0003] However, the related technologies have problems such as incomplete collection of mobile interaction behaviors and third-party market price data, resulting in limited accuracy of user portraits and product recommendation models; there are problems that static stratified sampling strategies are difficult to adapt to dynamic risk scenarios, resulting in a fraud undetected rate of high-value transactions in cross-border e-commerce exceeding 25% and a sharp increase in the consumption of full-scale review resources; there are problems that traditional encryption technologies conflict with privacy compliance requirements, resulting in a decrease in the availability of desensitized data by more than 40%, which cannot support real-time risk control decisions. Summary of the Invention

[0004] Based on this, it is necessary to provide an e-commerce trading platform based on big data and a method for collecting e-commerce trading information for the above-mentioned technical problems, so as to improve the accuracy of user portraits and product recommendation models, reduce the fraud undetected rate of transactions, and solve the problem of incompatibility between traditional encryption technologies and privacy compliance requirements.

[0005] In the first aspect, the present application provides a method for collecting e-commerce trading information based on big data, including:

[0006] Collecting user mobile interaction behavior data through an embedded SDK, synchronously calling the e-commerce platform API to obtain structured transaction data, and combining distributed crawlers to capture third-party public market data to generate an original transaction data set;

[0007] Based on a preset transaction amount threshold, product risk level, and user credit score, dynamically stratify and label the original transaction data through a hierarchical rule engine to obtain a transaction data set with hierarchical labels;

[0008] According to the initial sampling rate assigned by the weight of the hierarchical label, an improved isolation forest algorithm is used to detect abnormal patterns in the transaction data set, dynamically adjust the sampling weight of high-risk transactions, and generate an optimized sampling instruction set;

[0009] Performing text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, converting unstructured data into feature vectors, and constructing a time-series transaction knowledge graph;

[0010] Implement differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate an auditable statistical report through a homomorphic encryption algorithm, and output the desensitized multi-dimensional analysis results.

[0011] Furthermore, implement differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate an auditable statistical report through a homomorphic encryption algorithm, and output the desensitized multi-dimensional analysis results, including:

[0012] Add Laplace noise to the user location data in the time-series transaction knowledge graph to generate fuzzy geographical information;

[0013] Encrypt and sum the payment amounts through the Paillier algorithm to generate an auditable encrypted statistical result;

[0014] Integrate the fuzzy geographical information and the auditable encrypted statistical result into the desensitized multi-dimensional analysis result.

[0015] Furthermore, perform text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, convert the unstructured data into feature vectors, and construct a time-series transaction knowledge graph, including:

[0016] Generate sentiment feature vectors by semantically vectorizing user comments through the BERT model;

[0017] Parse the screenshots of product detail pages through the CRNN network to generate structured product description data;

[0018] Perform entity relationship mapping on the sentiment feature vectors, structured product description data, and transaction data to generate a time-series transaction knowledge graph.

[0019] Furthermore, allocate an initial sampling rate according to the weights of hierarchical labels, use an improved isolation forest algorithm to detect abnormal patterns in the transaction data set, dynamically adjust the sampling weights of high-risk transactions, and generate an optimized sampling instruction set, including:

[0020] Allocate a high basic sampling rate for high-risk, a medium basic sampling rate for medium-risk layers, and a low basic sampling rate for low-risk layers based on the weights of hierarchical labels;

[0021] Detect abnormal patterns in the transaction data set through an improved isolation forest algorithm to generate abnormal transaction identifiers;

[0022] According to the dynamic distribution of abnormal transaction identifiers, increase the sampling weights of the corresponding layers to generate an optimized sampling instruction set.

[0023] Further, based on a preset transaction amount threshold, product risk level, and user credit score, the original transaction data is dynamically stratified and labeled through a hierarchical rule engine to obtain a set of transaction data with hierarchical labels, including:

[0024] Stratify the transaction amounts in the original transaction dataset according to a preset low-risk threshold, medium-risk range, and high-risk threshold to generate amount stratification labels;

[0025] Generate product risk level labels based on the weighted sum result of the historical return times and fake goods complaint times of the product;

[0026] Train the features of order completion rate, average customer unit price, and social network correlation through the XGBoost algorithm to generate user credit score labels;

[0027] Merge the amount stratification labels, product risk level labels, and user credit score labels into a set of transaction data with hierarchical labels.

[0028] Further, collect user mobile terminal interaction behavior data through an embedded SDK, synchronously call the e-commerce platform API to obtain structured transaction data, and combine distributed crawlers to capture third-party public market data to generate an original transaction dataset, including:

[0029] Capture user touch trajectory data through the embedded SDK, record the contact point coordinate sequence and corresponding timestamps to generate mobile terminal interaction behavior data;

[0030] Call the payment interface to obtain the transaction amount, product category, and user ID, parse and verify the data integrity to generate structured transaction data;

[0031] Use the Scrapy-Redis cluster to capture competitor platform price data, set the dynamic request interval and IP proxy pool rotation strategy to generate third-party public market data;

[0032] Merge the mobile terminal interaction behavior data, structured transaction data, and third-party public market data into the original transaction dataset.

[0033] In the second aspect, the present application also provides an e-commerce trading platform based on big data, including:

[0034] A data acquisition module for collecting user mobile terminal interaction behavior data through an embedded SDK, synchronously calling the e-commerce platform API to obtain structured transaction data, and combining distributed crawlers to capture third-party public market data to generate an original transaction dataset;

[0035] The first processing module dynamically stratifies and labels the original transaction data through a hierarchical rule engine based on a preset transaction amount threshold, commodity risk level, and user credit score, obtaining a set of transaction data with hierarchical labels.

[0036] The second processing module is used to allocate an initial sampling rate according to the weight of the hierarchical label, detect abnormal patterns in the transaction data set using an improved isolation forest algorithm, dynamically adjust the sampling weight of high-risk transactions, and generate an optimized sampling instruction set.

[0037] The third processing module is used to perform text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, convert unstructured data into feature vectors, and construct a time-series transaction knowledge graph.

[0038] The result output module is used to perform differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate an auditable statistical report through a homomorphic encryption algorithm, and output the desensitized multi-dimensional analysis result.

[0039] The technical solution provided by this application includes the following technical effects: By providing a method for collecting e-commerce transaction information based on big data, including: collecting user mobile terminal interaction behavior data through an embedded SDK, synchronously calling the e-commerce platform API to obtain structured transaction data, combining distributed crawlers to capture third-party public market data, and generating an original transaction data set; dynamically stratifying and labeling the original transaction data through a hierarchical rule engine based on a preset transaction amount threshold, commodity risk level, and user credit score, obtaining a set of transaction data with hierarchical labels; allocating an initial sampling rate according to the weight of the hierarchical label, detecting abnormal patterns in the transaction data set using an improved isolation forest algorithm, dynamically adjusting the sampling weight of high-risk transactions, and generating an optimized sampling instruction set; performing text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, converting unstructured data into feature vectors, and constructing a time-series transaction knowledge graph; performing differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generating an auditable statistical report through a homomorphic encryption algorithm, and outputting the desensitized multi-dimensional analysis result to improve the accuracy of the user portrait and commodity recommendation model, reduce the undetected rate of transaction fraud, and solve the problem of incompatibility between traditional encryption technologies and privacy compliance requirements. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 Flow chart of the method for collecting e-commerce transaction information based on big data in an embodiment of the present invention;

[0042] Figure 2 Structural diagram of the e-commerce transaction platform based on big data in an embodiment of the present invention. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] As Figure 1 shown, the present application provides a method for collecting e-commerce transaction information based on big data, including:

[0045] S101: Collect user mobile terminal interaction behavior data through an embedded SDK, synchronously call the e-commerce platform API to obtain structured transaction data, and combine a distributed crawler to capture third-party public market data to generate an original transaction data set.

[0046] Specifically, first, by embedding an embedded SDK in a mobile application, various interaction behavior data of users on the mobile terminal can be captured in real time, such as clicks, swipes, and residence times. The above data helps to deeply understand the operation habits and preferences of users. At the same time, by synchronously calling the API interface of the e-commerce platform, structured transaction data can be obtained, including order details, payment records, product information, etc. These data have high accuracy and standardization and are important bases for building user portraits and product recommendation models. In addition, in order to obtain more comprehensive market information, a distributed crawler technology is also combined to capture relevant data from third-party public market platforms, such as product price trends, competitor information, user evaluations, etc. These data can enrich the data resources of the platform and provide a broader perspective for market analysis and decision-making. By combining these three data collection methods, a rich and wide-coverage original transaction data set can be generated, providing a solid data foundation for subsequent data analysis and business optimization.

[0047] S102: Based on a preset transaction amount threshold, product risk level, and user credit score, dynamically layer and label the original transaction data through a hierarchical rule engine to obtain a transaction data set with hierarchical labels.

[0048] Specifically, first, stratification bases such as a transaction amount threshold, a commodity risk level, and a user credit score are preset. The transaction amount threshold can be set according to the platform's historical transaction data and risk control requirements. For example, the transaction amount can be divided into three intervals: high, medium, and low. The commodity risk level is comprehensively evaluated and determined based on factors such as the category of the commodity, price fluctuations, and market popularity. For example, electronic products, luxury goods, etc. usually have a higher risk level. The user credit score is calculated by analyzing multi-dimensional data such as the user's purchase history, payment behavior, and evaluation records, reflecting the user's credit status. Then, using a stratification rule engine, the original transaction data is processed according to the preset rules. The stratification rule engine dynamically divides each transaction data into different levels according to the interval to which the transaction amount belongs, the risk level of the commodity, and the user's credit score, and labels the corresponding stratification tags, finally obtaining a set of transaction data with stratification tags, providing a basis for subsequent risk assessment and precision marketing, etc.

[0049] S103: According to the weights of the stratification tags, an initial sampling rate is allocated, and an improved isolation forest algorithm is used to detect abnormal patterns in the set of transaction data, dynamically adjust the sampling weights of high-risk transactions, and generate an optimized sampling instruction set.

[0050] Specifically, first, according to the weights of the stratification tags, an initial sampling rate is allocated, comprehensively considering the influence of factors such as the transaction amount, the commodity risk level, and the user credit score on the initial sampling rate. Then, an improved isolation forest algorithm is used to detect abnormal patterns in the set of transaction data. By introducing an attention mechanism, this algorithm can dynamically adjust the features and sample points to be concerned, assign greater weights to important features, and thus more effectively identify abnormal transaction patterns. During the detection process, the algorithm will automatically adjust the construction method of the isolation tree according to the data distribution and the importance of the features to improve the accuracy and efficiency of abnormal detection. Finally, the sampling weights of high-risk transactions are dynamically adjusted according to the results of the abnormal detection. This step means that if certain transactions are identified as having a higher risk, then in the subsequent sampling process, these transactions will be given higher weights and are thus more likely to be selected for detailed review. Through the above series of steps, an optimized sampling instruction set is finally generated, which will guide the platform on how to scientifically sample transaction data to ensure that both high-risk transactions can be covered and the review resources can be reasonably utilized, thereby effectively reducing the undetected rate of transaction fraud and improving the platform's risk control ability.

[0051] S104: Perform text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, convert unstructured data into feature vectors, and construct a time-series transaction knowledge graph.

[0052] Specifically, first, perform sentiment analysis on the text content in transaction data, including text information such as user comments and product evaluations. Through natural language processing techniques, such as deep learning-based models (e.g., BERT, LSTM, etc.), preprocess the text, extract sentiment-related features, and judge the sentiment tendency to convert the text data into sentiment feature vectors. Secondly, for the image information in transaction data, such as product pictures and invoices uploaded by users, use optical character recognition (OCR) technology for text recognition. The process of OCR technology includes image preprocessing (such as grayscale conversion, binarization, denoising, etc.), text recognition, and post-processing, and finally extract the text information in the image and convert it into text feature vectors. For video stream data, such as user operation videos and product display videos, process them through video stream parsing technology. Convert the video content into video feature vectors. Finally, integrate the feature vectors extracted from different data types, combine the time series information of the transactions, and construct a time-series transaction knowledge graph. The construction of the knowledge graph includes steps such as entity recognition, relationship extraction, and time attribute association. By using the transaction data at different time points and their related features as nodes and edges, a knowledge network with a time dimension is formed, so as to be able to more comprehensively understand and analyze transaction behaviors and their evolution trends.

[0053] S105: Implement differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate an auditable statistical report through a homomorphic encryption algorithm, and output the desensitized multi-dimensional analysis results.

[0054] Specifically, first, differential privacy processing is implemented on sensitive fields in the time series transaction knowledge graph. In this step, by selecting an appropriate privacy budget ε value, a differential privacy algorithm such as the Laplace mechanism or the Gaussian mechanism is determined to add noise to the sensitive fields, so that the privacy information of a single record cannot be inferred, while maintaining the overall statistical characteristics of the data set as much as possible. For example, for a numeric sensitive field, its sensitivity can be calculated, and then the corresponding noise can be added according to the sensitivity and the privacy budget ε; for a text sensitive field, it can be converted into numeric data using a suitable encoding method before adding noise. Then, the desensitized data is encrypted using a homomorphic encryption algorithm to generate an auditable statistical report. The homomorphic encryption algorithm allows specific calculation operations to be performed on encrypted data, and the results are still encrypted, which ensures the confidentiality of the data during the statistical and analytical process. For example, the Paillier homomorphic encryption algorithm is used to encrypt the desensitized data, and then statistical operations such as summation and counting are performed on the encrypted domain to generate a statistical report. Finally, the desensitized multidimensional analysis results are output. In this step, the desensitized data that has been processed with differential privacy and homomorphic encryption is subjected to multi-dimensional analysis, such as analysis by time, product category, user group, etc., to extract valuable information while ensuring that the analysis results do not leak the original sensitive data. Through the above series of steps, the privacy of users is protected, and the effective analysis and utilization of transaction data is achieved, providing support for the operation and decision-making of the platform.

[0055] Furthermore, differential privacy processing is implemented on sensitive fields in the time-series transaction knowledge graph, and auditable statistical reports are generated through homomorphic encryption algorithms to output desensitized multi-dimensional analysis results, including:

[0056] Add Laplace noise to the user location data in the time-series transaction knowledge graph to generate fuzzy geographic information;

[0057] The payment amount is encrypted and summed using the Paillier algorithm to generate auditable encrypted statistical results;

[0058] Integrate fuzzy geographic information and auditable encrypted statistical results into desensitized multidimensional analysis results.

[0059] Specifically, first, differential privacy processing is performed on the user location data in the time-series transaction knowledge graph, and Laplace noise is added to generate blurred geographical information. According to the preset privacy budget ε value, the sensitivity of the user location data is calculated, and then the corresponding noise is added using the Laplace mechanism, so that the true location information of the user is blurred, thereby protecting the privacy of the user. Second, the payment amounts are encrypted and summed using the Paillier homomorphic encryption algorithm to generate an auditable encrypted statistical result. The Paillier algorithm is an additive homomorphic encryption algorithm that allows summation operations to be performed on encrypted data without decrypting the data, which ensures the confidentiality of the payment amount data during the statistical process. Specifically, after each payment amount is encrypted, the summation operation is directly performed in the encrypted domain, and the generated encrypted statistical result can be audited without revealing the original payment amount information. Finally, the blurred geographical information and the auditable encrypted statistical result are integrated into a desensitized multi-dimensional analysis result. In this step, the blurred geographical information processed by differential privacy and the encrypted statistical result generated by homomorphic encryption are combined to form a multi-dimensional analysis result containing geographical location and payment amount statistical information, while ensuring that this information does not disclose the sensitive data of the user, thereby providing support for the operation and decision-making of the platform while protecting privacy.

[0060] Furthermore, text sentiment analysis, image OCR recognition, and video stream parsing are performed on the transaction data corresponding to the optimized sampling instruction set, and the unstructured data is converted into feature vectors to construct a time-series transaction knowledge graph, including:

[0061] The semantic vectorization of user comments is performed through the BERT model to generate sentiment feature vectors;

[0062] The screenshots of the product detail pages are parsed through the CRNN network to generate structured product description data;

[0063] The sentiment feature vectors, structured product description data, and transaction data are subjected to entity relationship mapping to generate a time-series transaction knowledge graph.

[0064] Specifically, first, perform sentiment analysis on the text content in transaction data, including text information such as user comments and product evaluations. Through natural language processing techniques, such as deep learning-based models (e.g., BERT, LSTM, etc.), preprocess the text, extract sentiment-related features, and judge the sentiment tendency, converting the text data into sentiment feature vectors. Secondly, for the image information in transaction data, such as product pictures and invoices uploaded by users, use optical character recognition (OCR) technology for text recognition. The process of OCR technology includes image preprocessing (such as grayscale conversion, binarization, denoising, etc.), text recognition, and post-processing, finally extracting the text information in the image and converting it into text feature vectors. For video stream data, such as user operation videos and product display videos, process them through video stream parsing technology. This involves steps such as key frame extraction, object detection and recognition, and feature extraction of the video, converting the video content into video feature vectors. Finally, integrate these feature vectors extracted from different data types, combine with the time series information of the transactions, and construct a time-series transaction knowledge graph. The construction of the knowledge graph includes steps such as entity recognition, relationship extraction, and time attribute association. By using transaction data and their related features at different time points as nodes and edges, a knowledge network with a time dimension is formed, enabling a more comprehensive understanding and analysis of transaction behaviors and their evolution trends.

[0065] Furthermore, according to the weight distribution of hierarchical labels, allocate the initial sampling rate, and use the improved isolation forest algorithm to detect abnormal patterns in the transaction data set, dynamically adjust the sampling weights of high-risk transactions, and generate an optimized sampling instruction set, including:

[0066] Allocate a high basic sampling rate for high risks, a medium basic sampling rate for the medium-risk layer, and a low basic sampling rate for the low-risk layer based on the weights of hierarchical labels;

[0067] Use the improved isolation forest algorithm to detect abnormal patterns in the transaction data set and generate abnormal transaction identifiers;

[0068] According to the dynamic distribution of abnormal transaction identifiers, increase the sampling weights of the corresponding layers to generate an optimized sampling instruction set.

[0069] Specifically, first, according to the weights of the hierarchical tags, initial sampling rates are assigned to layers of different risk levels. Specifically, a high basic sampling rate is assigned to the high-risk layer, a medium basic sampling rate is assigned to the medium-risk layer, and a low basic sampling rate is assigned to the low-risk layer. In this step, by analyzing the characteristics and risk distribution of the transaction data, the initial sampling ratios of each layer are determined to ensure that high-risk transactions can be reviewed more frequently. Then, an improved isolation forest algorithm is used to detect abnormal patterns in the transaction data set. By introducing an attention mechanism, the improved isolation forest algorithm can dynamically adjust the features and sample points of concern, giving greater weights to important features, thereby more effectively identifying abnormal transaction patterns. During the detection process, the algorithm will automatically adjust the construction method of the isolation tree according to the data distribution and the importance of the features, generating abnormal transaction identifiers to improve the accuracy and efficiency of anomaly detection. Finally, according to the dynamic distribution of the abnormal transaction identifiers, the sampling weights of the corresponding layers are increased to generate an optimized sampling instruction set. This means that if certain transactions are identified as having higher risks, then in subsequent sampling processes, these transactions will be given higher weights and thus are more likely to be selected for detailed review. Through the above series of steps, an optimized sampling instruction set is finally generated, which will guide the platform on how to scientifically sample the transaction data to ensure that both high-risk transactions can be covered and the review resources can be reasonably utilized, thereby effectively reducing the undetected rate of transaction fraud and improving the risk control ability of the platform.

[0070] Furthermore, based on a preset transaction amount threshold, commodity risk level, and user credit score, the original transaction data is dynamically hierarchically labeled through a hierarchical rule engine to obtain a transaction data set with hierarchical tags, including:

[0071] The transaction amounts in the original transaction data set are hierarchically divided according to a preset low-risk threshold, medium-risk interval, and high-risk threshold to generate amount hierarchical tags;

[0072] Based on the weighted sum result of the historical return times and fake goods complaint times of the commodity, a commodity risk level tag is generated;

[0073] The order completion rate, average customer unit price, and social network association degree features are trained through the XGBoost algorithm to generate user credit score tags;

[0074] The amount hierarchical tags, commodity risk level tags, and user credit score tags are combined into a transaction data set with hierarchical tags.

[0075] Specifically, first, the transaction amounts in the original transaction dataset are stratified according to the preset low-risk threshold, medium-risk range, and high-risk threshold to generate amount stratification labels. In this step, by analyzing the distribution and risk characteristics of the transaction amounts, the amount ranges for different risk levels are determined, and the amount of each transaction is mapped to the corresponding risk layer. Then, based on the weighted sum result of the historical return times and fake goods complaint times of the commodity, a commodity risk level label is generated. Specifically, by counting the historical return times and fake goods complaint times of the commodity and performing a weighted sum according to the preset weights, the risk score of the commodity is obtained, and then the risk level of the commodity is determined. Next, the order completion rate, average customer unit price, and social network correlation degree features are trained through the XGBoost algorithm to generate user credit score labels. The XGBoost algorithm is an efficient machine learning algorithm that can process a large amount of feature data and generate an accurate prediction model. During the training process, features such as the user's order completion rate, average customer unit price, and social network correlation degree are used as inputs to train a model that can predict the user's credit score, thereby generating credit score labels for the users of each transaction. Finally, the amount stratification labels, commodity risk level labels, and user credit score labels are merged into a transaction data set with stratification labels. In this step, the three labels of each transaction are integrated to form a stratification label containing multi-dimensional risk information, providing a basis for subsequent risk assessment and precision marketing, etc.

[0076] Furthermore, the user mobile terminal interaction behavior data is collected through an embedded SDK, the structured transaction data is obtained by synchronously calling the e-commerce platform API, and the third-party public market data is crawled by combining with a distributed crawler to generate an original transaction dataset, including:

[0077] The user touch trajectory data is captured through the embedded SDK, and the contact point coordinate sequence and the corresponding timestamps are recorded to generate mobile terminal interaction behavior data;

[0078] The payment interface is called to obtain the transaction amount, commodity category, and user ID, and the data integrity is parsed and verified to generate structured transaction data;

[0079] The Scrapy-Redis cluster is used to crawl the price data of the competitor platform, and the dynamic request interval and the IP proxy pool rotation strategy are set to generate third-party public market data;

[0080] The mobile terminal interaction behavior data, structured transaction data, and third-party public market data are merged into the original transaction dataset.

[0081] Specifically, first, by embedding an embedded SDK in the mobile application, it is possible to capture the user's interaction behavior data on the mobile side in real time, specifically including the user's touch trajectory data, recording the sequence of contact coordinates and the corresponding timestamps, and generating the mobile-side interaction behavior data. This step helps to deeply understand the user's operation habits and preferences. At the same time, synchronously call the API interface of the e-commerce platform to obtain structured transaction data. Specifically, call the payment interface to obtain information such as transaction amount, product category, and user ID, and parse and verify the obtained data to ensure data integrity, generating reliable structured transaction data. These data are highly accurate and standardized and are an important basis for building user portraits and product recommendation models. In addition, in order to obtain more comprehensive market information, distributed crawler technology is also combined. Use the Scrapy-Redis cluster to crawl the price data of competing product platforms, set the dynamic request interval and the IP proxy pool polling strategy, and generate third-party public market data. These data can enrich the platform's data resources and provide a broader perspective for market analysis and decision-making. Finally, merge the mobile-side interaction behavior data, structured transaction data, and third-party public market data into the original transaction data set, providing a solid data foundation for subsequent data analysis and business optimization.

[0082] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0083] In one embodiment, as Figure 2 shown, the present application also provides an e-commerce trading platform 200 based on big data, including:

[0084] A data acquisition module 201, configured to collect user mobile-side interaction behavior data through an embedded SDK, synchronously call the e-commerce platform API to obtain structured transaction data, and combine distributed crawlers to crawl third-party public market data to generate an original transaction data set;

[0085] A first processing module 202, based on a preset transaction amount threshold, product risk level, and user credit score, dynamically stratify and label the original transaction data through a hierarchical rule engine to obtain a set of transaction data with hierarchical labels;

[0086] The second processing module 203 is used to allocate an initial sampling rate according to the weights of the hierarchical tags, detect abnormal patterns in the transaction data set by using an improved isolation forest algorithm, dynamically adjust the sampling weights of high-risk transactions, and generate an optimized sampling instruction set;

[0087] The third processing module 204 is used to perform text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, convert unstructured data into feature vectors, and construct a time-series transaction knowledge graph;

[0088] The result output module 205 is used to perform differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate an auditable statistical report through a homomorphic encryption algorithm, and output the desensitized multi-dimensional analysis results.

[0089] Specifically, first, the data acquisition module 201 captures the interaction behavior data of users on the mobile terminal in real time through the embedded SDK, including information such as touch trajectories, contact point coordinate sequences, and timestamps. At the same time, it calls the e-commerce platform API to obtain structured transaction data, such as transaction amounts, product categories, user IDs, etc., and uses a distributed crawler to capture third-party public market data, such as competitor platform price data, and integrates them to generate an original transaction data set. Then, the first processing module 202 performs dynamic hierarchical annotation on the original transaction data by using a hierarchical rule engine based on a preset transaction amount threshold, product risk level assessment (weighted sum of historical return times and fake product complaint times of the product), and user credit score (generated by training features such as order completion rate, average customer unit price, and social network association degree through the XGBoost algorithm), to obtain a transaction data set with hierarchical tags.

[0090] Next, the second processing module 203 allocates an initial sampling rate according to the weights of the hierarchical tags, allocates a high basic sampling rate to the high-risk layer, a medium basic sampling rate to the medium-risk layer, and a low basic sampling rate to the low-risk layer, and uses an improved isolation forest algorithm (introducing an attention mechanism to adjust feature weights) to detect abnormal patterns in the transaction data set, generate abnormal transaction identifiers, and then, according to the dynamic distribution of these identifiers, increase the sampling weights of the corresponding layers to generate an optimized sampling instruction set. Next, the third processing module 204 performs multi-dimensional processing on the transaction data corresponding to the optimized sampling instruction set: uses the BERT model to generate sentiment feature vectors by semantic vectorization of user comments, parses the screenshots of product details pages through the CRNN network to generate structured product description data, and combines video stream parsing technology to extract video key frame feature vectors. Finally, map the feature vectors converted from these unstructured data to the transaction data to construct a time-series transaction knowledge graph containing time-series information.

[0091] Finally, the result output module 205 performs differential privacy processing on sensitive fields in the time-series transaction knowledge graph. For example, Laplace noise is added to user location data to generate fuzzy geographical information. At the same time, the payment amount is encrypted and summed through a homomorphic encryption algorithm (such as the Paillier algorithm) to generate an auditable encrypted statistical result. Finally, the fuzzy geographical information and the auditable encrypted statistical result are integrated to output the desensitized multi-dimensional analysis result, which not only protects user privacy but also provides comprehensive and accurate data support for platform operation and decision-making.

[0092] The result output module 205 is also used for:

[0093] Performing differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generating an auditable statistical report through a homomorphic encryption algorithm, and outputting the desensitized multi-dimensional analysis result, including:

[0094] Adding Laplace noise to user location data in the time-series transaction knowledge graph to generate fuzzy geographical information;

[0095] Encrypting and summing the payment amount through the Paillier algorithm to generate an auditable encrypted statistical result;

[0096] Integrating the fuzzy geographical information and the auditable encrypted statistical result into the desensitized multi-dimensional analysis result.

[0097] The third processing module 204 is also used for:

[0098] Performing text sentiment analysis, image OCR recognition, and video stream parsing on the transaction data corresponding to the optimized sampling instruction set, converting unstructured data into feature vectors, and constructing a time-series transaction knowledge graph, including:

[0099] Semantically vectorizing user comments through the BERT model to generate sentiment feature vectors;

[0100] Parsing screenshots of product detail pages through the CRNN network to generate structured product description data;

[0101] Performing entity relationship mapping on the sentiment feature vectors, structured product description data, and transaction data to generate a time-series transaction knowledge graph.

[0102] The second processing module 203 is also used for:

[0103] According to the weight distribution of hierarchical tags to assign an initial sampling rate, using an improved isolation forest algorithm to detect abnormal patterns in the transaction data set, dynamically adjusting the sampling weights of high-risk transactions, and generating an optimized sampling instruction set, including:

[0104] Assigning a high basic sampling rate to high-risk based on the weight of hierarchical tags, a medium basic sampling rate to the medium-risk layer, and a low basic sampling rate to the low-risk layer;

[0105] Detect abnormal patterns in the transaction data set by improving the Isolation Forest algorithm to generate abnormal transaction identifiers;

[0106] According to the dynamic distribution of the abnormal transaction identifiers, enhance the sampling weights of the corresponding layers to generate an optimized sampling instruction set.

[0107] The first processing module 202 is also used for:

[0108] Based on a preset transaction amount threshold, commodity risk level, and user credit score, perform dynamic hierarchical annotation on the original transaction data through a hierarchical rule engine to obtain a transaction data set with hierarchical labels, including:

[0109] Stratify the transaction amounts in the original transaction data set according to a preset low-risk threshold, medium-risk interval, and high-risk threshold to generate amount stratification labels;

[0110] Generate commodity risk level labels based on the weighted sum result of the historical return times and fake goods complaint times of the commodity;

[0111] Train features such as order completion rate, average customer unit price, and social network association degree through the XGBoost algorithm to generate user credit score labels;

[0112] Merge the amount stratification labels, commodity risk level labels, and user credit score labels into a transaction data set with hierarchical labels.

[0113] The data acquisition module 201 is also used for:

[0114] Collect user mobile terminal interaction behavior data through an embedded SDK, synchronously call the e-commerce platform API to obtain structured transaction data, and combine distributed crawlers to capture third-party public market data to generate an original transaction data set, including:

[0115] Capture user touch trajectory data through an embedded SDK, record the contact point coordinate sequence and the corresponding timestamps to generate mobile terminal interaction behavior data;

[0116] Call the payment interface to obtain the transaction amount, commodity category, and user ID, parse and verify the data integrity to generate structured transaction data;

[0117] Use the Scrapy-Redis cluster to capture competitor platform price data, set the dynamic request interval and IP proxy pool rotation strategy to generate third-party public market data;

[0118] Merge the mobile terminal interaction behavior data, structured transaction data, and third-party public market data into an original transaction data set.

[0119] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0120] The above embodiments only represent several implementation manners of the embodiments of the present application. The descriptions are relatively specific and detailed, but should not be construed as limiting the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A method for collecting e-commerce transaction information based on big data, characterized in that: The method comprises: Collect user mobile interactive behavior data through embedded SDK, synchronously call e-commerce platform API to obtain structured transaction data, and use distributed crawlers to capture third-party public market data to generate original transaction data sets; Based on the preset transaction amount threshold, commodity risk level and user credit score, the original transaction data is dynamically labeled by a hierarchical rule engine to obtain a transaction data set with hierarchical labels; Allocating an initial sampling rate according to the weights of the stratified labels, using an improved isolation forest algorithm to detect abnormal patterns on the transaction data set, dynamically adjusting the sampling weights of high-risk transactions, and generating an optimized sampling instruction set; Perform text sentiment analysis, image OCR recognition and video stream analysis on the transaction data corresponding to the optimized sampling instruction set, convert unstructured data into feature vectors, and construct a time-series transaction knowledge graph; Differential privacy processing is implemented on the sensitive fields in the time-series transaction knowledge graph, and auditable statistical reports are generated through a homomorphic encryption algorithm, and desensitized multi-dimensional analysis results are output.

2. The method for collecting e-commerce transaction information based on big data according to claim 1 is characterized in that: The sensitive fields in the time series transaction knowledge graph are differentially privacy processed, auditable statistical reports are generated through homomorphic encryption algorithms, and desensitized multi-dimensional analysis results are output, including: Adding Laplace noise to the user location data in the time-series transaction knowledge graph to generate fuzzy geographic information; The payment amount is encrypted and summed using the Paillier algorithm to generate auditable encrypted statistical results; The fuzzy geographic information and the auditable encrypted statistical results are integrated into the desensitized multi-dimensional analysis results.

3. The method for collecting e-commerce transaction information based on big data according to claim 1, characterized in that: The transaction data corresponding to the optimized sampling instruction set is subjected to text sentiment analysis, image OCR recognition and video stream parsing, unstructured data is converted into feature vectors, and a time series transaction knowledge graph is constructed, including: Use the BERT model to semantically vectorize user comments and generate sentiment feature vectors; Analyze product detail page screenshots through the CRNN network to generate structured product description data; The sentiment feature vector, the structured product description data and the transaction data are mapped into entity relationships to generate the time-series transaction knowledge graph.

4. The method for collecting e-commerce transaction information based on big data according to claim 1, characterized in that: The initial sampling rate is allocated according to the weight of the layered labels, an improved isolation forest algorithm is used to detect abnormal patterns on the transaction data set, the sampling weight of high-risk transactions is dynamically adjusted, and an optimized sampling instruction set is generated, including: Based on the weights of the stratified labels, a high basic sampling rate is assigned to the high risk layer, a medium basic sampling rate is assigned to the medium risk layer, and a low basic sampling rate is assigned to the low risk layer; Performing abnormal pattern detection on the transaction data set by using the improved isolation forest algorithm to generate an abnormal transaction identifier; According to the dynamic distribution of the abnormal transaction identifiers, the sampling weight of the corresponding layer is increased to generate the optimized sampling instruction set.

5. The method for collecting e-commerce transaction information based on big data according to claim 1, characterized in that: Based on the preset transaction amount threshold, commodity risk level and user credit score, the original transaction data is dynamically labeled by the hierarchical rule engine to obtain a transaction data set with hierarchical labels, including: stratifying the transaction amounts in the original transaction data set according to a preset low-risk threshold, a medium-risk interval, and a high-risk threshold, and generating amount stratification labels; Generate a product risk level label based on the weighted sum of the number of historical product returns and the number of counterfeit complaints; The XGBoost algorithm is used to train order completion rate, average customer unit price and social network association features to generate user credit score labels; The amount stratification label, the commodity risk level label and the user credit score label are combined into the transaction data set with stratification labels.

6. The method for collecting e-commerce transaction information based on big data according to claim 1, characterized in that: The embedded SDK is used to collect user mobile terminal interactive behavior data, and the e-commerce platform API is synchronously called to obtain structured transaction data. The distributed crawler is used to capture third-party public market data to generate the original transaction data set, including: Capturing user touch trajectory data through the embedded SDK, recording the contact point coordinate sequence and corresponding timestamp, and generating mobile terminal interaction behavior data; Call the payment interface to obtain the transaction amount, product category and user ID, parse and verify the data integrity, and generate structured transaction data; Use Scrapy-Redis cluster to crawl price data from competing platforms, set dynamic request intervals and IP proxy pool rotation strategies, and generate third-party public market data; The mobile terminal interaction behavior data, the structured transaction data and the third-party open market data are combined into the original transaction data set.

7. An e-commerce trading platform based on big data, characterized by: The platform comprises: The data acquisition module is used to collect user mobile terminal interaction behavior data through the embedded SDK, synchronously call the e-commerce platform API to obtain structured transaction data, and combine the distributed crawler to capture third-party public market data to generate the original transaction data set; The first processing module dynamically labels the original transaction data in layers through a layered rule engine based on a preset transaction amount threshold, commodity risk level, and user credit score, thereby obtaining a set of transaction data with layered labels; A second processing module is used to allocate an initial sampling rate according to the weights of the stratified labels, use an improved isolation forest algorithm to detect abnormal patterns on the transaction data set, dynamically adjust the sampling weights of high-risk transactions, and generate an optimized sampling instruction set; The third processing module is used to perform text sentiment analysis, image OCR recognition and video stream analysis on the transaction data corresponding to the optimized sampling instruction set, convert the unstructured data into feature vectors, and construct a time-series transaction knowledge graph; The result output module is used to implement differential privacy processing on sensitive fields in the time-series transaction knowledge graph, generate auditable statistical reports through a homomorphic encryption algorithm, and output desensitized multi-dimensional analysis results.

Citation Information

Cited By

  • Sampling method for forecasting rainfall sample under complex terrain condition of mountainous area

    CN121834114A