A cross-border e-commerce product recommendation system and method based on multi-source data fusion
By constructing a unified data warehouse for cross-border e-commerce and an improved BERT model, combined with dynamic weight adaptive algorithms and risk simulation technology, the problems of data integration and risk quantification in cross-border e-commerce product selection methods have been solved, enabling scientific decision-making and risk assessment for cross-border e-commerce product selection.
Patent Information
- Application Number
- CN202510881135.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Traditional cross-border e-commerce product selection methods cannot effectively integrate multi-source heterogeneous data, lack federated learning collaboration mechanisms and differential privacy protection, cannot deeply analyze the semantics of user feedback text, lack dynamic adaptive scoring mechanisms, and cannot accurately quantify cross-border trade risks.
By collecting cross-border e-commerce platform, user behavior, and external market data through distributed collectors, a unified cross-border data warehouse is built and federated learning collaboration is carried out. An improved BERT model is used for hierarchical text classification, and scoring is optimized by combining multi-dimensional feature extraction and dynamic weight adaptive algorithms. Risk assessment is carried out through real-time event data fusion and Monte Carlo risk simulation.
It has achieved comprehensive data integration and precise risk quantification for cross-border e-commerce product selection, improving the scientific nature and accuracy of product selection decisions and reducing the subjectivity and blindness of decision-making.
Smart Images

Figure CN120689120B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a cross-border e-commerce product recommendation system and method based on multi-source data fusion. Background Technology
[0002] With the acceleration of globalization and the rapid development of internet technology, cross-border e-commerce has become an important growth point for international trade. Cross-border e-commerce platforms offer consumers a wide variety of product choices, while also providing sellers with convenient channels to expand into overseas markets. Traditional cross-border e-commerce product selection methods mainly rely on manual experience, simple market research, and basic data statistical analysis. Sellers make product selection decisions by observing competitors' product performance, analyzing historical sales data, and referring to market reports. Existing product selection analysis technologies typically employ single-data source analysis methods, focusing primarily on basic product attributes such as price and sales volume. Some technical solutions incorporate simple user behavior data for analysis.
[0003] However, traditional product selection methods have significant technical limitations. First, existing technologies struggle to effectively integrate multi-source heterogeneous data, failing to fully utilize multi-dimensional information resources such as internal data from cross-border e-commerce platforms, external market data, and social media data. They also lack federated learning collaboration mechanisms and differential privacy protection technologies, resulting in a lack of comprehensiveness and accuracy in product selection decisions. Second, traditional methods lag behind in user feedback text analysis, lacking multilingual BERT variants and cross-lingual semantic alignment technologies. They cannot deeply analyze the complex sentiment tendencies and multi-layered preference information in cross-border user reviews, lack interpretable attention mechanisms and dynamic sub-classifier expansion capabilities, and can only perform simple sentiment polarity judgments, failing to capture users' nuanced preferences for different product attributes. Furthermore, existing product selection scoring mechanisms use a fixed-weight linear weighting method, lacking reinforcement learning frameworks and Q-learning optimization algorithms. They cannot dynamically adjust scoring strategies based on changes in the market environment and historical decision-making effects, and lack Pareto optimal solution search and weight convergence detection mechanisms, resulting in insufficient adaptability.
[0004] How to construct an effective mechanism for collecting and fusing multi-source heterogeneous data, breaking through the limitations of traditional single data sources, and realizing distributed collector deployment and federated learning collaborative data generation, has become a core technical problem that urgently needs to be solved. How to design advanced intelligent parsing technology for user feedback text, integrating multilingual processing capabilities and interpretable attention annotation functions, to deeply mine hierarchical semantic information and fine-grained sentiment preferences in user comments, surpassing the shallow text analysis capabilities of existing technologies. Simultaneously, how to establish a dynamic weight adaptive scoring optimization mechanism based on reinforcement learning, adjusting scoring weights in real time according to market changes and decision feedback to overcome the lack of adaptability of fixed-weight methods, and how to conduct accurate quantitative assessment of cross-border compliance risks through real-time event data fusion and Monte Carlo risk simulation, and achieve real-time sorting processing through lightweight preprocessing and streaming computing pipelines, are all problems that traditional technical solutions cannot effectively solve. Summary of the Invention
[0005] This application provides a cross-border e-commerce product recommendation system and method based on multi-source data fusion, which solves the technical problems of traditional cross-border e-commerce product selection methods being unable to effectively integrate multi-source heterogeneous data, deeply analyze the semantic information of user feedback text, dynamically and adaptively adjust scoring weights, and accurately quantify cross-border trade risks.
[0006] Firstly, this application provides a cross-border e-commerce product recommendation system based on multi-source data fusion, the cross-border e-commerce product recommendation system based on multi-source data fusion comprising:
[0007] The data collection module is used to collect and process multi-source heterogeneous data from cross-border e-commerce platforms, user behavior data, and external market data through a distributed collector to obtain a cross-border raw dataset. The cross-border raw dataset is then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data.
[0008] The classification module is used to perform hierarchical text classification processing based on user feedback text data in the cross-border unified data warehouse by improving the BERT model, so as to obtain user preference parsing results and interpretable attention annotations;
[0009] The quantification module is used to quantify the cross-border product selection features based on the cross-border unified data warehouse and the user preference analysis results, using a multi-dimensional feature extraction algorithm to obtain a cross-border product selection feature matrix.
[0010] The fusion module is used to perform weighted fusion processing on the product selection score based on the cross-border product selection feature matrix through a dynamic weight adaptive algorithm, so as to obtain the basic comprehensive score and the weight convergence detection result.
[0011] The assessment module is used to perform risk assessment processing based on the basic comprehensive score, through real-time event data fusion and Monte Carlo risk simulation, to obtain the risk probability distribution and cross-border compliance assessment results;
[0012] The sorting module is used to process the federated learning collaborative data and the interpretable attention annotations in real time through lightweight preprocessing and streaming computing pipelines to obtain a comprehensive score ranking list for cross-border product selection.
[0013] Secondly, this application provides a cross-border e-commerce product recommendation method based on multi-source data fusion, the method comprising:
[0014] The distributed collector performs multi-source heterogeneous data collection and processing on cross-border e-commerce platform data, user behavior data, and external market data to obtain cross-border raw datasets. The cross-border raw datasets are then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data.
[0015] Based on the user feedback text data in the cross-border unified data warehouse, hierarchical text classification is performed by improving the BERT model to obtain user preference parsing results and interpretable attention annotations;
[0016] Based on the cross-border unified data warehouse and the user preference analysis results, the cross-border product selection features are quantified using a multi-dimensional feature extraction algorithm to obtain a cross-border product selection feature matrix.
[0017] Based on the cross-border product selection feature matrix, the product selection scores are weighted and fused using a dynamic weight adaptive algorithm to obtain the basic comprehensive score and weight convergence detection results.
[0018] Based on the aforementioned comprehensive score, risk assessment is conducted through real-time event data fusion and Monte Carlo risk simulation to obtain the risk probability distribution and cross-border compliance assessment results.
[0019] Based on the federated learning collaboration data and the interpretable attention annotations, a comprehensive score ranking list for cross-border product selection is obtained through real-time processing using lightweight preprocessing and streaming computing pipelines.
[0020] Thirdly, a cross-border e-commerce product recommendation device based on multi-source data fusion is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the cross-border e-commerce product recommendation device based on multi-source data fusion to execute the aforementioned cross-border e-commerce product recommendation method based on multi-source data fusion.
[0021] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the aforementioned cross-border e-commerce product recommendation method based on multi-source data fusion.
[0022] The technical solution provided in this application uses a distributed collector to collect and process multi-source heterogeneous data from cross-border e-commerce platforms, user behavior data, and external market data. This effectively solves the technical problem of single data sources in traditional product selection methods. The distributed collection architecture can simultaneously obtain product information, user interaction data, and external market trend data from multiple cross-border e-commerce platforms, breaking through the limitations of existing technologies that rely solely on data from a single platform. The generation of federated learning collaborative data protects data privacy through differential privacy algorithms while enabling multi-party collaboration. The construction of a unified cross-border data warehouse provides a comprehensive and secure data foundation for subsequent analysis. The improved hierarchical text classification technology of the BERT model has played a key role in the field of cross-border e-commerce user feedback analysis. Multilingual BERT variants and cross-lingual semantic alignment layers can effectively handle multilingual reviews in cross-border scenarios. The interpretable attention annotation function improves the transparency and credibility of the model through keyword weight calculation. The online learning mechanism enables the dynamic updating of sub-classifiers to automatically adapt to the emergence of new product categories. Compared with traditional planar text classification methods, this hierarchical processing approach can more accurately identify users' preferences for different product attributes. Especially in the context of diversified user feedback content and complex expression in the cross-border e-commerce environment, the multilingual semantic understanding ability of the improved BERT model is significantly better than that of traditional text analysis techniques.
[0023] The multidimensional feature extraction algorithm is designed to meet the specific needs of cross-border e-commerce product selection. It quantifies features from multiple dimensions, including comprehensive user features, comprehensive product features, market trend features, and competitive features. The cross-border product selection feature matrix integrates key factors that affect the sales performance of cross-border products. Compared with existing technologies that only consider the basic attributes of products, the multidimensional feature extraction method of this application can comprehensively characterize the competitive position and development potential of products in the cross-border market.
[0024] The Q-learning optimization strategy achieves intelligent weight updates through Markov decision process modeling. The Pareto optimal solution search algorithm balances conflicting objectives such as profit maximization and risk minimization. The weight convergence detection mechanism effectively prevents overfitting, overcoming the technical shortcomings of traditional fixed-weight methods that cannot adapt to market changes. Real-time event data fusion and Monte Carlo risk simulation technology access geopolitical news and customs policy change events via API. The NLP risk event level extraction function dynamically corrects the risk index. The risk probability distribution generated by Monte Carlo simulation replaces the traditional standard deviation calculation method. The cross-border compliance assessment results integrate the target country's commodity access regulations as an independent risk dimension, significantly improving the accuracy of risk quantification. Lightweight preprocessing and streaming computing pipeline technology achieve real-time user profile updates through the Apache Flink framework. Combined with federated learning collaborative data and interpretable attention annotation, the generated cross-border product selection comprehensive score ranking list fully considers the complexity and dynamism of the cross-border trade environment, providing cross-border e-commerce sellers with scientific and accurate product selection guidance and significantly reducing the subjectivity and blindness of product selection decisions. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of an embodiment of a cross-border e-commerce product recommendation system based on multi-source data fusion in this application.
[0027] Figure 2 This is a schematic diagram of an embodiment of the cross-border e-commerce product recommendation method based on multi-source data fusion in this application.
[0028] Figure 3 This is a schematic block diagram of the cross-border e-commerce product recommendation device based on multi-source data fusion in an embodiment of the present invention. Detailed Implementation
[0029] This application provides a cross-border e-commerce product recommendation system and method based on multi-source data fusion. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the cross-border e-commerce product recommendation system based on multi-source data fusion in this application includes:
[0031] The data acquisition module 101 is used to collect and process multi-source heterogeneous data from cross-border e-commerce platform data, user behavior data, and external market data through a distributed collector to obtain cross-border raw datasets. The cross-border raw datasets are then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data.
[0032] Classification module 102 is used to perform hierarchical text classification processing based on user feedback text data in the cross-border unified data warehouse by improving the BERT model, and obtain user preference parsing results and interpretable attention annotations;
[0033] The quantification module 103 is used to quantify the cross-border product selection features based on the cross-border unified data warehouse and user preference analysis results, and obtain the cross-border product selection feature matrix through a multi-dimensional feature extraction algorithm.
[0034] The fusion module 104 is used to perform weighted fusion processing on the product selection score based on the cross-border product selection feature matrix through a dynamic weight adaptive algorithm to obtain the basic comprehensive score and weight convergence detection result.
[0035] Assessment module 105 is used to conduct risk assessment based on the basic comprehensive score, through real-time event data fusion and Monte Carlo risk simulation, to obtain risk probability distribution and cross-border compliance assessment results;
[0036] The sorting module 106 is used to process federated learning collaborative data and interpretable attention annotations in real time through lightweight preprocessing and streaming computing pipelines to obtain a comprehensive score ranking list for cross-border product selection.
[0037] It is understood that the executing entity of this application can be a cross-border e-commerce product recommendation system based on multi-source data fusion, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.
[0038] Specifically, the data collection module 101 uses a distributed collector to collect multi-source heterogeneous data. First, it deploys crawlers to various cross-border e-commerce platforms based on a pre-defined collection node configuration table. This table includes parameters such as platform domain names, data interface addresses, and crawling frequencies. Each collection node independently collects user browsing history, search keywords, purchase records, and other behavioral data according to its configuration. Simultaneously, it acquires external market data through search trend and social media interfaces. The internal behavioral data and external market data are then merged and processed using data source identification to form the original cross-border dataset. Subsequently, a rule engine and anomaly detection algorithm are used to clean the original data. The rule engine identifies duplicate records using predefined rules, and the anomaly detection algorithm uses the isolated forest method to detect outliers in numerical fields. After cleaning, product classification standards are unified, and currency units are converted. Finally, a star schema is used to construct a product fact table and a user dimension table to form a unified cross-border data warehouse. A differential privacy algorithm adds noise to the standardized dataset to generate a privacy-preserving dataset, and federated learning nodes generate federated learning collaborative data through model gradient sharing. The classification module 102 receives user feedback text data from a cross-border unified data warehouse and performs cross-language word vector conversion using a multilingual BERT variant. This multilingual BERT variant is a pre-trained model supporting multiple languages, capable of mapping words from different languages to a unified vector space. The cross-language semantic alignment algorithm maps word embedding vectors from different languages to the same semantic space through linear transformation, ensuring that semantically similar words have similar vector representations in different languages. The positional encoding algorithm uses sine and cosine functions to generate a unique code for each word's position; sine values are used for even-numbered dimensions, and cosine values for odd-numbered dimensions. These are added element-wise to the word embedding matrix to form the positional encoded text matrix. The multi-head self-attention mechanism calculates the association weights between words using 12 parallel attention heads. Each attention head independently calculates the attention score for the query matrix, key matrix, and value matrix. The interpretability attention layer normalizes the attention weights into a probability distribution using a softmax function and labels keywords that influence the classification results. The parent classifier categorizes text into six types: product quality evaluation, price sensitivity expression, logistics experience feedback, after-sales service evaluation, purchase intention inquiry, and competitor comparison opinions. The sub-classifier further subdivides each parent category. The online learning mechanism updates classifier parameters using a gradient descent algorithm and automatically expands the sub-classifier label system when a new product category is detected. Quantization module 103 performs clustering analysis on user geographic distribution data and purchasing power data based on the K-means clustering algorithm. The K-means algorithm iteratively optimizes and groups similar users into the same cluster, calculating the centroid of each cluster as the user feature vector. User preference weight quantification converts sentiment into numerical weights, assigning high weights to positive reviews and low weights to negative reviews. Product basic features are converted into binary vectors using one-hot encoding, and price and sales data are mapped to the zero-to-one interval using maximum-minimum value normalization.The seasonal fluctuation analysis algorithm uses time series decomposition to separate historical sales into trend, seasonal, and random components, and identifies periodic patterns through Fourier transform. Market trend characteristics are calculated using a moving average algorithm to determine the sliding average of keyword search popularity, and competition density is calculated using the Herfindahl-Hirschman Index to measure market concentration. Vector concatenation and fusion connects the features of each dimension column-wise to form a unified matrix representation.
[0039] The fusion module 104 employs an LSTM neural network as the time series prediction model. LSTM controls information flow through forget gates, input gates, and output gates, inputting historical sales data in a time series manner to predict future sales trends. A collaborative filtering algorithm constructs a user-product rating matrix, using matrix factorization to decompose the sparse matrix into user latent factor matrices and product latent factor matrices, calculating user matching scores using cosine similarity. The reinforcement learning framework models weight adjustment as a Markov decision process, with the state space composed of historical decision effects. The Q-learning algorithm updates and optimizes the weight strategy through a value function. The Pareto optimal solution search algorithm finds the non-dominated solution set, where a non-dominated solution is not strictly dominated by other solutions on any objective. Weight convergence detection determines whether convergence has been achieved by comparing the weight change rate with a preset threshold.
[0040] The assessment module 105 acquires geopolitical news and customs policy changes in real time via API. NLP algorithms employ named entity recognition and sentiment analysis to extract risk event levels. The exchange rate risk index is calculated using the standard deviation and coefficient of variation of historical exchange rate data; the logistics risk index is assessed based on cost fluctuations and timeliness ranges; and the competition risk index is quantified through market competition density and price competition intensity. The Monte Carlo simulation algorithm treats risk variables as random variables and performs extensive random sampling, establishing a risk distribution model through a probability density function. Cross-border compliance assessment detects compliance deficiencies by matching them to the target country's commodity access regulations, and the compliance risk quantification algorithm converts the severity of deficiencies into risk values.
[0041] The ranking module 106 inputs the federated learning collaborative data into a lightweight preprocessing model for feature extraction. This lightweight model is a neural network model with a small parameter size, suitable for operation in resource-constrained environments. Apache Flink is a distributed stream processing framework that supports real-time data stream computation, using window functions to group and aggregate data streams. The streaming computation pipeline organizes feature data into a continuous data stream, updating user profile information in real time. The ranking algorithm sorts the products in descending order based on their matching scores, generating a final ranking list of cross-border product selection based on comprehensive scores.
[0042] In one specific embodiment, the acquisition module 101 is used for:
[0043] Based on the preset data collection node configuration table, a distributed crawler is deployed on the cross-border e-commerce platform to obtain the platform data collection nodes. User browsing history, search keywords, and purchase records are input into the platform data collection nodes for behavioral data capture and processing to obtain raw user behavior data.
[0044] External data acquisition and processing are performed on search trend data and social media mention data based on search trend interface and social media interface to obtain raw external market data. The raw user behavior data and raw external market data are then merged by data source identification to obtain cross-border raw dataset.
[0045] Based on the rule engine and anomaly detection algorithm, the original cross-border dataset is processed to remove duplicate data and filter outout values to obtain a cleaned dataset. The cleaned dataset is then processed to unify the commodity classification standards and convert the currency units to obtain a standardized dataset.
[0046] Based on the star schema data model, the standardized dataset is processed to construct a product fact table and a user dimension table to obtain a cross-border unified data warehouse.
[0047] The standardized dataset is noise-added using a differential privacy algorithm to obtain a privacy-preserving dataset. The privacy-preserving dataset is then federated through federated learning nodes to share model gradients, resulting in federated learning collaborative data.
[0048] Specifically, the data collection module 101 implements distributed crawler deployment through a preset data collection node configuration table. This table includes configuration parameters such as the target platform's domain name, API endpoint, access frequency, and request headers. The distributed crawler deployment process creates independent data collection programs on multiple servers based on these parameters. Each program is responsible for data collection tasks on a specific platform, forming platform data collection nodes. Behavioral data collection takes user browsing history, search keywords, and purchase records as input. It extracts structured user interaction information by parsing HTML page structures, calling platform APIs, and analyzing user session logs, outputting raw user behavior data. External data acquisition processing obtains keyword search volume changes over different time periods through a search trend interface. The search trend interface returns time-series data in JSON format, including fields such as keywords, search volume, and timestamps. The social media interface provides social interaction data such as user product discussions, reposts, and likes on social platforms, also returning structured information in JSON format. The data source identification merging process adds source tags and timestamps to each data record, marking the original user behavior data as "internal" source and the original external market data as "external" source. Then, it performs association matching by product ID and time dimension, and combines related data into a unified record to form a cross-border original dataset.
[0049] The deduplication process identifies duplicate entries by calculating the similarity between data records. Similarity calculation is based on the textual similarity and numerical differences of key fields such as product titles, prices, and merchant information. When the similarity exceeds a threshold, the data is considered duplicate and deleted. The anomaly detection algorithm employs the Isolation Forest method. This algorithm constructs multiple random decision trees to segment data points. Normal data points cluster in the shallow layers of the trees, requiring fewer segmentations, while abnormal data points are distributed in the deeper layers, requiring more segmentations. Anomaly scores are calculated based on the average path length, and data points exceeding a threshold are identified as outliers and filtered out. The product classification standardization process establishes a cross-platform category mapping table, mapping classification tags from different platforms to a unified standard classification system. Currency conversion processing converts various currencies to US dollars based on real-time exchange rates.
[0050] The star schema data model employs a centralized data organization structure. The product fact table serves as the core, storing metrics such as product ID, name, price, and sales volume. The user dimension table stores attribute information such as user ID, geographic location, and purchasing power level. Foreign key relationships are established between the fact table and the dimension table, forming a unified cross-border data warehouse. Differential privacy algorithms protect individual privacy by adding carefully designed random noise to the original data. The noise intensity is controlled by a privacy budget parameter; a smaller privacy budget results in greater noise and a higher degree of privacy protection, generating a privacy-preserving dataset. Federated learning nodes are the computational units in a distributed machine learning system. Model gradient sharing refers to each node training its model locally and then uploading its gradient parameters to a central server for aggregation. The server calculates the global model parameters and distributes them to each node, achieving collaborative model optimization through multiple iterations.
[0051] In one specific embodiment, the classification module 102 is used for:
[0052] User feedback text data from the cross-border unified data warehouse is input into a multilingual BERT variant for cross-lingual word vector transformation to obtain a multilingual word embedding matrix. Based on a cross-lingual semantic alignment algorithm, the multilingual word embedding matrix is subjected to semantic space mapping to obtain an aligned word embedding matrix.
[0053] The positional information of the aligned word embedding matrix is fused according to the positional encoding algorithm to obtain the positional encoded text matrix. The positional encoded text matrix is then input into the multi-head self-attention mechanism for semantic association calculation to obtain the semantic feature vector.
[0054] The semantic feature vector is processed by calculating the keyword weights based on the interpretable attention layer to obtain the attention weight distribution. The keywords that affect the classification results are labeled according to the attention weight distribution to obtain interpretable attention labels.
[0055] The semantic feature vector is input into the parent classifier for six categories of parent category classification: product quality evaluation, price sensitivity expression, logistics experience feedback, after-sales service evaluation, purchase intention inquiry, and competitor comparison opinion. The parent category classification results are obtained. Based on the parent category classification results, the corresponding sub-class classifier is activated to perform sub-category recognition processing to obtain the sub-category classification results.
[0056] Based on the online learning mechanism, the new product category is expanded by sub-classifier to obtain a dynamically updated sub-classifier. The parent category classification results and the sub-category classification results are then fused by the dynamically updated sub-classifier to obtain the user preference analysis results.
[0057] Specifically, the classification module 102 first inputs user feedback text data from the cross-border unified data warehouse into a multilingual BERT variant for processing. The multilingual BERT variant is a pre-trained neural network model that supports multiple languages and can convert text in different languages into a unified numerical vector representation. The cross-lingual word vector transformation process maps each word to a 768-dimensional dense vector through the model's embedding layer, forming a multilingual word embedding matrix. The number of rows in the matrix corresponds to the number of words in the text, and the number of columns is fixed at 768. The cross-lingual semantic alignment algorithm maps word embedding vectors from different languages to the same semantic space through a linear transformation matrix, ensuring that semantically similar words have similar vector representations in different languages. The semantic space mapping process calculates the product of the transformation matrix and the word embedding matrix to obtain the aligned word embedding matrix.
[0058] The positional encoding algorithm generates a unique position vector for each word in the text. The algorithm uses sine and cosine functions to calculate the positional encoding value; the sine function is used for even dimensions, and the cosine function for odd dimensions. Positional information fusion processing adds the positional encoding vector element-wise to the word vector at the corresponding position in the aligned word embedding matrix to obtain a positional encoded text matrix containing positional information. The multi-head self-attention mechanism includes 12 parallel attention heads. Each attention head independently calculates the query matrix, key matrix, and value matrix. Semantic association calculation first calculates the matrix product of the query matrix and the transpose of the key matrix, then normalizes it by dividing by a scaling factor. The attention scores are converted into a probability distribution using the softmax function. Finally, the attention weights are multiplied by the value matrix to obtain a weighted feature representation. The outputs of multiple attention heads are concatenated to form a semantic feature vector.
[0059] The interpretable attention layer further processes the semantic feature vectors. Keyword weight calculation uses a fully connected neural network to calculate the contribution of each word to the final classification result, outputting an attention weight distribution. This distribution represents the importance of each word in the text to the classification decision. The annotation process identifies the words with the highest weight values as keywords based on the attention weight distribution. These keywords and their weight values are combined to form interpretable attention annotations, visually displaying the important words affecting the classification result.
[0060] The parent classifier employs a fully connected neural network structure, taking semantic feature vectors as input. After processing through linear transformations and activation functions, it outputs probability distributions for six categories. The parent category classification determines the primary category to which the text belongs based on the principle of maximizing probability. Subclass classifiers are activated based on the parent category classification results, with each parent category corresponding to a dedicated subclass classifier. The sub-category recognition process inputs the same semantic feature vectors into the activated subclass classifiers, resulting in a finer classification granularity. For example, the parent category of product quality evaluation can be subdivided into subcategories such as functional evaluation, appearance evaluation, and durability evaluation.
[0061] The online learning mechanism continuously updates classifier parameters through an incremental learning algorithm. When the system detects a new product category, the sub-classifier expansion process automatically adds new category labels and adjusts the network structure. The classifier weights are retrained using the new sample data through a gradient descent algorithm, resulting in a dynamically updated sub-classifier. The hierarchical label fusion process uses the parent category classification result as the root node and the sub-category classification result as the leaf node to construct a tree-like label structure. Each label contains a corresponding confidence score, ultimately outputting a user preference analysis result containing complete hierarchical information.
[0062] In one specific embodiment, the quantization module 103 is used for:
[0063] Based on user geographic distribution data and purchasing power data in the cross-border unified data warehouse, user feature clustering analysis is performed to obtain user feature vectors. Based on the user preference analysis results, the user's preference weight for product attributes is quantified and calculated to obtain user preference weight vectors.
[0064] The product category, price, and sales data in the cross-border unified data warehouse are standardized and encoded to obtain the basic feature vector of the product. Based on the seasonal fluctuation analysis algorithm, the historical sales data is periodically feature extracted to obtain the product time-series feature vector.
[0065] Based on exchange rate fluctuation data and consumption trend data, time series analysis is performed on the changes in the heat of the target market to obtain a market trend feature vector. Based on the competition density calculation algorithm, the market share data of similar products is quantified to obtain a competition feature vector.
[0066] The user feature vector and the user preference weight vector are concatenated and fused to obtain the user comprehensive feature vector. The product basic feature vector and the product time series feature vector are concatenated and fused to obtain the product comprehensive feature vector.
[0067] The user comprehensive feature vector, product comprehensive feature vector, market trend feature vector, and competition feature vector are matrix-concatenated and fused to obtain the cross-border product selection feature matrix.
[0068] Specifically, the quantification module 103 performs user feature clustering analysis on user geographic distribution data and purchasing power data based on the K-means clustering algorithm. The K-means algorithm iteratively calculates and groups similar users into the same cluster. The algorithm first randomly initializes K cluster centers, then calculates the Euclidean distance from each user to each cluster center, assigns the user to the nearest cluster center, and then recalculates the position of the center point of each cluster, repeating the iteration until the cluster centers no longer change significantly. User geographic distribution data includes geographical attributes such as the user's continent, economic development level, and consumption culture background, while purchasing power data includes economic attributes such as average order amount, purchase frequency, and consumer product category preference. Cluster analysis groups users with similar geographic and purchasing power characteristics into the same group, and the center point vector of each cluster serves as the representative user feature vector of that group. The user preference weight quantification process converts the sentiment tendency in the user preference analysis results into numerical weights. Positive sentiment evaluations are assigned high weight values of 0.8 to 1.0, neutral sentiment evaluations are assigned medium weight values of 0.4 to 0.6, and negative sentiment evaluations are assigned low weight values of 0.1 to 0.3. The weight calculation is based on the product operation of sentiment intensity and mention frequency to obtain the user preference weight vector.
[0069] The standardized coding process for basic product features uses one-hot encoding to convert the text tags of product categories into binary vectors. Each category corresponds to one dimension in the vector, with the value of that dimension being 1 and the values of the other dimensions being 0. Price data is mapped to a numerical range of 0 to 1 through maximum-minimum value normalization. The normalization formula is the current price minus the minimum price, divided by the price range. Sales data is processed using logarithmic transformation to address the long-tailed distribution problem. The logarithmic transformation takes the natural logarithm of the original sales values to compress the numerical range. The seasonal fluctuation analysis algorithm uses time series decomposition to separate historical sales data into trend components, seasonal components, and random components. The periodic feature extraction process uses Fourier transform to identify periodic patterns in sales data. The Fourier transform converts the time-domain signal into a frequency-domain signal, extracting the amplitude and phase information of different frequency components to identify time patterns such as monthly, quarterly, and annual cycles. The intensity coefficient and phase offset of each periodic component form the product time-series feature vector.
[0070] Time series analysis arranges exchange rate fluctuation data and consumption trend data chronologically. A moving average algorithm is used to calculate the average value for different time windows to smooth short-term fluctuations. The moving average algorithm takes the arithmetic mean of values from N consecutive time points as the representative value for that period. The standard deviation of the exchange rate fluctuation data measures exchange rate stability; a larger standard deviation indicates more volatile fluctuations. The first difference of the consumption trend data reflects the rate of market change; the first difference is the difference between values at adjacent time points, with positive values indicating an upward trend and negative values indicating a downward trend. The competition density calculation algorithm counts the number of sellers of similar products and their respective market shares. The Herfindahl-Hirschman Index (HHI) is used to calculate market concentration; the HHI is the sum of the squares of each seller's market share. An index value close to 0 indicates intense market competition, while an index value close to 1 indicates a highly concentrated market. The competition intensity quantification is based on the reciprocal of market concentration to calculate the degree of competition, forming a competition feature vector.
[0071] Vector concatenation and fusion processing connects the user feature vector and the user preference weight vector column-wise to form a comprehensive user feature vector. The concatenation operation then arranges all elements of the two vectors sequentially to form a new, longer vector. Similarly, the product basic feature vector and the product time-series feature vector are concatenated column-wise to form a comprehensive product feature vector, ensuring the complete preservation of both static attributes and dynamic time-series information. Matrix concatenation and fusion processing connects the comprehensive user feature vector, comprehensive product feature vector, market trend feature vector, and competition feature vector column-wise to form a unified feature matrix. The number of rows in the matrix equals the number of candidate products, and the number of columns equals the sum of the dimensions of all feature vectors. Each row represents a complete feature description of a product.
[0072] In one specific embodiment, the fusion module 104 is used for:
[0073] Based on the time series prediction model, the sales-related features in the cross-border product selection feature matrix are predicted and calculated to obtain the sales prediction score. The user preference features and product features are matched and processed according to the collaborative filtering algorithm to obtain the user matching score. The market trend features are input into the trend analysis model to calculate the trend fit and obtain the market trend score.
[0074] Based on the reinforcement learning framework, the weight adjustment is modeled as a Markov decision process. The state space is constructed by analyzing the historical product selection decision effects to obtain the decision state matrix. Based on the decision state matrix, the weight update strategy is optimized by the Q-learning algorithm to obtain the dynamic weight coefficient set.
[0075] The sales forecast score, user matching score and market trend score are used as the objective function input. Multi-objective optimization is performed based on the Pareto optimal solution search algorithm to obtain the non-dominated solution set. The non-dominated solution set is then weighted and summed according to the dynamic weight coefficient group to obtain the basic comprehensive score.
[0076] Based on the weight fluctuation threshold, the convergence detection process of the dynamic weight coefficient group is performed to obtain the weight change rate. The weight change rate is compared with the preset convergence threshold to obtain the weight convergence detection result.
[0077] The sales prediction score, user matching score, and market trend score are finally weighted and fused using a dynamic weight adaptive algorithm to obtain the basic comprehensive score and weight convergence detection results.
[0078] Specifically, the fusion module 104 processes sales-related features in the cross-border product selection feature matrix based on the LSTM time series prediction model. The LSTM neural network controls information flow through three gating mechanisms: the forget gate, the input gate, and the output gate. The forget gate determines which information is discarded from the cell state, the input gate controls the storage degree of new information, and the output gate determines which parts of the cell state are output. The prediction calculation process inputs historical sales data into the model as a time series. A sliding window mechanism is used to predict the sales trend for the next 7 days using the sales data of the past 30 days as the input sequence. The LSTM learns the time dependency relationship through the backpropagation algorithm. The sales prediction score is calculated by the ratio of the predicted sales to the historical average sales. The collaborative filtering algorithm constructs a user-product rating matrix, where rows represent users and columns represent products, and element values represent the implicit ratings of users for products. The matching degree calculation process uses matrix factorization to decompose the sparse rating matrix into a user latent factor matrix and a product latent factor matrix. Matrix factorization minimizes the reconstruction error through the gradient descent algorithm. The user matching score is obtained by calculating the cosine similarity between the user latent factors and the product latent factors. The trend analysis model uses linear regression to fit the time-varying patterns of market trend characteristics. The trend fit calculation process inputs features such as changes in market popularity and shifts in consumer preferences into the model, and estimates the slope and intercept parameters of the trend line using the least squares method. The market trend score is calculated based on the positive or negative value and significance level of the trend slope.
[0079] The reinforcement learning framework models weight adjustment as a Markov decision process. The state space consists of factors such as the current weight configuration, historical decision effects, and changes in the market environment. The state space construction process encodes these factors into state vectors. The historical product selection decision effects are measured by indicators such as the error between actual and predicted sales, user satisfaction, and market response, forming a decision state matrix. The Q-learning algorithm is a model-free reinforcement learning method that optimizes decision-making strategies by learning a state-action value function. The weight update strategy optimization process uses the Bellman equation to update the Q-value, which represents the long-term expected return of taking a specific action in a specific state. The algorithm uses an exploration-exploitation balancing mechanism to learn the optimal strategy while continuing to explore new strategies, resulting in a dynamic set of weight coefficients.
[0080] The Pareto optimality search algorithm seeks solutions that are not strictly dominated by other solutions across multiple objective functions. A non-dominated solution is one where no other solution is superior or equal to it across all objectives and is strictly superior to it in at least one objective. Multi-objective optimization processes use genetic algorithms or particle swarm optimization to search for the Pareto front. The algorithm maintains a solution set and generates new solutions through operations such as crossover and mutation. Non-dominated sorting stratifies the solutions according to their dominance relationships, resulting in a non-dominated solution set. Weighted summation processes multiply the weights in the dynamic weight coefficient group by the sales prediction score, user matching score, and market trend score respectively, then sum the results. The total weight sum is normalized to 1 to ensure the rationality of the score, yielding a basic comprehensive score.
[0081] The convergence detection process calculates the change range of the dynamic weight coefficient group within a continuous time step. The weight change rate is calculated by dividing the difference between the current weight and the weight at the previous time step by the weight at the previous time step. The absolute value of the change rate reflects the drasticness of the weight adjustment. The comparison and judgment process compares the weight change rate with a preset convergence threshold. When the change rate is less than the threshold, the weight is determined to have stabilized. The weight stability indicator records the convergence state, forming the weight convergence detection result.
[0082] The dynamic weight adaptive algorithm comprehensively considers factors such as the current market environment, historical decision-making effects, and weight convergence status. The final weighted fusion processing determines whether to continue adjusting the weights based on the weight convergence detection results. When the weights converge, stable weights are used for score calculation. When the weights do not converge, Q-learning is used to optimize the weights. The three scores are then weighted and summed according to the final determined weights, and the basic comprehensive score and weight convergence detection results are output simultaneously.
[0083] In one specific embodiment, the evaluation module 105 includes:
[0084] Real-time data acquisition and processing of geopolitical news and customs policy change events are performed through API interfaces to obtain raw real-time event data. The raw real-time event data is then correlated and matched with the basic comprehensive score to obtain the event impact score matrix.
[0085] The risk event level is extracted from the raw real-time event data based on the NLP algorithm to obtain the risk event classification identifier. The event impact scoring matrix is then fused based on the risk event classification identifier to obtain the fused event data.
[0086] Exchange rate risk, logistics risk, and competition risk are input as random variables into the Monte Carlo simulation algorithm for random sampling to obtain a risk variable sample set. Based on the risk variable sample set, the probability density function is used to calculate the distribution and obtain the risk probability distribution.
[0087] The integrated event data is processed for compliance testing based on the target country's commodity access regulations to obtain compliance defect identifiers. Based on these compliance defect identifiers, an assessment is conducted using a compliance risk quantification algorithm to obtain the cross-border compliance assessment results.
[0088] The merged event data is then processed through Monte Carlo risk simulation for final risk assessment, yielding risk probability distribution and cross-border compliance assessment results.
[0089] Specifically, the assessment module 105 acquires geopolitical news and customs policy change events in real time via API interfaces. API interfaces are application programming interfaces that allow the system to access external data sources in a standardized manner. Real-time data acquisition and processing calls the news aggregation API and government policy release API via HTTP requests to obtain structured information such as news titles, release times, event types, and impact scope in JSON format, forming raw real-time event data. The correlation and matching processing matches real-time events with basic comprehensive scores according to dimensions such as product category, target market, and time window. The matching algorithm identifies the relevance between events and products through keyword extraction and semantic similarity calculation, quantifies the impact of related events into numerical values, and combines them with the basic scores of the corresponding products to form an event impact score matrix. In this matrix, rows represent products, columns represent event types, and element values represent the intensity of the impact of a specific event on a specific product.
[0090] The NLP algorithm uses named entity recognition (NER) and sentiment analysis to process raw real-time event data. NER uses a pre-trained BERT model to identify key entities in the text, such as country names, policy types, and product categories. Sentiment analysis uses a sentiment dictionary and a deep learning model to determine whether an event has a positive or negative impact on the market. Risk event level extraction categorizes events into low-risk, medium-risk, and high-risk levels based on factors such as the scope, duration, and severity of the impact. The grading criteria are based on historical event data and an expert knowledge base, resulting in risk event level labels. Data fusion processing weights the values in the event impact scoring matrix according to the risk level. The impact coefficient is set to 1.5 for high-risk events, 1.0 for medium-risk events, and 0.5 for low-risk events. The adjusted scoring matrix is then linearly combined with the basic product scores to obtain the fused event data.
[0091] Monte Carlo simulation is a numerical computation method based on random sampling. It treats exchange rate risk, logistics risk, and competition risk as random variables following specific probability distributions. Random sampling is used to fit probability distribution functions to historical data for each risk variable. Exchange rate risk follows a normal distribution, logistics risk follows an exponential distribution, and competition risk follows a uniform distribution. A pseudo-random number generator produces a large number of random samples conforming to each distribution, resulting in a risk variable sample set. The probability density function is calculated based on kernel density estimation, using a Gaussian kernel function to smooth the sample points. Distribution modeling represents the joint distribution of multiple risk variables as a multidimensional probability density function. The function value represents the probability density of a specific risk combination. Numerical integration is used to calculate the cumulative probability of different risk levels, forming a risk probability distribution.
[0092] The compliance inspection process checks the integrated event data according to the target country's commodity access regulations, including clauses on safety standards, environmental requirements, and certification procedures. The inspection algorithm identifies the degree of compliance between commodity attributes and regulatory requirements through rule matching and text analysis. Compliance defect identification records the commodity's non-compliance with various regulatory requirements, including defect type, severity, and difficulty of remediation. The compliance risk quantification algorithm converts defect information into numerical risk scores, assigning high-risk weights to severe defects and low-risk weights to minor defects. The algorithm considers the cost and time requirements for defect remediation and calculates a comprehensive compliance risk value through weighted summation to obtain the cross-border compliance assessment result.
[0093] The final risk assessment process uses the fused event data as input parameters for the Monte Carlo simulation. The simulation algorithm considers the impact of real-time events on various risks, adjusts the distribution parameters of risk variables to reflect the event impact, re-sampling and probability calculation, generates an updated risk probability distribution, and outputs cross-border compliance assessment results that include the impact of the event.
[0094] In one specific embodiment, the sorting module 106 is used for:
[0095] Federated learning collaborative data is input into a lightweight preprocessing model for feature extraction to obtain collaborative feature vectors. Key product attributes are weighted based on interpretable attention annotations to obtain attention-weighted features.
[0096] Feature fusion processing is performed based on collaborative feature vectors and attention-weighted features to obtain a fused feature matrix. The fused feature matrix is then input into a streaming computing pipeline for real-time data stream processing to obtain a real-time feature stream.
[0097] Based on the Apache Flink framework, the real-time feature stream is processed by streaming computation to obtain dynamically updated user profiles. The product matching degree is then recalculated based on the dynamically updated user profiles to obtain real-time matching scores.
[0098] The real-time matching scores are compared and analyzed with historical score data to obtain score change trends. Based on the score change trends, the products are prioritized using a sorting algorithm to obtain preliminary sorting results.
[0099] Based on the preliminary sorting results, a final sorting optimization process is performed to obtain a comprehensive score-ranked list of cross-border product selections.
[0100] Specifically, the sorting module 106 inputs the federated learning collaborative data into a lightweight preprocessing model for processing. This lightweight preprocessing model is a neural network model with a small parameter size, employing depthwise separable convolution and parameter quantization techniques to reduce computational complexity, making it suitable for operation in resource-constrained environments. Feature extraction processing maps high-dimensional data such as user behavior patterns, product attribute information, and market trend changes in the federated learning collaborative data into low-dimensional dense vector representations through the model's encoder. The encoder uses a multilayer perceptron structure, with each layer extracting features at different abstraction levels through linear transformations and activation functions, ultimately outputting a fixed-dimensional collaborative feature vector. Weight allocation processing uses the keyword weights marked in interpretable attention annotation as an importance indicator for product attributes. Different weight coefficients are assigned to attributes such as price, quality, brand, and function. Weight normalization ensures that the sum of all attribute weights is 1, and attention-weighted features are obtained through weighted averaging.
[0101] Feature fusion processing performs element-level fusion operations on collaborative feature vectors and attention-weighted features. The fusion method employs a weighted combination: collaborative feature vectors reflect the overall preference patterns of a user group, while attention-weighted features highlight key attributes influencing user decisions. The two are linearly combined according to preset weight ratios to form a fused feature vector. The fused feature vectors of multiple products are arranged in rows to form a fused feature matrix. The streaming computing pipeline is a computational framework for processing continuous data streams. Real-time data stream processing converts the fused feature matrix into a time-sequential data stream. Each element in the data stream contains a timestamp and feature vector information. Streaming processing supports incremental updates and sliding window computation to obtain a real-time feature stream.
[0102] Apache Flink is a distributed stream processing framework that supports low-latency real-time data processing and complex event processing. Streaming computation uses event-time semantics and watermarking mechanisms to handle out-of-order data. Dynamically updating user profiles involves aggregating recent user behavior data through a sliding window, with the window size set to the past 24 hours. Aggregation operations include statistical functions such as summation, averaging, and maximum values, calculating user profile features such as preference strength for different product categories, changes in purchase frequency, and price sensitivity. Recalculation processing then performs similarity calculations based on the updated user profile and product features. The similarity algorithm uses cosine similarity to measure the degree of matching between the user preference vector and the product feature vector; a higher similarity value indicates that the product better matches the user's preferences, resulting in a real-time matching score.
[0103] The comparative analysis process compares the current real-time matching score with the score data within the historical time window. Historical score data is stored in a time-series database. Comparison methods include calculating indicators such as score difference, rate of change, and trend slope. The score change trend is identified through time-series analysis, identifying upward, downward, and stable trend patterns. The trend detection algorithm uses moving averages and linear regression to fit the change pattern of scores over time. The positive or negative value and significance level of the trend slope reflect the dynamic changes in product popularity. The ranking algorithm prioritizes products based on score change trends. Priority calculation comprehensively considers the current score and the magnitude of trend change; products with upward trends receive a priority boost, while products with downward trends have a lower priority. The ranking uses a quicksort algorithm to sort products from highest to lowest priority score, yielding a preliminary ranking result.
[0104] The final ranking optimization process fine-tunes the initial ranking results. The optimization algorithm considers additional factors such as product inventory status, supply chain stability, and diverse user needs. It balances recommendation accuracy and diversity through multi-objective optimization. The optimization process uses a local search algorithm to make minor adjustments to the ranking results to ensure that the recommendation list satisfies user preferences while maintaining a reasonable product distribution. The final output is a ranking list of cross-border product selection based on comprehensive scores.
[0105] The above describes the cross-border e-commerce product recommendation system based on multi-source data fusion in the embodiments of this application. The following describes the cross-border e-commerce product recommendation method based on multi-source data fusion in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the cross-border e-commerce product recommendation method based on multi-source data fusion in this application includes:
[0106] S201. Through distributed collectors, multi-source heterogeneous data collection and processing are performed on cross-border e-commerce platform data, user behavior data, and external market data to obtain cross-border raw datasets. The cross-border raw datasets are then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data.
[0107] S202. Based on user feedback text data in the cross-border unified data warehouse, hierarchical text classification is performed by improving the BERT model to obtain user preference parsing results and interpretable attention annotations.
[0108] S203. Based on the cross-border unified data warehouse and user preference analysis results, the cross-border product selection features are quantified using a multi-dimensional feature extraction algorithm to obtain a cross-border product selection feature matrix.
[0109] S204. Based on the cross-border product selection feature matrix, the product selection score is weighted and fused using a dynamic weight adaptive algorithm to obtain the basic comprehensive score and weight convergence detection results.
[0110] S205. Based on the basic comprehensive score, risk assessment is carried out through real-time event data fusion and Monte Carlo risk simulation to obtain the risk probability distribution and cross-border compliance assessment results.
[0111] S206. Based on federated learning collaborative data and interpretable attention annotations, a comprehensive score ranking list for cross-border product selection is obtained through real-time processing using lightweight preprocessing and streaming computing pipelines.
[0112] above Figure 2 The cross-border e-commerce product recommendation method based on multi-source data fusion in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The cross-border e-commerce product recommendation device based on multi-source data fusion in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0113] Reference Figure 3 This invention also provides a method where the cross-border e-commerce product recommendation device based on multi-source data fusion can be a server, and its internal structure can be as follows: Figure 3As shown, the cross-border e-commerce product recommendation device based on multi-source data fusion includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor in this computer design provides computing and control capabilities. The memory of the cross-border e-commerce product recommendation device based on multi-source data fusion includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the cross-border e-commerce product recommendation device based on multi-source data fusion is used to store the data corresponding to this embodiment. The network interface of the cross-border e-commerce product recommendation device based on multi-source data fusion is used for communication with external terminals via network connection. When the computer program is executed by the processor, it implements the above-described method.
[0114] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the cross-border e-commerce product recommendation device based on multi-source data fusion applied thereto.
[0115] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the cross-border e-commerce product recommendation system based on multi-source data fusion.
[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a cross-border e-commerce product recommendation device (which may be a personal computer, server, or network device, etc.) based on multi-source data fusion to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-border e-commerce product recommendation system based on multi-source data fusion, characterized in that, The system includes: The data collection module is used to collect and process multi-source heterogeneous data from cross-border e-commerce platforms, user behavior data, and external market data through a distributed collector to obtain a cross-border raw dataset. The cross-border raw dataset is then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data. The classification module is used to perform hierarchical text classification processing based on user feedback text data in the cross-border unified data warehouse by improving the BERT model, so as to obtain user preference parsing results and interpretable attention annotations; The quantification module is used to quantify the cross-border product selection features based on the cross-border unified data warehouse and the user preference analysis results, using a multi-dimensional feature extraction algorithm to obtain a cross-border product selection feature matrix. The fusion module is used to perform weighted fusion processing on the product selection score based on the cross-border product selection feature matrix through a dynamic weight adaptive algorithm to obtain a basic comprehensive score and a weight convergence detection result. The assessment module is used to perform risk assessment processing based on the basic comprehensive score, through real-time event data fusion and Monte Carlo risk simulation, to obtain the risk probability distribution and cross-border compliance assessment results; The sorting module is used to process the federated learning collaborative data and the interpretable attention annotations in real time through lightweight preprocessing and streaming computing pipelines to obtain a comprehensive score ranking list for cross-border product selection.
2. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The acquisition module is used for: Based on the preset data collection node configuration table, a distributed crawler is deployed on the cross-border e-commerce platform to obtain the platform data collection nodes. User browsing history, search keywords, and purchase records are input into the platform data collection nodes for behavioral data capture and processing to obtain raw user behavior data. External data acquisition and processing are performed on search trend data and social media mention data based on search trend interface and social media interface to obtain external market raw data. The user behavior raw data and the external market raw data are then merged by data source identification to obtain the cross-border raw dataset. Based on the rule engine and anomaly detection algorithm, the original cross-border dataset is processed to remove duplicate data and filter outout values to obtain a cleaned dataset. The cleaned dataset is then processed to unify the commodity classification standards and convert the currency units to obtain a standardized dataset. Based on the star schema data model, the standardized dataset is processed to construct a product fact table and a user dimension table to obtain the cross-border unified data warehouse; The standardized dataset is noise-added using a differential privacy algorithm to obtain a privacy-preserving dataset. The privacy-preserving dataset is then processed through federated learning nodes to share model gradients, resulting in federated learning collaborative data.
3. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The classification module is used for: User feedback text data from the cross-border unified data warehouse is input into a multilingual BERT variant for cross-lingual word vector conversion to obtain a multilingual word embedding matrix. The multilingual word embedding matrix is then subjected to semantic space mapping based on a cross-lingual semantic alignment algorithm to obtain an aligned word embedding matrix. The alignment word embedding matrix is fused with positional information according to the positional encoding algorithm to obtain a positional encoded text matrix. The positional encoded text matrix is then input into a multi-head self-attention mechanism for semantic association calculation to obtain a semantic feature vector. The semantic feature vector is processed by keyword weight calculation based on the interpretable attention layer to obtain the attention weight distribution. The keywords that affect the classification result are labeled according to the attention weight distribution to obtain the interpretable attention label. The semantic feature vector is input into the parent classifier for six categories of parent category classification processing: product quality evaluation, price sensitivity expression, logistics experience feedback, after-sales service evaluation, purchase intention inquiry, and competitor comparison opinion. The parent category classification results are obtained. Based on the parent category classification results, the corresponding sub-classifier is activated to perform sub-category recognition processing to obtain the sub-category classification results. Based on the online learning mechanism, the new product category is expanded by a sub-classifier to obtain a dynamically updated sub-classifier. The parent category classification result and the sub-category classification result are then fused by the dynamically updated sub-classifier to obtain the user preference parsing result.
4. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The quantization module is used for: Based on the user geographic distribution data and purchasing power data in the cross-border unified data warehouse, user feature clustering analysis is performed to obtain user feature vectors. Based on the user preference parsing results, the user's preference weight for product attributes is quantitatively calculated to obtain user preference weight vectors. The commodity category, price, and sales data in the cross-border unified data warehouse are standardized and encoded to obtain the commodity basic feature vector. Based on the seasonal fluctuation analysis algorithm, the historical sales data are periodically feature extracted to obtain the commodity time-series feature vector. Based on exchange rate fluctuation data and consumption trend data, time series analysis is performed on the changes in the heat of the target market to obtain a market trend feature vector. Based on the competition density calculation algorithm, the market share data of similar products is quantified to obtain a competition feature vector. The user feature vector and the user preference weight vector are concatenated and fused to obtain the user comprehensive feature vector. The product basic feature vector and the product time series feature vector are concatenated and fused to obtain the product comprehensive feature vector. The user comprehensive feature vector, the product comprehensive feature vector, the market trend feature vector, and the competition feature vector are matrix-concatenated and fused to obtain the cross-border product selection feature matrix.
5. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The fusion module is used for: Based on the time series prediction model, the sales-related features in the cross-border product selection feature matrix are predicted and calculated to obtain the sales prediction score. The user preference features and product features are matched and processed according to the collaborative filtering algorithm to obtain the user matching score. The market trend features are input into the trend analysis model to calculate the trend fit and obtain the market trend score. Based on the reinforcement learning framework, the weight adjustment is modeled as a Markov decision process. The state space is constructed by analyzing the historical product selection decision effects to obtain a decision state matrix. Based on the decision state matrix, the weight update strategy is optimized by the Q-learning algorithm to obtain a dynamic weight coefficient set. The sales forecast score, the user matching score, and the market trend score are used as objective function inputs. Multi-objective optimization is performed based on the Pareto optimal solution search algorithm to obtain a non-dominated solution set. The non-dominated solution set is then weighted and summed according to the dynamic weight coefficient group to obtain the basic comprehensive score. The convergence detection process of the dynamic weight coefficient group is performed based on the weight fluctuation threshold to obtain the weight change rate. The weight change rate is compared with the preset convergence threshold to obtain the weight convergence detection result. The sales prediction score, user matching score, and market trend score are finally weighted and fused using a dynamic weight adaptive algorithm to obtain the basic comprehensive score and the weight convergence detection result.
6. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The evaluation module is used for: Real-time data acquisition and processing of geopolitical news and customs policy change events are performed through API interfaces to obtain raw real-time event data. The raw real-time event data is then correlated and matched with the basic comprehensive score to obtain an event impact score matrix. The real-time event raw data is processed by extracting risk event levels based on NLP algorithms to obtain risk event classification identifiers. The event impact scoring matrix is then fused based on the risk event classification identifiers to obtain fused event data. Exchange rate risk, logistics risk, and competition risk are input as random variables into the Monte Carlo simulation algorithm for random sampling to obtain a risk variable sample set. Based on the risk variable sample set, distribution modeling is performed by calculating the probability density function to obtain the risk probability distribution. The fused event data is processed for compliance detection according to the target country's commodity access regulations to obtain compliance defect identifiers. Based on the compliance defect identifiers, an evaluation is performed using a compliance risk quantification algorithm to obtain the cross-border compliance evaluation result. The fused event data is then processed through Monte Carlo risk simulation for final risk assessment to obtain the risk probability distribution and the cross-border compliance assessment results.
7. The cross-border e-commerce product recommendation system based on multi-source data fusion according to claim 1, characterized in that, The sorting module is used for: The federated learning collaborative data is input into a lightweight preprocessing model for feature extraction to obtain a collaborative feature vector. Based on the interpretable attention annotation, the key product attributes are weighted to obtain attention-weighted features. The collaborative feature vector and the attention-weighted feature are used to perform feature fusion processing to obtain a fused feature matrix. The fused feature matrix is then input into a streaming computing pipeline for real-time data stream processing to obtain a real-time feature stream. The real-time feature stream is processed using the Apache Flink framework to obtain a dynamically updated user profile. The product matching degree is then recalculated based on the dynamically updated user profile to obtain a real-time matching score. The real-time matching score is compared and analyzed with historical score data to obtain the score change trend. Based on the score change trend, the product priority is sorted using a sorting algorithm to obtain a preliminary sorting result. Based on the preliminary sorting results, a final sorting optimization process is performed to obtain the comprehensive score sorting list for cross-border product selection.
8. A method for recommending cross-border e-commerce products based on multi-source data fusion, characterized in that, The cross-border e-commerce product recommendation system based on multi-source data fusion as described in any one of claims 1-7 is implemented, wherein the cross-border e-commerce product recommendation method based on multi-source data fusion includes: The distributed collector performs multi-source heterogeneous data collection and processing on cross-border e-commerce platform data, user behavior data, and external market data to obtain cross-border raw datasets. The cross-border raw datasets are then cleaned and standardized to obtain a cross-border unified data warehouse and federated learning collaborative data. Based on the user feedback text data in the cross-border unified data warehouse, hierarchical text classification is performed by improving the BERT model to obtain user preference parsing results and interpretable attention annotations; Based on the cross-border unified data warehouse and the user preference analysis results, the cross-border product selection features are quantified using a multi-dimensional feature extraction algorithm to obtain a cross-border product selection feature matrix. Based on the cross-border product selection feature matrix, the product selection scores are weighted and fused using a dynamic weight adaptive algorithm to obtain the basic comprehensive score and weight convergence detection results. Based on the aforementioned comprehensive score, risk assessment is conducted through real-time event data fusion and Monte Carlo risk simulation to obtain the risk probability distribution and cross-border compliance assessment results. Based on the federated learning collaboration data and the interpretable attention annotations, a comprehensive score ranking list for cross-border product selection is obtained through real-time processing using lightweight preprocessing and streaming computing pipelines.
9. A cross-border e-commerce product recommendation device based on multi-source data fusion, characterized in that, It includes a memory and a processor, the memory storing a computer program that can run on the processor, and the processor executing the computer program to implement the cross-border e-commerce product recommendation method based on multi-source data fusion as described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the cross-border e-commerce product recommendation method based on multi-source data fusion as described in claim 8.
Citation Information
Patent Citations
Multi-dimensional cross-border commodity matching method and system
CN118503807A
Cross-border e-commerce market demand prediction system based on big data
CN119863271A