High-efficiency multi-source heterogeneous data processing algorithm
By acquiring and processing data from multiple platforms, we establish user portraits and market prediction models, solving the problem of data isolation in the e-commerce industry and achieving accurate user portraits and product recommendations.
Patent Information
- Application Number
- CN202510798600.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-10
AI Technical Summary
The isolation of data in various links of the e-commerce industry makes it difficult to maximize the value of data mining. Existing technologies are unable to provide accurate data support and user portraits, affecting the effectiveness of product promotion.
Obtain multi-source e-commerce data through multi-platform login, perform data normalization and classification, establish data relationships, build personalized user portraits, and use ARIMA models to predict market demand and recommend corresponding products.
It achieves accurate prediction of user purchasing habits and effective analysis of market conditions, improves the accuracy of user portraits and the pertinence of product recommendations, and helps merchants update their products.
Smart Images

Figure CN120765282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-source heterogeneous data processing, and in particular to a high-efficiency multi-source heterogeneous data processing algorithm. Background Art
[0002] As a vital part of the economy, the e-commerce industry connects multiple entities, including products, sellers, platforms, consumers, and warehouses. Each entity carries a vast amount of data, which flows along with the transportation of goods. In the era of big data, information is value, and the effective utilization of this data has become a new opportunity within the industry.
[0003] Currently, data from various stakeholders in the e-commerce industry is isolated, and data mining is often conducted independently. With insufficient and disconnected data, it's difficult for each entity to maximize its value. To more accurately target marketing recommendations for target users, how can we bridge the data silos between products, sellers, platforms, consumers, and warehouses, thereby creating a more precise user profile?
[0004] A big data-based e-commerce data processing method and device, with publication number CN117575657A, comprises a controller that obtains e-commerce data of a target user, the e-commerce data including waybill express data, purchased product data, e-commerce platform data, merchant data, and consumer data; pre-processes the e-commerce data to obtain pre-processed target e-commerce data; identifies corresponding key information from various types of target e-commerce data based on a text recognition model; obtains a user portrait of the target user based on the key information, and finally updates the fully mined data and the connections between the data to the target user's portrait system, so that the target user's portrait system can play a greater role in data value in each link, thereby more accurately portraying the user portrait.
[0005] However, in the process of e-commerce customer information statistics, a separate data information acquisition solution cannot provide accurate data support for e-commerce, and a simple user portrait is not conducive to e-commerce's targeted product promotion, so improvements are needed. Summary of the Invention
[0006] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a high-efficiency multi-source heterogeneous data processing algorithm.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions: A high-efficiency multi-source heterogeneous data processing algorithm includes the following steps: S1. Multi-platform login: Staff members log in to multiple platforms, including but not limited to: social networking, e-commerce, sports, dating, and second-hand; S2. Acquisition of multi-source e-commerce data: Multi-source e-commerce data acquisition needs to be collected from the corresponding e-commerce platform. The corresponding multi-source e-commerce data collects user multi-source data based on user authorization; S3 data classification and data relationship establishment: Multi-source data is normalized, missing values and outliers are processed, and data is classified according to user categories. At the same time, category-based relationships are established between data. S4. Original data storage: back up the acquired data and mark them according to the acquisition date; S5. User portrait establishment: used to segment the acquired multi-source data in order to clearly define the portraits of users in each branch area. The more you know about the user, the more you can overlap multiple portrait areas to recommend corresponding products to the user.
[0008] Compared with the existing technology, this application can obtain user data through multiple platforms, and fully process and classify the data, while establishing associations between the data, so as to predict users' purchasing habits and recommend corresponding products to users. At the same time, it can also make market predictions so that merchants can replace products and facilitate consumers' purchases.
[0009] Preferably, the data classification includes an age module, a purchased item module, a consumer price module, a gender module and an after-sales module; The data relationship establishment includes the following steps: Gender: used to confirm the user's gender; Age: Specify the user's age; Shopping preferences: clearly used for shopping categories, follow categories, search categories, and collection categories; Shopping price: clearly define the shopping price range; After-sales feedback: clearly indicate user reviews, products, comments, exchanges, and return rates; Consumption prediction: predict the price of products purchased by users and the trend of product categories and make recommendations.
[0010] Furthermore, data can be effectively classified to facilitate faster user profiling.
[0011] Preferably, the multi-source user data includes but is not limited to user gender, user age, user location, user frequently purchased product categories, user spending amount, user purchase history, user collection history, user browsing history, user complaint history, user return history, and multi-source market data of user purchased items; The multi-source market data includes but is not limited to product categories, product inventory, product sales, product praise rate, product positive demand trend, and product market growth rate.
[0012] Furthermore, the data related to user purchases can be fully obtained, the data can be classified, and the market situation can be fully understood in order to remind merchants to update their products.
[0013] Preferably, the user profile is formed by normalizing the user's multi-source data, processing missing values and outliers, and then using the Pierre coefficient to calculate the correlation between the shopping behavior and other data in the user's multi-source data; Compare the correlation with the correlation threshold and filter out multi-source data with correlation lower than the correlation threshold; Integrate multi-source user data with correlation greater than or equal to the correlation threshold, normalize and label them into a personalized user portrait feature set, and build a personalized user portrait based on the personalized user portrait feature set.
[0014] Furthermore, the data is processed sufficiently through corresponding data to ensure the continuity of the data so that the data can be processed.
[0015] Preferably, the multi-source market data are arranged in chronological order to form a time series {Yt}, t=1, 2, ..., T, where Yt represents the sales volume of the product at the tth moment, and T is the total duration of the historical data; The ARIMA model is trained using historical sales data, and the model parameters are estimated by minimizing the error between the predicted value and the true value. The trained ARIMA model is combined with future seasonal information and promotional activity plans to predict the demand for goods in the next h moments. During the forecasting process, the model parameters, known historical data, and future input features are substituted into the model formula for calculation to predict the future demand for goods and complete the inventory forecast.
[0016] Furthermore, multi-source market data is processed to effectively complete the forecast of inventory data and understand market changes.
[0017] Preferably, the consumption prediction is to calculate the similarity through the user's historical behavior; for example, the purchase price of the product, the time interval for purchasing similar products, the delivery address of the purchased products, recommend the products that the user has purchased to the user, and extract the description keywords in the product attributes and product description text to recommend matching products based on the user portrait.
[0018] Furthermore, analyze and understand users' purchasing habits.
[0019] Preferably, the recommended corresponding product is to extract the commodity attribute and the description keyword in the commodity description text, convert the user's historical purchase and browsing record into a user preference vector, then the user preference keyword set is P={p1, p2,..., ps}, for each keyword pj, according to the user's behavior data of the goods containing the keyword, the user preference vector Vp is formed; for each commodity, according to the keyword set extracted from it, the commodity feature vector Vg is formed, the similarity of the user preference vector and the commodity feature vector is calculated using the cosine similarity, which is represented as: is the TF-IDF value corresponding to the user preference keyword pj in the commodity feature vector, the higher the similarity, the more matched the commodity and the user preference, the commodities are sorted according to the similarity from high to low, and the commodities with high similarity are recommended to the user.
[0020] Further, in order to screen the goods, it is convenient to recommend similar products to the user.
[0021] The beneficial effects of the present application are: 1. The user data is obtained through multiple platforms, and the data is fully processed and classified, and the correlation between the data is established, so as to predict the user's purchase habit and recommend corresponding products to the user, and also can predict the market, so that the merchant can replace the goods, and the consumer can purchase conveniently; 2. The data can be effectively classified to quickly generate user portraits, and the data involved in the user's purchase is fully obtained, and the data is classified, and the market situation is fully understood, so as to remind the merchant to update the goods; 3. The data is processed to ensure the continuity of the data, so that the data can be processed; the market multi-source data is processed to effectively complete the prediction of the inventory data, and the market trend change can also be understood; the user's purchase habit is analyzed and understood. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 A flow chart of a high-efficiency multi-source heterogeneous data processing algorithm is provided for the present application; Figure 2 A data classification block diagram in a high-efficiency multi-source heterogeneous data processing algorithm is provided for the present application; Figure 3 A data relationship establishment step block diagram in a high-efficiency multi-source heterogeneous data processing algorithm is provided for the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0024] Reference Figure 1-3 , an efficient multi-source heterogeneous data processing algorithm, comprising the following steps: S1. Multi-platform login: Staff log in to multiple platforms, including but not limited to: social networking, e-commerce, sports, dating, and second-hand goods; as many platforms as possible are involved in multiple categories, while facilitating the acquisition of corresponding user information from the platforms; relevant data is obtained through user authorization, and platforms registered with real names can quickly aggregate user information in the later stage to facilitate data processing and the creation of user portraits; S2. Acquisition of multi-source e-commerce data: Multi-source e-commerce data acquisition requires collection from the corresponding e-commerce platforms. The corresponding multi-source e-commerce data is collected based on user authorization. This can expand the scope of data acquisition, thereby increasing the data categories and data volume involved in the user's end. At the same time, the real-name system can ensure that the corresponding data is accurately classified to the corresponding user end, so that the corresponding algorithms can be used to complete and calculate the connections between the corresponding data. S3 data classification and data relationship establishment: Multi-source data is normalized, missing values and outliers are processed, and data is classified according to user categories. At the same time, category-based relationships are established between data. S4. Original data storage: back up the acquired data and mark it according to the acquisition date to effectively ensure the security of the original data and enable continuous replenishment; S5. User portrait establishment: used to segment the acquired multi-source data in order to clearly define the portraits of users in each branch area. The more you know about the user, the more you can overlap multiple portrait areas to recommend corresponding products to the user.
[0025] In the present invention, data classification includes age module, purchased item module, consumption price module, gender module and after-sales module; The establishment of data relationships includes the following steps: Gender: used to confirm the user's gender; Age: Specify the user's age; Shopping preferences: clearly used for shopping categories, follow categories, search categories, and collection categories; Shopping price: clearly define the shopping price range; After-sales feedback: clearly indicate user reviews, products, comments, exchanges, and return rates; Consumption prediction: predict the price of products purchased by users and the trend of product categories and make recommendations.
[0026] In the present invention, user multi-source data includes but is not limited to user gender, user age, user location, user frequently purchased product categories, user spending amount, user purchase history, user collection history, user browsing history, user complaint history, user return history, and multi-source data on the market where the user purchased items; Multi-source market data includes but is not limited to product categories, product inventory, product sales, product praise rate, product positive demand trends, and product market growth rate.
[0027] In the present invention, user profiling is performed by normalizing the user's multi-source data, processing missing values and outliers, and then using the Pierre coefficient to calculate the correlation between the shopping behavior and other data in the user's multi-source data; Compare the correlation with the correlation threshold and filter out multi-source data with correlation lower than the correlation threshold; Integrate multi-source user data with correlation greater than or equal to the correlation threshold, normalize and label them into a personalized user portrait feature set, and build a personalized user portrait based on the personalized user portrait feature set.
[0028] In the present invention, multi-source market data are arranged in chronological order to form a time series {Yt}, t = 1, 2, ..., T, where Yt represents the sales volume of the product at the tth moment, and T is the total duration of the historical data; The ARIMA model is trained using historical sales data, and the model parameters are estimated by minimizing the error between the predicted value and the true value. The trained ARIMA model is combined with future seasonal information and promotional activity plans to predict the demand for goods in the next h moments. During the forecasting process, the model parameters, known historical data, and future input features are substituted into the model formula for calculation to predict the future demand for goods and complete the inventory forecast.
[0029] In the present invention, consumption prediction is to calculate the similarity through the user's historical behavior; for example, the purchase price of the product, the time interval for purchasing similar products, the delivery address of the purchased products, recommend the products that the user has purchased to the user, and extract the description keywords in the product attributes and product description text to recommend matching products based on the user portrait.
[0030] In the present invention, the corresponding product recommendation is to extract the descriptive keywords from the product attributes and product description text, and convert the user's historical purchase and browsing records into a user preference vector. The user preference keyword set is P = {p1, p2, ..., ps}. For each keyword pj, the user preference vector Vp is formed based on the user's behavioral data on the product containing the keyword. For each product, the product feature vector Vg is formed based on the extracted keyword set. The cosine similarity is used to calculate the similarity between the user preference vector and the product feature vector, which is expressed as: ; in, It is the TF-IDF value corresponding to the user preference keyword pj in the product feature vector. The higher the similarity, the more the product matches the user preference. Products are sorted from high to low according to the similarity, and products with high similarity are recommended to users.
[0031] In the present invention, staff log in to multiple platforms, including but not limited to: social networking, e-commerce, sports, dating, and second-hand; as many platforms as possible involving multiple categories, while facilitating the acquisition of corresponding user information from the platforms; the corresponding data is obtained through user authorization, and the platform registered through the real-name system can quickly aggregate user information in the later stage to facilitate data processing and the creation of user portraits; Multi-source e-commerce data acquisition needs to be collected from the corresponding e-commerce platforms. The corresponding multi-source e-commerce data collects multi-source user data based on user authorization; it can expand the scope of data acquisition to increase the data categories and data volume involved in the user's end. At the same time, through the role of real-name registration, it can ensure that the corresponding data is accurately classified to the corresponding user end, so that the corresponding algorithms can be used to complete and calculate the relationship between the corresponding data; Multi-source data is normalized, missing values and outliers are processed, and the data is classified according to user categories. At the same time, category-based connections are established between the data. Back up the acquired data and mark it according to the acquisition date to effectively ensure the security of the original data and enable continuous replenishment; It is used to segment the acquired multi-source data in order to clearly define the portraits of users in each branch area. The more we know about users, the more we can overlap multiple portrait areas to recommend corresponding products to users.
[0032] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A high-performance multi-source heterogeneous data processing algorithm, characterized by: The following steps are involved: S1. Multi-platform login: Staff members log in to multiple platforms, including but not limited to: social networking, e-commerce, sports, dating, and second-hand; S2. Acquisition of multi-source e-commerce data: Multi-source e-commerce data acquisition needs to be collected from the corresponding e-commerce platform. The corresponding multi-source e-commerce data collects user multi-source data based on user authorization; S3 data classification and data relationship establishment: Multi-source data is normalized, missing values and outliers are processed, and data is classified according to user categories. At the same time, category-based relationships are established between data. S4. Original data storage: back up the acquired data and mark them according to the acquisition date; S5. User portrait establishment: used to segment the acquired multi-source data in order to clearly define the portraits of users in each branch area. The more you know about the user, the more you can overlap multiple portrait areas to recommend corresponding products to the user.
2. A high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The data classification includes age module, purchased item module, consumption price module, gender module and after-sales module; The data relationship establishment includes the following steps: Gender: used to confirm the user's gender; Age: Specify the user's age; Shopping preferences: clearly used for shopping categories, follow categories, search categories, and collection categories; Shopping price: clearly define the shopping price range; After-sales feedback: clearly indicate user reviews, products, comments, exchanges, and return rates; Consumption prediction: predict the price of products purchased by users and the trend of product categories and make recommendations.
3. The high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The multi-source user data includes but is not limited to user gender, user age, user location, user frequently purchased product categories, user spending amount, user purchase history, user collection history, user browsing history, user complaint history, user return history, and multi-source market data of user purchased items; The multi-source market data includes but is not limited to product categories, product inventory, product sales, product praise rate, product positive demand trend, and product market growth rate.
4. The high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The user profile is created by normalizing the user's multi-source data, processing missing values and outliers, and then using the Pierre coefficient to calculate the correlation between the user's shopping behavior and other data in the multi-source data; Compare the correlation with the correlation threshold and filter out multi-source data with correlation lower than the correlation threshold; Integrate multi-source user data with correlation greater than or equal to the correlation threshold, normalize and label them into a personalized user portrait feature set, and build a personalized user portrait based on the personalized user portrait feature set.
5. The high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The multi-source market data are arranged in chronological order to form a time series {Yt}, t = 1, 2, ..., T, where Yt represents the sales volume of the product at time t, and T is the total duration of the historical data; The ARIMA model is trained using historical sales data, and the model parameters are estimated by minimizing the error between the predicted value and the true value. The trained ARIMA model is combined with future seasonal information and promotional activity plans to predict the demand for goods in the next h moments. During the forecasting process, the model parameters, known historical data, and future input features are substituted into the model formula for calculation to predict the future demand for goods and complete the inventory forecast.
6. The high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The consumption prediction is to calculate the similarity through the user's historical behavior; for example, the purchase price of the product, the time interval for purchasing similar products, and the delivery address of the purchased products, recommend the products that the user has purchased to the user, and extract the description keywords in the product attributes and product description text to recommend matching products based on the user portrait.
7. The high-performance multi-source heterogeneous data processing algorithm according to claim 1, characterized in that: The recommended product is to extract the descriptive keywords from the product attributes and product description text, and convert the user's historical purchase and browsing history into a user preference vector. The user preference keyword set is P = {p1, p2, ..., ps}. For each keyword pj, a user preference vector Vp is formed based on the user's behavioral data for the product containing the keyword. For each product, a product feature vector Vg is formed based on the extracted keyword set. The cosine similarity is used to calculate the similarity between the user preference vector and the product feature vector, which is expressed as: ; in, It is the TF-IDF value corresponding to the user preference keyword pj in the product feature vector. The higher the similarity, the more the product matches the user preference. Products are sorted from high to low according to the similarity, and products with high similarity are recommended to users.
Citation Information
Patent Citations
E-commerce data processing method and device based on big data
CN117575657A