E-commerce user portrait construction method based on reinforcement and increase learning

Through the reinforcement learning-based e-commerce user portrait construction method, labels and strategies are dynamically adjusted, which solves the problem of lack of real-time response and optimization in traditional e-commerce user portrait construction and achieves higher accuracy and optimization effects.

CN120687908AActive Publication Date: 2025-09-23MOUTAI INST

Patent Information

Application Number
CN202510801232.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-23
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Traditional e-commerce user portrait construction methods lack a dynamic adjustment mechanism, making it difficult to reflect user changes. Push strategies lack real-time response and optimization, and are insufficiently accurate.

Method used

A reinforcement learning-based method is used to generate and grade labels through data collection, processing and analysis. A reward mechanism is built based on user feedback. The reinforcement learning algorithm is used to evaluate the push strategy and dynamically adjust the label weights and strategies.

Benefits of technology

It improves the accuracy of user portraits and the optimization effect of push strategies, and enhances user experience and precision marketing effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687908A_ABST
    Figure CN120687908A_ABST
Patent Text Reader

Abstract

The invention discloses an e-commerce user portrait construction method based on reinforcement and increase learning, and the system comprises a data collection module which collects the information of a user, and provides a data source for the comprehensive analysis of user behaviors; the data processing and analyzing module is used for cleaning and preprocessing the data and analyzing the behavior pattern and preference of the user according to the processed data; and the label generating and grading module is used for automatically generating labels for the user according to the result output by the data processing and analyzing module in combination with a preset label rule base, and calculating the comprehensive score of the user portrait. According to the method, browsing behavior characteristics are calculated through a specific formula, interest preferences are analyzed by applying a text mining technology and a deep learning model, and user behavior modes and interest points can be accurately grasped; a reward mechanism is constructed based on feedback, a push strategy is evaluated by using a reinforcement learning algorithm, an optimal strategy is continuously selected, and the push effect and the accuracy of a user portrait are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user portrait construction, and in particular to an e-commerce user portrait construction method based on reinforcement learning. Background Art

[0002] In today's digital age, e-commerce is booming, and e-commerce platforms face massive user growth and fierce market competition. To better meet user needs, enhance user experience, increase user stickiness, and achieve targeted marketing, e-commerce companies need to gain a deep understanding of user characteristics and behaviors, thereby providing personalized services and recommendations.

[0003] Traditional e-commerce user profile building methods often have limitations. For example, in data collection, they may only focus on superficial information, resulting in an incomplete understanding of users. In data analysis, simple statistical analysis methods are often used, making it difficult to deeply explore users' potential behavior patterns and interests. The tag generation process can be crude, lacking dynamic adjustment mechanisms and failing to reflect user changes in a timely manner. Push strategies are often based on static rules, lacking real-time response to user feedback and strategy optimization.

[0004] After searching, the application scheme of Chinese patent application number CN201510944668.6 discloses a user portrait establishment method and user portrait management system based on big data, which uses user behavior and / or content within the effective time limit to establish a temporary user portrait, and makes the temporary user portrait inherit the descriptive label attributes that match the user behavior and / or content within the effective time limit from the user portrait. When the user behavior and / or content within the effective time limit does not match the descriptive label attributes of the user portrait, a new descriptive label attribute is created in the temporary user portrait. The portrait establishment method in the above patent has the following deficiencies: there is a lack of feedback construction reward mechanism, the accuracy of user portrait construction is insufficient, and there is room for improvement. Summary of the Invention

[0005] The purpose of the present invention is to address the shortcomings of the existing technology and propose an e-commerce user portrait construction method based on reinforcement learning.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for constructing e-commerce user portraits based on reinforcement learning, the system comprising:

[0008] Data collection module, the data collection module collects user information and provides a data source for comprehensive analysis of user behavior;

[0009] Data processing and analysis module: The data processing and analysis module performs data cleaning and preprocessing, and analyzes user behavior patterns and preferences based on the processed data;

[0010] The label generation and grading module automatically generates labels for users based on the output of the data processing and analysis module and combines it with the preset label rule library, and calculates the comprehensive score of the user profile;

[0011] Push test module: The push test module identifies users with unclear label classification by setting multi-dimensional screening conditions based on the comprehensive score S of the user portrait and the matching degree of the label;

[0012] Strengthen the learning module and build a refined reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy. Use the reinforcement learning algorithm to compare and evaluate the cumulative reward values ​​of different push strategies, and select the strategy with the highest cumulative reward value as the current optimal strategy.

[0013] Preferably, the data processing and analysis module performs in-depth user behavior analysis and mines browsing behavior characteristics, as follows:

[0014] Calculate the user's average browsing frequency F avg , the formula is:

[0015]

[0016] Among them, n represents the number of views within the set time period, F i Indicates the frequency of the i-th browsing; at the same time, analyze the fluctuation of browsing frequency and calculate the standard deviation σ F To measure the stability of user browsing behavior, the formula is:

[0017]

[0018] Calculate the user's average browsing time T avg , the formula is:

[0019]

[0020] Among them, m represents the number of views within the statistical time, T j Indicates the time of the jth browsing; analyze the distribution of user browsing time and the browsing time change curve to discover user behavior patterns and preferences.

[0021] Preferably, the data processing and analysis module performs user interest preference analysis, specifically as follows:

[0022] Using text mining technology and the deep learning model BERT, we can extract semantic information from text and identify keywords, topics, and sentiments that users are interested in.

[0023] By calculating the TF-IDF value of keywords in the user's browsing content, the formula is:

[0024]

[0025] Among them, TF is the word frequency, N is the total number of documents, and DF is the number of documents containing the keyword. Keywords that strongly represent user interests are screened out;

[0026] When constructing a user interest vector, we comprehensively consider the various attributes of the product and the user's behavior weights; specifically, as follows:

[0027] For product category c k , user's interest The calculation formula is:

[0028]

[0029] in, is the number of times users browse products in this category, N total is the total number of views by users, is the inherent weight of the category’s products, is the user's rating of the product in this category, and β is the adjustment coefficient, which is used to control the impact of the rating on interest. In this way, the user's interest in different product categories is quantified.

[0030] Preferably: when the label generation and grading module automatically generates labels for users, in the results output by the data processing and analysis module, when the user meets the matching conditions of the corresponding label, the corresponding label is generated; when the user does not meet the matching conditions of the corresponding label, the matching degree data of the label is obtained according to the proportion.

[0031] Preferably, the tag generation and grading module, when generating tags, automatically generates rich tags for users based on the results output by the data processing and analysis module and in combination with a preset tag rule library; and adjusts tag weights or forms combined tags based on tag association.

[0032] Preferably, the specific manner in which the label generation and grading module performs label grading and dynamic weight adjustment is:

[0033] User tags are divided into three levels: primary tags, secondary tags, and tertiary tags;

[0034] Among them, the first-level tag reflects the user's basic information and macro-behavioral characteristics, including age group, gender, and regional distribution. Its initial weight W1 is set in the range of 0.5-0.7;

[0035] Among them, the secondary tags reflect the user's interest preferences and common behavioral pattern characteristics, including: the major categories of products of interest and peak browsing time. The initial weight W2 is between 0.2-0.3;

[0036] Among them, the third-level tags reflect the user's detailed behavioral characteristics and special preference features, including: promotion matching type, brand preference, and the initial weight W3 is 0.1-0.2;

[0037] The calculation formula for the comprehensive score S of the user portrait is:

[0038]

[0039] Among them, n1, n2, and n3 are the number of first-level, second-level, and third-level tags respectively, and T 1i 、T 2j 、T 3k are the scores of the corresponding labels;

[0040] Regularly re-evaluate the weight and classification of tags based on the latest user behavior data and feedback information.

[0041] Preferably, the push test module filters users whose label classification is unclear, as follows:

[0042] Based on the comprehensive score S of the user profile and the degree of matching of the label, we set multi-dimensional screening conditions to identify users with unclear label classification, as follows:

[0043] For the comprehensive score S threshold , S threshold is the set threshold, and the difference between multiple label scores D <D threshold , D thresholf For users who exceed the set threshold, they will be included in the category of users with unclear label classification;

[0044] Analyze the browsing history and purchase records of these users;

[0045] The push testing module customizes a personalized push content combination for each user based on their historical behavior and interest preferences, combined with tag relevance information, for users with unclear tag classifications;

[0046] Based on the product categories or keywords that frequently appear in their browsing history, we filter out products with high relevance and push them together, as follows:

[0047] ​Calculate the weight W of each product or keyword in the user's browsing history hist , the formula is:

[0048]

[0049] Among them, N hist N is the number of times a user browses a certain type of product or a certain keyword. total_hist It is the total number of browsing history records of the user. Then, the top K product categories or keywords are selected from high to low according to the weight, and the popular products corresponding to these categories or keywords are mixed with the overall popular products of the platform to form a personalized push content combination.

[0050] Preferably, the reinforcement learning module builds a reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy; specifically, as follows:

[0051] For push strategies where users click on push content and make purchases, a higher reward score R is given. buy , the calculation formula of reward points is:

[0052] R buy =α×P buy +β×M profit +γ×P inventroy

[0053] Among them, P buy , α is the weight coefficient of the purchase amount, M profit is the commodity profit, β is the weight coefficient of profit, P inventroy is the inventory turnover rate improvement index, γ is the weight coefficient of inventory turnover rate;

[0054] For push strategies where users click but do not purchase, a lower reward score R is given clicl The reward score is based on the contribution of the user's click behavior to the platform traffic and user activity. The formula is:

[0055] R ckick =δ×T stay +∈×P page

[0056] Among them, T stay is the dwell time after the user clicks on the pushed content, δ is the weight coefficient of the dwell time, P page is the number of pages viewed after the user clicks, ∈ is the weight coefficient of the number of pages;

[0057] For push strategies that users do not click, a certain penalty score R is given noclickThe penalty score is determined based on the expected effect of the push strategy and the actual number of non-clicks. It is adjusted based on the set basic penalty value and combined with the attractiveness of the push content and the accuracy of the target users;

[0058] According to the reward mechanism, the cumulative reward value Q of each push strategy is calculated as follows:

[0059] Q = ∑(R × γ t )

[0060] Where R is the reward score for each push, t is the time step, and γ is the discount factor, 0<γ<1, which is used to take into account the uncertainty of future rewards and time value decay;

[0061] The reinforcement learning algorithm Q-learning is used to compare and evaluate the cumulative reward values ​​of different push strategies, and the strategy with the highest cumulative reward value is selected as the current optimal strategy.

[0062] Preferably, the method comprises the following steps:

[0063] S1: The data collection module collects users' browsing behavior data, transaction data, and other related information in real time during the operation of the e-commerce platform, and performs preliminary organization and packaging;

[0064] S2: timely transmit the collected data to the data processing and analysis module on the server side;

[0065] S3: After receiving the data, the data processing and analysis module immediately starts the data cleaning program to eliminate invalid data and perform normalization and feature extraction on the valid data;

[0066] S4: Use text mining technology to analyze product and user review texts, extract keywords and semantic information; use statistical analysis methods to calculate user browsing behavior characteristic indicators;

[0067] S5: Generate initial tags for users based on the data processing results and the tag generation rule library;

[0068] S6: Calculate the correlation between tags and build a tag association network;

[0069] S7: Based on the label classification rules and initial weight settings, the generated labels are graded and assigned values, and the comprehensive score of the user profile is calculated to form a preliminary user profile;

[0070] S8: Improve user portrait.

[0071] Preferably, the e-commerce user portrait construction method improves the user portrait in the following specific ways:

[0072] S81: The push test module filters out user groups with unclear label classification based on the comprehensive score of the user portrait and the degree of label certainty;

[0073] S82: For these users, an algorithm is generated based on the personalized push content combination to tailor push content for them, and appropriate push timing and channels are selected for push.

[0074] S83: During the push process, the user's feedback information is tracked in real time; the feedback information is promptly transmitted to the reinforcement learning module;

[0075] S84: The reinforcement learning module calculates the cumulative reward value of each push strategy according to the received feedback information and the reward mechanism, and uses the reinforcement learning algorithm to evaluate and optimize the strategy;

[0076] S85: Select the strategy with the highest cumulative reward value as the optimal strategy, and adjust the weights and parameters of the push strategy according to the actual situation;

[0077] S86: Update user portrait tags based on the optimal strategy and the latest user feedback. Adjust the weight or delete tags that are verified to be inconsistent with the user's actual interests during push testing. Add new tags and assign them grades in a timely manner for newly discovered user interests or behavioral characteristics to improve the user portrait.

[0078] The beneficial effects of the present invention are:

[0079] 1. This invention calculates browsing behavior characteristics through a specific formula and uses text mining technology and deep learning models to analyze interest preferences, which can more accurately grasp user behavior patterns and points of interest. It builds a reward mechanism based on feedback and uses a reinforcement learning algorithm to evaluate push strategies, which can continuously select the optimal strategy, thereby improving push effects and the accuracy of user portraits.

[0080] 2. The present invention automatically generates tags based on the processing and analysis results and the rule base, and can also adjust weights or form combined tags based on the degree of association to make the tags richer and accurately reflect user characteristics.

[0081] 3. This invention can identify users with unclear label classifications and customize personalized push content combinations for them, thereby improving the accuracy and effectiveness of push notifications. It builds a reward mechanism based on feedback and uses reinforcement learning algorithms to evaluate push notification strategies, which can continuously select the optimal strategy and improve push notification effects and the accuracy of user portraits. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 This is a flowchart of a method for constructing e-commerce user portraits based on reinforcement learning proposed by the present invention;

[0083] Figure 2This is a flowchart of a method for improving user portraits in an e-commerce user portrait construction method based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION

[0084] The technical solution of the present invention will be further described in detail below in conjunction with specific implementation methods.

[0085] Example 1:

[0086] A method for constructing e-commerce user portraits based on reinforcement learning is implemented based on an e-commerce system. The system includes:

[0087] Data collection module, the data collection module collects user information and provides a data source for comprehensive analysis of user behavior;

[0088] Data processing and analysis module: The data processing and analysis module performs data cleaning and preprocessing, and analyzes user behavior patterns and preferences based on the processed data;

[0089] The label generation and grading module automatically generates labels for users based on the output of the data processing and analysis module and combines it with the preset label rule library, and calculates the comprehensive score of the user profile;

[0090] Push test module: The push test module identifies users with unclear label classification by setting multi-dimensional screening conditions based on the comprehensive score S of the user portrait and the matching degree of the label;

[0091] Strengthen the learning module and build a refined reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy. Use the reinforcement learning algorithm to compare and evaluate the cumulative reward values ​​of different push strategies, and select the strategy with the highest cumulative reward value as the current optimal strategy.

[0092] The data collection module performs multi-channel collection, specifically as follows:

[0093] For e-commerce platform web pages, JavaScript code is embedded to capture user browsing behavior in real time, recording in detail the URL, product ID, stay timestamp, scroll depth and other information of the page the user browses, and comprehensively reflecting the user's operation details on the web page;

[0094] For mobile applications, we use the sensors and log systems of mobile devices to collect data such as user touch operations, screen browsing history, and application usage time, and transmit it to the server via the network to ensure the complete collection of mobile user data.

[0095] Integrate the backend transaction data of the e-commerce platform, including user purchase records, order amount, payment method, delivery address and other information.

[0096] The data processing and analysis module identifies and eliminates duplicate, erroneous, or incomplete data records. For example, this ensures data accuracy and reliability by comparing timestamp logic and checking the legality of page URL formats. Data of different dimensions is normalized to bring all data into a uniform range of magnitude, facilitating subsequent analysis and calculations. For example, data such as browsing frequency and browsing time are mapped to the [0, 1] interval to eliminate the impact of dimensional differences on analysis results.

[0097] Conduct in-depth user behavior analysis and explore browsing behavior characteristics, as follows:

[0098] Calculate the user's average browsing frequency F avg , the formula is:

[0099]

[0100] Among them, n represents the number of views within the set time period, F i Indicates the frequency of the i-th browsing. At the same time, analyze the fluctuation of browsing frequency and calculate the standard deviation σ F To measure the stability of user browsing behavior, the formula is:

[0101]

[0102] Calculate the user's average browsing time T avg , the formula is:

[0103]

[0104] Among them, m represents the number of views within the statistical time, T j represents the time of the jth browsing. Further analysis of the distribution of user browsing time can be done, such as plotting a distribution histogram of browsing time in different time periods (e.g., early morning, morning, noon, afternoon, and evening), as well as a browsing time variation curve for each day of the week, to discover user behavior patterns and preferences.

[0105] User interest preference analysis: Text mining technology is used to conduct in-depth analysis of text data such as product descriptions, category information, and user reviews. The deep learning model BERT is used to extract semantic information from the text and identify keywords, topics, and sentiments that users are interested in.

[0106] For example, by calculating the TF-IDF value of a keyword in the content browsed by the user, the formula is:

[0107]

[0108] Among them, TF is the word frequency, N is the total number of documents, and DF is the number of documents containing the keyword, which is used to filter out keywords that have a strong representation of user interests.

[0109] When constructing the user interest vector, the various attributes of the product and the user's behavior weight are comprehensively considered. k , user's interest The calculation formula is:

[0110]

[0111] in, is the number of times users browse products in this category, N total is the total number of views by users, It is the inherent weight of the products in this category (determined by factors such as product sales and popularity). is the user's rating of the product in that category, and β is an adjustment coefficient used to control the impact of the rating on interest. In this way, the user's interest in different product categories can be more accurately quantified.

[0112] Among them, when the label generation and grading module automatically generates labels for users, in the results output by the data processing and analysis module, when the user meets the matching conditions of the corresponding label, the corresponding label is generated; when the user does not meet the matching conditions of the corresponding label, the matching degree data of the label is obtained according to the proportion; for example, when the condition for obtaining label A is to purchase sneakers 10 times, if the user has only purchased sneakers 7 times, the matching conditions of the corresponding label are not met and the corresponding label cannot be generated, but the matching degree of label A is 0.7 or 70%.

[0113] Among them, the label generation and grading module, when generating labels, automatically generates rich labels for users based on the results output by the data processing and analysis module and combined with the preset label rule library. Labels cover multiple dimensions, including basic information of users (such as age, gender, region, etc.), behavioral characteristics (such as browsing frequency level, browsing time preference, etc.), interest preferences (such as product categories and brands of interest, etc.), and consumption habits (such as purchase frequency, average order value, etc.). For example, when a user's browsing frequency and purchase times for a certain brand exceed a certain threshold, a label of "Brand Loyalty-[Brand Name]" is added to it;

[0114] Adjust tag weights or create combined tags based on tag associations. Specifically, this involves analyzing user data to uncover potential associations between tags. For example, if a strong association is found between the tag "sports enthusiast" and tags like "sports shoes," "sportswear," and "fitness equipment," when a user is labeled "sports enthusiast," the weights of these associated tags are increased or a corresponding combined tag is generated, such as "sports equipment preference - shoes, clothing, and equipment."

[0115] The specific method of the label generation and grading module to perform label grading and dynamic weight adjustment is as follows:

[0116] User tags are divided into three levels: primary tags, secondary tags, and tertiary tags;

[0117] Among them, the first-level label reflects the user's basic information and macro-behavioral characteristics, such as "age group", "gender", "geographical distribution", etc. Its initial weight W1 is set in the range of 0.5-0.7, and is fine-tuned according to the universality and stability of the user group.

[0118] Among them, the secondary tags reflect the user's interest preferences and common behavioral pattern characteristics, such as "interest-product category", "peak browsing time", etc. The initial weight W2 is between 0.2-0.3, and is adjusted with the dynamic changes of user behavior and the influence of tag correlation.

[0119] Among them, the third-level labels reflect the user's detailed behavioral characteristics and special preference features, such as "sensitive to specific promotional activities" and "niche brand preference". The initial weight W3 is 0.1-0.2, and it is flexibly adjusted according to the user's personalized feedback and rare behaviors.

[0120] The calculation formula for the comprehensive score S of the user portrait is:

[0121]

[0122] Among them, n1, n2, and n3 are the number of first-level, second-level, and third-level tags respectively, and T 1i 、T 2j 、T 3k are the scores of the corresponding tags (calculated dynamically based on the importance of the tags and the user's performance).

[0123] Regularly re-evaluate the weight and grading of tags based on the latest user behavior data and feedback. For example, if a secondary tag shows a high degree of match with a user's actual interests in multiple push tests, it will be upgraded to a primary tag and its weight will be adjusted accordingly. Conversely, if a tag is consistently inconsistent with user behavior, its weight will be reduced or merged into other relevant tags.

[0124] The push test module filters users whose label classification is unclear, as follows:

[0125] Based on the comprehensive score S of the user profile and the matching degree of the label, we set multi-dimensional screening conditions to identify users with unclear label classification. The details are as follows:

[0126] For those with lower comprehensive scores (such as S threshold , where S​threshold is the set threshold), and multiple label scores are similar (by calculating the difference D between label scores, if D <D threshold , then users with similar tag scores are considered to be classified as users with unclear tag classification. For example, set S threshold =0.3, D threshold =0.3, when the user's comprehensive score is lower than 0.3 and the difference between any two tag scores is less than 0.1, the user tag classification is determined to be unclear.

[0127] Further analyze the browsing history, purchase history and other data of these users, deeply explore the potential characteristics and preference clues in their behavior patterns, and provide a more accurate basis for subsequent push testing.

[0128] The push testing module customizes a personalized push content combination for each user based on their historical behavior and interest preferences and tag association information for users with unclear tag classification.

[0129] For example, for users with a certain browsing history but unclear interests, in addition to pushing popular products, we also filter out highly relevant products based on the product categories or keywords that frequently appear in their browsing history and push them in combination. The details are as follows:

[0130] Calculate the weight W of each product or keyword in the user's browsing history hist , the formula is:

[0131]

[0132] Among them, N hist N is the number of times a user browses a certain type of product or a certain keyword. total_hist is the total number of browsing history of the user. Then, the top K product categories or keywords are selected from high to low weights, and the popular products corresponding to these categories or keywords are mixed with the overall popular products of the platform to form a personalized push content combination.

[0133] Consider the diversity and novelty of pushed content, avoid over-reliance on users' historical behavior, and appropriately introduce new products or categories that may be potentially relevant to users' current interests but have not yet been viewed by them to stimulate their interest and desire to explore. For example, based on the tag relevance model, a certain percentage (e.g., 20%) of products associated with the user's existing tags but not yet viewed by the user can be recommended for inclusion in the pushed content mix.

[0134] The push test module, when formulating and executing push strategies, is specifically as follows:

[0135] Choose the appropriate push timing and channel based on contextual information such as the user's active time, device type, and geographic location. For example, for users who frequently browse e-commerce platforms on mobile devices at night, personalized push content can be sent via mobile push notifications during their leisure time. For users who typically browse on computers during work hours, push content can be sent via email or web pop-up ads during lunch breaks or before leaving get off work.

[0136] Embed tracking codes or identifiers in push content to accurately collect user feedback information, such as key indicators such as click-through rate, purchase conversion rate, and dwell time.

[0137] The enhanced learning module builds a reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy. The details are as follows:

[0138] For push strategies where users click on push content and make purchases, a higher reward score R is given. buy , the calculation formula of reward points is:

[0139] R buy =α×P buy +β×M profit +γ×P inventroy

[0140] Among them, P buy , α is the weight coefficient of the purchase amount, M profit is the commodity profit, β is the weight coefficient of profit, P inventroy is the inventory turnover rate improvement index (calculated based on inventory changes), and γ is the weight coefficient of inventory turnover rate. For push strategies where users click but do not purchase, a lower reward score R is given. click The reward score is mainly based on the contribution of the user's click behavior to the platform traffic and user activity. The formula is:

[0141] R click =δ×T stay +∈×P page

[0142] Among them, T stay is the dwell time after the user clicks on the pushed content, δ is the weight coefficient of the dwell time, P page is the number of pages the user browses after clicking, and ∈ is the weight coefficient of the number of pages. For push strategies that users do not click, a certain penalty score R is given. noclick The penalty score is determined based on the expected effect of the push strategy and the actual number of non-clicks. It is adjusted based on the set basic penalty value and combined with the attractiveness of the pushed content and the accuracy of the target users.

[0143] According to the reward mechanism, the cumulative reward value Q of each push strategy is calculated as follows:

[0144] Q = ∑(R × γ t )

[0145] Where R is the reward score for each push, t is the time step, and γ is the discount factor (0<γ<1) to take into account the uncertainty of future rewards and time value decay.

[0146] The reinforcement learning algorithm Q-learning is used to compare and evaluate the cumulative reward values ​​of different push strategies, and the strategy with the highest cumulative reward value is selected as the current optimal strategy.

[0147] Based on the optimal strategy and the latest user feedback, push strategies and user profile tags are updated in real time. For example, if a push strategy performs well over a period of time, not only will the weight of the strategy be increased, but the reasons for its success, such as the specific product combination, push timing, or copywriting style, will also be analyzed, and these success factors will be incorporated into similar push scenarios. If a tag is proven to be inconsistent with the user's actual interests after multiple push tests, in addition to adjusting the tag's weight or deleting it, an in-depth analysis of the causes of misjudgment, such as data collection bias, inaccurate tag definition, or sudden changes in user behavior, will be conducted to improve the data collection method and tag definition rules in a targeted manner, thereby improving the accuracy and stability of user profiles.

[0148] The e-commerce user portrait construction method includes the following steps:

[0149] S1: The data collection module collects users' browsing behavior data, transaction data, and other related information in real time during the operation of the e-commerce platform, and performs preliminary organization and packaging;

[0150] S2: timely transmit the collected data to the data processing and analysis module on the server side;

[0151] S3: After receiving the data, the data processing and analysis module immediately starts the data cleaning program to eliminate invalid data and perform normalization and feature extraction on the valid data;

[0152] S4: Use text mining technology to analyze product and user review texts, extract keywords and semantic information; use statistical analysis methods to calculate user browsing behavior characteristic indicators;

[0153] S5: Generate initial tags for users based on the data processing results and the tag generation rule library;

[0154] S6: Calculate the correlation between tags and build a tag association network;

[0155] S7: Based on the label classification rules and initial weight settings, the generated labels are graded and assigned values, and the comprehensive score of the user profile is calculated to form a preliminary user profile;

[0156] S8: Improve user portrait.

[0157] The specific method of constructing the e-commerce user portrait and improving the user portrait is as follows:

[0158] S81: The push test module filters out user groups with unclear label classification based on the comprehensive score of the user portrait and the degree of label certainty;

[0159] S82: For these users, an algorithm is generated based on the personalized push content combination to tailor push content for them, and appropriate push timing and channels are selected for push.

[0160] S83: During the push process, real-time tracking of user feedback information, including key indicators such as click-through rate, purchase conversion rate, dwell time, and page views, is performed; this feedback information is promptly transmitted to the enhanced learning module;

[0161] S84: The reinforcement learning module calculates the cumulative reward value of each push strategy according to the received feedback information and the reward mechanism, and uses the reinforcement learning algorithm to evaluate and optimize the strategy;

[0162] S85: Select the strategy with the highest cumulative reward value as the optimal strategy, and adjust the weights and parameters of the push strategies based on actual conditions. For example, if a push strategy performs well in the test, increase its weight and appropriately increase its frequency of use in subsequent pushes. If a strategy performs poorly, reduce its weight or modify and optimize it before retesting.

[0163] S86: Update user portrait tags based on the optimal strategy and the latest user feedback. Adjust the weight or delete tags that are verified to be inconsistent with the user's actual interests during push testing. Add new tags and assign them grades in a timely manner for newly discovered user interests or behavioral characteristics to improve the user portrait.

[0164] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for constructing e-commerce user portraits based on reinforcement learning, characterized in that: Based on the e-commerce system, the system includes: Data collection module, the data collection module collects user information and provides a data source for comprehensive analysis of user behavior; Data processing and analysis module: The data processing and analysis module performs data cleaning and preprocessing, and analyzes user behavior patterns and preferences based on the processed data; The label generation and grading module automatically generates labels for users based on the output of the data processing and analysis module and combines it with the preset label rule library, and calculates the comprehensive score of the user profile; Push test module: The push test module identifies users with unclear label classification by setting multi-dimensional screening conditions based on the comprehensive score S of the user portrait and the matching degree of the label; Strengthen the learning module and build a refined reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy. Use the reinforcement learning algorithm to compare and evaluate the cumulative reward values ​​of different push strategies, and select the strategy with the highest cumulative reward value as the current optimal strategy.

2. The method for constructing an e-commerce user profile based on reinforcement learning according to claim 1, characterized in that: The data processing and analysis module conducts in-depth user behavior analysis and mines browsing behavior characteristics, as follows: Calculate the user's average browsing frequency F avg , the formula is: Among them, n represents the number of views within the set time period, F i Indicates the frequency of the i-th browsing; at the same time, analyze the fluctuation of browsing frequency and calculate the standard deviation σ F To measure the stability of user browsing behavior, the formula is: Calculate the user's average browsing time T avg , the formula is: Among them, m represents the number of views within the statistical time, T j Indicates the time of the jth browsing; analyze the distribution of user browsing time and the browsing time change curve to discover user behavior patterns and preferences.

3. The method for constructing an e-commerce user profile based on reinforcement learning according to claim 1, characterized in that: The data processing and analysis module performs user interest preference analysis, specifically as follows: Using text mining technology and the deep learning model BERT, we can extract semantic information from text and identify keywords, topics, and sentiments that users are interested in. By calculating the TF-IDF value of keywords in the user's browsing content, the formula is: Among them, TF is the word frequency, N is the total number of documents, and DF is the number of documents containing the keyword. Keywords that strongly represent user interests are screened out; When constructing a user interest vector, we comprehensively consider the various attributes of the product and the user's behavior weights; specifically, as follows: For product category c k , user's interest The calculation formula is: in, is the number of times users browse products in this category, N total is the total number of views by users, is the inherent weight of the category’s products, is the user's rating of the product in this category, and β is the adjustment coefficient, which is used to control the impact of the rating on interest. In this way, the user's interest in different product categories is quantified.

4. The method for constructing an e-commerce user profile based on reinforcement learning according to claim 1, characterized in that: When the label generation and grading module automatically generates labels for users, in the results output by the data processing and analysis module, when the user meets the matching conditions of the corresponding label, the corresponding label is generated; when the user does not meet the matching conditions of the corresponding label, the matching degree data of the label is obtained according to the proportion.

5. The method for constructing an e-commerce user profile based on reinforcement learning according to claim 4, characterized in that: The label generation and grading module, when generating labels, automatically generates rich labels for users based on the results output by the data processing and analysis module and in combination with a preset label rule library; and adjusts label weights or forms combined labels based on label relevance.

6. The method for constructing an e-commerce user profile based on reinforcement learning according to claim 4, characterized in that: The specific method of the label generation and grading module to perform label grading and dynamic weight adjustment is as follows: User tags are divided into three levels: primary tags, secondary tags, and tertiary tags; Among them, the first-level tag reflects the user's basic information and macro-behavioral characteristics, including age group, gender, and regional distribution. Its initial weight W1 is set in the range of 0.5-0.7; Among them, the secondary tags reflect the user's interest preferences and common behavioral pattern characteristics, including: the major categories of products of interest and peak browsing time. The initial weight W2 is between 0.2-0.3; Among them, the third-level tags reflect the user's detailed behavioral characteristics and special preference features, including: promotion matching type, brand preference, and the initial weight W3 is 0.1-0.2; The calculation formula for the comprehensive score S of the user portrait is: Among them, n1, n2, and n3 are the number of first-level, second-level, and third-level tags respectively, and T 1i 、T 2j 、T 3k are the scores of the corresponding labels; Regularly re-evaluate the weight and classification of tags based on the latest user behavior data and feedback information.

7. The method for constructing e-commerce user portraits based on reinforcement learning according to claim 6, characterized in that: The push test module filters users whose label classification is unclear, as follows: Based on the comprehensive score S of the user profile and the degree of matching of the label, we set multi-dimensional screening conditions to identify users with unclear label classification, as follows: For the comprehensive score D threshold , S threshold is the set threshold, and the difference between multiple label scores S <D threshold , D threshold For users who exceed the set threshold, they will be included in the category of users with unclear label classification;​ Analyze the browsing history and purchase records of these users; The push testing module customizes a personalized push content combination for each user based on their historical behavior and interest preferences, combined with tag relevance information, for users with unclear tag classifications; Based on the product categories or keywords that frequently appear in their browsing history, we filter out products with high relevance and push them together, as follows: Calculate the weight W of each product or keyword in the user's browsing history hist , the formula is: Among them, N hist N is the number of times a user browses a certain type of product or a certain keyword. total_hist It is the total number of browsing history records of the user. Then, the top K product categories or keywords are selected from high to low according to the weight, and the popular products corresponding to these categories or keywords are mixed with the overall popular products of the platform to form a personalized push content combination.

8. The method for constructing e-commerce user portraits based on reinforcement learning according to claim 7, characterized in that: The enhanced learning module builds a reward mechanism based on user feedback to comprehensively evaluate the effectiveness of the push strategy; the details are as follows: For push strategies where users click on push content and make purchases, a higher reward score R is given. buy , the calculation formula of reward points is: R buy =α×P buy +β×M profit +γ×P inventroy Among them, P buy , α is the weight coefficient of the purchase amount, M profit is the commodity profit, β is the weight coefficient of profit, P inventroy is the inventory turnover rate improvement index, γ is the weight coefficient of inventory turnover rate; For push strategies where users click but do not purchase, a lower reward score R is given click The reward score is based on the contribution of the user's click behavior to the platform traffic and user activity. The formula is: R click =δ×T stay +∈×P page Among them, T stay is the dwell time after the user clicks on the pushed content, δ is the weight coefficient of the dwell time, P page is the number of pages viewed after the user clicks, ∈ is the weight coefficient of the number of pages; For push strategies that users do not click, a certain penalty score R is given noclick The penalty score is determined based on the expected effect of the push strategy and the actual number of non-clicks. It is adjusted based on the set basic penalty value and combined with the attractiveness of the push content and the accuracy of the target users; According to the reward mechanism, the cumulative reward value Q of each push strategy is calculated as follows: Q=Σ(R×γ t ) Where R is the reward score for each push, t is the time step, and γ is the discount factor, 0<γ<1, which is used to take into account the uncertainty of future rewards and time value decay; The reinforcement learning algorithm Q-learning is used to compare and evaluate the cumulative reward values ​​of different push strategies, and the strategy with the highest cumulative reward value is selected as the current optimal strategy.

9. The method for constructing e-commerce user portraits based on reinforcement learning according to claim 8, characterized in that: The steps include: S1: The data collection module collects user information in real time during the operation of the e-commerce platform and performs preliminary sorting and packaging; S2: timely transmit the collected data to the data processing and analysis module on the server side; S3: After receiving the data, the data processing and analysis module immediately starts the data cleaning program to eliminate invalid data and perform normalization and feature extraction on the valid data; S4: Use text mining technology to analyze product and user review texts, extract keywords and semantic information; use statistical analysis methods to calculate user browsing behavior characteristic indicators; S5: Generate initial tags for users based on the data processing results and the tag generation rule library; S6: Calculate the correlation between tags and build a tag association network; S7: Based on the label classification rules and initial weight settings, the generated labels are graded and assigned values, and the comprehensive score of the user profile is calculated to form a preliminary user profile; S8: Improve user portrait.

10. The method for constructing e-commerce user portraits based on reinforcement learning according to claim 9, characterized in that: The specific method of constructing the e-commerce user portrait and improving the user portrait is as follows: S81: The push test module filters out user groups with unclear label classification based on the comprehensive score of the user portrait and the degree of label certainty; S82: For these users, an algorithm is generated based on the personalized push content combination to tailor push content for them, and appropriate push timing and channels are selected for push. S83: During the push process, real-time tracking of user feedback information; Transmit these feedback information to the reinforcement learning module in a timely manner; S84: The reinforcement learning module calculates the cumulative reward value of each push strategy according to the received feedback information and the reward mechanism, and uses the reinforcement learning algorithm to evaluate and optimize the strategy; S85: Select the strategy with the highest cumulative reward value as the optimal strategy, and adjust the weights and parameters of the push strategy according to the actual situation; S86: Update user portrait tags based on the optimal strategy and the latest user feedback. Adjust the weight or delete tags that are verified to be inconsistent with the user's actual interests during push testing. Add new tags and assign them grades in a timely manner for newly discovered user interests or behavioral characteristics to improve the user portrait.

Citation Information

Patent Citations

  • A method for building user profiles and a user profile management system based on big data

    CN105574159B

  • Big data-based user portrayal establishing method and user portrayal management system

    CN105574159A

  • User portrait optimization method

    CN115248896A

  • Intelligent fuzzy test method, device and system based on reinforcement learning

    CN115309628A

  • Content recommendation method and system based on industry knowledge graph and reinforcement learning

    CN117493687A

Cited By

  • Label list automatic generation method and system for user portraits

    CN121051282A