Product recommendation method and device, equipment, storage medium and program product
By using reinforcement learning algorithms and multimodal user interaction information to dynamically adjust the recommendation strategy network, the problem of low accuracy in insurance product recommendation methods is solved, achieving efficient and accurate product recommendations and improved user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-28
AI Technical Summary
Existing insurance product recommendation methods lack optimization mechanisms, resulting in low accuracy in understanding recommended products and impacting the efficiency of recommendation services and user satisfaction.
A recommendation strategy network based on reinforcement learning algorithm is adopted to match and recommend products through multimodal user interaction information (voice, text, and images). The parameters of the recommendation strategy network are optimized by using a reward mechanism, and the recommendation strategy is dynamically adjusted by combining product database and user feedback information.
It has achieved efficient and accurate product recommendations, improved the efficiency of insurance recommendation services and user satisfaction, and continuously optimized the performance of the recommendation strategy network.
Smart Images

Figure CN121937231A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recommendation technology, specifically to a product recommendation method, apparatus, device, storage medium, and program product. Background Technology
[0002] Currently, insurance product recommendations largely rely on traditional manual sales or intelligent recommendation systems based on simple text interaction. Manual sales are inefficient and heavily influenced by the professional level and subjective factors of sales personnel; intelligent recommendation systems based on simple text interaction struggle to accurately understand complex and implicit user needs, requiring users to describe their situation in lengthy texts, making the interaction process cumbersome. With the application of artificial intelligence technology in the insurance field, the lack of optimization mechanisms in existing insurance product recommendation methods leads to low accuracy in understanding recommended products, thus affecting the efficiency of insurance recommendation services and user satisfaction. Summary of the Invention
[0003] At least one embodiment of the present invention provides a product recommendation method, apparatus, device, storage medium, and program product to address the problem that the lack of optimization mechanisms in existing insurance product recommendation methods leads to low accuracy in understanding the recommended products.
[0004] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0005] In a first aspect, embodiments of the present invention provide a product recommendation method, comprising:
[0006] Based on the user's immediate product demand interaction information, multiple candidate recommended products are matched for the user in the product database;
[0007] Based on the recommendation strategy network and the user's second-time recommended product information, a target recommended product for the user at the first time is determined from among multiple candidate recommended products; wherein, the parameters of the recommendation strategy network are updated based on a first reward value obtained from the user's first feedback information on the recommended product at the second time; the second time is before the first time.
[0008] Optionally, in the product recommendation method, if the first feedback information is used to indicate that the user has not rejected the recommended product at the second time, and / or the type of the recommended product at the second time matches the user's risk preference type, then the first reward value is positive.
[0009] If the first feedback information is used to indicate that the user did not accept the recommended product at the second time, and / or the type of the recommended product at the second time is the same as the type of recommended product that the user repeatedly did not accept, then the first reward value is negative.
[0010] Optionally, the product recommendation method further includes:
[0011] Obtain the user's second feedback information on the target recommended product at the first time.
[0012] Based on the second feedback information, the second reward value is obtained;
[0013] The parameters of the recommendation strategy network are updated according to the second reward value to obtain the updated recommendation strategy network;
[0014] Based on the updated recommendation strategy network and the user's product demand interaction information at a third time, a target recommended product for the user at the third time is determined; wherein, the third time is after the first time.
[0015] Optionally, the product recommendation method includes matching multiple candidate recommended products for the user in the product database based on the user's immediate product demand interaction information, including:
[0016] Based on the user's immediate product demand interaction information, obtain the demand feature vector;
[0017] Based on the user demand analysis model and the demand feature vector, the initial product demand analysis results of the user at the first moment are obtained;
[0018] Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained.
[0019] Based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database.
[0020] Optionally, in the product recommendation method, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0021] Based on the user's immediate product request interaction information, obtain the demand feature vector, including:
[0022] Feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, which includes at least two of the following: voice feature vector, text feature vector, and image feature vector.
[0023] The initial feature vector is mapped to the target dimension to obtain the mapped initial feature vector;
[0024] Based on the correlation between the initial feature vector and the context vector at the first time, the attention weight corresponding to the initial feature vector is obtained;
[0025] Feature fusion is performed based on the attention weights and the initial feature vector to obtain the required feature vector.
[0026] Optionally, in the product recommendation method, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0027] Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained, including:
[0028] Based on the risk-product type mapping library in the product demand knowledge base, as well as the product demand analysis results and product demand interaction information, the user's first-time product recommendation type is obtained; wherein, the risk-product type mapping library is used to indicate the mapping relationship between different risks and product types existing in the user;
[0029] Based on the facial feature-product limit mapping library in the product demand knowledge base, as well as the product demand analysis results and the product demand interaction information, the user's first-time product recommendation limit is obtained; wherein, the facial feature-product limit mapping library is used to indicate the mapping relationship between the user's different facial features and product limits;
[0030] Based on the requirement-product usage period mapping library in the product requirement knowledge base, as well as the product requirement analysis results and product requirement interaction information, the user's first-time recommended product usage period is obtained; wherein, the requirement-product usage period mapping library is used to indicate the mapping relationship between the user's different product requirements and product usage periods;
[0031] Based on the product recommendation type, the product recommendation limit, and the product recommendation usage period, the user's first-time product demand analysis results are obtained.
[0032] Optionally, the product recommendation method, wherein, based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database, including:
[0033] Based on the product demand analysis results, the product demand analysis feature vector of the user is obtained;
[0034] In terms of the target dimension, obtain the product feature vector of each product in the product database, where the target dimension is related to the product demand analysis results;
[0035] Based on the matching degree between the product demand analysis feature vector and the product feature vector, multiple candidate recommended products are matched for the user in the product database.
[0036] Optionally, the product recommendation method, wherein determining the target recommended product for the user in the first time from among a plurality of candidate recommended products based on the recommendation strategy network and the user's second-time recommended product information, includes:
[0037] Based on the recommendation strategy network, product recommendation rules, and the user's second-time recommended product information, the probability distribution corresponding to each candidate recommended product is predicted; wherein, the product recommendation rules are formulated according to different product demand scenarios;
[0038] Based on the probability distribution corresponding to the candidate recommended products, the target recommended product for the user at the first time is determined from among the multiple candidate recommended products.
[0039] Secondly, embodiments of the present invention also provide a product recommendation device, comprising:
[0040] The matching module is used to match multiple candidate recommended products for the user in the product database based on the user's first-time product demand interaction information.
[0041] The determining module is configured to determine the target recommended product for the user at the first time from among a plurality of candidate recommended products based on the recommendation strategy network and the user's recommended product information at the second time; wherein, the parameters of the recommendation strategy network are obtained by updating a first reward value based on the user's first feedback information on the recommended product at the second time; the second time is before the first time.
[0042] Thirdly, embodiments of the present invention also provide a product recommendation device, comprising: a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the processor executes the program or instructions to implement the product recommendation method as described in the first aspect.
[0043] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the product recommendation method as described in the first aspect.
[0044] Fifthly, embodiments of the present invention also provide a computer program product, including computer instructions, which, when executed by a processor, implement the product recommendation method as described in the first aspect.
[0045] Compared with existing technologies, embodiments of the present invention provide a product recommendation method, apparatus, device, storage medium, and program product. Based on a user's first-time product demand interaction information, multiple candidate recommended products are matched for the user in a product database. Based on a recommendation strategy network and the user's second-time recommended product information, a target recommended product for the user in the first-time period is determined from among the multiple candidate recommended products. The parameters of the recommendation strategy network are updated based on a first reward value obtained from the user's first feedback information on the recommended product in the second-time period. The second time period is prior to the first time period. This achieves linkage between user feedback and the optimization of the recommendation strategy network, solving the problem of low accuracy in understanding recommended products due to the lack of optimization mechanisms in insurance product recommendation methods. It can continuously improve the performance of the recommendation strategy network, achieving efficient and accurate product recommendations, and improving the efficiency of insurance recommendation services and user satisfaction. Attached Figure Description
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0047] Figure 1 This is a flowchart illustrating the product recommendation method described in an embodiment of the present invention;
[0048] Figure 2 This is a flowchart illustrating one embodiment of the product recommendation method described in this invention.
[0049] Figure 3 This is a schematic diagram of the product recommendation device according to an embodiment of the present invention;
[0050] Figure 4 This is a hardware block diagram of the product recommendation device described in an embodiment of the present invention. Detailed Implementation
[0051] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, the "or" in this invention indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0052] See Figure 1 This invention provides a product recommendation method that can be applied to insurance product recommendation scenarios or customer service scenarios.
[0053] Furthermore, the method includes:
[0054] Step 101: Based on the user's immediate product demand interaction information, match multiple candidate recommended products for the user in the product database;
[0055] For example, the first time could be the current time.
[0056] Since the product recommendation method described in this embodiment of the invention can be applied to insurance product recommendation scenarios, the product database can be an insurance product database that stores information on various insurance products, including but not limited to: product terms; coverage; premiums; and claims conditions.
[0057] Step 102: Based on the recommendation strategy network and the user's second-time recommended product information, determine the target recommended product for the user from among the multiple candidate recommended products; wherein, the parameters of the recommendation strategy network are obtained by updating the first reward value based on the user's first feedback information on the second-time recommended product; the second time is before the first time.
[0058] Since the first time can be the current time and the second time is before the first time, the second time can be a historical time.
[0059] In this embodiment of the invention, the recommendation strategy network is a neural network based on a reinforcement learning algorithm. The input of the recommendation strategy network includes the user's second-time recommended product information and information on multiple candidate recommended products. The output of the recommendation strategy network includes the user's first-time product demand interaction information.
[0060] It should be noted that this reinforcement learning algorithm includes:
[0061] Status (S): The user's second-time recommended product information and information on multiple candidate recommended products;
[0062] Action (A): Select and recommend the first target recommended product to the user from among the multiple candidate recommended products;
[0063] Reward (R): A numerical signal set based on the user's feedback on the recommendation action (i.e., the target recommended product), which guides the optimization parameters of the recommendation strategy network.
[0064] This reinforcement learning algorithm employs a reward mechanism, which quantifies the evaluation of the user's feedback on the recommendation action by defining a reward value (R), directly influencing the optimization direction of the recommendation strategy network. Specifically, the algorithm optimizes the recommendation strategy network based on the user's second-time recommended product and first-time feedback information. By effectively utilizing the user's feedback information through the reward value, the recommendation strategy network can evolve from fixed rules to adaptive learning, continuously improving the accuracy of product recommendations and user acceptance.
[0065] In one implementation, optionally, if the first feedback information is used to indicate that the user did not reject the recommended product at the second time, and / or the type of the recommended product at the second time matches the user's risk preference type, then the first reward value is positive;
[0066] If the first feedback information is used to indicate that the user did not accept the recommended product at the second time, and / or the type of the recommended product at the second time is the same as the type of recommended product that the user repeatedly did not accept, then the first reward value is negative.
[0067] In this embodiment of the invention, the reward value includes a positive reward value (i.e., the reward value is positive) and a negative reward value (i.e., the reward value is negative); wherein, the positive reward value is used to encourage the recommendation policy network; and the negative reward value is used to penalize the recommendation policy network.
[0068] Specifically, if the first reward value is a positive reward value, then the first feedback information is used to indicate that the user did not reject the recommended product at the second time, and / or that the type of the recommended product at the second time matches the user's risk preference type. Wherein, the user's failure to reject the recommended product at the second time is determined by setting different first reward values based on the type and intensity of the first feedback information, and can be categorized into at least one of the following types:
[0069] Strong positive reward type: The user explicitly accepts the recommended product at the second time, for example, by clicking the "buy" or "insure" button, and is given a higher positive reward value, which directly strengthens the association between the recommended action and the current state; Assuming that the user explicitly accepts the recommended product at the second time, which is behavioral event A, if behavioral event A occurs, the calculation formula for the first reward value R can be expressed as R = +10;
[0070] Medium positive reward type: When the user shows further interest in the recommended product at the second time, such as clicking on the product details page, adding the product to favorites, or asking for more details, a medium positive reward value is given. Assuming the user shows further interest in the recommended product at the second time, which is behavioral event B, if behavioral event B occurs, the formula for calculating the first reward value can be expressed as R = +3 to +5, encouraging the recommendation strategy network to continue exploring similar products; the first reward value can be fine-tuned by the number of times behavioral event B occurs within a certain time window, n, such as R = +3 + n / 10;
[0071] Weak positive reward type: If the user does not explicitly reject the recommended product at the second time but shows interest, such as listening to the voice recommendation or browsing the text description for more than 3 seconds, a weak positive reward value is given. Assuming the user does not explicitly reject the recommended product at the second time but shows interest, this is behavioral event C. If behavioral event C occurs, the calculation formula for the first reward value can be expressed as R = +1, preventing the recommendation strategy network from prematurely abandoning potentially suitable products.
[0072] The type of the recommended product at the second time point matches the user's risk preference type. For example, if the user's risk preference type is risk-averse, and the type of the recommended product at the second time point is conservative, the user's first feedback information indicates that the user did not refuse. Assuming that the type of the recommended product at the second time point highly matches the user's risk preference type and the user did not refuse, this is considered a behavioral event G. If this behavioral event G occurs, the formula for calculating the first reward value can be expressed as R = +2.
[0073] Specifically, if the first reward value is a negative reward value, then the first feedback information is used to indicate that the user did not accept the recommended product at the second time, and / or, the type of the recommended product at the second time is the same as the type of recommended product that the user repeatedly did not accept. Wherein, the user not accepting the recommended product at the second time includes at least one of the following:
[0074] Strong negative reward type: When the user explicitly rejects the recommended product at the second time, for example by clicking the "Not interested" or "Change" button, a higher negative reward value is given. Assuming the user explicitly rejects the recommended product at the second time, which is behavioral event D, if behavioral event D occurs, the formula for calculating the first reward value can be expressed as R=-8, which inhibits the occurrence of this recommendation action in similar states;
[0075] Medium negative reward type: When the user exhibits resistance to the recommended product at the second time, such as quickly skipping the recommendation or closing the recommendation window, a medium negative reward value is given. Assuming the user exhibits resistance to the recommended product at the second time, which is behavioral event E, if behavioral event E occurs, the calculation formula for the first reward value can be expressed as R = -3 to -5, prompting the recommendation strategy network to adjust the product type or recommendation method;
[0076] Weak negative reward type: If the user has no interaction with the recommended product at the second time, for example, no action within 10 seconds after the recommendation, a weak negative reward value is given. Assuming the user has no interaction with the recommended product at the second time, this is behavioral event F. If behavioral event F occurs, the calculation formula for the first reward value can be expressed as R = -1, guiding the recommendation strategy network to optimize the timing of recommendations or the attractiveness of content.
[0077] The type of the recommended product at the second time point is the same as the type of recommended product that the user has repeatedly rejected. For example, if the user previously explicitly rejected "dividend insurance," it may be recommended again. Assuming that the type of recommended product at the second time point is the same as the type of recommended product that the user has historically rejected, this is considered a behavioral event H. If this behavioral event H occurs, the formula for calculating the first reward value can be expressed as R = -5, thus preventing the recommendation strategy network from repeating invalid actions.
[0078] In one embodiment, optionally, the method further includes:
[0079] Obtain the user's second feedback information on the target recommended product at the first time.
[0080] Based on the second feedback information, the second reward value is obtained;
[0081] The parameters of the recommendation strategy network are updated according to the second reward value to obtain the updated recommendation strategy network;
[0082] Based on the updated recommendation strategy network and the user's product demand interaction information at a third time, a target recommended product for the user at the third time is determined; wherein, the third time is after the first time.
[0083] In this embodiment of the invention, the parameters of the recommendation strategy network can be updated based on the second reward value obtained from the user's second feedback information on the target recommended product at the first time. Then, based on the updated recommendation strategy network and the user's product demand interaction information at the third time, the target recommended product for the user at the third time is determined. Since the first time can be the current time, and the third time is before the first time, the third time can be a future time.
[0084] It should be noted that the target recommended products in the first instance can be in multimodal form, including text, voice, or image. Specifically, the information of the target recommended products in the first instance is in text form, providing a detailed introduction to the recommended products; text-to-speech technology is used to convert the text information of the target recommended products in the first instance into voice for broadcast; based on the text information of the target recommended products in the first instance, images or charts are generated to intuitively display the comparative information, profit predictions, etc. of the target recommended products. It is ensured that the information of the target recommended products in multiple modalities such as text, voice, and images is consistent in content and time, and presented to the user in a natural and clear manner.
[0085] Furthermore, the second feedback information from the user regarding the target recommended product at the first time includes at least one of the following: satisfaction evaluation information, such as star rating; further demand information, such as "wanting a product with a lower price"; and behavioral feedback information, such as skipping the recommendation or adding the product to favorites. Based on the second feedback information, the recommendation evaluation effect is obtained, such as click-through rate, conversion rate, or satisfaction score, and the shortcomings of the recommendation strategy network are analyzed, such as failing to accurately capture the user's sensitivity to insurance premiums.
[0086] Based on the second feedback information, not only can the parameters of the recommendation strategy network be updated, but the user demand analysis model can also be updated, such as adjusting feature weights, and the insurance product matching algorithm can be updated, such as optimizing feature dimensions, so as to continuously improve the accuracy of product recommendations and user satisfaction.
[0087] The reinforcement learning algorithm in this embodiment of the invention employs a reward mechanism, the working principle of which includes at least one of the following:
[0088] (1) Real-time feedback guidance: Through the reward value of real-time feedback, the recommendation strategy network can quickly learn which products are more easily accepted in which user states. For example, if positive reward values are obtained multiple times when recommending family liability insurance to a 30-year-old married user, the recommendation strategy network will gradually increase the probability of recommending such products in this scenario.
[0089] (2) Long-term cumulative reward optimization: The goal of reinforcement learning algorithms is to maximize long-term cumulative rewards, rather than single rewards. The reward mechanism balances immediate rewards and long-term benefits to prevent the recommendation strategy network from falling into the trap of short-term optimality but long-term ineffectiveness. For example, for products with low initial interest but potential long-term needs, such as long-term life insurance, a weak positive reward value can guide the recommendation strategy network to appropriately retain recommendation opportunities, rather than completely abandoning them due to a lack of feedback once.
[0090] (3) Adapting to multimodal interaction characteristics: The reward mechanism combines user multimodal feedback information, such as negative emotions like "too expensive" in voice, requests for cheaper prices in text input, or frowning gestures in facial expressions, and dynamically adjusts the reward value. For example, when a user's voice tone shows impatience, the negative reward value for the target recommended product will be increased, such as from -3 to -5.
[0091] Furthermore, the reward mechanism drives the recommendation strategy network to update in the following ways:
[0092] (1) Recommendation Strategy Network ( (S) Output the probability distribution of the recommended action A based on the current state S;
[0093] (2) After performing the recommended action A, obtain the reward value R corresponding to the user feedback;
[0094] (3) Using reinforcement learning algorithms, such as Q-Learning (a classic value-based reinforcement learning algorithm) or PPO (Proximal Policy Optimization, an advanced policy-based algorithm), update the parameters of the recommendation policy network based on the reward value R, so that the probability of the recommendation action with high reward value increases in similar states and the probability of the recommendation action with low reward value decreases.
[0095] (4) As the user's product demand interaction information accumulates, the reward mechanism will gradually teach the recommendation strategy network to understand the preferences of different user groups, such as young people paying more attention to premiums and middle-aged people paying more attention to coverage, and finally realize the dynamic optimization of the personalized recommendation strategy network.
[0096] In one implementation, optionally, based on the recommendation strategy network and the user's second-time recommended product information, determining the target recommended product for the user from among a plurality of candidate recommended products for the first time includes:
[0097] Based on the recommendation strategy network, product recommendation rules, and the user's second-time recommended product information, the probability distribution corresponding to each candidate recommended product is predicted; wherein, the product recommendation rules are formulated according to different product demand scenarios;
[0098] Based on the probability distribution corresponding to the candidate recommended products, the target recommended product for the user at the first time is determined from among the multiple candidate recommended products.
[0099] The product recommendation rules can be formulated based on product demand scenarios. For example, for users with children's education needs, education savings insurance is recommended first; employer liability insurance is recommended for business owners; and traffic accident insurance is recommended for business people.
[0100] In this embodiment of the invention, the product recommendation rule is used to indicate recommended products corresponding to the user's risk preference information. This risk preference information can be obtained based on the user's immediate product demand interaction information. For example, if the user's risk preference information is risk-averse, then the recommended product for the user is a stable insurance product, such as critical illness insurance or medical insurance; if the user's risk preference information is risk-averse, then the recommended product for the user is an insurance product with certain investment attributes, such as participating insurance or universal life insurance.
[0101] In one implementation method, optionally, based on the user's immediate product demand interaction information, multiple candidate recommended products are matched for the user in the product database, including:
[0102] Based on the user's immediate product demand interaction information, obtain the demand feature vector;
[0103] Based on the user demand analysis model and the demand feature vector, the initial product demand analysis results of the user at the first moment are obtained;
[0104] Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained.
[0105] Based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database.
[0106] In this embodiment of the invention, the method for collecting the product demand interaction information at the first time includes at least one of the following:
[0107] The microphone is used to collect product demand interaction information at the first moment, including the user's verbal description or questions about product demand;
[0108] Receive the product demand interaction information manually input by the user at the first time, including personal basic information or existing product information;
[0109] The system captures the user's facial image using a camera and extracts product demand interaction information from the facial image at the first moment, including the user's emotional state or age.
[0110] Optionally, the user demand analysis model is pre-built based on a deep learning model (such as Transformer), and the user demand analysis model is used to analyze the demand feature vector. The system processes and identifies user intent to obtain the user's initial product demand analysis results in real time.
[0111] Furthermore, to address the semantic misunderstanding of the user demand analysis model regarding the demand feature vectors, such as misjudging facial features like "20-25 years old" and "relaxed smile" as indicating no insurance demand, a knowledge graph dynamic reasoning module is introduced to correct the initial product demand analysis results output by the user demand analysis model. That is, based on the product demand knowledge base, the initial product demand analysis results, and the product demand interaction information, the user's first-time product demand analysis results are obtained.
[0112] For example, the initial product demand analysis result output by the user demand analysis model is no insurance demand. The product demand knowledge base is then invoked, and combined with the product demand interaction information, to obtain the user's first-time product demand analysis result.
[0113] The product demand knowledge base includes a risk-product type mapping library. If the product demand interaction information is a demand of young users, then the risk of the user is determined to be no family responsibility but with accidental risk. The corresponding protection type for no family responsibility but with accidental risk is obtained from the risk-product type mapping library, which is accident insurance.
[0114] Then, based on the above product demand analysis results, supporting information is obtained from the user's first-time product demand interaction information. This supporting information includes, for example, voice interaction information indicating fear of accidents on the road, text interaction information indicating frequent business trips, and image interaction information indicating age 20 to 25. If such supporting information exists, the above product demand analysis results are further triggered to trigger the user demand analysis model to accurately calculate the coverage amount and coverage period.
[0115] For example, the image interaction information in the user's first-time product demand interaction information is 22 years old and has a student-like appearance. The user demand analysis model initially judges that there is no demand for critical illness insurance. However, combined with the text interaction information in the user's first-time product demand interaction information, which is that the user is studying in another city, and the voice-text interaction information, which is that the user is afraid of getting sick and having no one to take care of them, the product demand knowledge base verifies that students in other cities have medical reimbursement needs. The insurance type is obtained as a demand for million-dollar medical insurance. The output result is corrected to a demand for million-dollar medical insurance, with a coverage amount of 3 million and a coverage period of 1 year, thereby reducing the rate of missed judgment of the needs of young users.
[0116] In one embodiment, optionally, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0117] Based on the user's immediate product request interaction information, obtain the demand feature vector, including:
[0118] Feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, which includes at least two of the following: voice feature vector, text feature vector, and image feature vector.
[0119] The initial feature vector is mapped to the target dimension to obtain the mapped initial feature vector;
[0120] Based on the correlation between the initial feature vector and the context vector at the first time, the attention weight corresponding to the initial feature vector is obtained;
[0121] Feature fusion is performed based on the attention weights and the initial feature vector to obtain the required feature vector.
[0122] In this embodiment of the invention, the product demand interaction information is multimodal product demand interaction information, including at least two of voice interaction information, text interaction information, and image interaction information.
[0123] Since the product demand interaction information is voice interaction information, feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, including:
[0124] By using the Voice Activity Detection (VAD) algorithm, the effective area of speech is accurately located, silent segments and non-speech interference (such as breathing sounds and background noise) are removed, and redundant data is reduced.
[0125] Algorithms such as spectral subtraction and Wiener filtering are used to eliminate environmental noise (such as conference room echoes and street noise) and improve the signal-to-noise ratio of speech signals.
[0126] The Mel-Frequency Cepstral Coefficients (MFCC) algorithm is used to convert the preprocessed speech interaction information into an initial feature vector that conforms to the characteristics of human hearing.
[0127] Assuming the voice interaction information is a continuous time series x(t), the initial feature vector obtained after extraction by the MFCC algorithm is Fvoice=[c1,c2,...,ci], where ci is the i-th MFCC, reflecting the energy distribution and spectral characteristics of the speech in a specific frequency band, and is the core input for speech recognition and emotion analysis.
[0128] Since the product demand interaction information is text-based, feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, including:
[0129] The text interaction information is processed by word segmentation (e.g., Chinese is segmented using the "jieba" word segmentation component, and English is segmented by spaces), part-of-speech tagging (e.g., nouns and verbs are classified), and named entity recognition (extracting key entities such as personal names, place names, organization names, and insurance types).
[0130] To unify the text format of the text interaction information (such as case conversion, traditional Chinese to simplified Chinese, and removal of special symbols), and eliminate interference caused by differences in expression;
[0131] The core content (such as keywords, topic phrases, sentiment words, and demand description words) is filtered from the processed text interaction information to obtain the initial feature vector.
[0132] Suppose the text interaction information is T=[w1,w2,...,wm], where wm represents the m-th character (such as a Chinese character, letter, or punctuation mark); after word segmentation, a word sequence is obtained, and then the initial feature vector Ftext=[k1,k2,...,ks] is extracted through the above processing, where ks represents the s-th key information unit (such as "30 years old", "health insurance", "annual income of 500,000", etc., which are linguistic units with practical significance).
[0133] Since the product demand interaction information is image interaction information, feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, including:
[0134] The quality of the image interaction information is optimized by operations such as contrast adjustment, noise reduction (such as Gaussian filtering), or sharpening, so as to solve problems such as insufficient light and blur.
[0135] Multi-task Cascaded Convolutional Networks (MTCNN) are used to locate facial regions, and image segmentation or keypoint detection is combined to extract fine features such as facial contours and keypoints.
[0136] The cropped or aligned face region image is input into a convolutional neural network (CNN), and abstract visual features are extracted through multiple convolution and pooling operations to obtain an initial feature vector.
[0137] Assuming the image interaction information is I, after feature extraction by CNN, an initial feature vector Fimg=[f1,f2,...,fp] is obtained, where fp represents the p-th initial feature vector extracted from the image interaction information I, corresponding to the numerical representation of the image in abstract dimensions (such as texture patterns, contour features, facial expression features (such as the emotion value corresponding to smiling / frowning), and age).
[0138] Next, the initial feature vectors are fused using an attention-based mechanism to achieve dynamic integration of cross-modal features. The core logic is to adaptively allocate weights according to task requirements to highlight key information, including the following steps:
[0139] S1: Input initial feature vectors, including the initial feature vector corresponding to voice interaction information, referred to as voice feature vector Fvoice; the initial feature vector corresponding to text interaction information, referred to as text feature vector Ftext; and the initial feature vector corresponding to image interaction information, referred to as image feature vector Fimg.
[0140] S2: Map the initial feature vectors to the same target dimension to obtain the mapped initial feature vectors. Since the dimensions and distributions of the initial feature vectors of the three modalities are very different (e.g., Ftext is 40-dimensional, Fvoice is 768-dimensional, and Fimg is 2048-dimensional), they cannot be directly fused. Therefore, it is necessary to use a linear transformation (e.g., using a fully connected layer) to map the initial feature vectors of the three modalities to the same target dimension (assuming the target dimension is d) to ensure feature additivity or splicing.
[0141] For the initial feature vector of each modality, a corresponding fully connected layer can be designed to complete the mapping:
[0142] The speech feature vector Fvoice∈Rn1 is mapped to Fvoice'∈Rd:
[0143] Fvoice' = W1·Fvoice + b1;
[0144] The text feature vector Ftext∈Rn2 is mapped to Ftext'∈Rd:
[0145] Ftext' = W2·Ftext + b2;
[0146] Image feature vectors Fimg∈Rn3 are mapped to Fimg'∈Rd:
[0147] Fimg' = W3·Fimg + b3;
[0148] Where W1∈Rd×n1, W2∈Rd×n2, W3∈Rd×n3 are attention weights, and b1,b2,b3∈Rd are bias terms, all of which are learnable parameters.
[0149] Here, the selection of the target dimension d needs to balance expressive power and computational efficiency. If the target dimension d is too small (e.g., 32, 64), important information may be lost, resulting in poor fusion performance; if the target dimension d is too large (e.g., 1024, 2048), it may increase model parameters and the risk of overfitting, increasing computational costs. Therefore, the size of the target dimension d can be selected based on the complexity of the product recommendation task. For example, for sentiment analysis tasks, the target dimension d = 128 or 256; for complex multimodal understanding tasks, the target dimension d = 512.
[0150] S3: Calculate the correlation degree between the initial feature vector and the context vector at the first time step using a Multi-Layer Perceptron (MLP). Here, we use... These represent the correlation between the initial feature vector corresponding to voice interaction information, the initial feature vector corresponding to text interaction information, and the initial feature vector corresponding to image interaction information and the context vector at the first time point, respectively.
[0151] It should be noted that, through data-driven supervised training, the correlation strength between each modality and the task objective (relevance or intent) is automatically learned under the goal of minimizing prediction loss. During the training process, the MLP will continuously try and fail: if the initial feature vector of a certain modality can help the MLP to more accurately judge the relevance or user intent, the MLP will increase the weight of that modality through backpropagation; otherwise, it will decrease the weight. Finally, the trained MLP can output the relevance reflecting the true importance of each modality based on the initial feature vector of the input multimodal modalities, as shown in the following formula (1):
[0152] (1);
[0153] in, Perform calculations for voice, text, and img respectively; This refers to the context vector, such as the hidden state of the context vector at the second time step. Indicates feature splicing;
[0154] S4: The above correlations are transformed into attention weights using the Softmax function (a commonly used activation function in deep learning), thus obtaining the attention weights corresponding to each modality. , , Furthermore, the sum of the attention weights for the three modalities is 1, i.e. The magnitude of each attention weight reflects the importance of the corresponding modality to the current task, as shown in formulas (2) to (4) below:
[0155] (2);
[0156] (3);
[0157] (4);
[0158] Here, the Softmax function is mainly used in the output layer of multi-class tasks. It can convert the correlation of the MLP output into a numerical value representing the probability distribution, so that the attention weight corresponding to each modality of the output is in the range of [0,1], and the sum of all attention weights is 1.
[0159] S5: Perform feature fusion based on the attention weights and initial feature vectors corresponding to each modality to obtain the required feature vector. As shown in the following formula (5):
[0160] (5).
[0161] In one embodiment, optionally, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0162] Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained, including:
[0163] Based on the risk-product type mapping library in the product demand knowledge base, as well as the product demand analysis results and product demand interaction information, the user's first-time product recommendation type is obtained; wherein, the risk-product type mapping library is used to indicate the mapping relationship between different risks and product types existing in the user;
[0164] Based on the facial feature-product limit mapping library in the product demand knowledge base, as well as the product demand analysis results and the product demand interaction information, the user's first-time product recommendation limit is obtained; wherein, the facial feature-product limit mapping library is used to indicate the mapping relationship between the user's different facial features and product limits;
[0165] Based on the requirement-product usage period mapping library in the product requirement knowledge base, as well as the product requirement analysis results and product requirement interaction information, the user's first-time recommended product usage period is obtained; wherein, the requirement-product usage period mapping library is used to indicate the mapping relationship between the user's different product requirements and product usage periods;
[0166] Based on the product recommendation type, the product recommendation limit, and the product recommendation usage period, the user's first-time product demand analysis results are obtained.
[0167] In this embodiment of the invention, the user's immediate product demand analysis results include the product recommendation type, the product recommendation limit, and the product recommendation usage period, thereby enabling quantifiable product demand output. Since the product recommendation method described in this embodiment can be applied to insurance product recommendation scenarios, the product recommendation type can be the insurance product's coverage type, the product recommendation limit can be the insurance product's coverage limit, and the product recommendation usage period can be the insurance product's coverage period.
[0168] In the insurance product recommendation scenario, the risk-product type mapping library can indicate the mapping relationship between different risks and protection types of the user; the facial feature-product limit mapping library can indicate the mapping relationship between different facial features and protection limits of the user; and the demand-product usage period mapping library can indicate the mapping relationship between different product demands and protection periods of the user.
[0169] For example, if the user's initial product demand interaction information includes voice interaction indicating fear of accidents, text interaction indicating frequent business trips, and image interaction indicating being 25 years old and energetic (i.e., without obvious health problems), then the user's risk is determined to be that of a young business traveler without obvious health problems. The corresponding insurance type for this young business traveler without obvious health problems is then obtained from the risk-product type mapping library, which is accident insurance. Alternatively, the corresponding insurance type for this young business traveler without obvious health problems can be obtained from the risk-product type mapping library, sorted by priority from high to low as accident insurance, comprehensive medical insurance, and short-term critical illness insurance.
[0170] For example, if the text interaction information in the product demand interaction information at the first time is "want to report an accident of 1 million", and the image interaction information is "a business travel user aged 25 to 30", then it is determined that the user's facial features are those of a business travel user aged 25 to 30, and the corresponding protection limit for the accident insurance product demand is obtained from the facial feature and product limit mapping library, which is 1 million for accident insurance and 3 million for medical insurance.
[0171] For example, if the image interaction information in the product demand interaction information at the first time is that of a business travel user aged 25 to 30, then the user's product demand is determined to be an accident insurance demand, and the corresponding coverage period for the accident insurance demand is obtained from the demand-product usage period mapping library, which is 3 months; if the image interaction information in the product demand interaction information at the first time is that of a student user, then the user's different product demand is determined to be a medical insurance demand, and the corresponding coverage period for the medical insurance demand is obtained from the demand-product usage period mapping library, which is 1 year.
[0172] In one implementation method, optionally, based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database, including:
[0173] Based on the product demand analysis results, the product demand analysis feature vector of the user is obtained;
[0174] In terms of the target dimension, obtain the product feature vector of each product in the product database, where the target dimension is related to the product demand analysis results;
[0175] Based on the matching degree between the product demand analysis feature vector and the product feature vector, multiple candidate recommended products are matched for the user in the product database.
[0176] In this embodiment of the invention, since the product recommendation method described in this embodiment can be applied to insurance product recommendation scenarios, the product database can be an insurance product database. This insurance product database stores information on various insurance products, including but not limited to: product terms; coverage; premium; and claims conditions. Each insurance product in this insurance product database is represented as a product feature vector Pi=[pi1,pi2,...,pit], where pit represents the feature value of the i-th insurance product in the t-th target dimension. The target dimension includes but is not limited to: premium amount; coverage period; coverage matching score; and claims success rate.
[0177] Based on the product demand analysis results, the user's product demand analysis feature vector U=[u1,u2,...,ut] is obtained, where ut represents the user's feature value in the t-th demand dimension, which includes, but is not limited to: expected coverage amount; expected coverage period; and weight of the type of coverage preferred.
[0178] The cosine similarity algorithm is used to calculate the matching degree between each insurance product in the insurance product database and the user's product demand analysis feature vector. As shown in the following formula (6):
[0179] (6);
[0180] Among them, matching degree The result range is [-1, 1], based on the matching degree. Sort the insurance products in the insurance product database and filter them based on their matching degree. Higher-priced insurance products are recommended as candidate products.
[0181] Figure 2 This is a flowchart illustrating one embodiment of the product recommendation method described in this invention. Figure 2 As shown, the process involves collecting multimodal product demand interaction information, preprocessing the product demand interaction information, fusing the initial feature vectors corresponding to each modality to obtain demand feature vectors, initial user demand analysis results, matching candidate recommended products, formulating the recommendation strategy network, presenting the target recommended product in a multimodal manner, and finally updating and optimizing the recommendation strategy network based on user feedback and re-analyzing user demands to form a complete closed-loop recommendation process, ensuring the accuracy and personalization of product recommendations.
[0182] In summary, the product recommendation method described in this invention is essentially a product recommendation method based on multimodal human-computer dialogue. This method deeply integrates multimodal product demand interaction information: it deeply integrates voice, text, and image multimodal product demand interaction information through linear transformation to unify dimensions and dynamic attention weighting. This fully utilizes the complementarity of different modal product demand interaction information, such as voice emotion, text requirements, and image expressions, to more accurately analyze user product needs. Compared to existing single-modal or simple fusion methods, it can obtain more comprehensive and accurate product demand analysis results, improving the comprehensiveness and accuracy of product demand analysis result extraction. Combined with user needs... By analyzing models and product demand knowledge bases, we gain a deep understanding of users' potential needs. Simultaneously, we employ a recommendation strategy network based on product recommendation rules and reinforcement learning, dynamically considering user risk preferences to achieve personalized product recommendations. This improves recommendation accuracy and user acceptance, significantly enhancing the personalization of recommendations. Through the linkage between user feedback and recommendation strategy network optimization, we continuously adjust the user demand analysis model, matching algorithm, and recommendation strategy network, adapting to changes in different user groups and business scenarios, achieving sustainable optimization, and maintaining high recommendation accuracy in the long term. Explicit algorithms are used in the integration of attention weights, matching degrees, and reward mechanisms, making the product recommendation process interpretable and reproducible, reducing the influence of subjective factors, and enhancing its scientific rigor.
[0183] See Figure 3 This invention also provides a product recommendation device, comprising:
[0184] Matching module 301 is used to match multiple candidate recommended products for the user in the product database based on the user's first-time product demand interaction information;
[0185] The determining module 302 is used to determine the target recommended product for the user in the first time from among a plurality of candidate recommended products based on the recommendation strategy network and the user's second time-time recommended product information; wherein, the parameters of the recommendation strategy network are obtained by updating a first reward value obtained from the user's first feedback information on the recommended product in the second time-time; the second time-time is before the first time-time.
[0186] Optionally, in the product recommendation device, if the first feedback information is used to indicate that the user has not rejected the recommended product at the second time, and / or the type of the recommended product at the second time matches the user's risk preference type, then the first reward value is positive.
[0187] If the first feedback information is used to indicate that the user did not accept the recommended product at the second time, and / or the type of the recommended product at the second time is the same as the type of recommended product that the user repeatedly did not accept, then the first reward value is negative.
[0188] Optionally, the product recommendation device further includes:
[0189] The acquisition module is used to acquire the user's second feedback information on the target recommended product at the first time.
[0190] The reward module is used to obtain a second reward value based on the second feedback information;
[0191] An update module is used to update the parameters of the recommendation strategy network according to the second reward value, so as to obtain the updated recommendation strategy network;
[0192] The recommendation module is used to determine the target recommended product for the user in the third time period based on the updated recommendation strategy network and the user's product demand interaction information in the third time period; wherein the third time period is after the first time period.
[0193] Optionally, in the product recommendation device, the matching module 301 includes:
[0194] The first acquisition unit is used to obtain the demand feature vector based on the user's immediate product demand interaction information;
[0195] The second obtaining unit is used to obtain the user's initial product demand analysis results at the first moment based on the user demand analysis model and the demand feature vector;
[0196] The third obtaining unit is used to obtain the user's first-time product requirement analysis results based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information.
[0197] The matching unit is used to match multiple candidate recommended products for the user in the product database based on the product demand analysis results.
[0198] Optionally, in the product recommendation device, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0199] The first obtaining unit is specifically used for:
[0200] Feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, which includes at least two of the following: voice feature vector, text feature vector, and image feature vector.
[0201] The initial feature vector is mapped to the target dimension to obtain the mapped initial feature vector;
[0202] Based on the correlation between the initial feature vector and the context vector at the first time, the attention weight corresponding to the initial feature vector is obtained;
[0203] Feature fusion is performed based on the attention weights and the initial feature vector to obtain the required feature vector.
[0204] Optionally, in the product recommendation device, the product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information.
[0205] The second obtaining unit is specifically used for:
[0206] Based on the risk-product type mapping library in the product demand knowledge base, as well as the product demand analysis results and product demand interaction information, the user's first-time product recommendation type is obtained; wherein, the risk-product type mapping library is used to indicate the mapping relationship between different risks and product types existing in the user;
[0207] Based on the facial feature-product limit mapping library in the product demand knowledge base, as well as the product demand analysis results and the product demand interaction information, the user's first-time product recommendation limit is obtained; wherein, the facial feature-product limit mapping library is used to indicate the mapping relationship between the user's different facial features and product limits;
[0208] Based on the requirement-product usage period mapping library in the product requirement knowledge base, as well as the product requirement analysis results and product requirement interaction information, the user's first-time recommended product usage period is obtained; wherein, the requirement-product usage period mapping library is used to indicate the mapping relationship between the user's different product requirements and product usage periods;
[0209] Based on the product recommendation type, the product recommendation limit, and the product recommendation usage period, the user's first-time product demand analysis results are obtained.
[0210] Optionally, in the product recommendation device, the matching unit is specifically used for:
[0211] Based on the product demand analysis results, the product demand analysis feature vector of the user is obtained;
[0212] In terms of the target dimension, obtain the product feature vector of each product in the product database, where the target dimension is related to the product demand analysis results;
[0213] Based on the matching degree between the product demand analysis feature vector and the product feature vector, multiple candidate recommended products are matched for the user in the product database.
[0214] Optionally, in the product recommendation device, the determining module 302 is specifically used for:
[0215] Based on the recommendation strategy network, product recommendation rules, and the user's second-time recommended product information, the probability distribution corresponding to each candidate recommended product is predicted; wherein, the product recommendation rules are formulated according to different product demand scenarios;
[0216] Based on the probability distribution corresponding to the candidate recommended products, the target recommended product for the user at the first time is determined from among the multiple candidate recommended products.
[0217] It should be noted that the apparatus provided in the embodiments of the present invention can implement all the method steps implemented in the above-mentioned product recommended method embodiments and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.
[0218] This invention also provides a product recommendation device, such as... Figure 4 As shown, it includes:
[0219] The processor 401, memory 402, transceiver 403, and a program or instructions stored in the memory 402 and executable on the processor 401; when the processor 401 executes the program or instructions, it implements the various processes of the above-described product recommended method embodiments and achieves the same technical effect. To avoid repetition, these will not be described again here.
[0220] The transceiver 403 is used to receive and send data under the control of the processor 401.
[0221] Among them, Figure 4 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically connecting various circuits of one or more processors represented by processor 401 and memory represented by memory 402. The bus architecture can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. Transceiver 403 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. For different user equipment, the user interface 404 can also be an interface capable of connecting external or internal devices, including but not limited to keypads, displays, speakers, microphones, joysticks, etc.
[0222] The processor 401 is responsible for managing the bus architecture and general processing, while the memory 402 can store the data used by the processor 401 when performing operations.
[0223] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the product recommendation method embodiments described above and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0224] This invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described product recommended method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0225] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0226] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0227] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these modifications are within the protection scope of the present invention.
Claims
1. A product recommendation method, characterized in that, include: Based on the user's immediate product demand interaction information, multiple candidate recommended products are matched for the user in the product database; Based on the recommendation strategy network and the user's second-time recommended product information, a target recommended product for the user at the first time is determined from among multiple candidate recommended products; wherein, the parameters of the recommendation strategy network are updated based on a first reward value obtained from the user's first feedback information on the recommended product at the second time; the second time is before the first time.
2. The method according to claim 1, characterized in that, If the first feedback information is used to indicate that the user did not reject the recommended product at the second time, and / or the type of the recommended product at the second time matches the user's risk preference type, then the first reward value is positive; If the first feedback information is used to indicate that the user did not accept the recommended product at the second time, and / or the type of the recommended product at the second time is the same as the type of recommended product that the user repeatedly did not accept, then the first reward value is negative.
3. The method according to claim 1, characterized in that, The method further includes: Obtain the user's second feedback information on the target recommended product at the first time. Based on the second feedback information, the second reward value is obtained; The parameters of the recommendation strategy network are updated according to the second reward value to obtain the updated recommendation strategy network; Based on the updated recommendation strategy network and the user's product demand interaction information at a third time, a target recommended product for the user at the third time is determined; wherein, the third time is after the first time.
4. The method according to claim 1, characterized in that, Based on the user's immediate product request interaction information, multiple candidate recommended products are matched for the user in the product database, including: Based on the user's immediate product demand interaction information, obtain the demand feature vector; Based on the user demand analysis model and the demand feature vector, the initial product demand analysis results of the user at the first moment are obtained; Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained. Based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database.
5. The method according to claim 4, characterized in that, The product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information; Based on the user's immediate product request interaction information, obtain the demand feature vector, including: Feature extraction is performed on the user's first-time product demand interaction information to obtain an initial feature vector, which includes at least two of the following: voice feature vector, text feature vector, and image feature vector. The initial feature vector is mapped to the target dimension to obtain the mapped initial feature vector; Based on the correlation between the initial feature vector and the context vector at the first time, the attention weight corresponding to the initial feature vector is obtained; Feature fusion is performed based on the attention weights and the initial feature vector to obtain the required feature vector.
6. The method according to claim 4, characterized in that, The product demand interaction information includes at least two of the following: voice interaction information, text interaction information, and image interaction information; Based on the product requirement knowledge base, the initial product requirement analysis results, and the product requirement interaction information, the user's first-time product requirement analysis results are obtained, including: Based on the risk-product type mapping library in the product demand knowledge base, as well as the product demand analysis results and product demand interaction information, the user's first-time product recommendation type is obtained; wherein, the risk-product type mapping library is used to indicate the mapping relationship between different risks and product types existing in the user; Based on the facial feature-product limit mapping library in the product demand knowledge base, as well as the product demand analysis results and the product demand interaction information, the user's first-time product recommendation limit is obtained; wherein, the facial feature-product limit mapping library is used to indicate the mapping relationship between the user's different facial features and product limits; Based on the requirement-product usage period mapping library in the product requirement knowledge base, as well as the product requirement analysis results and product requirement interaction information, the user's first-time recommended product usage period is obtained; wherein, the requirement-product usage period mapping library is used to indicate the mapping relationship between the user's different product requirements and product usage periods; Based on the product recommendation type, the product recommendation limit, and the product recommendation usage period, the user's first-time product demand analysis results are obtained.
7. The method according to claim 6, characterized in that, Based on the product demand analysis results, multiple candidate recommended products are matched for the user in the product database, including: Based on the product demand analysis results, the product demand analysis feature vector of the user is obtained; In terms of the target dimension, obtain the product feature vector of each product in the product database, where the target dimension is related to the product demand analysis results; Based on the matching degree between the product demand analysis feature vector and the product feature vector, multiple candidate recommended products are matched for the user in the product database.
8. The method according to claim 1, characterized in that, Based on the recommendation strategy network and the user's second-time recommended product information, the target recommended product for the user in the first time is determined from among multiple candidate recommended products, including: Based on the recommendation strategy network, product recommendation rules, and the user's second-time recommended product information, the probability distribution corresponding to each candidate recommended product is predicted; wherein, the product recommendation rules are formulated according to different product demand scenarios; Based on the probability distribution corresponding to the candidate recommended products, the target recommended product for the user at the first time is determined from among the multiple candidate recommended products.
9. A product recommendation device, characterized in that, include: The matching module is used to match multiple candidate recommended products for the user in the product database based on the user's first-time product demand interaction information. The determining module is configured to determine the target recommended product for the user at the first time from among a plurality of candidate recommended products based on the recommendation strategy network and the user's recommended product information at the second time; wherein, the parameters of the recommendation strategy network are obtained by updating a first reward value based on the user's first feedback information on the recommended product at the second time; the second time is before the first time.
10. A product recommendation device, characterized in that, include: A processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the processor, when executing the program or instructions, implements the product recommendation method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the product recommendation method as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the product recommendation method as described in any one of claims 1 to 8.