Commodity recommendation model training method, commodity recommendation method and related device

By generating semantic-similar incremental product recommendation sessions and product recommendation logics with limited real product recommendation data, the problem of insufficient generalization ability of the recommendation model is solved and stronger generalization ability of the product recommendation model is achieved.

CN120180134AActive Publication Date: 2025-06-20ALIBABA HEALTH TECH (CHINA) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510527862.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-20
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

When the amount of real product recommendation data is limited, the recommended model trained has the problem of insufficient generalization ability.

Method used

By obtaining real product recommendation sessions, generate incremental product recommendation sessions with similar semantics and their corresponding product recommendation logic, and use these data as the basis for model training to improve the generalization ability of the model.

Benefits of technology

When the real product recommendation session is limited, through incremental data generation and logical derivation, the trained product recommendation model can have stronger generalization capabilities and can more accurately recommend products that meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180134A_ABST
    Figure CN120180134A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a commodity recommendation model training method, a commodity recommendation method and a related device. The commodity recommendation model training method specifically comprises the steps of obtaining a real commodity recommendation session; generating an incremental commodity recommendation session based on the real commodity recommendation session, and generating a commodity recommendation logic corresponding to the incremental commodity recommendation session; wherein the semantic similarity between the incremental commodity recommendation session and the real commodity recommendation session conforms to expectation; the commodity recommendation logic is used for releasing recommendation reasons for commodities in the incremental commodity recommendation session according to a logic derivation mode; and training a commodity recommendation model based on the real commodity recommendation session, the incremental commodity recommendation session and the commodity recommendation logic. By implementing the embodiment of the invention, the generalization ability of the commodity recommendation model can be improved under the condition that the data volume for model training is limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in the present application relate to the field of artificial intelligence technology, and particularly to a method for training a product recommendation model, a product recommendation method, and related devices. Background Art

[0002] In order to meet diverse user needs, e-commerce platforms provide product recommendation services. When users are unsure which product to purchase, the product recommendation service can be triggered to recommend products that meet the user's needs according to the user's needs.

[0003] In traditional product recommendation methods, when a recommendation requirement is detected, a pre-trained recommendation model is called to generate product recommendation information. The generalization ability of the recommendation model is very important for the accuracy of model recommendations. The generalization ability refers to whether the recommendation model can also achieve accurate recommendations for user needs when faced with user needs that have not been learned from the samples. In order to enable the recommendation model to have stronger generalization ability, it is usually necessary to train the recommendation model using a large amount of real product recommendation data.

[0004] However, in the case where the amount of real product recommendation data is limited, the trained recommendation model has the problem of insufficient generalization ability. Summary of the Invention

[0005] In view of this, multiple embodiments of the present application are committed to providing a method for training a product recommendation model and related devices, which can improve the generalization ability of the product recommendation model when the amount of data used for model training is limited.

[0006] An embodiment of the present application provides a method for training a product recommendation model, the method including: obtaining real product recommendation sessions; generating incremental product recommendation sessions based on the real product recommendation sessions, and generating product recommendation logics corresponding to the incremental product recommendation sessions; wherein, the semantic similarity between the incremental product recommendation sessions and the real product recommendation sessions meets the expectation; the product recommendation logics are used to explain the recommendation reasons for the products in the incremental product recommendation sessions in a logical derivation manner; training a product recommendation model based on the real product recommendation sessions, the incremental product recommendation sessions, and the product recommendation logics.

[0007] An embodiment of the present application further provides a product recommendation method, the method including: when receiving consultation information, calling the product recommendation model in an embodiment of the present application, so that the product recommendation model generates product recommendation information and product recommendation logics corresponding to the consultation information; wherein, the product recommendation logics are used to explain the recommendation reasons for the products in the product recommendation information in a logical derivation manner.

[0008] An embodiment of the present application further provides a training device for a commodity recommendation model. The device includes: a session acquisition module, configured to acquire real commodity recommendation sessions; an incremental data generation module, configured to generate incremental commodity recommendation sessions based on the real commodity recommendation sessions, and generate commodity recommendation logics corresponding to the incremental commodity recommendation sessions; wherein, the semantic similarity between the incremental commodity recommendation sessions and the real commodity recommendation sessions meets the expectation; the commodity recommendation logics are used to explain the recommendation reasons for the commodities in the incremental commodity recommendation sessions in a logical derivation manner; a model training module, configured to train a commodity recommendation model based on the real commodity recommendation sessions, the incremental commodity recommendation sessions and the commodity recommendation logics.

[0009] An embodiment of the present application further provides a commodity recommendation device. The device includes: a commodity recommendation module, configured to, when receiving consultation information, call the commodity recommendation model in the training device, so that the commodity recommendation model generates commodity recommendation information and commodity recommendation logics corresponding to the consultation information; wherein, the commodity recommendation logics are used to explain the recommendation reasons for the commodities in the commodity recommendation information in a logical derivation manner.

[0010] An embodiment of the present application further provides a computer device. The computer device includes a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method as described above.

[0011] An embodiment of the present application further provides a computer-readable storage medium. At least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, it can implement the method as described above.

[0012] An embodiment of the present application further provides a computer program product, which is used to implement the method as described above.

[0013] In multiple embodiments provided by the present application, incremental commodity recommendation sessions semantically similar to real commodity recommendation sessions and commodity recommendation logics corresponding to the incremental commodity recommendation sessions can be generated based on the real commodity recommendation sessions. Since the commodity recommendation logics explain the recommendation reasons for the commodities in the incremental commodity recommendation sessions, therefore, using the real commodity recommendation sessions, the incremental commodity recommendation sessions and the commodity recommendation logics as the basis for model training can enable the commodity recommendation model to be sufficiently trained and learn the commodity recommendation logics from them when the real commodity recommendation sessions are limited, thereby improving the generalization ability of the commodity recommendation model. Description of the Drawings

[0014] Figure 1Schematic diagram of a system architecture for implementing a training method of a product recommendation model provided for an embodiment of the present application.

[0015] Figure 2 Flowchart of a training method of a product recommendation model provided for an embodiment of the present application.

[0016] Figure 3 Schematic diagram of modules of a training device of a product recommendation model provided for an embodiment of the present application.

[0017] Figure 4 Schematic diagram of modules of a product recommendation device provided for an embodiment of the present application.

[0018] Figure 5 Schematic diagram of a computer device provided for an embodiment of the present application. Specific embodiments

[0019] The following will clearly and completely describe the information to be retrieved in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0020] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0021] An e-commerce platform provides users with the function of purchasing products online. However, in some cases, users have a purchasing need but no clear purchasing direction. To meet the diverse purchasing needs of users, the e-commerce platform provides a product recommendation function. Among them, the product recommendation function can exist in the form of an intelligent customer service, a shopping guide assistant, etc., and can recommend products that can meet their needs (such as XX brand simple swivel chairs) based on the consultation information input by users (such as, which office chair is better to choose in a narrow work station). The core of the product recommendation function is the recommendation model. The generalization ability of the recommendation model depends on the data volume of the training samples. The larger the data volume, the easier it is to train a recommendation model with strong generalization ability. A recommendation model with strong generalization ability can also use the learned recommendation logic when dealing with consultation information that has not been learned, accurately understand the semantics of the consultation information, and recommend products that match the consultation information. In the case where the data volume of real product recommendation data is limited, the trained recommendation model has the problem of insufficient generalization ability.

[0022] The means adopted by related technologies to improve the generalization ability of the model lies in: increasing the data volume of training samples. To obtain a sufficient amount of samples, related technologies use at least one of the following methods to collect data as samples: intelligent customer service repeatedly asks for information related to the user's needs, requests information on the user's expressed needs from other platforms, and obtains historical user search records / historical user shopping records / historical user browsing records. Obviously, the above methods require consuming a large amount of computing resources or communication resources to obtain a sufficient amount of samples to train the model, so that the generalization ability of the model meets the expectations.

[0023] Therefore, it is necessary to provide a training method for a commodity recommendation model, which can generate an incremental commodity recommendation session semantically similar to a real commodity recommendation session based on the real commodity recommendation session, as well as a commodity recommendation logic corresponding to the incremental commodity recommendation session. Since the commodity recommendation logic explains the reasons for recommending the commodities in the incremental commodity recommendation session, therefore, using the real commodity recommendation session, the incremental commodity recommendation session and the commodity recommendation logic as the basis for model training can enable the commodity recommendation model to be sufficiently trained and learn the commodity recommendation logic from it under the condition that the real commodity recommendation session is limited, thereby improving the generalization ability of the commodity recommendation model.

[0024] Please refer to Figure 1 . In multiple implementation manners provided in this application, the training method for the commodity recommendation model can be applied to a training device for the commodity recommendation model. The training device for the commodity recommendation model can be an electronic device with certain computing capabilities and network access capabilities. The electronic device can be a desktop computer, a laptop computer, a tablet computer, or a server. The electronic device can be connected to the server through a network. The server can be a distributed server, including multiple processors, memories, network communication modules, etc., which cooperate to realize various functions. Or, the server can also be a server cluster formed by several servers, with higher computing and data processing capabilities. With the development of science and technology, the server can also be realized by new forms of technical means, such as a new type of "server" based on quantum computing. Of course, in some implementation manners, the training device for the commodity recommendation model can also be a program module running in the electronic device.

[0025] Specifically, the server is used to receive the real commodity recommendation session sent by the electronic device, then generate an incremental commodity recommendation session from the real commodity recommendation session, and generate a commodity recommendation logic corresponding to the incremental commodity recommendation session, and train the commodity recommendation model based on the real commodity recommendation session, the incremental commodity recommendation session and the commodity recommendation logic. Furthermore, when the electronic device receives the consultation information, it sends the consultation information to the server, so that the server calls the commodity recommendation model, so that the commodity recommendation model generates commodity recommendation information and commodity recommendation logic corresponding to the consultation information.

[0026] In some embodiments, an electronic device / server is used to obtain a real product recommendation session, generate an incremental product recommendation session based on the real product recommendation session, generate product recommendation logic corresponding to the incremental product recommendation session, and train a product recommendation model based on the real product recommendation session, the incremental product recommendation session, and the product recommendation logic. On this basis, the electronic device / server can call the product recommendation model when receiving consultation information, so that the product recommendation model generates product recommendation information and product recommendation logic corresponding to the consultation information.

[0027] An application scenario example of a method for training a product recommendation model is provided in an embodiment of the present application. The method for training the product recommendation model can be applied to the scenario of an intelligent customer service recommending products.

[0028] After obtaining the real product recommendation session {User: "What mop mops the floor clean?"; Intelligent customer service: "Have you ever used a floor washer?"; User: "I've used the AA brand. The water tank is too small and it's very troublesome to add water several times when mopping the floor once."; Intelligent customer service: "I recommend you use the BB brand floor washer. This floor washer is equipped with a 1L large water tank. Or, the CC brand floor washer. This floor washer is equipped with an external backpack water tank in addition to its own water tank, and can carry a total of 2L of water."}

[0029] Furthermore, an incremental product recommendation session {User: "Recommend a floor washer. I have experience using a floor washer and need a floor washer with a larger water-carrying capacity."; Intelligent customer service: "It is recommended to buy the BB brand floor washer equipped with a 1L large water tank, and the CC brand floor washer equipped with its own water tank and an external backpack water tank that can carry 2L of water."} can be generated based on the real product recommendation session. And, generate product recommendation logic corresponding to the incremental product recommendation session to show the logical derivation relationship between the consultation information "Recommend a floor washer. I have experience using a floor washer and need a floor washer with a larger water-carrying capacity." entered by the user and the final output product recommendation information "It is recommended to buy the BB brand floor washer equipped with a 1L large water tank, and the CC brand floor washer equipped with its own water tank and an external backpack water tank that can carry 2L of water."

[0030] Training the product recommendation model based on the real product recommendation session, the incremental product recommendation session, and the product recommendation logic can obtain a product recommendation model that meets the expected generalization ability.

[0031] Deploy the product recommendation model in the intelligent customer service. When receiving the consultation information input by the user, "Which vacuum cleaner is easy to use and has more replacement heads?", the intelligent customer service calls the product recommendation model, so that the product recommendation model analyzes the user's needs in the consultation information, "a vacuum cleaner with multiple replacement heads", based on the learned product recommendation logic, and queries the product "SS brand multi-head vacuum cleaner" that matches it from the product description information, and then can push the "SS brand multi-head vacuum cleaner" to the user to meet the user's needs.

[0032] Please refer to Figure 2 An embodiment of the present application provides a method for training a product recommendation model. The method for training a product recommendation model can be applied to a training device for a product recommendation model. The method for training the product recommendation model may include the following steps.

[0033] Step S110: Obtain a real product recommendation session.

[0034] Step S120: Generate an incremental product recommendation session based on the real product recommendation session, and generate a product recommendation logic corresponding to the incremental product recommendation session; wherein, the semantic similarity between the incremental product recommendation session and the real product recommendation session meets the expectation; the product recommendation logic is used to explain the recommendation reason for the product in the incremental product recommendation session in a logical derivation manner.

[0035] Step S130: Train the product recommendation model based on the real product recommendation session, the incremental product recommendation session, and the product recommendation logic.

[0036] In this embodiment, a real product recommendation session refers to the actual interaction record between a user and an intelligent customer service / human customer service. The real product recommendation session contains consultation information for characterizing user needs (such as, "recent nasal congestion and headache", "which mop mops more cleanly") and product recommendation information for meeting user needs (such as, "recommend using drugs A and B in combination", "recommend trying out a new floor washer"). In some embodiments, the real product recommendation session may also contain communication information for mining user needs (such as, "considering that it is spring recently and the pollen content in the air is high, do you have a history of pollen allergy causing nasal congestion and headache", "what model of mop are you currently using, have you used a floor washer"). The training device of the product recommendation model can obtain real product recommendation sessions from at least one e-commerce platform through a data interface or a database connection method. The real product recommendation session can be used as a basis for model training and can also be used as a basis for generating an incremental product recommendation session.

[0037] In some embodiments, after obtaining a real product recommendation session and before generating an incremental product recommendation session based on the real product recommendation session, the training device of the product recommendation model may also preprocess the real product recommendation session to standardize the format of each real product recommendation session, or remove the noise in the real product recommendation session. Herein, the noise refers to the information considered invalid in the real product recommendation session (such as, "Thank you for your affirmation, dear", "Are you still there, dear").

[0038] In this embodiment, the training device of the product recommendation model may call a pre-trained large language model based on the real product recommendation session, so that the pre-trained large language model rewrites the information in the real product recommendation session (such as, "Suffering from rhinitis for ten years, long-term conditioning is required") to generate an incremental product recommendation session whose semantic similarity with the real product recommendation session meets the expectation (such as, an incremental product recommendation session containing "Suffering from rhinitis for a long time, are there any products suitable for continuous use"), or so that the pre-trained large language model performs semantic understanding on the real product recommendation session and generates an incremental product recommendation session similar to the real product recommendation session according to the product knowledge learned online. Herein, the incremental product recommendation session is a fictional product recommendation session, and its form is the same as that of the real product recommendation session, that is, the incremental product recommendation session may also include any one of the following information: consultation information for characterizing user needs, product recommendation information for meeting user needs, communication information for mining user needs, etc.

[0039] In some embodiments, the training device of the product recommendation model may also call a semantic similarity model to verify the semantic similarity between the real product recommendation session and the incremental product recommendation session to ensure that the semantic similarity between the real product recommendation session and the incremental product recommendation session meets the expectation. Herein, the expectation may include a preset similarity (such as, 0.8). When the semantic similarity is greater than or equal to the preset similarity, it can be considered that the semantic similarity meets the expectation; wherein, the semantic similarity can be understood as the similarity between the full texts of the real product recommendation session and the incremental product recommendation session, or can be understood as the similarity of the main ideas between the real product recommendation session and the incremental product recommendation session. The main idea is used to reflect the core topic discussed in the session.

[0040] In some embodiments, the training device of the product recommendation model may generate multiple incremental product recommendation sessions based on one real product recommendation session to increase the sample size for model training, or may combine multiple real product recommendation sessions to generate an incremental product recommendation session that integrates the user needs in multiple real product recommendation sessions to enrich the sample diversity for model training.

[0041] In this embodiment, in order to improve the usability of the incremental product recommendation session as a model training sample, the training device of the product recommendation model can also call a pre-trained logical reasoning model based on the incremental product recommendation session, so that the logical reasoning model generates a product recommendation logic corresponding to the incremental product recommendation session. The product recommendation logic can be used to represent the process by which the large model thinks from the consultation information in the incremental product recommendation session to obtain the product recommendation information. Optionally, the form of the product recommendation logic can be embodied as text, and this text can be displayed on the user interface when triggered to be shown. The product recommendation logic explains the recommendation reasons for the products in the incremental product recommendation session according to the logical derivation method. Optionally, the logical derivation method defines a user demand analysis process (for example, the user has long-term rhinitis and may need long-term treatment), a product knowledge retrieval process (for example, finding a multi-course product set suitable for long-term conditioning of rhinitis), and a recommendation reason generation process (for example, recommending product A because it contains a 3-month course and is suitable for long-term use). The product recommendation logic executes the above processes in the above order to obtain the recommendation reasons.

[0042] In this embodiment, the product recommendation logic serves as a model training sample, and the model can learn how to extract user demands from the incremental product recommendation session and how to use product knowledge to search for products that can meet the user demands from the product library. In the case where the number of real product recommendation sessions is limited, incremental product recommendation sessions can be generated as model training samples to increase the model training volume. In order to enable the model to have the expected generalization ability, the product recommendation logic corresponding to the incremental product recommendation session is also used as a model training sample. This can enable the model to not only receive sufficient training but also learn the product recommendation logic from the training. Compared with the related technology that generally improves the model generalization ability by training the model with a large amount of labeled data, this application provides corresponding product recommendation logic and incremental product recommendation sessions as model training samples based on real product recommendation sessions, which can supplement the model training volume while ensuring that a product recommendation model with the expected generalization ability is obtained.

[0043] In some embodiments, the pre-trained logical reasoning model can also be called based on real product recommendation sessions, so that the logical reasoning model generates product recommendation logic corresponding to the real product recommendation sessions as model training samples to further improve the model generalization ability.

[0044] In this embodiment, the trained product recommendation model can be deployed in an intelligent customer service scenario. When detecting the consulting information input by the user (e.g., "Which contact lenses are better for dry eyes?"), the product recommendation model can be invoked based on this consulting information. Since the product recommendation model has learned the product recommendation logic, it can, based on the product recommendation logic, infer the user's needs corresponding to the consulting information and the products that meet the user's needs, and generate product recommendation information (e.g., "It is recommended to use XX brand contact lenses. This type of contact lens has a relatively high oxygen permeability and a relatively low water content, and will not aggravate the symptoms of dry eyes.") and provide it to the user.

[0045] In this embodiment, an incremental product recommendation session semantically similar to the real product recommendation session and the product recommendation logic corresponding to the incremental product recommendation session can be generated. Since the product recommendation logic explains the reasons for recommending products in the incremental product recommendation session, using the real product recommendation session, the incremental product recommendation session, and the product recommendation logic as the basis for model training can, when the real product recommendation sessions are limited, enable the product recommendation model to be sufficiently trained and learn the product recommendation logic from them, thereby ensuring that the product recommendation model has the expected generalization ability.

[0046] In some embodiments, the training device of the product recommendation model can extract product description information from product images; wherein, the product images include at least one of the following: product display pictures, product detail page images, product application scenario pictures, product usage effect pictures; and train the product recommendation model based on the product recommendation logic, the product description information, the real product recommendation session, and the incremental product recommendation session.

[0047] In this embodiment, in order to improve the recommendation accuracy of the product recommendation model, the training device of the product recommendation model can extract product description information as the input for model training, so that the product recommendation model can use comprehensive product description information as the basis when making product recommendations, which is conducive to improving the product recommendation accuracy. Specifically, there are a large number of product images stored in the e-commerce platform. One product corresponds to at least one product image. The product image can be a product display picture for showing the standardized appearance of the product (e.g., the outer packaging of drugs, the physical picture of medical devices), a product detail page image for characterizing the functions of the product (e.g., the ingredient list of drugs, the screenshot of the user manual), a product application scenario picture for indicating the product in an actual usage scenario (e.g., a schematic diagram of a user using a nebulizer to treat rhinitis), or a product usage effect picture for showing the comparison of the product before and after use (e.g., a comparison picture of the skin condition before and after using skin care products).

[0048] In this embodiment, the training device of the product recommendation model can preprocess the product image (e.g., denoising, adjusting the resolution, etc.), and then call the optical character recognition model to extract the text information in the product image as the product description information (e.g., "Suitable for chronic rhinitis"), or call the multimodal model to extract the visual features of the image (e.g., the visual style of the drug packaging, the user behavior features in the scene graph) as the product description information. Among them, the product description information can characterize the product from one or more dimensions (e.g., efficacy, treatment course, applicable population).

[0049] In some embodiments, the training device of the product recommendation model can associate the product description information with the product SKU to form a product knowledge base for the product recommendation model to call. For example, the product knowledge base includes: {Product A: Ingredients (Cetirizine Hydrochloride), Applicable population (Patients with chronic rhinitis), Treatment course (3 months)}.

[0050] In this embodiment, the training device of the product recommendation model can input the product recommendation logic, the product description information, the real product recommendation session, and the incremental product recommendation session into the product recommendation model, so that the product recommendation model can learn the product recommendation logic and query the product that can be used to meet the user's needs from the product description information, and push the product to the user in the form of product recommendation information.

[0051] In some embodiments, the training device of the product recommendation model can perform text recognition on the product image to obtain a text recognition result; based on the text recognition result, call a multimodal evaluation model to enable the multimodal evaluation model to generate an evaluation result for characterizing the relationship between the text recognition result and the product image; in the case where the evaluation result indicates that the product image contains the text recognition result, recognize the text recognition result as the product description information.

[0052] In this embodiment, in order to ensure the accuracy of the product description information as the model input and to combat model hallucinations, the training device of the product recommendation model can use the OCR algorithm to extract text from the product image, and call a pre-trained large model to perform semantic parsing on the non-text information (such as charts, symbols) in the product image to obtain a text recognition result. Furthermore, based on the text recognition result and the product image, call a multimodal evaluation model to enable the multimodal evaluation model to generate an evaluation result characterizing the relationship between the two. Among them, the evaluation result is used to determine whether the text recognition result truly reflects the content of the product image, and based on the evaluation result, high-confidence product description information can be screened out from the text recognition result.

[0053] In some embodiments, the training device of the product recommendation model may call a multimodal evaluation model based on the text recognition result, so that the multimodal evaluation model generates a mapping of the text recognition result to text semantic features and a mapping of the product image to image semantic features, generates a cross-vector space representing the feature matching relationship between the text semantic features and the image semantic features, and combines the fusion result of the text semantic features, the image semantic features, and the cross-vector space to generate an evaluation result of the relationship between the text recognition result and the product image.

[0054] In this embodiment, in order to obtain an accurate evaluation of the relationship between the text recognition result and the product image, the training device of the product recommendation model may call a multimodal evaluation model based on the text recognition result, so that the multimodal evaluation model maps / encodes the text recognition result into text semantic features through a feature mapping unit, which can also be referred to as a text semantic space. For example, the text recognition result "3-month course of treatment" is mapped / encoded into a vector [0.2, 0.5, 0.8]. Also, the product image is mapped / encoded into image semantic features through the feature mapping unit, which can also be referred to as an image semantic space. For example, the product image is mapped / encoded into [0.3, 0.6, 0.7].

[0055] In this embodiment, the multimodal evaluation model may also use a cross-attention mechanism (i.e., the CrossAttention mechanism) to calculate the feature matching relationship between the text semantic features and the image semantic features to generate a cross-vector space. The cross-vector space is used to represent the correlation between the features of the text recognition result and the product image. For example, if "3-month course of treatment" in the text recognition result is highly correlated with the "course of treatment description" area in the product image, the cross-vector space is used to highlight the strong correlation between the two in the form of a feature vector.

[0056] In this embodiment, the multimodal evaluation model may also obtain the fusion result of the text semantic features, the image semantic features, and the cross-vector space, and generate an evaluation result that can evaluate the relationship between the text recognition result and the product image through a linear regression layer. The linear regression layer includes a self-attention layer (Self-Attention Layer) for highlighting key features, a pooling layer (Pooling Layer) for compressing the feature dimension and retaining the core semantics, and a normalization layer (Normalization Layer) for normalizing the feature distribution. The evaluation result can be expressed as a confidence score between 0 and 1, and the larger the confidence score, the higher the correlation between the text recognition result and the product image.

[0057] In some cases, when the evaluation result is greater than a preset value (e.g., 0.8), it is determined that the product image contains the text recognition result.

[0058] In some embodiments, the training device of the product recommendation model may call a quality evaluation model based on the product recommendation logic and the incremental product recommendation session, so that the quality evaluation model generates a quality evaluation result based on the relevance between the context in the incremental product recommendation session and the product recommendation logic; wherein, the quality evaluation result is used to characterize the rationality of the product recommendation logic; the quality evaluation result of the product recommendation logic used to train the product recommendation model meets the expectations.

[0059] In this embodiment, in order to make the product recommendation logic of the input for model training have a certain degree of accuracy, the training device of the product recommendation model may call a quality evaluation model, so that the quality evaluation model calculates the text similarity between the incremental product recommendation session and the product recommendation logic, verifies the coverage of the key elements representing user needs in the incremental product recommendation session by the product recommendation logic, checks the consistency between the products in the product recommendation logic and the product description information, and fuses the text similarity, coverage, and consistency to obtain a quality evaluation result for characterizing the product recommendation logic. The quality evaluation result can be expressed as a numerical value, and the higher the numerical value, the higher the rationality of the product recommendation logic. Furthermore, the product recommendation logic that meets the expected quality evaluation result can be selected as the input for model training to improve the model training accuracy.

[0060] In some embodiments, the training device of the product recommendation model may train an initial model based on pre-annotated sample data and a prompt instruction for indicating the model task, so as to obtain the product recommendation model to be trained.

[0061] In this embodiment, in order to achieve model warm-up, the initial model can be first driven to learn the tasks to be completed, thinking paths, answer styles, etc. based on the pre-annotated sample data and the prompt instruction for indicating the model task, so as to obtain the product recommendation model. The pre-annotated sample data may come from real product recommendation sessions. The pre-annotated sample data includes consultation information, corresponding product recommendation information, and the recommendation logic between the two. The initial model is an artificial intelligence model with natural language understanding ability, which can be understood as a base model. By driving the initial model to perform parameter iteration through the pre-annotated sample data, the product recommendation model to be trained can be obtained. The product recommendation model to be trained already has preliminary product recommendation capabilities. Training the product recommendation model based on the product recommendation logic, incremental product recommendation sessions, and real product recommendation sessions can obtain a product recommendation model with expected generalization capabilities.

[0062] In some embodiments, the training device of the product recommendation model may call the product recommendation model based on the sample session information, so that the product recommendation model generates target product recommendation information and target product recommendation logic corresponding to the sample consultation information; drive the dynamic tuning of the recommendation logic parameters of the product recommendation model based on the semantic similarity between the target product recommendation logic and the standard recommendation logic; and drive the dynamic tuning of the product recommendation parameters of the product recommendation model based on the consistency between the target product recommendation information and the standard recommended products.

[0063] In this embodiment, in order to improve the model training accuracy, supervised training based on positive feedback may also be performed on the product recommendation model. Among them, the sample session information may be a real product recommendation session or other product recommendation sessions, and the embodiments of the present application do not limit this. The sample session information includes standard recommended products used as the basis for model tuning, and the sample session information corresponds to the standard recommendation logic used as the basis for model tuning. Specifically, the training device of the product recommendation model may call the product recommendation model based on the sample session information, so that the product recommendation model generates target product recommendation information and target product recommendation logic corresponding to the sample consultation information. The product recommendation model includes product recommendation parameters related to product recommendation and recommendation logic parameters participating in logical reasoning. The Regular Reward Model (ORM) can be used to verify the consistency between the target product recommendation information and the standard recommended products. When the two are consistent, reward information for motivating the dynamic tuning of the product recommendation parameters of the product recommendation model can be output. And the Process Reward Model (PRM) can be used to determine the similarity between the standard recommendation logic and the target product recommendation logic. When the similarity meets the expectation, reward information for motivating the dynamic tuning of the recommendation logic parameters of the product recommendation model can be output.

[0064] In this embodiment, the cooperation of the Regular Reward Model (ORM) and the Process Reward Model (PRM) can make the product recommendation logic of the product recommendation model more reasonable, and can also make the matching degree between the products recommended by the product recommendation model and the user needs higher.

[0065] The embodiments of the present application also provide a product recommendation method. The product recommendation method includes: when receiving consultation information, calling the product recommendation model in an embodiment of the present application, so that the product recommendation model generates product recommendation information and product recommendation logic corresponding to the consultation information; where the product recommendation logic is used to explain the recommendation reasons for the products in the product recommendation information in a logical derivation manner.

[0066] In this embodiment, for the specific functions and effects achieved by the training method of the product recommendation model, reference may be made to the explanations in other embodiments of the present application, and details will not be elaborated here.

[0067] Please refer to Figure 3 Figure 3 This embodiment of the present application also provides a training device for a commodity recommendation model. The training device for the commodity recommendation model may include: a session acquisition module, configured to acquire real commodity recommendation sessions; an incremental data generation module, configured to generate incremental commodity recommendation sessions based on the real commodity recommendation sessions, and generate commodity recommendation logics corresponding to the incremental commodity recommendation sessions; wherein, the semantic similarity between the incremental commodity recommendation sessions and the real commodity recommendation sessions meets the expectation; the commodity recommendation logics are used to explain the recommendation reasons for the commodities in the incremental commodity recommendation sessions in a logical derivation manner; a model training module, configured to train a commodity recommendation model based on the real commodity recommendation sessions, the incremental commodity recommendation sessions, and the commodity recommendation logics. Figure 3

[0068] In this embodiment, for the specific functions and effects achieved by the training device for the commodity recommendation model, reference may be made to other embodiments of the present application for comparative explanation, which will not be elaborated herein.

[0068]

[0069] Please refer to Figure 4 Figure 4 This embodiment of the present application also provides a commodity recommendation device, which includes: a commodity recommendation module, configured to, when receiving consultation information, call the commodity recommendation model in the training device, so that the commodity recommendation model generates commodity recommendation information and commodity recommendation logics corresponding to the consultation information; wherein, the commodity recommendation logics are used to explain the recommendation reasons for the commodities in the commodity recommendation information in a logical derivation manner. Figure 4

[0070] In this embodiment, for the specific functions and effects achieved by the training device for the commodity recommendation model, reference may be made to other embodiments of the present application for comparative explanation, which will not be elaborated herein.

[0070]

[0071] Please refer to Figure 5 Figure 5 This embodiment of the present application also provides a computer device, which includes: a memory and a processor, where at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method as described above. Figure 5

[0072] This embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the processor implements the method as described above.

[0072]

[0073] This embodiment of the present application also provides a computer program product including instructions, and when the computer program product is executed by a processor, the method as described above is implemented.

[0073]

[0074] The user information or user account information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, etc.) involved in multiple embodiments of this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws and regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0075] It can be understood that the specific examples in this article are only to help those skilled in the art better understand the embodiments of this application, rather than limiting the scope of the present invention.

[0076] It can be understood that in various embodiments of this application, the magnitudes of the sequence numbers of each process do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0077] It can be understood that the various embodiments described in this application can be implemented alone or in combination, and the embodiments of this application do not limit this.

[0078] Unless otherwise specified, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the technical field of this application. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the scope of this application. The term "and / or" used in this application includes any and all combinations of one or more of the related listed items. The singular forms "a", "above", and "the" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0079] It can be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0080] It can be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0081] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the information to be retrieved. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0082] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0083] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0084] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0086] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the information to be retrieved in the present application, or the part that contributes to the prior art, or the part of the information to be retrieved can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0087] The above is only the specific embodiment of the present application, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for training a product recommendation model, characterized in that: The method comprises: Get real product recommendation session; An incremental product recommendation session is generated based on the real product recommendation session, and a product recommendation logic corresponding to the incremental product recommendation session is generated; wherein the semantic similarity between the incremental product recommendation session and the real product recommendation session is as expected; and the product recommendation logic is used to explain the reasons for recommending the products in the incremental product recommendation session in a logical deduction manner; A product recommendation model is trained based on the real product recommendation session, the incremental product recommendation session and the product recommendation logic.

2. The method according to claim 1, characterized in that The step of training a product recommendation model based on the real product recommendation session, the incremental product recommendation session and the product recommendation logic includes: Extracting product description information from a product image; wherein the product image includes at least one of the following: a product display image, a product details page image, a product application scenario image, and a product use effect image; A product recommendation model is trained based on the product recommendation logic, the product description information, the real product recommendation session, and the incremental product recommendation session.

3. The method according to claim 2, characterized in that The steps of extracting product description information from the product image include: Perform text recognition on the product image to obtain the text recognition result; Invoking a multimodal evaluation model based on the text recognition result, so that the multimodal evaluation model generates an evaluation result for characterizing the relationship between the text recognition result and the product image; When the evaluation result indicates that the product image includes the text recognition result, the text recognition result is recognized as product description information.

4. The method according to claim 3, characterized in that The step of calling a multimodal evaluation model based on the text recognition result so that the multimodal evaluation model generates an evaluation result for characterizing the relationship between the text recognition result and the product image includes: A multimodal evaluation model is called based on the text recognition result, so that the multimodal evaluation model generates a map of the text recognition result to a text semantic feature and maps the product image to an image semantic feature, generates a cross vector space representing the feature matching relationship between the text semantic feature and the image semantic feature, and combines the text semantic feature, the image semantic feature, and the fusion result of the cross vector space to generate an evaluation result of the relationship between the text recognition result and the product image.

5. The method according to claim 1, characterized in that The method further comprises: Based on the product recommendation logic and the incremental product recommendation session, a quality evaluation model is called so that the quality evaluation model generates a quality evaluation result based on the correlation between the context in the incremental product recommendation session and the product recommendation logic; wherein the quality evaluation result is used to characterize the rationality of the product recommendation logic; and the quality evaluation result of the product recommendation logic used to train the product recommendation model is in line with expectations.

6. The method according to claim 1, characterized in that The method further comprises: Calling the product recommendation model based on the sample session information so that the product recommendation model generates target product recommendation information and target product recommendation logic corresponding to the sample consultation information; Based on the semantic similarity between the target product recommendation logic and the standard recommendation logic, driving the product recommendation model to dynamically tune the recommendation logic parameters; Based on the consistency between the target product recommendation information and the standard recommended products, the product recommendation model is driven to dynamically optimize product recommendation parameters.

7. A product recommendation method, characterized in that: The method comprises: When consulting information is received, the product recommendation model described in any one of claims 1 to 6 is called so that the product recommendation model generates product recommendation information and product recommendation logic corresponding to the consulting information; wherein the product recommendation logic is used to explain the reasons for recommending the products in the product recommendation information in a logical deduction manner.

8. A training device for a product recommendation model, characterized in that: The device comprises: A session acquisition module is used to acquire real product recommendation sessions; An incremental data generation module, configured to generate an incremental product recommendation session based on the real product recommendation session, and to generate a product recommendation logic corresponding to the incremental product recommendation session; wherein the semantic similarity between the incremental product recommendation session and the real product recommendation session is as expected; and the product recommendation logic is configured to explain the reasons for recommending the products in the incremental product recommendation session in a logically deductive manner; A model training module is used to train a product recommendation model based on the real product recommendation session, the incremental product recommendation session and the product recommendation logic.

9. A commodity recommendation device, characterized in that: The device comprises: A product recommendation module, for, upon receiving consulting information, calling the product recommendation model described in any one of claims 1 to 6, so that the product recommendation model generates product recommendation information and product recommendation logic corresponding to the consulting information; wherein the product recommendation logic is used to explain the reasons for recommending the products in the product recommendation information in a logical deduction manner.

10. A computer device, characterized in that: The computer device includes a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the method according to any one of claims 1 to 7 can be implemented.

12. A computer program product, characterized in that The computer program product is used to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for providing commodity information and electronic equipment

    CN116308682A

  • Commodity recommendation reason generation method and device and electronic equipment

    CN116894711A

  • Optimization method and device for recommendation reason generation model

    CN117540078A

  • Information generation method and device, computer equipment and medium

    CN117593069A

  • Commodity recommendation model training method, commodity recommendation method and electronic equipment

    CN117974276A