Commodity recommendation method and system based on multi-modal information, electronic equipment and storage medium

Through multimodal information pre-training and recoding technology, the problem of data isolation of a single e-commerce platform is solved, the accuracy and diversity of cross-platform product recommendations are achieved, the effect of new users is improved, and the calculation and maintenance costs are reduced.

CN120338911APending Publication Date: 2025-07-18阿里巴巴(中国)网络技术有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510318370.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing product recommendation system based on a single e-commerce platform has cross-platform data isolation problems, resulting in poor accuracy and diversity, and high calculation and maintenance costs, especially poor recommendations for new users.

Method used

Multimodal information is used for model pre-training, multimodal information of products is encoded and autoregressively learned through image encoder and large language model, multimodal pre-trained large model is generated, combined with residual quantization variational autoencoder for recoding, and a double tower recall model is built to realize the unified representation of cross-platform product features and cross-domain migration of user interests.

Benefits of technology

It improves the accuracy, diversity and efficiency of product recommendations, improves the recommendation effect in the cold start scenario of new users, and realizes the unified understanding of cross-platform product information and cross-domain acceptance of user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338911A_ABST
    Figure CN120338911A_ABST
Patent Text Reader

Abstract

The invention provides a commodity recommendation method and system based on multi-modal information, electronic equipment and a storage medium. The commodity recommendation method comprises the steps of performing model pre-training based on multi-modal information of a target commodity of a first e-commerce platform to obtain a multi-modal pre-training large model; determining a high-dimensional multi-modal representation of the target commodity according to the multi-modal pre-training large model; performing recoding processing on the high-dimensional multi-modal representation to obtain a hierarchical code and a low-dimensional multi-modal representation of the target commodity; determining a user feature vector of a user tower in the double-tower recall model according to the user historical behavior sequence of the user on the second e-commerce platform, the hierarchical code and the sliding window representation of the hierarchical code; determining commodity feature vectors of commodity towers in the double-tower recall model according to the hierarchical codes and the sliding window representation thereof; and determining a commodity feature vector matched with the user feature vector, so as to obtain a target recommended commodity according to the determined commodity feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product recommendation based on an e-commerce platform. Specifically, it relates to a product recommendation method and system, an electronic device, and a storage medium based on multi-modal information. Background Art

[0002] With the rapid development of e-commerce, users' shopping behaviors have become increasingly diversified, and their interests have become more scattered. Traditional product recommendation systems mainly rely on user behavior data within a single e-commerce platform, such as browsing history, purchase records, etc., for personalized recommendation.

[0003] However, the inventors of the present application found that there are problems of data isolation between multiple e-commerce platforms in this product recommendation method based on data from a single e-commerce platform. For example, product attribute data between various e-commerce platforms are isolated from each other (such as data like product IDs and store IDs cannot be directly interconnected), which makes it difficult for a single e-commerce platform to comprehensively understand users' cross-platform shopping interests and behaviors, restricting the accuracy and diversity of the product recommendation system.

[0004] In addition, the inventors also found that for new users, due to the lack of historical behavior data, it is difficult for traditional product recommendation systems to accurately capture their interest preferences based on the above-mentioned data from a single e-commerce platform, which also leads to poor recommendation effects. Moreover, traditional cross-platform joint training methods require integrating data from different platforms and performing complex model training and updates (such as knowledge transfer), which results in relatively high computational costs and maintenance costs.

[0005] The content in the background art section is only the technology known to the applicant and does not necessarily represent the prior art in this field. Summary of the Invention

[0006] The present invention provides a product recommendation method and system, an electronic device, and a storage medium based on multi-modal information, aiming to solve the problems of poor accuracy and diversity, as well as relatively high computational costs and maintenance costs, existing in the current product recommendation method based on data from a single e-commerce platform.

[0007] According to one aspect of the present invention, the present invention provides a commodity recommendation method based on multimodal information, including: performing model pre-training based on the multimodal information of target commodities on a first e-commerce platform to obtain a large multimodal pre-trained model; determining high-dimensional multimodal representations of target commodities according to the large multimodal pre-trained model; performing re-encoding processing on the high-dimensional multimodal representations to obtain hierarchical encodings and low-dimensional multimodal representations of target commodities; determining user feature vectors of the user tower in a two-tower recall model according to the user's historical behavior sequence on a second e-commerce platform, as well as the hierarchical encodings and their sliding window representations; determining commodity feature vectors of the commodity tower in the two-tower recall model according to the hierarchical encodings and their sliding window representations; determining the commodity feature vectors that match the user feature vectors, so as to obtain target recommended commodities according to the determined commodity feature vectors.

[0008] According to some embodiments of the present invention, performing model pre-training based on the multimodal information of target commodities on a first e-commerce platform to obtain a large multimodal pre-trained model includes: encoding the image modality information of target commodities based on a preset image encoder to obtain visual representations of target commodities; performing alignment processing on the visual representations and the text representations of target commodities; determining description information of target commodities based on a large language model according to the aligned visual representations and text encoding instructions; performing autoregressive learning according to the description information and the text modality information of target commodities to obtain a large multimodal pre-trained model.

[0009] According to some embodiments of the present invention, performing re-encoding processing on the high-dimensional multimodal representations to obtain hierarchical encodings and low-dimensional multimodal representations of target commodities includes: reducing the dimension of the high-dimensional multimodal representations to obtain residual initial representations; performing multiple residual quantizations based on the residual initial representations to obtain hierarchical encodings corresponding to target commodities and low-dimensional multimodal representations after cumulative recombination of hierarchical encodings.

[0010] According to some embodiments of the present invention, the commodity recommendation method further includes: performing sliding window processing on the hierarchical encodings to obtain sliding window representations.

[0011] According to another aspect of the present invention, the present invention provides a commodity recommendation system based on multimodal information. The commodity recommendation system based on multimodal information includes a multimodal model training module, a multimodal representation processing module, a user tower processing module, a commodity tower processing module, and a commodity recommendation processing module. The multimodal model training module performs model pre-training based on the multimodal information of target commodities on the first e-commerce platform to obtain a multimodal pre-trained large model. The multimodal representation processing module determines the high-dimensional multimodal representation of the target commodity according to the multimodal pre-trained large model, and performs re-encoding processing on the high-dimensional multimodal representation to obtain the hierarchical encoding and low-dimensional multimodal representation of the target commodity. The user tower processing module determines the user feature vector of the user tower in the two-tower recall model according to the user's historical behavior sequence on the second e-commerce platform, the hierarchical encoding, and its sliding window representation. The commodity tower processing module determines the commodity feature vector of the commodity tower in the two-tower recall model according to the hierarchical encoding and its sliding window representation. The commodity recommendation processing module determines the commodity feature vector that matches the user feature vector, and obtains the target recommended commodity according to the determined commodity feature vector.

[0012] According to some embodiments of the present invention, the multimodal model training module encodes the image modality information of the target commodity based on a preset image encoder to obtain the visual representation of the target commodity; the multimodal model training module performs alignment processing on the visual representation and the text representation of the target commodity; the multimodal model training module determines the description information of the target commodity based on the large language model according to the aligned visual representation and text encoding instructions; the multimodal model training module performs autoregressive learning according to the description information and the text modality information of the target commodity to obtain a multimodal pre-trained large model.

[0013] According to some embodiments of the present invention, the multimodal representation processing module reduces the dimension of the high-dimensional multimodal representation to obtain the residual initial representation; the multimodal representation processing module performs multiple residual quantizations based on the residual initial representation to obtain the hierarchical encoding corresponding to the target commodity and the low-dimensional multimodal representation after cumulative recombination of the hierarchical encoding.

[0014] According to some embodiments of the present invention, the multimodal representation processing module performs a sliding window process on the hierarchical encoding to obtain a sliding window representation.

[0015] According to another aspect of the present invention, the present invention also provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the commodity recommendation method as described above.

[0016] According to another aspect of the present invention, the present invention further provides a non-volatile computer-readable storage medium having a computer program stored thereon, which can implement the commodity recommendation method as described above when executed by a processor.

[0017] According to another aspect of the present invention, the present invention further provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the commodity recommendation method as described above.

[0018] Beneficial Effects

[0019] The present invention can effectively process the multimodal information of the target product through a multimodal pre-trained large model, and can extract rich modal features, thereby realizing information understanding of products across e-commerce platforms. The present invention can map product information of different platforms and different modalities to a unified vector space based on high-dimensional vector recoding technology, and can realize unified representation and comparison of cross-platform product features. And the present invention can also construct a dual-tower recall model based on multimodal representation, which can realize cross-domain migration and inheritance of user interests on different e-commerce platforms. The present invention can improve the accuracy, diversity and efficiency of product recommendations, and can also improve the product recommendation effect in the new user cold start scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A schematic diagram showing a process flow of a commodity recommendation method according to an embodiment of the present invention;

[0022] Figure 2 Another schematic diagram showing a flow chart of a commodity recommendation method according to an embodiment of the present invention;

[0023] Figure 3 Another schematic diagram showing a flow chart of a commodity recommendation method according to an embodiment of the present invention;

[0024] Figure 4 A schematic diagram showing a re-encoding process according to an embodiment of the present invention;

[0025] Figure 5 A schematic diagram showing a sliding window process according to an embodiment of the present invention;

[0026] Figure 6Schematic structural diagram of the commodity recommendation system according to an embodiment of the present invention.

[0027] Description of reference numerals:

[0028] Commodity recommendation system 1; Multimodal model training module 10; Multimodal feature processing module 20; User tower processing module 30; Commodity tower processing module 40; Commodity recommendation processing module 50. Detailed implementation manners

[0029] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0031] The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of these specific details, or other methods, components, materials, devices, etc. can be adopted. In these cases, well-known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0032] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0033] The terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order.

[0034] Combined with the accompanying drawings in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0035] Traditional cross-platform collaborative training methods may include building a basic dual-tower recall structure, multi-modal representation modeling, knowledge transfer, etc.

[0036] However, the inventors found that in the way of building a basic dual-tower recall structure, since the training parameters of the item tower (the product tower in the dual-tower model) are fixed, the model training is insufficient, which affects the training results. In the multi-modal representation modeling method, since it is not the multi-modal representation produced by the large model, there is a problem of poor expected effect. The knowledge transfer method needs to introduce more out-of-domain data for training, and both the training data volume and the model complexity are relatively high, resulting in high computational costs and maintenance costs.

[0037] Based on this, according to one aspect of the present invention, the present invention provides a product recommendation method based on multi-modal information.

[0038] Figure 1 A flowchart showing the product recommendation method according to an embodiment of the present invention is as follows. Figure 1 As shown, the product recommendation method may include steps S100 - S600. Exemplarily, the product recommendation method may be executed by a product recommendation system with computing capabilities.

[0039] According to an exemplary embodiment, in step S100, the product recommendation system performs model pre-training based on the multi-modal information of the target product on the first e-commerce platform to obtain a multi-modal pre-trained large model.

[0040] Multi-modal information refers to information that simultaneously includes multiple modalities. Exemplarily, the multi-modal information of the target product may include text modality information (such as product title, detailed description, etc.), image modality information (such as product main image, etc.), video modality information (such as product display video, etc.), and structured data information (such as price, sales volume, rating, decision-making, merchant level, and category name, etc.).

[0041] It can be understood here that although there is data isolation between ID-based commodities on different e-commerce platforms, the multi-modal information of commodities can be cross-platform understood. For example, in step S100, the commodity recommendation system collects the multi-modal information of the target commodity within the first e-commerce platform and uses it as a data set to pre-train the model, which can enable it to learn general feature representations. Further, the commodity recommendation system fine-tunes the pre-trained model on the data set of a specific task based on autoregressive learning, so that the obtained multi-modal pre-trained large model can adapt to the specific task.

[0042] Exemplarily, the multi-modal pre-trained large model can be used to perform the task of extracting the multi-modal representation of commodities.

[0043] Figure 2 Another flowchart showing the commodity recommendation method according to an embodiment of the present invention.

[0044] Optionally, as Figure 2 shown, step S100 may further include steps S110-S140.

[0045] In step S110, the commodity recommendation system encodes the image modal information of the target commodity based on a preset image encoder to obtain the visual representation of the target commodity.

[0046] For example, the image modal information may be the main image of the target commodity (such as the core picture for commodity display). The commodity recommendation system encodes the input main image of the commodity through a pre-trained and parameter-frozen image encoder (such as EVA-CL IP-G, EvolvingAttention-Contrastive Language Image Pretraining-G, a model for multi-modal tasks of processing images and texts), and can extract the visual representation of the main image of the commodity (which is represented in vector form).

[0047] In step S120, the commodity recommendation system aligns the visual representation with the text representation of the target commodity.

[0048] For example, the text representation of the target commodity may be generated by encoding the text modal information of the target commodity through a large language model, and the text representation is represented in vector form.

[0049] Since the visual representation and the text representation are usually in different vector spaces, the commodity recommendation system can align the visual representation and the text representation in the vector space based on Q-Former (Querying Transformer, a model for converting image features from an image encoder into a representation suitable for processing by a language model) and a linear projection layer.

[0050] For example, the Q-Former can extract features related to text tasks from visual representations through self-attention and cross-attention mechanisms; the linear projection layer can map the features extracted by the Q-Former to a space with the same input dimension as the large language model. Thus, the commodity recommendation system can achieve the alignment of visual and text representations in the vector space based on the attention mechanism of the Q-Former and the dimension mapping of the linear projection layer.

[0051] In step S130, the commodity recommendation system determines the description information of the target commodity based on the aligned visual representation and text encoding instructions and the large language model.

[0052] For example, the commodity recommendation system can generate the description information of the target commodity by concatenating the visual representation and the text encoding instruction (such as "Please describe this commodity") and inputting them into the large language model.

[0053] In step S140, the commodity recommendation system performs autoregressive learning based on the description information and the text modality information of the target commodity to obtain a multi-modal pre-trained large model.

[0054] For example, the commodity recommendation system can compare the description information with the text modality information (such as the title, comments, and other text information of the target commodity itself) to achieve the learning of the autoregressive task, so as to fine-tune the multi-modal pre-trained large model based on the autoregressive learning and finally obtain the trained multi-modal pre-trained large model.

[0055] In step S200, the commodity recommendation system determines the high-dimensional multi-modal representation of the target commodity based on the multi-modal pre-trained large model.

[0056] For example, the commodity recommendation system can process various modality information (such as text, image, or video modality information) of the target commodity by using the multi-modal pre-trained large model to extract a unified high-dimensional multi-modal representation that can represent the target commodity (which is represented in the form of a high-dimensional vector).

[0057] Exemplarily, after the fine-tuning task of the above multi-modal pre-trained large model is completed, the commodity recommendation system extracts the feature vector from the intermediate layer of the multi-modal pre-trained large model to obtain the high-dimensional multi-modal representation of the target commodity.

[0058] As an embodiment, the high-dimensional multi-modal representation can be a 4096-dimensional high-dimensional vector.

[0059] In step S300, the commodity recommendation system performs re-encoding processing on the high-dimensional multi-modal representation to obtain a hierarchical encoding and a low-dimensional multi-modal representation.

[0060] For example, since it is relatively difficult to directly use high-dimensional multi-modal representations offline or online, the commodity recommendation system performs semantic re-encoding processing on high-dimensional multi-modal representations based on RQ-VAE (Residual Quantized Variational Autoencoder), and can obtain low-dimensional multi-modal representations (such as low-dimensional vectors of 512 dimensions), and can quantize the target commodity into a set of hierarchical ID codes, and can obtain a set of hierarchical codes.

[0061] Figure 3 Another schematic flowchart of the commodity recommendation method according to an embodiment of the present invention is shown; Figure 4 A schematic diagram of the re-encoding process according to an embodiment of the present invention is shown.

[0062] Optionally, as Figure 3 shown, step S300 may include steps S310-S320.

[0063] In step S310, the commodity recommendation system reduces the dimension of the high-dimensional multi-modal representation to obtain an initial residual representation.

[0064] For example, the commodity recommendation system maps high-dimensional input data (such as high-dimensional multi-modal representations of 4096 dimensions) to a low-dimensional latent representation space (such as low-dimensional vectors of 512 dimensions) through an encoder network, and generates an initial residual representation. Exemplarily, the initial residual representation can be represented by r0.

[0065] The encoder network is usually composed of multiple fully connected layers or multiple convolutional layers, and has the characteristics of high efficiency and flexibility. With such a setting, while reducing the dimension and compressing the data, key semantic information can be retained, providing guarantee for subsequent residual quantization and task processing.

[0066] In step S320, the commodity recommendation system performs multiple residual quantizations based on the initial residual representation to obtain a hierarchical code corresponding to the target commodity and a low-dimensional multi-modal representation after cumulative recombination of the hierarchical codes.

[0067] For example, in the residual quantization unit, there are D sub-modules, such as d = 0, 1,..., D-1. Each sub-module corresponds to a codebook, such as:

[0068]

[0069] Among them, e k is the embedding vector of each code in the codebook, and K is the size of the code value space.

[0070] As Figure 4 shown, the commodity recommendation system maps r0 to the embedding vector e closest to it in the C0 codebook c0where the corresponding code index C0 = argmin k ||r o - e k ||. Then the residual of the next level is r1 = r0 - e co . Repeating the above operation D times, D ID code combinations can be obtained, that is, a set of hierarchical codes (i.e., hierarchical codes) corresponding to the target commodity.

[0071] Exemplarily, Figure 4 the DNN Encoder in is the encoder of the deep neural network DNN, the DNN Decoder is the decoder of the deep neural network DNN, the codebook is the codebook (which can be abbreviated as code), the Residual Quantization is the residual quantization unit, the Semantic codes are the semantic codes, the Quantized representation is the quantization representation, and the Embedding is the embedding.

[0072] Optionally, the commodity recommendation system can also perform residual reconstruction.

[0073] For example, the commodity recommendation system sets up a neural network symmetric to the encoder structure to convert the quantized residual representation back to the original high-dimensional data space and trains the model by minimizing the vector error before and after reconstruction. Such a setting can ensure that the quantized low-dimensional representation can reconstruct the original data as accurately as possible, thereby retaining the semantic information of the data.

[0074] Through the above embodiments, the present invention can establish a unified representation of target commodities under different categories and platform systems by re-encoding high-dimensional vectors using RQ-VAE. The hierarchical codes and low-dimensional multi-modal representations can be applied as generalization features to the commodity recommendation link.

[0075] In step S400, the commodity recommendation system determines the user feature vector of the user tower in the two-tower recall model according to the user's historical behavior sequence on the second e-commerce platform, the hierarchical code, and its sliding window representation.

[0076] The two-tower recall model includes a user tower and a commodity tower. The two-tower recall model can map users and commodities into a vector space respectively. By calculating the similarity between the user feature vector and the commodity feature vector, a candidate set of commodities most relevant to the user's interests can be quickly screened out.

[0077] For example, the two-tower recall model is pre-trained by the commodity recommendation system. The training data set of the two-tower recall model can include {user behavior sequence, positive and negative samples}, and the training data set can include various combinations.

[0078] As an example, the training dataset may include three combination forms, namely {user extra-domain behavior sequence, extra-domain positive / negative samples}, {user extra-domain behavior sequence, in-domain positive / negative samples}, and {user in-domain behavior sequence, in-domain positive / negative samples}. This can enable the training model to be adapted to more application scenarios.

[0079] Here, it can be understood that "in-domain" refers to the current e-commerce platform, and "extra-domain" refers to other e-commerce platforms except the current one. Positive samples are the product samples actually clicked by users; negative samples include simple negative samples and hard negative samples. Simple negative samples are randomly selected samples globally (such as 3 samples), and hard negative samples are randomly selected samples within the same category (such as 10 samples), and simple negative samples can be shared within the same batch.

[0080] The user's historical behavior sequence can be the historical behavior data of the user on e-commerce platforms other than the first e-commerce platform (such as the second e-commerce platform), including but not limited to behavior sequences such as clicks, collections, add-to-carts, and orders.

[0081] In step S400, the product recommendation system uses the user's historical behavior sequence as the input of the user tower, and uses information such as hierarchical encoding and its sliding window representation as the main features. Based on the Embedding Layer, a sequence vector is obtained. Then, the interaction relationship between features is captured through FM / Concat (interaction method), and the internal dependencies within the sequence are captured based on Self-Attention (self-attention mechanism), thereby generating a context-aware fused feature vector. Finally, high-order features are extracted through a multi-layer perceptron to generate a final 128-dimensional user feature vector.

[0082] In step S500, the product recommendation system determines the product feature vector of the product tower in the dual-tower recall model according to the hierarchical encoding and its sliding window representation.

[0083] For example, taking the hierarchical encoding and its sliding window representation as the input, after being processed by the embedding layer and the fully connected layer, the product recommendation system can obtain a product feature vector of the same size as the user feature vector (such as 128 dimensions).

[0084] In step S600, the product recommendation system determines the product feature vector that matches the user feature vector, and obtains the target recommended products according to the determined product feature vector.

[0085] For example, the product recommendation system can determine the product feature vector that matches the user feature vector based on vector retrieval technology to complete the matching and recall of users and products, and thus can correspondingly obtain the target recommended products (such as the TopN products with higher similarity).

[0086] Exemplarily, the vector retrieval technology includes, but is not limited to, calculating the inner product between two vectors, cosine similarity, etc., and the present invention does not limit this.

[0087] Optionally, the product recommendation system can also perform a sliding window process on the hierarchical encoding to obtain a sliding window representation.

[0088] For example, in the process of processing low-dimensional multi-modal representations, the multi-modal re-encoding features of the target product are not independent of each other. Each dimension of the encoding is obtained by residual fitting of the upper level. Therefore, the single-dimensional encoding is not interpretable, and only the products with continuous encoding segments will show certain regularity.

[0089] The complete D-dimensional encoding combination can achieve the effect of product recommendation based on the product ID in terms of product discrimination (the value space of each dimension is 0-255). However, the number of parameters of this encoding combination is very large (such as 255**16), which poses great difficulties for model training.

[0090] Therefore, the present invention can divide the multi-dimensional multi-modal representation into multiple sliding window segment features by performing a sliding window process on the hierarchical encoding, and then send these multiple sliding window segment features into the model for training.

[0091] Figure 5 A schematic diagram showing the sliding window process of an embodiment of the present invention is shown.

[0092] As an embodiment, the window size of the sliding window process can be 3. As Figure 5 shown, a group of 16-dimensional hierarchical encodings can obtain 14 sliding window segment features after the sliding window process.

[0093] Exemplarily, in Figure 5 , window is the sliding window and window Feature is the window feature.

[0094] Through the above embodiments, the present invention can perform model pre-training and fine-tuning processing based on the multi-modal information of the target product to obtain a multi-modal pre-trained large model, and then can determine the high-dimensional multi-modal representation of the target product based on the multi-modal pre-trained large model. And the present invention can obtain a hierarchical encoding and a low-dimensional multi-modal representation by performing a re-encoding process on the high-dimensional multi-modal representation, and then can enable the dual-tower recall model to determine the corresponding user feature vector and product feature vector according to the hierarchical encoding and its sliding window representation, and finally can determine the target recommended product according to the product feature vector matching the user feature vector.

[0095] The present invention can effectively process the multimodal information of the target product through a multimodal pre-trained large model, and can extract rich modal features, thereby realizing information understanding of products across e-commerce platforms. The present invention can map product information of different platforms and different modalities to a unified vector space based on high-dimensional vector recoding technology, and can realize unified representation and comparison of cross-platform product features. And the present invention can also construct a dual-tower recall model based on multimodal representation, which can realize cross-domain migration and inheritance of user interests on different e-commerce platforms. The present invention can improve the accuracy, diversity and efficiency of product recommendations, and can also improve the product recommendation effect in the new user cold start scenario.

[0096] According to yet another aspect of the present invention, the present invention further provides a commodity recommendation system based on multimodal information.

[0097] Figure 6 FIG. 2 is a schematic diagram showing the structure of a commodity recommendation system according to an embodiment of the present invention. Figure 6 As shown, the product recommendation system 1 may include a multimodal model training module 10 , a multimodal representation processing module 20 , a user tower processing module 30 , a product tower processing module 40 and a product recommendation processing module 50 .

[0098] According to an exemplary embodiment, the multimodal model training module 10 performs model pre-training based on the multimodal information of the target product of the first e-commerce platform to obtain a multimodal pre-trained large model.

[0099] Multimodal information refers to information that contains multiple modalities at the same time. For example, the multimodal information of the target product may include text modal information (such as product title, detailed description, etc.), image modal information (such as product main picture, etc.), video modal information (such as product display video, etc.) and structured data information (such as price, sales volume, rating, decision, merchant level and category name, etc.).

[0100] It can be understood here that although there is data isolation between ID-type commodities on different e-commerce platforms, the multimodal information of commodities can be understood across platforms. For example, the multimodal model training module 10 collects the multimodal information of the target commodity on the first e-commerce platform, and uses it as a data set to pre-train the model, so that it can learn common feature representations. Furthermore, the multimodal model training module 10 fine-tunes the pre-trained model on a dataset of a specific task based on autoregressive learning, so that the obtained multimodal pre-trained large model can adapt to specific tasks.

[0101] Exemplarily, the multimodal pre-trained large model can be used to perform the task of extracting multimodal representations of commodities.

[0102] Optionally, the multimodal model training module 10 encodes the image modality information of the target commodity based on a preset image encoder to obtain the visual representation of the target commodity.

[0103] For example, the image modality information can be the main image of the target commodity (such as the core picture used for commodity display). The multimodal model training module 10 encodes the input main commodity image through a pre-trained image encoder with frozen parameters (such as the EVA-CLIP-G model), and can extract the visual representation of the main commodity image (which is represented in vector form).

[0104] The multimodal model training module 10 aligns the visual representation with the text representation of the target commodity.

[0105] For example, the text representation of the target commodity can be generated by encoding the text modality information of the target commodity through a large language model, and the text representation is represented in vector form.

[0106] Since the visual representation and the text representation are usually in different vector spaces, the multimodal model training module 10 can achieve the alignment of the visual representation and the text representation in the vector space based on Q-Former and a linear projection layer.

[0107] For example, Q-Former can extract features related to text tasks from the visual representation through self-attention mechanism and cross-attention mechanism; the linear projection layer can map the features extracted by Q-Former to a space consistent with the input dimension of the large language model. Thus, the multimodal model training module 10 can achieve the alignment of the visual representation and the text representation in the vector space based on the attention mechanism of Q-Former and the dimension mapping of the linear projection layer.

[0108] The multimodal model training module 10 determines the description information of the target commodity based on the aligned visual representation and text encoding instructions, based on the large language model.

[0109] For example, the multimodal model training module 10 splices the visual representation and the text encoding instructions (such as "Please describe this commodity") and inputs them into the large language model, and can generate the description information of the target commodity.

[0110] The multimodal model training module 10 performs autoregressive learning based on the description information and the text modality information of the target commodity to obtain a multimodal pre-trained large model.

[0111] For example, the multi-modal model training module 10 can compare the description information with the text modal information (such as the title, comments, and other text information of the target commodity itself), and can implement the learning of the autoregressive task to fine-tune the multi-modal pre-trained large model based on the autoregressive learning, and finally obtain the trained multi-modal pre-trained large model.

[0112] According to the example embodiment, the multi-modal representation processing module 20 determines the high-dimensional multi-modal representation of the target commodity according to the multi-modal pre-trained large model.

[0113] For example, by using the multi-modal pre-trained large model to process various modal information (such as text, image, or video modal information) of the target commodity, the multi-modal representation processing module 20 can extract a unified high-dimensional multi-modal representation that can represent the target commodity (which is represented in the form of a high-dimensional vector).

[0114] Exemplarily, after the fine-tuning task of the above multi-modal pre-trained large model is completed, the multi-modal representation processing module 20 extracts feature vectors in the intermediate layer of the multi-modal pre-trained large model, and can obtain the high-dimensional multi-modal representation of the target commodity.

[0115] As an embodiment, the high-dimensional multi-modal representation can be a 4096-dimensional high-dimensional vector.

[0116] According to the example embodiment, the multi-modal representation processing module 20 performs re-encoding processing on the high-dimensional multi-modal representation to obtain hierarchical encoding and low-dimensional multi-modal representation.

[0117] For example, since it is relatively difficult to directly use the high-dimensional multi-modal representation offline or online, the multi-modal representation processing module 20 performs semantic re-encoding processing on the high-dimensional multi-modal representation based on RQ-VAE, and can obtain a low-dimensional multi-modal representation (such as a 512-dimensional low-dimensional vector), and can quantize the target commodity into a set of hierarchical ID encodings, and can obtain a set of hierarchical encodings.

[0118] Optionally, the multi-modal representation processing module 20 reduces the dimension of the high-dimensional multi-modal representation to obtain a residual initial representation.

[0119] For example, the multi-modal representation processing module 20 maps the high-dimensional input data (such as a 4096-dimensional high-dimensional multi-modal representation) to a low-dimensional latent representation space (such as a 512-dimensional low-dimensional vector) through an encoder network and generates a residual initial representation. Exemplarily, the residual initial representation can be represented by r0.

[0120] The encoder network is usually composed of multiple fully connected layers or multiple convolutional layers, and has the characteristics of high efficiency and flexibility. With such a setting, while reducing the dimension and compressing the data, the key semantic information can be retained, providing guarantee for subsequent residual quantization and task processing.

[0121] The multimodal representation processing module 20 performs multiple residual quantizations based on the residual initial representation to obtain a hierarchical encoding corresponding to the target commodity and a low-dimensional multimodal representation after cumulative recombination of the hierarchical encodings.

[0122] For example, in the residual quantization unit, there are D sub-modules, such as d = 0, 1,..., D - 1. Each sub-module corresponds to a codebook, such as:

[0123]

[0124] where e k is the embedding vector of each code in the codebook, and K is the size of the code value space.

[0125] The multimodal representation processing module 20 maps r0 to the embedding vector e c0 in the C0 codebook that is closest, and the corresponding code index C0 = argmin k ||r o - e k ||. Then the next-level residual is r1 = r0 - e co . Repeating the above operation D times can obtain D ID code combinations, that is, a set of hierarchical encodings (i.e., hierarchical encoding) corresponding to the target commodity.

[0126] Optionally, the multimodal representation processing module 20 can also perform residual reconstruction.

[0127] For example, the multimodal representation processing module 20 sets up a neural network symmetric to the encoder structure to convert the quantized residual representation back to the original high-dimensional data space and trains the model by minimizing the vector error before and after reconstruction. Such a setting can ensure that the quantized low-dimensional representation can reconstruct the original data as accurately as possible, thereby retaining the semantic information of the data.

[0128] Through the above embodiments, the present invention can establish a unified expression for target commodities under different categories and platform systems by re-encoding and processing high-dimensional vectors using RQ-VAE. The hierarchical encoding and the low-dimensional multimodal representation can be applied as generalizable features to the link of commodity recommendation.

[0129] According to the exemplary embodiment, the user tower processing module 30 determines the user feature vector of the user tower in the two-tower recall model according to the user's historical behavior sequence on the second e-commerce platform, as well as the hierarchical encoding and its sliding window representation.

[0130] The two-tower recall model includes a user tower and a product tower. The two-tower recall model can map users and products into a vector space respectively. By calculating the similarity between the user feature vector and the product feature vector, a candidate set of products most relevant to the user's interests can be quickly screened out.

[0131] For example, the two-tower recall model is pre-trained for a product recommendation system. The training dataset of the two-tower recall model can include {user behavior sequences, positive and negative samples}, and the training dataset can include various combinations.

[0132] As an embodiment, the training dataset can include three combination forms, namely {user out-of-domain behavior sequences, out-of-domain positive / negative samples}, {user out-of-domain behavior sequences, in-domain positive / negative samples}, and {user in-domain behavior sequences, in-domain positive / negative samples}. This can enable the training model to be adapted to more application scenarios.

[0133] Here, it can be understood that in-domain refers to the current e-commerce platform, and out-of-domain refers to other e-commerce platforms except the current one. Positive samples are product samples actually clicked by users; negative samples include simple negative samples and hard negative samples. Simple negative samples are globally randomly selected samples (such as 3 samples), and hard negative samples are randomly selected samples under the same category (such as 10 samples), and simple negative samples can be shared within the same batch.

[0134] The user historical behavior sequence can be the historical behavior data of the user on an e-commerce platform other than the first e-commerce platform (such as the second e-commerce platform), including but not limited to behavior sequences such as clicks, collections, add-to-carts, and orders.

[0135] The user tower processing module 30 uses the user historical behavior sequence as the input of the user tower, and takes information such as hierarchical encoding and its sliding window representation as the main features. Based on the Embedding Layer, a sequence vector is obtained. Then, the interaction relationship between features is captured through FM / Concat, and the internal dependence of the sequence is captured based on Self-Attention. Furthermore, a context-aware fusion feature vector is generated. Finally, high-order features are extracted through a multi-layer perceptron to generate a final 128-dimensional user feature vector.

[0136] According to the example embodiment, the product tower processing module 40 determines the product feature vector of the product tower in the two-tower recall model according to the hierarchical encoding and its sliding window representation.

[0137] For example, the product tower processing module 40 takes the hierarchical encoding and its sliding window representation as the input, and after processing by the embedding layer and the fully connected layer, a product feature vector of the same size as the user feature vector (such as 128 dimensions) can be obtained.

[0138] According to the exemplary embodiment, the product recommendation processing module 50 determines the product feature vector that matches the user feature vector, so as to obtain the target recommended product based on the determined product feature vector.

[0139] For example, the product recommendation processing module 50 can determine the product feature vector that matches the user feature vector based on vector retrieval technology to complete the matching and recall of users and products, so as to obtain the target recommended product accordingly (such as the TopN products with higher similarity can be selected).

[0140] Exemplarily, the vector retrieval technology includes but is not limited to calculating the inner product, cosine similarity, etc. between two vectors, and the present invention does not limit this.

[0141] Optionally, the multi-modal representation processing module 20 can also perform a sliding window process on the hierarchical encoding to obtain a sliding window representation.

[0142] For example, in the process of processing low-dimensional multi-modal representations, the multi-modal re-encoded features of the target product are not independent of each other, and each dimension of the encoding is obtained by residual fitting of the upper level. Therefore, a single-dimensional encoding is not interpretable, and only the products with continuous encoding segments will show certain regularity.

[0143] The complete D-dimensional encoding combination can achieve the effect of product recommendation based on the product ID in terms of product discrimination (the value range of each dimension is 0 to 255). However, the number of parameters of this encoding combination is very large (such as 255**16), which is very difficult for model training.

[0144] Therefore, in the present invention, by performing a sliding window process on the hierarchical encoding, the multi-dimensional multi-modal representation can be divided into multiple sliding window segment features, and these multiple sliding window segment features are sent into the model for training.

[0145] Through the above embodiments, the present invention performs model pre-training and fine-tuning processing based on the multi-modal information of the target product, and can obtain a multi-modal pre-trained large model. Furthermore, the present invention can determine the high-dimensional multi-modal representation of the target product based on the multi-modal pre-trained large model. And by performing re-encoding processing on the high-dimensional multi-modal representation, the present invention can obtain hierarchical encoding and low-dimensional multi-modal representation. Furthermore, based on the two-tower recall model, the corresponding user feature vector and product feature vector can be determined according to the hierarchical encoding and its sliding window representation. Finally, the target recommended product can be determined according to the product feature vector that matches the user feature vector.

[0146] The present invention can effectively process the multimodal information of the target product through a multimodal pre-trained large model, and can extract rich modal features, thereby realizing information understanding of products across e-commerce platforms. The present invention can map product information of different platforms and different modalities to a unified vector space based on high-dimensional vector recoding technology, and can realize unified representation and comparison of cross-platform product features. And the present invention can also construct a dual-tower recall model based on multimodal representation, which can realize cross-domain migration and inheritance of user interests on different e-commerce platforms. The present invention can improve the accuracy, diversity and efficiency of product recommendations, and can also improve the product recommendation effect in the new user cold start scenario.

[0147] According to another aspect of the present invention, the present invention further provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the commodity recommendation method as described above.

[0148] According to another aspect of the present invention, the present invention further provides a non-volatile computer-readable storage medium having a computer program stored thereon, which can implement the commodity recommendation method as described above when executed by a processor.

[0149] According to another aspect of the present invention, the present invention further provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the commodity recommendation method as described above.

[0150] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions of the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A commodity recommendation method based on multimodal information, characterized in that, Including: Performing model pre-training based on the multi-modal information of the target commodity on the first e-commerce platform to obtain a multi-modal pre-trained large model; Determining the high-dimensional multi-modal representation of the target commodity according to the multi-modal pre-trained large model; Performing re-encoding processing on the high-dimensional multi-modal representation to obtain the hierarchical encoding and low-dimensional multi-modal representation of the target commodity; Determining the user feature vector of the user tower in the two-tower recall model according to the user's historical behavior sequence on the second e-commerce platform and the hierarchical encoding and its sliding window representation; Determining the commodity feature vector of the commodity tower in the two-tower recall model according to the hierarchical encoding and its sliding window representation; Determining the commodity feature vector that matches the user feature vector, and obtaining the target recommended commodity according to the determined commodity feature vector.

2. The product recommendation method according to claim 1, wherein The performing model pre-training based on the multi-modal information of the target commodity on the first e-commerce platform to obtain a multi-modal pre-trained large model includes: Encoding the image modal information of the target commodity based on a preset image encoder to obtain the visual representation of the target commodity; Performing alignment processing on the visual representation and the text representation of the target commodity; Determining the description information of the target commodity based on the large language model according to the aligned visual representation and text encoding instructions; Performing autoregressive learning according to the description information and the text modal information of the target commodity to obtain the multi-modal pre-trained large model.

3. The commodity recommendation method according to claim 1, characterized in that, The performing re-encoding processing on the high-dimensional multi-modal representation to obtain the hierarchical encoding and low-dimensional multi-modal representation of the target commodity: Reducing the dimension of the high-dimensional multi-modal representation to obtain the residual initial representation; Performing multiple residual quantizations based on the residual initial representation to obtain the hierarchical encoding corresponding to the target commodity and the low-dimensional multi-modal representation after cumulative recombination of the hierarchical encoding.

4. The commodity recommendation method according to claim 3, wherein The commodity recommendation method further includes: Performing a sliding window process on the hierarchical encoding to obtain the sliding window representation.

5. A commodity recommendation system based on multimodal information, characterized in that, Including: A multi-modal model training module that performs model pre-training based on the multi-modal information of the target commodity on the first e-commerce platform to obtain a multi-modal pre-trained large model; A multi-modal representation processing module that determines the high-dimensional multi-modal representation of the target commodity according to the multi-modal pre-trained large model, and performs re-encoding processing on the high-dimensional multi-modal representation to obtain the hierarchical encoding and low-dimensional multi-modal representation of the target commodity; A user tower processing module that determines the user feature vector of the user tower in the two-tower recall model according to the user's historical behavior sequence on the second e-commerce platform and the hierarchical encoding and its sliding window representation; A commodity tower processing module that determines the commodity feature vector of the commodity tower in the two-tower recall model according to the low-dimensional hierarchical encoding and its sliding window representation; A commodity recommendation processing module that determines the commodity feature vector that matches the user feature vector, and obtains the target recommended commodity according to the determined commodity feature vector.

6. The product recommendation system according to claim 5, wherein The multi-modal model training module encodes the image modal information of the target commodity based on a preset image encoder to obtain the visual representation of the target commodity; The multi-modal model training module aligns the visual representation with the text representation of the target commodity; The multi-modal model training module determines the description information of the target commodity based on the large language model according to the aligned visual representation and the text encoding instruction; The multi-modal model training module performs autoregressive learning according to the description information and the text modality information of the target commodity to obtain the multi-modal pre-trained large model.

7. The merchandise recommendation system according to claim 5, wherein The multi-modal representation processing module reduces the dimension of the high-dimensional multi-modal representation to obtain an initial residual representation; The multi-modal representation processing module performs multiple residual quantizations based on the initial residual representation to obtain the hierarchical encoding corresponding to the target commodity and the low-dimensional multi-modal representation after the hierarchical encoding cumulative recombination.

8. The commodity recommendation system according to claim 7, characterized in that, The multi-modal representation processing module performs a sliding window process on the hierarchical encoding to obtain the sliding window representation.

9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the commodity recommendation method according to any one of claims 1-4.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the commodity recommendation method according to any one of claims 1-4.

Citation Information

Cited By

  • Generation method and recommendation method of commodity reasoning identifier and related device

    CN121544353A