Item recommendation method and device, item training method and device, electronic device, and storage medium

CN122779949APending Publication Date: 2026-09-18BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611240496.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]然而,上述大型语言模型对用户偏好的理解不够全面,也就导致了推荐系统的准确性很低,因此,亟需一种准确性更高的物品推荐系统

Benefits of technology

[0023]The technical solution provided in this disclosure uses an item recommendation model to process text and visual data to obtain multimodal features. These multimodal features fully integrate the characteristics of various modalities, comprehensively capture user preferences, and predict target title features and target image features based on the multimodal features. Based on these features, the item recommendation model selects items to be recommended from a candidate item library. Since the predicted titles and images are mutually constrained, the prediction bias of a single modality can be reduced, thereby improving the accuracy of the recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779949A_ABST
    Figure CN122779949A_ABST
Patent Text Reader

Abstract

The present disclosure provides an article recommendation method and device, an article training method and device, an electronic device and a storage medium, and belongs to the technical field of Internet. The technical scheme provided by the embodiments of the present disclosure processes text data and visual data based on an article recommendation model to obtain multi-modal features. The multi-modal features fully fuse the characteristics of multiple different modalities, comprehensively capture user preferences, and predict target title features and target image features based on the multi-modal features by the article recommendation model. Based on these features, the article to be recommended is selected from a candidate article library. Since the predicted title and image constrain each other, the single-modal prediction deviation can be reduced, thereby improving the recommendation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a method and apparatus for recommending items, a method and apparatus for training items, an electronic device, and a storage medium. Background Technology

[0002] In recent years, Large Language Models (LLMs) have driven a paradigm shift in the recommender system field due to their powerful semantic understanding, world knowledge, and reasoning capabilities. LLM-based recommendation methods mainly fall into two categories: one utilizes LLMs to extract rich feature representations of historically interacted items to enhance the recommender system's understanding of user preferences and item attributes; the other directly fine-tunes the base LLM to generate information about the next item the user might interact with, i.e., generative recommendation.

[0003] However, the aforementioned large-scale language models do not have a comprehensive understanding of user preferences, which leads to low accuracy in the recommendation system. Therefore, there is an urgent need for a more accurate item recommendation system. Summary of the Invention

[0004] This disclosure provides a method and apparatus for recommending items, a method and apparatus for training items, an electronic device, and a storage medium, which can ensure the accuracy of the recommended items.

[0005] The technical solution disclosed herein is as follows: According to one aspect of the embodiments of this disclosure, an item recommendation method is provided, including: Obtain the historical interaction item sequence of the user to be predicted, the historical interaction item sequence including multiple items that the user has interacted with in the past; The feature extraction unit in the item recommendation model processes the text and visual data of each item in the historical interaction item sequence to obtain multimodal features, which indicate the user's preference for items in terms of text and visual aspects. The multimodal features are processed by the title prediction unit in the item recommendation model to obtain the target title features; The target image features are obtained by reconstructing the image based on the multimodal features of the item using the visual reconstruction unit in the item recommendation model; Based on the target title features and the target image features, joint features are obtained, and based on the joint features, items to be recommended are determined from the candidate item library.

[0006] According to another aspect of the embodiments of this disclosure, an item training method is provided, the method comprising: Obtain the historical interaction item sequence of the sample user, the historical interaction item sequence including multiple sample items that the sample user has interacted with; The feature extraction unit in the item recommendation model processes the text and visual data of each sample item in the historical interaction item sequence to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects. The title prediction unit in the item recommendation model processes the multimodal features of the sample items to obtain the predicted title features. Using the multimodal features of the sample item and multiple sub-image features of the sample item's image as input, the visual reconstruction unit in the item recommendation model performs image reconstruction based on the multimodal features of the sample item to obtain the reconstructed image features of the sample item. The item recommendation model is trained based on the title features of the next sample item of the sample item, the predicted title features, the image features of the next sample item of the sample item, and the reconstructed image features.

[0007] According to another aspect of the embodiments of this disclosure, an item recommendation device is provided, comprising: The sequence acquisition module is configured to acquire the historical interaction item sequence of the user to be predicted, the historical interaction item sequence including multiple items that the user has previously interacted with; The feature extraction module is configured to process the text and visual data of each item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain multimodal features, which indicate the sample users' preferences for items in terms of text and visual aspects. The title feature acquisition module is configured to process the multimodal features through the title prediction unit in the item recommendation model to obtain the target title features; The image feature acquisition module is configured to perform image reconstruction based on the multimodal features of the item through the visual reconstruction unit in the item recommendation model to obtain the target image features; The item determination module is configured to obtain joint features based on the target title features and the target image features, and determine the item to be recommended from the candidate item library based on the joint features.

[0008] In some embodiments, the feature extraction module is configured to: convert the text data and prompt statements of each item in the historical interactive item sequence using the feature extraction unit in the item recommendation model to obtain the text embedding of each item; process the visual data of each item in the historical interactive item sequence using the feature extraction unit in the item recommendation model to obtain the visual features of each item, and project the visual features onto the embedding space of the title prediction unit to obtain the visual embedding of each item; concatenate the text embedding and the visual embedding of each item, and encode the concatenation result to obtain multimodal features.

[0009] In some embodiments, the title feature acquisition module is configured to perform autoregressive prediction on the multimodal features through the title prediction unit in the item recommendation model, and obtain each word in the word sequence of the target title one by one.

[0010] In some embodiments, the image feature acquisition module includes a first reconstruction unit or a second reconstruction unit; The first reconstruction unit includes a denoising subunit and a reconstruction subunit. The denoising subunit is configured to denoise the sub-image features of a preset image with added random noise by using a diffusion module with the multimodal features of the item as a condition. The diffusion module is a visual reconstruction unit in the item recommendation model. The reconstruction subunit is configured to reconstruct the target image features based on the denoising result. The second reconstruction unit generates the target image features based on the multimodal features through the VQ-VAE module, where the VQ-VAE module is the visual reconstruction unit in the item recommendation model.

[0011] In some embodiments, the denoising result includes the denoising result of the plurality of sub-image features, and the reconstruction sub-unit is configured to stitch together the denoising results of the plurality of sub-image features to obtain the target image feature.

[0012] In some embodiments, the item determination module is configured to determine a plurality of first items from a plurality of candidate items in the candidate item library based on the joint features, wherein the similarity between the joint features of each first item and the joint features meets a target condition; and output the first items ranked in the top preset position of similarity as the items to be recommended.

[0013] According to another aspect of the embodiments of this disclosure, an item training device is provided, comprising: The sample sequence acquisition module is configured to acquire the historical interaction item sequence of a sample user, the historical interaction item sequence including multiple sample items that the sample user has interacted with; The sample feature extraction module is configured to process the text and visual data of each sample item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects. The sample title feature acquisition module is configured to process the multimodal features of the sample items through the title prediction unit in the item recommendation model to obtain predicted title features; The sample image feature acquisition module is configured to perform image reconstruction based on the multimodal features of the sample item through the visual reconstruction unit in the item recommendation model, and obtain the reconstructed image features of the sample item. The training module is configured to train the item recommendation model based on the title features of the next sample item of the sample item and the predicted title features, the image features of the next sample item of the sample item and the reconstructed image features.

[0014] In some embodiments, the sample feature extraction module is configured to: convert the text data and prompt statements of each sample item in the historical interactive item sequence using the feature extraction unit in the item recommendation model to obtain the text embedding of each sample item; process the visual data of each sample item in the historical interactive item sequence using the feature extraction unit in the item recommendation model to obtain the visual features of each sample item, and project the visual features onto the embedding space of the title prediction unit to obtain the visual embedding of each sample item; and concatenate the text embedding and the visual embedding of each sample item to obtain the multimodal features of each sample item.

[0015] In some embodiments, the sample title feature acquisition module is configured to perform autoregressive prediction on the multimodal features of the sample item and the word sequence of the title of the sample item in the historical interaction item sequence through the title prediction unit in the item recommendation model, and obtain each word in the word sequence of the title of the predicted item one by one.

[0016] In some embodiments, the sample image feature acquisition module includes a third reconstruction unit or a fourth reconstruction unit; The third reconstruction unit includes a sample denoising subunit and a sample reconstruction subunit. The sample denoising subunit is configured to denoise the sub-image features of a preset image with added random noise by using a diffusion module with the multimodal features of the sample item as a condition. The diffusion module is a visual reconstruction unit in the item recommendation model. The sample reconstruction subunit is configured to reconstruct the sample item based on the denoising result to obtain the reconstructed image features of the sample item. The fourth reconstruction unit is configured to generate reconstructed image features of the sample item based on the multimodal features through the VQ-VAE module, wherein the VQ-VAE module is a visual reconstruction unit in the item recommendation model.

[0017] In some embodiments, the apparatus further includes a feature acquisition module configured to segment the image of the preset image to obtain multiple consecutive sub-images, extract features from each sub-image to obtain multiple sub-image features, and randomly add noise to the multiple sub-image features to obtain multiple sub-image features with added random noise.

[0018] In some embodiments, the denoising result includes the denoising result of sub-image features, and the sample reconstruction subunit is configured to stitch together the denoising results of multiple sub-image features to obtain the reconstructed image features of the sample item.

[0019] In some embodiments, the training module is configured to: obtain a first loss value based on the title features of the next sample item of the sample item and the title features of the next sample item of the sample item, the first loss value indicating the difference between the title features of the next sample item of the sample item and the predicted title features; obtain a second loss value based on the image features of the next sample item of the sample item and the reconstructed image features, the second loss value indicating the difference between the image features of the next sample item of the sample item and the reconstructed image features; obtain a joint loss value based on the first loss value and the second loss value, the joint loss value indicating the textual and visual differences between the next sample item of the sample item and the predicted result; and adjust the model parameters of the item recommendation model based on the joint loss value.

[0020] According to another aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing processor-executable program code; wherein the processor is configured to execute the program code to implement the above-described item recommendation method or item training method.

[0021] According to another aspect of the present disclosure, a computer-readable storage medium is provided that, when program code in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to perform the above-described item recommendation method or item training method.

[0022] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the above-described item recommendation method or item training method.

[0023] The technical solution provided in this disclosure uses an item recommendation model to process text and visual data to obtain multimodal features. These multimodal features fully integrate the characteristics of various modalities, comprehensively capture user preferences, and predict target title features and target image features based on the multimodal features. Based on these features, the item recommendation model selects items to be recommended from a candidate item library. Since the predicted titles and images are mutually constrained, the prediction bias of a single modality can be reduced, thereby improving the accuracy of the recommendation.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0026] Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment.

[0027] Figure 2 This is an architecture diagram of an item recommendation model illustrated according to an exemplary embodiment.

[0028] Figure 3 This is a flowchart illustrating an item recommendation method according to an exemplary embodiment.

[0029] Figure 4 This is a flowchart illustrating another item recommendation method according to an exemplary embodiment.

[0030] Figure 5 This is a flowchart illustrating pre-training based on another visual reconstruction unit according to an exemplary embodiment.

[0031] Figure 6 This is a block diagram illustrating an item recommendation device according to an exemplary embodiment.

[0032] Figure 7 This is a block diagram illustrating an article training device according to an exemplary embodiment.

[0033] Figure 8 This is a block diagram illustrating a terminal according to an exemplary embodiment.

[0034] Figure 9 This is a block diagram illustrating a server according to an exemplary embodiment. Detailed Implementation

[0035] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0036] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0037] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the task execution progress and objects involved in this disclosure were obtained with full authorization.

[0038] LLMs (Large Language Models) are trained on massive amounts of text data and possess powerful semantic understanding, knowledge storage, and reasoning capabilities, enabling them to handle various natural language-related tasks.

[0039] MLLMs (Multi-modal Large Language Models) are a class of artificial intelligence models that can receive and process multiple modal inputs such as text and images, and output corresponding results, possessing the ability to understand and fuse cross-modal information.

[0040] VLMs (Vision-Language Models) integrate visual perception and language understanding capabilities, enabling them to process both image and text information simultaneously and achieve cross-modal semantic alignment and interaction.

[0041] MMSR (Multi-modal Sequential Recommendation) adaptively fuses multimodal features through dual attention map propagation, balancing intra-modal and inter-modal relationships.

[0042] VQ-VAE (Vector Quantized-Variational Autoencoder) is an autoencoder model based on discrete latent representations. It converts continuous data into discrete words through vector quantization and is suitable for tasks such as visual reconstruction and speech synthesis.

[0043] Diffusion models (Denoising Diffusion Probabilistic Models) are a class of generative artificial intelligence models that transform noise into structured output by simulating a stepwise denoising process, and have shown excellent performance in fields such as image generation and visual reconstruction.

[0044] Recommendation Systems (RS) are information filtering systems used to predict users' preferences for items and provide personalized recommendations.

[0045] NDCG@10 (Normalized Discounted Cumulative Gain at 10) is a metric for evaluating the ranking quality of a recommender system. It considers both the relevance and location of the recommended items, with higher scores indicating better performance.

[0046] HR@10 (Hit Rate@10) is a metric for evaluating the coverage of a recommendation system. It calculates the percentage of users whose next recommended item appears in the top 10 recommendation results.

[0047] SASRec (Self-attentive Sequential Recommendation) uses a self-attention mechanism to capture long-term dependencies in user behavior sequences, thereby improving the accuracy of sequence recommendations.

[0048] GRU4Rec (Gated Recurrent Unit for Sequential Recommendation) uses a GRU network to capture short-term behavioral sequence features of users for conversational recommendation tasks.

[0049] DIN (Deep Interest Network) is a deep learning model for click-through rate prediction that improves recommendation performance by capturing users' immediate interests.

[0050] UniMP (Unified Multi-modal Personalization) is a framework that fine-tunes large visual language models to handle diverse personalization tasks, including preference prediction, interpretation generation, and image selection.

[0051] MLLM-MSR (Multi-modal Large Language Model for Multi-modal Sequential Recommendation) dynamically summarizes user preferences through a two-stage approach to achieve multi-modal sequence recommendation.

[0052] ViT (Vision Transformer) uses the Transformer architecture to process image data, segmenting the image into a sequence of patches before extracting and processing features.

[0053] AdamW is an optimizer, an improved version of the Adam optimizer, which enhances the generalization ability of the model through a weight decay mechanism.

[0054] DeepSpeed ​​is a deep learning optimization library that provides features such as distributed training and mixed-precision training to improve the training efficiency of large models. ZeRO Stage 2 is a zero-redundancy optimization strategy in DeepSpeed ​​that improves distributed training performance by optimizing parameter storage and communication.

[0055] Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment. Taking an electronic device provided as a terminal as an example, see [link to example]. Figure 1 The implementation environment specifically includes: terminal 101 and server 102. Terminal 101 and server 102 can be connected directly or indirectly through wired or wireless communication, which is not limited herein.

[0056] Terminal 101 is at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player, MP4 player, and laptop computer. Terminal 101 runs an application that supports item recommendation. Users can log in to this application through Terminal 101 to access the services provided by the application. For example, users can browse and interact with items through the application on Terminal 101. Alternatively, Terminal 101 can also run an application that supports item recommendation model calls, enabling item recommendations by invoking the item recommendation model.

[0057] Terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be several terminals, or dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.

[0058] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 provides backend services for item recommendation, such as model training and model updates. Server 102 can connect to terminal 101 via a wireless or wired network. Server 102 can provide database services to the terminal, and of course, it can also provide other types of services, such as data preprocessing. In some embodiments, the number of servers may be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In other embodiments, server 102 can also serve as a carrier for model training to implement the model training process, so as to provide the trained item recommendation model to the terminal for use.

[0059] This disclosure proposes a multimodal generative item recommendation method. By integrating text generation and visual reconstruction tasks and introducing diffusion units to enhance visual generation capabilities, it achieves deep fusion and efficient utilization of multimodal information, thereby improving recommendation accuracy. The overall architecture of the item recommendation model designed in this disclosure is as follows: Figure 2 As shown, the model mainly consists of three core parts: a feature extraction unit for modeling user preferences; a title prediction unit for predicting item titles; and a visual reconstruction unit for visually reconstructing the predicted items based on fine-grained diffusion. These three parts work together to simultaneously capture both the user's textual and visual preferences, outputting recommendation results that are both consistent with user interests and visually coherent. This model architecture can be viewed as a type of MMGR (Multi-modal Generative Recommendation), which enhances the recommendation system's ability to capture user preferences by integrating text generation and visual reconstruction tasks.

[0060] The following section explains the item recommendation process based on the model structure described above. (See also...) Figure 3 The subject executing this item recommendation method can be either a terminal or a server, and the method includes the following steps.

[0061] In step 301, the historical interaction item sequence of the user to be predicted is obtained. The historical interaction item sequence includes multiple items that the user has interacted with.

[0062] In this step, to recommend items to a specific user, data collection and preprocessing for that user are necessary. This includes collecting the user's historical item interaction sequence, which contains items the user has interacted with on the current online platform (or an authorized online platform), such as purchased or favorited items. Each item's information includes textual data, such as the title and word segmentation results of comments, as well as visual data, such as images or videos of the item.

[0063] It should be noted that the above data is all data that has been fully authorized by the user, and the data can be anonymized. That is, the above data does not include sensitive information such as the user's personal information and financial information.

[0064] In step 302, the text and visual data of each item in the historical interaction item sequence are processed by the feature extraction unit in the item recommendation model to obtain multimodal features. The multimodal features indicate the user's preference for items in terms of text and visual aspects.

[0065] By analyzing the aforementioned textual and visual data, we can understand users' preferences for items in terms of both text and visuals. Since text and visuals are important factors in attracting users to interact, we can process the textual and visual data of items to provide a reference for subsequent recommendations, thereby giving us a feature that can comprehensively describe user preferences from multiple dimensions.

[0066] The following describes each processing step in the feature extraction unit. Step 302 includes steps 3021 to 3023.

[0067] In step 3021, the text data and prompts of each item in the historical interactive item sequence are transformed by the feature extraction unit in the item recommendation model to obtain the text embedding of each item.

[0068] For item recommendation models, they can perform corresponding recommendation steps based on prompts, and can also perform word segmentation based on text data and prompts, and convert the segmentation results into word embedding sequences, that is, obtain the text embedding of each item.

[0069] In step 3022, the visual data of each item in the historical interactive item sequence is processed by the feature extraction unit in the item recommendation model to obtain the visual features of each item, and the visual features are projected into the embedding space of the title prediction unit to obtain the visual embedding of each item.

[0070] This feature extraction unit includes multiple visual encoders that process visual data to obtain video features for each item. This visual data includes both images and videos; video processing can be achieved by manipulating video frames. To further facilitate the subsequent title prediction unit, a linear layer is used to process the visual features, projecting them into the embedding space of the title prediction unit to obtain a visual embedding.

[0071] In step 3023, the text embedding and the visual embedding of each item are concatenated, and the concatenation result is encoded to obtain multimodal features.

[0072] This splicing process can fuse text embeddings and visual embeddings to obtain the multimodal features of the object. This splicing can be direct, that is, using text embedding h... t and visual embedding h v By concatenating the components, we obtain the multimodal embedding h=[h t ;h v Furthermore, the multimodal embedding is encoded using an encoder (containing a self-attention mechanism and a feedforward network) to generate hidden states z that contain user textual and visual preferences, i.e., multimodal features z. This step comprehensively characterizes user preferences from both textual and visual perspectives, providing more complete information and capturing user preferences more accurately than unimodal methods.

[0073] In step 303, the multimodal features are processed by the title prediction unit in the item recommendation model to obtain the target title features.

[0074] In step 303, the title prediction unit in the item recommendation model performs autoregressive prediction on the multimodal features to obtain each word in the word sequence of the target title features.

[0075] The title prediction unit in step 303 above can be regarded as a text generation model. It can use an autoregressive prediction method to generate title text word by word based on the input multimodal features. This title text is the title of the next item that the user may buy, predicted by the item recommendation model.

[0076] This step is based on multimodal features that contain user textual and visual preferences for prediction. Multimodal features cover more comprehensive information, which improves the accuracy of prediction. Furthermore, the fused multimodal model can better reflect the underlying semantic and visual attributes.

[0077] In some embodiments, the item recommendation model may include MLLMs, which include feature extraction units and title prediction units, which may be language model heads (LM heads). A language model head (LM head) is connected to the output of the MLLMs. The language model head (LM head) is used to generate the title words of the next item in an autoregressive manner based on the multimodal features obtained from the MLLMs, until a complete title sequence is generated.

[0078] For example, the LM Head includes a linear layer and a normalization layer (Softmax). The linear layer is used to map multimodal features to the size of the pre-trained vocabulary of the language model head to obtain the score of each token in the vocabulary. The normalization layer is used to convert the score of each token into a probability between 0 and 1, and the sum of the probabilities of all tokens is 1.

[0079] Based on the aforementioned LM Head, the steps to obtain each token in the token sequence of the target title feature include: mapping the multimodal features to the size of the pre-trained vocabulary of the language model head through the LM Head, and obtaining the score of each token in the vocabulary; converting the score of each token into a probability through the LM Head; outputting the first token w1 in the token sequence of the target title feature based on the probability of each token through the LM Head (for example, outputting the token with the highest probability as w1); fusing w1 and the multimodal features with MLLMs to obtain z1, and predicting the second token w2 in the token sequence of the target title feature based on z1 through the LM Head; repeating the above steps, that is, predicting the next token in the token sequence of the target title feature based on the multimodal features and the obtained tokens each time, until the end symbol is predicted or the token sequence of the target title feature reaches a preset value, stopping the loop, concatenating all the predicted tokens, and obtaining the token sequence of the target title feature.

[0080] In step 304, the visual reconstruction unit in the item recommendation model performs image reconstruction based on the multimodal features of the item to obtain the target image features.

[0081] In this step, the target image features obtained through image reconstruction, which are also the item images that match the user's potential preferences, achieve a precise mapping from abstract user preferences to concrete visual items. This has stronger personalized expression capabilities, generalization capabilities, and robustness, and can effectively improve the accuracy and novelty of item recommendations.

[0082] The visual reconstruction unit in this step can be of various models, such as a diffusion module or a VQ-VAE module. The following will use a diffusion module or a VQ-VAE module as the visual reconstruction unit to describe the image reconstruction process based on the visual reconstruction unit in detail. The following describes each processing step in the diffusion module. In some embodiments, step 304 above includes steps 30411 to 30412.

[0083] In step 30411, the diffusion module uses the multimodal features of the item as a condition to denoise the sub-image features of the preset image with added random noise, and obtains the denoising result.

[0084] In this step, the sub-image features of a preset image with added random noise are denoised based on the multimodal features of the item. The random noise is a noise tensor randomly sampled from a standard normal distribution. For example, a preset image is divided into 256 sub-image features. Based on the multimodal features of the item, the 256 sub-image features with added random noise are denoised to obtain multiple denoising results.

[0085] In some embodiments, the diffusion module includes a diffusion head and an output layer, with the diffusion head being the core component of image reconstruction. The diffusion head is based on the SimpleMLPAdaLN architecture, containing three MLP layers (256 to 1536 dimensions), a linear projection layer (1536 to 1536 dimensions), and three AdaLN-modulated residual blocks. Each residual block contains layer normalization (LayerNorm), an extended MLP, and a modulation network (1536 to 4608 dimensions). The output features of the diffusion head ultimately generate multiple denoising results through the output layer (1536 to 3072 dimensions). The diffusion head employs a simplified and lightweight architecture, adding only a small number of parameters and a small amount of inference time. In this embodiment, the diffusion head contains only 50M parameters, and the overall inference time of the item recommendation model increases by only 11.21% compared to basic MLLMs (single-user inference time is 0.2837s), resulting in minimal loss of inference efficiency. It can meet the needs of large-scale batch processing, and the diffusion head can be seamlessly integrated into existing item recommendation models. By organically integrating the diffusion head with MLLMs, the item recommendation model possesses fine-grained visual generation capabilities, avoiding the quantization loss of traditional visual reconstruction methods and significantly improving the generation quality of target image features and their matching degree with user preferences. The network structure of this diffusion head is an implementable architecture; the diffusion head can be other architectures, such as different numbers of MLP layers or residual blocks included in the diffusion head, etc., which are not specifically limited in this embodiment.

[0086] Based on the aforementioned diffusion module, the specific steps for denoising include: performing linear transformation and nonlinear activation on the 256-dimensional spliced ​​features of multiple sub-image features and multimodal features of the object image using a 3-layer MLP, thereby increasing the dimensionality of the 256-dimensional spliced ​​features to 1536 dimensions, and fusing the residual blocks with the multimodal features; performing a linear transformation on the output features of the 3-layer MLP using a linear projection layer to correct the feature distribution; sequentially inputting the output features of the linear projection layer into three residual blocks, performing layer normalization, feature enhancement, and preference modulation on the input features through each residual block to denoise and extract visual features; and performing a linear transformation on the output features of the residual blocks using an output layer to map the 1536-dimensional features to 3072 dimensions, obtaining the denoised result. In this step, the MLP and linear projection layer ensure that the feature dimensions and numerical ranges are adapted to the processing requirements of the subsequent residual blocks, and then the three cascaded residual blocks are used for denoising and to extract clear visual features.

[0087] The processing of input features through each residual block includes: performing a linear transformation on the input features through a linear projection layer to calibrate their feature distribution and eliminate the distribution shift after dimensionality upscaling through a 3-layer MLP; performing layer normalization on the output features of the linear projection layer through layer normalization of the residual block to stabilize the feature distribution; expanding the output features of the layer normalization from 1536 dimensions to 4608 dimensions through the extended MLP of the residual block, and then compressing them back to 1536 dimensions to capture the fine-grained details (such as texture and shape) of the normalized features; and generating normalized scaling and shift parameters through the modulation network of the residual block conditioned on multimodal features, and processing the output features of the extended MLP according to the scaling and shift parameters. In this step, each residual block completes a round of feature denoising and detail enhancement (i.e. visual feature extraction) through layer normalization, extended MLP and modulation network. Through three cascaded residual blocks, noise in the input features is gradually removed, and clear target image features that meet user preferences are reconstructed based on multimodal features.

[0088] For example, the denoising result is calculated using the following formula: Formula 1:

[0089] in, For the tth Denoising sub-image features at time step 1; t is the preset time step; σ t The noise level is set for the preset time step t; Add sub-image features with random noise to a preset time step t; The cumulative noise attenuation coefficient for a preset time step t; This refers to random sampling noise added during the back diffusion process; The prediction noise for the item recommendation model (achieved through a diffusion head), that is, based on... Noise predicted by t and z; α t Let be the noise attenuation coefficient at time step t; δ ~ N(0, I); random noise x T ~N(0,I).

[0090] In this step, denoising and reconstruction are performed based on the multimodal features of the items. The resulting target image features are highly aligned with the user's interests, enabling personalized recommendations for the user with high accuracy.

[0091] In step 30412, reconstruction is performed based on the denoising results to obtain the target image features.

[0092] In step 30412, the denoising result includes the denoising results of multiple sub-image features. The denoising results of multiple sub-image features are concatenated to obtain the target image feature V. u n+1 .

[0093] In this step, the denoising results are stitched together according to the original position order of the sub-image features of the preset image to obtain the target image features. By denoising and stitching together the sub-image features, fine-grained visual feature reconstruction can be achieved while reducing computational complexity, allowing for more accurate control over user preferences. Furthermore, the denoising error of a single sub-image feature only affects the local area and does not contaminate the entire target image features. Compared to global modeling, it is less sensitive to data noise and training fluctuations, thus improving robustness.

[0094] The diffusion module enables fine-grained target image feature generation, resulting in higher visual reconstruction accuracy. Compared to discretization methods such as VQ-VAE, it avoids quantization information loss, generates visual features with higher fidelity, and significantly improves the matching degree with user visual preferences. The diffusion model enhances the fineness of visual reconstruction, further improving the accuracy and effectiveness of item recommendation models.

[0095] The following describes the various processing steps in the VQ-VAE module. In some embodiments, step 304 includes generating target image features based on multimodal features using the visual reconstruction unit in the item recommendation model. This method can generate target image features more quickly.

[0096] In some embodiments, the VQ-VAE module includes a decoder for generating target image features based on multimodal features. The decoder can be a convolutional neural network (CNN) or a U-Net structure.

[0097] In step 305, joint features are obtained based on the target title features and target image features, and items to be recommended are determined from the candidate item library based on the joint features.

[0098] In this step, the joint features include information on the target title features and the target image features, which can more completely cover the title and image information of the next item that the user may buy. Based on this, the identified items to be recommended are more accurate.

[0099] The following describes each processing step in the feature extraction unit. Step 305 includes steps 3051 to 3052.

[0100] In step 3051, based on joint features, multiple first items are determined from multiple candidate items in the candidate item library, and the joint features of each first item and the similarity of the joint features meet the target conditions.

[0101] In this step, the higher the similarity between the joint features of candidate items, the more characteristics the first item represents. The first item is determined based on the similarity meeting the target condition, making the first item very likely to be the next item the user will purchase.

[0102] In step 3052, the first item ranked first in the similarity ranking is output as the item to be recommended.

[0103] In this step, the first item ranked in the top preset position based on similarity is output as the recommended item, which improves the accuracy of the output recommended items.

[0104] The steps of the item recommendation method have been described above using the prediction process. Below, using an e-commerce toy recommendation scenario as an example, the item recommendation method of this embodiment will be explained. The method includes: 1. Obtain user A's historical interactive toy sequence. The historical interactive toy sequence includes 5 items, namely: Item 1: Title "Melissa & Doug See & Spell", corresponding image; Item 2: Title "Fisher-Price Kid-Tough See Yourself Camera, Purple", corresponding image; Item 3: Title "KidKraft Large Kitchen", corresponding image; Item 4: Title "Melissa & Doug Sunny Patch Blaze Firefly Flashlight", corresponding image; Item 5: Title "Learning Resources Healthy Dinner Basket", corresponding image.

[0105] 2. The title of each item is segmented using the feature extraction unit in the item recommendation model to generate a text embedding h. t The image of each item is preprocessed and segmented, and visual features are generated using a ViT encoder. After projection, the visual embedding h is obtained. v The text embedding and visual embedding are concatenated to obtain the multimodal input h. The multimodal input h is then input into the Qwen2.5-VL-2B model to encode the hidden state z of user preferences.

[0106] 3. The title prediction unit (LM Head) in the item recommendation model generates the predicted title embedding of the next item that the user may buy based on the hidden state z of the user's preferences through autoregression. The text corresponding to the title embedding is "LearningResources Healthy Lunch Basket". The visual embedding v1 of an item is generated by the visual reconstruction unit (diffusion head) in the item recommendation model based on the user preference hidden state z through a backdiffusion process. 6 v2 6 , ..., v k 6 .

[0107] 4. Combine the predicted visual embedding and title embedding to obtain joint features. Calculate the similarity between the joint features and all items in the toy candidate library. After sorting, select the top 10 items to recommend to user A.

[0108] The technical solution provided in this disclosure uses an item recommendation model to process text and visual data to obtain multimodal features. These multimodal features fully integrate the characteristics of various modalities, comprehensively capture user preferences, and predict target title features and target image features based on the multimodal features. Based on these features, the item recommendation model selects items to be recommended from a candidate item library. Since the predicted titles and images are mutually constrained, the prediction bias of a single modality can be reduced, thereby improving the accuracy of the recommendation.

[0109] The method in this embodiment is applicable to various scenarios such as e-commerce, short videos, and content recommendation, and has a good recommendation effect on different types of items (toys, sporting goods, videos, etc.).

[0110] This disclosure addresses the lack of visual reconstruction capabilities in existing technologies by performing image reconstruction based on the multimodal features of items. It empowers item recommendation models with the ability to generate high-fidelity target image features, thereby improving the matching degree between recommendation results and user visual preferences.

[0111] Experimental results on multiple public datasets demonstrate that the recommendation performance of the method in this embodiment is significantly improved. For example, on the first dataset, HR@10 reaches 0.0963 (a 6.9% improvement over the best baseline) and NDCG@10 reaches 0.0535 (a 16.8% improvement over the best baseline); on the second dataset, HR@10 reaches 0.1020 (a 7.9% improvement over the best baseline) and NDCG@10 reaches 0.0587 (a 7.9% improvement over the best baseline); and on the third dataset, HR@10 reaches 0.0385 (a 4.6% improvement over the best baseline) and NDCG@10 reaches 0.0204 (a 6.8% improvement over the best baseline).

[0112] Figure 4 This is a block diagram illustrating an item training method according to an exemplary embodiment. See also... Figure 4 The training methods for this item include: In step 401, a training set is obtained, which includes multiple training samples. Each training sample includes a sequence of historical interaction items of the sample user. The sequence of historical interaction items includes multiple sample items that the sample user has interacted with.

[0113] In this step, the sample items are the items actually purchased by the sample users. The sample items are sorted by purchase time to form a historical interaction item sequence. The last item or the last preset number of items in the historical interaction item sequence are directly used as labels, so that the distribution learned by the item recommendation model is completely aligned with the real recommendation scenario, avoiding the bias caused by false labels and weak supervision.

[0114] The sequence of historical interactive items can be represented as S. u =[I u 1 ,I u 2 ,…,I u n ], where each item I u i (i=1,2,…,n) contains text data W u i =[w1 i w2 i ,…,w m i (e.g., word segmentation results for titles and comments) and visual data V ui =[v1 i v2 i ,…,v k i (e.g., patch lexical units of an image), where w j i This refers to the word segmentation results of the text data of the item (such as title, review), where m is the total number of tokens after segmentation, and v is the word segmentation result. l i Sub-image features are the visual data (images) of an item, such as k being the total number of patch words or sub-image features of the image.

[0115] In some embodiments, the sample items in the historical interaction item sequence of each sample user are arranged in chronological order. The length of the historical interaction item sequence is 5-20, that is, the historical interaction item sequence includes 5-20 sample items. When the length of the historical interaction item sequence is insufficient, it is filled with placeholders, and the part that exceeds the length of the historical interaction item sequence is truncated.

[0116] In some embodiments, the method further includes: obtaining prompt words, the prompt words including at least one of a task description and a task-specific prompt, wherein the task description includes predicted title features and reconstructed image features of an item recommendation model predicting the next item a sample user will purchase based on a sequence of historically interacted items by the sample user, and the task-specific prompt indicates the items included in the output of the item recommendation model. For example, referring to Table 1, for a historically interacted item sequence of length 20, the task description could be: the item recommendation model predicts the next item a sample user will purchase based on the first 19 sample items in the historically interacted item sequence, where the 20th sample item can be used as a ground truth label.

[0117] Table 1 Examples of prompt words

[0118] In step 402, the text and visual data of each sample item in the historical interaction item sequence are processed by the feature extraction unit in the item recommendation model to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects.

[0119] The principle for obtaining multimodal features is the same as that for obtaining multimodal features in the above embodiments, and will not be repeated here.

[0120] In step 403, the image of the sample item is segmented to obtain multiple consecutive sub-images. Feature extraction is performed on each sub-image to obtain multiple sub-image features. Noise is added to the multiple sub-image features at a preset time step to obtain multiple noisy sub-image features.

[0121] In this step, the image of the sample item is divided into multiple sub-images, and noise is added to the features of the sub-images. This allows the visual reconstruction unit in the item recommendation model to process the features of the noisy sub-images as the smallest unit. This reduces the computational cost of word processing in the visual reconstruction unit. Furthermore, the denoising error of a single sub-image feature only affects the local area. Compared with global modeling, it is less sensitive to data noise and training fluctuations, thus improving robustness.

[0122] For example, the following Formula 2 is used to calculate the noisy sub-image feature x at a preset time step t. t : Formula 2: x t =

[0123] Where t is the preset time step; x is the sub-image feature; For the added random noise, ε ~ N(0,I); The cumulative noise attenuation coefficient is the preset time step t, which is also the noise scheduling table.

[0124] In step 404, the visual reconstruction unit in the item recommendation model performs image reconstruction based on multiple noisy image features to obtain the predicted target item features.

[0125] In this step, the visual reconstruction unit in the item recommendation model is a diffusion module. The principle of image reconstruction is the same as in the above embodiment, and will not be repeated here.

[0126] In step 405, the true loss value is obtained based on the multimodal features and the predicted target item features. The true loss value indicates the difference between the multimodal features and the predicted target item features.

[0127] In this step, the smaller the actual loss value, the closer the multimodal features and the predicted target item features are, indicating that the denoising process of the visual reconstruction unit is closer to the user's preferences.

[0128] For example, the true loss value L(z,x) is calculated using the following formula three: Formula 3: L(z,x)=E ε,t [‖ε-ε θ (x t |t,z)_‖ 2 ] Where, ε θ ε represents the prediction noise of the visual reconstruction unit; x represents random noise; t For noisy image features; t is the preset time step; z is the multimodal feature; E ε,t [ The instruction is to take the mathematical expectation of the random noise ε and the diffusion time step t. || 2The square of the indicator norm.

[0129] In step 406, the parameters of the visual reconstruction unit in the item recommendation model are adjusted based on the true loss value.

[0130] In this step, by adjusting the parameters in a direction that makes the actual loss value smaller and smaller, the noise predicted by the visual reconstruction unit becomes closer and closer to the actual added random noise, and the visual reconstruction unit can learn a stable and accurate denoising capability.

[0131] In step 407, the iteration is performed according to steps 403 to 406 above until the number of iterations reaches a preset value, or the reduction in the actual loss value in a consecutive preset number of rounds is less than a preset magnitude.

[0132] Steps 402 to 407 above are the pre-training process. Through steps 402 to 405, the visual reconstruction unit learns to remove noise from the noisy image features, that is, it learns the basic visual denoising and reconstruction capabilities. Then, subsequent joint training is carried out to ensure stable training, high generation quality, and more reliable modal alignment.

[0133] In step 408, the multimodal features of the sample items are processed by the title prediction unit in the item recommendation model to obtain the predicted title features.

[0134] The principle for obtaining the predicted title features is the same as the principle for obtaining the target title features in the above embodiments, and will not be repeated here.

[0135] In step 409, the visual reconstruction unit in the item recommendation model performs image reconstruction based on the multimodal features of the sample items to obtain the reconstructed image features of the sample items.

[0136] For example, the image reconstruction method in step 409 includes the following two methods: The first method involves using a diffusion module to denoise sub-image features of a preset image with added random noise, based on the multimodal features of the item. Reconstruction is then performed based on the denoising results to obtain the reconstructed image features of the sample item. The diffusion module is the visual reconstruction unit in the item recommendation model. The principle behind obtaining the reconstructed image features in this embodiment is the same as the principle behind obtaining the target image features in the above embodiments, and will not be elaborated upon here.

[0137] In some embodiments, the method further includes: segmenting the preset image to obtain multiple consecutive sub-images; extracting features from each sub-image to obtain multiple sub-image features; and randomly adding noise to the multiple sub-image features to obtain multiple sub-image features with added random noise. The preset image can be a blank image of the same size as the image corresponding to the reconstructed image features. By randomly adding noise to the preset image, a denoising base image is provided for the visual reconstruction unit.

[0138] The second method generates reconstructed image features of sample items based on multimodal features using the VQ-VAE module, which is the visual reconstruction unit in the item recommendation model. The principle of obtaining reconstructed image features in this embodiment is the same as that of obtaining target image features in the above embodiments, and will not be repeated here.

[0139] In step 410, the parameters of the item recommendation model are adjusted based on the title features and predicted title features of the next sample item of the sample item, and the image features and reconstructed image features of the next sample item of the sample item.

[0140] The next sample item in the sample item refers to the sample item that was used as the real label in the historical interaction item sequence.

[0141] In this step, the parameters of the item recommendation model are adjusted through text generation (title features of the next sample item and predicted title features) and visual reconstruction (image features of the next sample item and reconstructed image features). The item recommendation model can simultaneously capture the user's textual and visual preferences, make fuller use of multimodal information, and make the recommendation results more in line with the user's actual preferences.

[0142] The training process of the item recommendation model will be described below. Step 410 above includes steps 4101 to 4104.

[0143] In step 4101, a first loss value is obtained based on the title features of the next sample item and the predicted title features of the sample item. The first loss value indicates the difference between the title features of the next sample item and the predicted title features of the sample item.

[0144] In this step, the title feature of the next sample item is used as the true label. The first loss value reflects the gap between the predicted result (predicted title feature) and the true label. This gap can guide the training of the model more accurately.

[0145] For example, the first loss value L is calculated using the following formula four. T : Formula 4: L T =

[0146] Among them, w t The t-th term of the title feature of the next sample item; w <t h represents the target title feature; h represents the multimodal embedding. Based on target title features w <t When using multimodal embedding h, the predicted probability of the t-th term of the title feature of the next sample item.

[0147] For example, if the target title (i.e., the title of the real tag) is "red dress, new summer style", the real word sequence after word segmentation is: w1=red, w2=dress, w3=summer, w4=new style; at step t, the item recommendation model will output the probability distribution of the entire word sequence: P(red)=0.7, P(blue)=0.1, P(dress)=0.05…, only the real word w is taken. t The probability of the first loss value is calculated as follows: at t=1, only P(w1=red), at t=2, only P(w2=dress), at t=3, only P(w3=summer), and at t=4, only P(w4=new style). Substituting w1, w2, w3, and w4 into Formula 4 yields the first loss value. The higher the probability of the true word element predicted by the item recommendation model, the smaller the first loss value; conversely, the lower the probability of the true word element predicted by the item recommendation model, the larger the first loss value. Therefore, maximizing the prediction probability of the true word element sequence by the item recommendation model can improve the accuracy of the predicted target title features.

[0148] In step 4102, a second loss value is obtained based on the image features of the next sample item and the reconstructed image features of the sample item. The second loss value indicates the difference between the image features of the next sample item and the reconstructed image features of the sample item.

[0149] In this step, the image features of the next sample item are used as the true label, and the second loss value reflects the gap between the predicted result (reconstructed image features) and the true label. This gap can guide the training of the model more accurately.

[0150] For example, the second loss value L is calculated using the following formula five. D : Formula 5: L D =-

[0151] Where, x i x is the i-th sub-image feature of the image features of the next sample item; <iis a sub-image feature in the target image features; n is the number of sub-image features included in the image features of the next sample item of the sample item; For sub-image features x in the target image features <i The i-th sub-image feature (x) of the image features of the next sample item predicted from the sample item. i The probability of ). This involves positionally summing the logarithms of the probabilities of each sub-image feature for the target image features.

[0152] In step 4103, a joint loss value is obtained based on the first loss value and the second loss value. The joint loss value indicates the textual and visual differences between the sample item and the prediction result.

[0153] In this step, joint training of dual tasks is achieved through joint loss. The item recommendation model can simultaneously capture users' textual and visual preferences, making fuller use of multimodal information and solving the problem of using only one type of relevant technical information. The recommendation results are more in line with users' actual preferences.

[0154] For example, the joint loss value L is calculated using the following formula six: Formula 6: L=L T +L D Among them, L T This is the first loss value; L D This is the second loss value.

[0155] For example, the joint loss value L is calculated using the following formula seven: Formula 7: L = λ × L T +(1-λ)×L D Among them, L T This is the first loss value; L D λ is the second loss value; λ is the weight, λ∈[0,1].

[0156] In step 4104, the model parameters of the item recommendation model are adjusted based on the joint loss value.

[0157] In this step, by adjusting the model parameters of the item recommendation model in the direction that reduces the joint loss value, the accuracy of the item recommendation model can be gradually improved.

[0158] In step 411, the iteration is performed according to steps 408 to 410 above until the number of iterations reaches a preset value, or the reduction of the joint loss value in a consecutive preset number of rounds is less than a preset magnitude.

[0159] For example, during training, the AdamW optimizer is used to jointly fine-tune all parameters of the item recommendation model, with a learning rate of lr=1e-5 (i.e., the step size for each parameter update is 0.00001), a batch size of 1 per GPU (i.e., a single GPU processes only 1 training sample at a time), training for 3 epochs (i.e., the number of iterations is 3, and one complete traversal of the training set is considered one iteration), and 4 gradient accumulation steps (i.e., accumulating the gradients of 4 training samples before updating the parameters all at once); distributed training is performed using DeepSpeed's ZeRO Stage 2 optimization to improve training efficiency.

[0160] Steps 403 to 407 above describe the pre-training process using the visual reconstruction unit as a diffusion module as an example. Below, the pre-training steps will be explained using the visual reconstruction unit as a VQ-VAE module, referring to... Figure 5 ,include: In step 501, the visual data of the sample items is converted into discrete words using the VQ-VAE module in the item recommendation model.

[0161] In some embodiments, the VQ-VAE module includes an encoder, a quantization layer, and a decoder. The encoder is used to extract features from the visual data of the sample items to obtain an m×n feature map, which includes m×n target visual features at different locations. The quantization layer is used to perform a linear transformation on each target visual feature, and then calculate the Euclidean distance between the linearly transformed target visual feature and all vectors in the codebook of the VQ-VAE module. The discrete word is output based on the vector with the smallest Euclidean distance. The encoder is used to generate target image features based on the visual word sequence.

[0162] In step 502, discrete lexical units are added to the vocabulary of the item recommendation model. The item recommendation model generates a visual lexical sequence through autoregression based on multimodal features and the vocabulary.

[0163] In this step, the original vocabulary of the item recommendation model includes the lexical units corresponding to the text data of each item, but does not include the lexical units corresponding to the visual data of each item. By adding discrete lexical units to the vocabulary of the item recommendation model, the visual lexical unit sequence generated by the item recommendation model can contain the user's textual preferences and visual preferences.

[0164] In step 503, target image features are generated based on visual lexical sequences using residual blocks in the item recommendation model.

[0165] In this step, the VQ-VAE module in the item recommendation model retrieves corresponding discrete vectors from the discrete codebook based on the visual word sequence. These discrete vectors are then used to reconstruct the visual word sequence into a continuous sequence of visual feature vectors. The VQ-VAE module then concatenates these continuous feature vector sequences to obtain the target image features. Both the vocabulary and the discrete codebook are trainable parameters of the item recommendation model.

[0166] In step 504, a reconstruction loss value is obtained based on the visual features of the sample item and the features of the target image. The reconstruction loss value indicates the difference between the visual features of the sample item and the features of the target image. The visual features of the sample item are the stitched features of multiple target visual features. Based on the reconstruction loss value, the parameters of the visual reconstruction unit in the item recommendation model are adjusted.

[0167] In this step, by adjusting the parameters in a direction that makes the reconstruction loss value smaller and smaller, the parameters of the VQ-VAE module can be adjusted, which can gradually improve the reconstruction accuracy of the VQ-VAE module.

[0168] For example, the reconstruction loss value L is calculated using the following formula eight. recon : Formula 8: L recon =‖x- || 2 Where x represents the visual feature of the next sample item; Features of the target image.

[0169] In step 505, the process is iterated according to steps 501 to 504 until the number of iterations reaches a preset value, or the reduction in reconstruction loss value is less than a preset value in a consecutive preset number of rounds.

[0170] Steps 501 to 505 above constitute the pre-training process. During this pre-training process, the VQ-VAE module learns to encode the visual data of the sample items into discrete tokens more accurately, and then decodes and restores them to ensure high-quality reconstruction.

[0171] This pre-training process enables the vocabulary of the item recommendation model to include more comprehensive visual data terms, improving the accuracy of discrete terms converted by the VQ-VAE module and the generated target image features.

[0172] The training method in the above embodiments involves jointly fine-tuning all parameters of the item recommendation model. In addition to this method, a two-stage training method can also be used, which includes the following steps: in the first stage, only the feature extraction unit and the title prediction unit are trained; in the second stage, the parameters of the feature extraction unit and the title prediction unit are frozen, and only the visual reconstruction unit is trained; the total loss function adopts an adaptive weighting strategy, which is Equation 7 above.

[0173] The above training methods can reduce the risk of gradient conflicts during training and have higher training stability, but the training cycle is longer. The adaptive weighting strategy can dynamically adjust the optimization focus according to the task difficulty. Its performance on some datasets (such as Amazon-Sport) is close to that of the original scheme, but it is slightly inferior to the joint training scheme in terms of the depth of multimodal information fusion.

[0174] In some embodiments, the MLLMs in the item recommendation model can be replaced with other pre-trained multimodal language models, such as LLaVA-NeXT and Rec-GPT4V, while keeping the structure of the title prediction unit and visual reconstruction unit unchanged. Only the pre-trained weights and input-output adaptation layers of the item recommendation model are adjusted. The item recommendation model after the MLLMs are replaced can achieve similar item recommendation functions as the original item recommendation model. The performance differences between different base models mainly stem from their pre-training data and architecture design. LLaVA-NeXT may have a slight advantage in inference accuracy, but the training cost is higher. Rec-GPT4V performs well in visual understanding and is suitable for scenarios with higher visual requirements.

[0175] In some embodiments, the item recommendation model uses Qwen2.5-VL-2B as a pre-trained MLLM, and its visual encoder is a ViT architecture. The maximum input sequence length of the item recommendation model is set to 6144, and the maximum output sequence length is set to 512. The hidden layer dimension of the diffuse head of the item recommendation model is 1536, the number of residual blocks is 3, the output dimension of the modulation network is 4608, and the final output dimension is 3072; a noise scheduling table is used. A linear scheduling strategy is adopted, with a time step T=1000.

[0176] Based on any training method in this embodiment, the hardware environment can be: a Linux server equipped with 8 NVIDIA A800 80GB GPUs, supporting distributed training and inference; the software environment can be: model training based on the PyTorchLightning framework, using the Python programming language, and relying on open source libraries such as DeepSpeed ​​and Transformers.

[0177] Three public datasets were used to train and evaluate the item recommendation model. The results are as follows: The first dataset contains 25,411 users, 20,276 items, and 223,263 interaction records, with a data sparsity of 99.96%, and includes multimodal information such as item titles and cover images; the second dataset contains 24,314 users, 18,906 items, and 209,281 interaction records, with a data sparsity of 99.95%; and the third dataset contains 40,358 users, 24,766 items, and 334,238 interaction records, with a data sparsity of 99.97%.

[0178] In some embodiments, the method further includes: preprocessing the text data. For example, the text data (item title) is segmented and stop word removed using the BPE (Byte Pair Encoding) segmentation algorithm, with a vocabulary size of 32,000.

[0179] In some embodiments, the method further includes: adjusting the visual data (object image) to 224×224 pixels, and using the ViT image segmentation strategy to segment the image into 14×14 patches, each patch having a dimension of 768.

[0180] This embodiment integrates the dual tasks of text generation (obtaining target title features) and visual reconstruction (obtaining target image features) into the training of the item recommendation model. Through the dual paradigm of "text data input - text data output" and "visual data input - visual data output", the item recommendation model is guided to deeply mine multimodal information, that is, item recommendation, which solves the core problems of single output modality and insufficient utilization of visual information in related recommendation methods.

[0181] The method in this embodiment constructs a text generation loss L. T With diffusion loss L D The combined training objectives ensure that the model's recommendation performance (such as recommendation accuracy) improves simultaneously in both textual and visual dimensions. Furthermore, the item recommendation model can capture both the textual and visual preferences of sample users, resulting in recommendation results that better match the actual preferences of users.

[0182] By jointly training the dual tasks of text generation and visual reconstruction, the item recommendation model can simultaneously capture users' textual and visual preferences, making fuller use of multimodal information, solving the problem of using only one type of relevant technical information, and making the recommendation results more in line with users' actual needs.

[0183] Figure 6 This is a block diagram illustrating an item recommendation device according to an exemplary embodiment. See also: Figure 6 The recommended device for this item includes: The sequence acquisition module 601 is configured to acquire the historical interaction item sequence of the user to be predicted, which includes multiple items that the user has interacted with in the past. The feature extraction module 602 is configured to process the text and visual data of each item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain multimodal features, which indicate the user's preference for items in terms of text and visual aspects. The title feature acquisition module 603 is configured to process multimodal features through the title prediction unit in the item recommendation model to obtain the target title features; The image feature acquisition module 604 is configured to perform image reconstruction based on the multimodal features of the item through the visual reconstruction unit in the item recommendation model, and to reconstruct the image based on the denoising result to obtain the target image features; The item determination module 605 is configured to obtain joint features based on target title features and target image features, and determine the item to be recommended from the candidate item library based on the joint features.

[0184] In some embodiments, the feature extraction module is configured to: convert the text data and prompt statements of each item in the historical interactive item sequence through the feature extraction unit in the item recommendation model to obtain the text embedding of each item; process the visual data of each item in the historical interactive item sequence through the feature extraction unit in the item recommendation model to obtain the visual features of each item, and project the visual features onto the embedding space of the title prediction unit to obtain the visual embedding of each item; and concatenate the text embedding and the visual embedding of each item to obtain multimodal features.

[0185] In some embodiments, the title feature acquisition module is configured to perform autoregressive prediction on multimodal features through the title prediction unit in the item recommendation model, and obtain each word in the word sequence of the title of the target title feature one by one.

[0186] In some embodiments, the image feature acquisition module includes a first reconstruction unit or a second reconstruction unit; The first reconstruction unit includes a denoising subunit and a reconstruction subunit. The denoising subunit is configured to denoise the sub-image features of the preset image with added random noise by using the visual reconstruction unit in the item recommendation model with the multimodal features of the item as a condition. The reconstruction subunit is configured to reconstruct the target image features based on the denoising result. The second reconstruction unit generates target image features based on multimodal features through the visual reconstruction unit in the item recommendation model.

[0187] In some embodiments, the denoising result includes the denoising result of multiple sub-image features, and the reconstruction sub-unit is configured to stitch together the denoising results of multiple sub-image features to obtain the target image features.

[0188] In some embodiments, the item determination module is configured to determine multiple first items from multiple candidate items in a candidate item library based on joint features, wherein the joint features of each first item and the similarity of the joint features meet the target conditions; and output the first items ranked in the top preset position of similarity as items to be recommended.

[0189] It should be noted that the item recommendation device provided in the above embodiments is only illustrated by the division of the above functional units when recommending items. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the item recommendation device and the item recommendation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0190] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0191] According to another aspect of embodiments of the present disclosure, an item training device is provided, with reference to... Figure 7 The device includes: The sample sequence acquisition module 701 is configured to acquire the historical interaction item sequence of the sample user, which includes multiple sample items that the sample user has interacted with. The sample feature extraction module 702 is configured to process the text and visual data of each sample item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects. The sample title feature acquisition module 703 is configured to process the multimodal features of the sample items through the title prediction unit in the item recommendation model to obtain the predicted title features; The sample image feature acquisition module 704 is configured to perform image reconstruction based on the multimodal features of the item through the visual reconstruction unit in the item recommendation model, and to perform reconstruction based on the denoising result to obtain the reconstructed image features of the item. Training module 705 is configured to train the item recommendation model based on the title features and predicted title features of the next sample item of the sample item, the image features and reconstructed image features of the next sample item of the sample item.

[0192] In some embodiments, the sample feature extraction module is configured to: convert the text data and prompt statements of each sample item in the historical interactive item sequence through the feature extraction unit in the item recommendation model to obtain the text embedding of each sample item; process the visual data of each sample item in the historical interactive item sequence through the feature extraction unit in the item recommendation model to obtain the visual features of each sample item, and project the visual features onto the embedding space of the title prediction unit to obtain the visual embedding of each sample item; and concatenate the text embedding and the visual embedding of each sample item to obtain the multimodal features of each sample item.

[0193] In some embodiments, the sample title feature acquisition module is configured to perform autoregressive prediction on the multimodal features of the sample item and the word sequence of the title of the sample item in the historical interaction item sequence through the title prediction unit in the item recommendation model, and obtain each word in the word sequence of the title of the predicted item one by one.

[0194] In some embodiments, the sample image feature acquisition module includes a third reconstruction unit or a fourth reconstruction unit; The third reconstruction unit includes a sample denoising subunit and a sample reconstruction subunit. The sample denoising subunit is configured to denoise the sub-image features of the preset image with added random noise by using the diffusion module with the multimodal features of the sample item as a condition. The diffusion module is the visual reconstruction unit in the item recommendation model. The sample reconstruction subunit is configured to reconstruct based on the denoising result to obtain the reconstructed image features of the sample item. The fourth reconstruction unit is configured to generate reconstructed image features of sample items based on multimodal features through the VQ-VAE module, which is the visual reconstruction unit in the item recommendation model.

[0195] In some embodiments, the apparatus further includes a feature acquisition module configured to segment a preset image to obtain multiple consecutive sub-images, extract features from each sub-image to obtain multiple sub-image features, and randomly add noise to the multiple sub-image features to obtain sub-image features with added random noise.

[0196] In some embodiments, the denoising result includes denoising results of multiple noisy sub-image features, and the sample reconstruction sub-unit is configured to stitch together the denoising results of multiple noisy sub-image features to obtain the reconstructed image features of the sample item.

[0197] In some embodiments, the training module is configured to obtain a first loss value based on the title features of the next sample item and the predicted title features of the sample item, the first loss value indicating the difference between the title features of the next sample item and the predicted title features of the sample item; Based on the image features of the next sample item and the reconstructed image features of the sample item, a second loss value is obtained, which indicates the difference between the image features of the next sample item and the reconstructed image features of the sample item; Based on the first loss value and the second loss value, a joint loss value is obtained, which indicates the textual and visual differences between the next sample item and the prediction result. The model parameters of the item recommendation model are adjusted based on the joint loss value.

[0198] It should be noted that the item training device provided in the above embodiments is only illustrated by the division of the above functional units when training the item recommendation model. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the item training device and the item training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0199] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0200] When an electronic device is provided as a terminal, Figure 8 This is a block diagram illustrating a terminal 800 according to an exemplary embodiment. The diagram shows a structural block diagram of a terminal 800 provided in an exemplary embodiment of this disclosure. The terminal 800 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 800 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0201] Typically, terminal 800 includes a processor 801 and a memory 802.

[0202] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0203] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one computer program, which is executed by the processor 801 to implement the method provided in the method embodiments of this application.

[0204] In some embodiments, the terminal 800 may also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0205] Peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0206] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0207] Display screen 805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, disposed on the front panel of terminal 800; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 800 or in a folded design; in other embodiments, display screen 805 may be a flexible display screen, disposed on a curved or folded surface of terminal 800. Furthermore, display screen 805 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0208] The camera assembly 806 is used to acquire images or videos. In some embodiments, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0209] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0210] Power supply 808 is used to supply power to the various components in terminal 800. Power supply 808 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 808 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0211] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0212] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, which can be executed by a processor of an electronic device, such as a processor 801 of a terminal 800, to complete the task execution method for the aforementioned network activities. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0213] The aforementioned computer equipment can also be implemented as a server. The structure of a server is described below: Figure 9This is a schematic diagram of a server structure provided in an embodiment of this application. The server 900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 901 and one or more memories 902. The one or more memories 902 store at least one computer program, which is loaded and executed by the one or more processors 901 to implement the methods provided in the various method embodiments described above. Of course, the server 900 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 900 may also include other components for implementing device functions, which will not be elaborated upon here.

[0214] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program that can be executed by a processor to perform the methods in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0215] In an exemplary embodiment, a computer program product or computer program is also provided, which includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the method described above.

[0216] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0217] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for recommending items, characterized in that, The method includes: Obtain the historical interaction item sequence of the user to be predicted, the historical interaction item sequence including multiple items that the user has interacted with in the past; The feature extraction unit in the item recommendation model processes the text and visual data of each item in the historical interaction item sequence to obtain multimodal features, which indicate the user's preference for items in terms of text and visual aspects. The multimodal features are processed by the title prediction unit in the item recommendation model to obtain the target title features; The target image features are obtained by reconstructing the image based on the multimodal features of the item using the visual reconstruction unit in the item recommendation model. Based on the target title features and the target image features, joint features are obtained, and based on the joint features, items to be recommended are determined from the candidate item library.

2. The item recommendation method according to claim 1, characterized in that, The feature extraction unit in the item recommendation model processes the text and visual data of each item in the historical interaction item sequence to obtain multimodal features, including: The feature extraction unit in the item recommendation model is used to transform the text data and prompts of each item in the historical interaction item sequence to obtain the text embedding of each item. The visual data of each item in the historical interactive item sequence is processed by the feature extraction unit in the item recommendation model to obtain the visual features of each item, and the visual features are projected into the embedding space of the title prediction unit to obtain the visual embedding of each item. The text embedding and visual embedding of each item are concatenated, and the concatenation result is encoded to obtain multimodal features.

3. The item recommendation method according to claim 1, characterized in that, The multimodal features are processed by the title prediction unit in the item recommendation model to obtain the target title features, including: The title prediction unit in the item recommendation model performs autoregressive prediction on the multimodal features to obtain each word in the word sequence of the target title features.

4. The item recommendation method according to claim 1, characterized in that, The image reconstruction performed by the visual reconstruction unit in the item recommendation model based on the multimodal features of the item yields the target image features, including: The diffusion module uses the multimodal features of the item as a condition to denoise the sub-image features of the preset image with added random noise, and reconstructs the target image features based on the denoising results. The diffusion module is the visual reconstruction unit in the item recommendation model. Alternatively, the target image features can be generated based on the multimodal features using the VQ-VAE module, where the VQ-VAE module is a visual reconstruction unit in the item recommendation model.

5. The item recommendation method according to claim 4, characterized in that, The reconstruction based on the denoising results to obtain the target image features includes: The denoising result includes the denoising results of the multiple sub-image features. The denoising results of the multiple sub-image features are concatenated to obtain the target image features.

6. The item recommendation method according to claim 1, characterized in that, The process of determining the items to be recommended from the candidate item library based on the joint features includes: Based on the joint features, multiple first items are determined from multiple candidate items in the candidate item library, and the similarity between the joint features of each first item and the joint features meets the target condition. The first item ranked first by similarity is output as the item to be recommended.

7. A method for training objects, characterized in that, The method includes: Obtain the historical interaction item sequence of the sample user, the historical interaction item sequence including multiple sample items that the sample user has interacted with; The feature extraction unit in the item recommendation model processes the text and visual data of each sample item in the historical interaction item sequence to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects. The title prediction unit in the item recommendation model processes the multimodal features of the sample items to obtain the predicted title features. The visual reconstruction unit in the item recommendation model performs image reconstruction based on the multimodal features of the sample items to obtain the reconstructed image features of the sample items. The item recommendation model is trained based on the title features of the next sample item of the sample item, the predicted title features, the image features of the next sample item of the sample item, and the reconstructed image features.

8. The method for training items according to claim 7, characterized in that, The feature extraction unit in the item recommendation model processes the text and visual data of each sample item in the historical interaction item sequence to obtain multimodal features for each sample item. These multimodal features indicate the sample user's preferences for the sample item in terms of both text and visual representation, including: The feature extraction unit in the item recommendation model is used to convert the text data and prompts of each sample item in the historical interaction item sequence to obtain the text embedding of each sample item. The visual data of each sample item in the historical interactive item sequence is processed by the feature extraction unit in the item recommendation model to obtain the visual features of each sample item, and the visual features are projected into the embedding space of the title prediction unit to obtain the visual embedding of each sample item. The text embedding and visual embedding of each sample item are concatenated to obtain the multimodal features of each sample item.

9. The method for training items according to claim 7, characterized in that, The process of processing the multimodal features of the sample items through the title prediction unit in the item recommendation model to obtain the predicted title features includes: The title prediction unit in the item recommendation model performs autoregressive prediction on the multimodal features of the sample items, and obtains each word in the word sequence of the predicted title features one by one.

10. The method for training items according to claim 7, characterized in that, The step of reconstructing the image based on the multimodal features of the sample items using the visual reconstruction unit in the item recommendation model to obtain the reconstructed image features of the sample items includes: The diffusion module uses the multimodal features of the sample item as a condition to denoise the sub-image features of the preset image with added random noise, and reconstructs the image features of the sample item based on the denoising results. The diffusion module is the visual reconstruction unit in the item recommendation model. Alternatively, the VQ-VAE module can be used to generate reconstructed image features of the sample items based on the multimodal features. The VQ-VAE module is a visual reconstruction unit in the item recommendation model.

11. The method for training items according to claim 10, characterized in that, The method further includes: The preset image is segmented to obtain multiple consecutive sub-images, and features are extracted from each sub-image to obtain multiple sub-image features; Random noise is added to the multiple sub-image features to obtain multiple sub-image features with added random noise.

12. The method for training items according to claim 10, characterized in that, The reconstructed image features of the sample items obtained by reconstructing based on the denoising results include: The denoising result includes the denoising results of the multiple sub-image features. The denoising results of the multiple sub-image features are stitched together to obtain the reconstructed image features of the sample item.

13. The method for training items according to claim 7, characterized in that, The step of training the item recommendation model based on the title features of the next sample item of the sample item, the predicted title features, the image features of the next sample item of the sample item, and the reconstructed image features includes: Based on the title features of the next sample item of the sample item and the predicted title features, a first loss value is obtained, wherein the first loss value indicates the difference between the title features of the next sample item of the sample item and the predicted title features; Based on the image features of the next sample item of the sample item and the reconstructed image features, a second loss value is obtained, which indicates the difference between the image features of the next sample item of the sample item and the reconstructed image features; Based on the first loss value and the second loss value, a joint loss value is obtained, which indicates the textual and visual differences between the next sample item and the prediction result of the sample item. Based on the joint loss value, the model parameters of the item recommendation model are adjusted.

14. An item recommendation device, characterized in that, include: The sequence acquisition module is configured to acquire the historical interaction item sequence of the user to be predicted, the historical interaction item sequence including multiple items that the user has previously interacted with; The feature extraction module is configured to process the text and visual data of each item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain multimodal features, which indicate the user's preference for items in terms of text and visual aspects. The title feature acquisition module is configured to process the multimodal features through the title prediction unit in the item recommendation model to obtain the target title features; The image feature acquisition module is configured to take the multimodal features of the item and multiple sub-image features of the image of the item as input, and perform image reconstruction based on the multimodal features of the item through the visual reconstruction unit in the item recommendation model to obtain the target image features; The item determination module is configured to obtain joint features based on the target title features and the target image features, and determine the item to be recommended from the candidate item library based on the joint features.

15. A training device for an item, characterized in that, include: The sample sequence acquisition module is configured to acquire the historical interaction item sequence of a sample user, the historical interaction item sequence including multiple sample items that the sample user has interacted with; The sample feature extraction module is configured to process the text and visual data of each sample item in the historical interaction item sequence through the feature extraction unit in the item recommendation model to obtain the multimodal features of each sample item. The multimodal features indicate the sample user's preference for the sample item in terms of text and visual aspects. The sample title feature acquisition module is configured to process the multimodal features of the sample items through the title prediction unit in the item recommendation model to obtain predicted title features; The sample image feature acquisition module is configured to take the multimodal features of the sample item and multiple sub-image features of the image of the sample item as input, and perform image reconstruction based on the multimodal features of the sample item through the visual reconstruction unit in the item recommendation model to obtain the reconstructed image features of the sample item. The training module is configured to train the item recommendation model based on the title features of the next sample item of the sample item and the predicted title features, the image features of the next sample item of the sample item and the reconstructed image features.

16. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the item recommendation method as described in any one of claims 1 to 6 or the item training method as described in any one of claims 7 to 13.

17. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the item recommendation method as described in any one of claims 1 to 6 or the item training method as described in any one of claims 7 to 13.