A collaborative generative recommendation method based on multimodal semantically enhanced tags
By constructing multimodal semantic enhancement identifiers and adopting cross-modal semantic alignment and collaborative semantic alignment mechanisms, the problem of insufficient multimodal information fusion in generative recommendation models is solved, and more accurate and personalized recommendation content generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- COMMUNICATION UNIVERSITY OF CHINA
- Filing Date
- 2025-11-19
- Publication Date
- 2026-05-26
AI Technical Summary
Existing generative recommendation models lack deep cross-modal semantic fusion and alignment when processing multimodal information, and fail to effectively utilize collaborative signals, resulting in a lack of semantic understanding and interactive modes of collaborative behavior in the generation process, and thus failing to achieve deep joint expression of multimodal information.
By constructing multimodal semantically enhanced identifiers, adopting cross-modal semantic alignment, interaction-aware modeling and joint coding optimization, integrating multimodal content such as text descriptions and product images, and embedding user-item interaction graphs and context co-occurrence patterns, and using a collaborative semantic alignment mechanism for explicit supervision, we can ensure that the distribution of semantic quantization representation and collaborative embedding is aligned.
It achieves synergistic enhancement of semantic richness and structural sensitivity in a unified semantic space, improving the accuracy and personalization of generative recommendations, breaking through the technical bottlenecks of modality fragmentation and signal isolation, and providing more accurate recommended content generation.
Smart Images

Figure CN121502084B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of generative recommendation systems, and more particularly to a collaborative generative recommendation method based on multimodal semantic enhancement tagging. Background Technology
[0002] Recommender systems play a crucial role in exploring personalized content across various scenarios, such as video platforms, e-commerce shopping, and movie recommendations. With the rapid development of large language models and recommender system technologies, generative recommendation, as a new generation of recommendation paradigm, is gradually evolving from the traditional "retrieval-ranking" model to an "understanding-generation" model. Compared to traditional methods that rely solely on user-item interaction history for matching, generative recommendation, through a constrained generation mechanism, can dynamically generate personalized recommendation results and improve content distribution quality in cold-start scenarios, significantly enhancing user experience and system flexibility. Encoding the multimodal information of items (including goods, video entities, objects, etc.) is the core of the "understanding-generation" model, effectively addressing the representational limitations of discriminative methods in the traditional "retrieval-ranking" model.
[0003] Currently, there are still some problems to be solved in the research on semantic understanding depth, multi-source information fusion capability, and utilization of collaborative signals. First, in real-world scenarios, multimodal information (such as text data, image data, user preferences, and item features) is often presented in multimodal forms such as text and images. However, the semantic identifier construction of current generative recommendation models relies on unimodal information. Some studies only process multimodal information at the stage of simple splicing or independent encoding, lacking deep fusion and alignment of cross-modal semantics. The isolation of different modalities hinders the model from learning a unified semantic representation that utilizes complementary information. Second, existing research only focuses on the semantic attributes of unimodality, ignoring collaborative signals derived from user interaction history, resulting in the inability to retain the interaction patterns that exist simultaneously between semantic understanding and collaborative behavior. Existing methods usually model collaborative signals and semantic content separately, failing to achieve their collaborative construction at the semantic enhancement level, resulting in a lack of deep joint guidance of multimodal semantics in the generation process. Against the backdrop of the explosive growth of multimodal content and the vigorous development of generative artificial intelligence, there is an urgent need for a semantically enhanced collaborative alignment fusion expression method that can deeply integrate multimodal semantics and collaborative signals to construct a new type of multimodal generative expression, thereby breaking through the technical bottlenecks of modal fragmentation and signal isolation.
[0004] Based on the aforementioned technical challenges and application requirements, this invention proposes a generative recommendation (Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation, MusicRec) method and model. To better integrate multimodal item semantics, it first extracts shared quantization codes across different modalities, and then learns specific quantization codes for each modality based on shared representation features, enabling the generated identifiers to capture richer cross-modal semantic representations. To effectively utilize collaborative information, a collaborative semantic alignment mechanism is introduced during the identifier construction learning process. This method ensures that the semantic quantization representation and collaborative embedding maintain distributional alignment, thereby maintaining the proximity between items with similar interaction patterns in the quantization space while effectively driving the generative recommendation model to achieve more accurate and semantically consistent recommendation content generation. Summary of the Invention
[0005] The purpose of this invention is to provide a collaborative generative recommendation method based on multimodal semantic enhancement identifiers. By cross-modal semantic alignment, interaction-aware modeling, and joint encoding optimization, a multimodal semantic enhancement identifier explicitly supervised by collaborative signals is constructed. As the core input representation of the generative recommendation model, the identifier not only integrates deep semantic information from multimodal content such as text descriptions and product images, but also embeds high-order collaborative signals such as user-item interaction graphs and context co-occurrence patterns, thereby achieving collaborative enhancement of semantic richness and structural sensitivity in a unified semantic space.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] A collaborative generative recommendation method based on multimodal semantically enhanced identifiers, the method comprising:
[0008] S1. Construct an entity database, where entities contain both text and image data; extract text semantic embeddings from each entity in the entity database. Image semantic embedding and collaborative embedding Collaborative embedding This includes associated information and / or text-image collaborative attributes, where associated information is the information associated with the text and / or image;
[0009] S2. Construct a semantically enhanced collaborative alignment and fusion representation method model. This model utilizes entity-related textual semantic embeddings. With image semantic embedding Semantic enhancement, co-alignment, and fusion processing are performed to obtain quantized text embeddings. and quantized image embedding And the corresponding data is stored in the entity database;
[0010] S3. The user inputs an entity of interest. Text semantic embedding and image semantic embedding are extracted for the entity of interest and input into the semantic enhancement collaborative alignment fusion representation model to obtain the quantized text embedding, quantized image embedding and collaborative embedding representation of the entity of interest. Taking the quantized text embedding, quantized image embedding and / or collaborative embedding representation of the entity of interest as the search task target, the entity database is searched for entities or / and combinations of entities associated with the entity of interest to generate recommendation results.
[0011] To better implement the present invention, in method S1, if the entity database is an e-commerce goods and services database, then the entities in the entity database are e-commerce goods or services, and the generated recommendation result is the associated recommended e-commerce goods or services; the entities in the entity database are associated with and stored as sub-entities with the same semantics or similarity greater than a threshold, and all sub-entities with the same semantics or similarity greater than a threshold are stored in the entity database with the same entity and the corresponding sub-entities are associated with and stored in the entity attributes.
[0012] Preferably, in method S3, the user inputs multimodal data of interest, which includes text data and / or image data, and the core entities in the multimodal data of interest are extracted by large language model recognition.
[0013] Preferably, the semantic enhancement collaborative alignment fusion expression method of the present invention has the following model processing method:
[0014] S11. Embedding text semantics With image semantic embedding Compressed into latent semantic embeddings of the same dimension and Embedding latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. ;
[0015] S12. The last shared layer input to the modality-specific quantizer is processed by modal shared residual quantization for modal separation, and then... Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences ;
[0016] S13, Modal sharing of code sequences With text-specific code sequences Jointly constructing quantized text embeddings ; Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings ;
[0017] S14. Semantic Enhancement Collaborative Alignment Fusion Representation Method Model Output Quantized Text Embedding of Entities and quantized image embedding .
[0018] Preferably, in method S14, the quantized text embedding is extracted. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers; image embeddings are extracted. Shared codewords and image-specific codewords constitute an image lexical sequence and serve as an image identifier; the cooperative embedding representation It also records quantified text embeddings With quantized image embedding Collaboratively perceive alignment information.
[0019] Preferably, in method S11, the latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. ; Each level in the layer shared codebook Each is equipped with a shared codebook , , The codeword with sequence number n is represented by N, where N is the size of the codebook; the residual quantization expression for modal shared residual quantization is as follows:
[0020] ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals For the selected codeword; modal shared code sequence .
[0021] Preferably, in method S12, a modality-specific quantizer is used to perform modality-shared quantization on the residual embedding of the last shared layer. Modality separation is performed to extract text semantics for text modality quantization. and image semantics for image modality quantization ; on text semantics pass The layered text's specific codebook is quantized into a sequence of codewords. Layer of a specific codebook The corresponding text modal codebook is obtained. Text-specific code sequences ; Image semantics pass The layer image is quantized into a codeword sequence based on a specific codebook. Layer image specific codebook level The corresponding image modality codebook is obtained. Image-specific code sequence .
[0022] Preferably, in quantized text embedding Reconstructing the original input features and encoding them yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The following loss function is constructed based on the residual quantization processing of methods S11 and S12:
[0023] ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, This represents a coefficient that balances the strength of code embedding and encoder optimization.
[0024] Preferably, a total loss function for quantization embedding is constructed, and the expression for the total loss function for quantization embedding is as follows:
[0025] ,in This represents the total loss of quantization embedding. This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.
[0026] Preferably, the total loss function expression of the semantic enhancement collaborative alignment fusion representation method model is as follows:
[0027] , For the collaborative alignment loss of collaborative embedding representation, To preserve the loss for the cooperative relationships in the cooperative embedding representation, , They represent control , Hyperparameters of capability strength.
[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0029] (1) This invention constructs a multimodal semantic enhancement identifier explicitly supervised by collaborative signals through cross-modal semantic alignment, interaction-aware modeling and joint coding optimization. The identifier serves as the core input representation of the generative recommendation model. It not only integrates deep semantic information from multimodal content such as text descriptions and product images, but also embeds high-order collaborative signals such as user-item interaction graphs and context co-occurrence patterns. Thus, it achieves synergistic enhancement of semantic richness and structural sensitivity in a unified semantic space. This invention breaks through the technical bottlenecks of modal fragmentation and signal isolation, and achieves synergistic enhancement of semantic understanding and structural awareness. Thus, it provides a unified representation foundation for generative recommendation that combines semantic richness, structural awareness and generative guidance capabilities.
[0030] (2) The application of this invention in multiple application scenarios demonstrates its effective utilization of multimodal semantic information and collaborative-semantic alignment mechanisms in the identifier generation process. The experimental results also highlight the complementarity and irreplaceability of multimodal semantics and collaborative signals in generative recommendation systems, emphasizing the importance of their deep integration. This invention provides key technical support for building next-generation semantically driven, structure-aware, and generatively controllable intelligent recommendation systems, and has broad application prospects and significant engineering promotion value in multimodal intensive recommendation scenarios such as e-commerce, short videos, and content platforms.
[0031] (3) This invention can better integrate multimodal item semantics, extract shared quantization codes between different modalities, and learn specific quantization codes for each modality based on shared representation features, enabling the generated identifiers to capture richer cross-modal semantic representations. In order to effectively utilize collaborative information, a collaborative semantic alignment mechanism is introduced in the identifier construction learning process. This invention ensures that the semantic quantization representation and collaborative embedding maintain distributional alignment, thereby maintaining the proximity between items with similar interaction patterns in the quantization space while effectively driving the generative recommendation model to achieve more accurate and semantically consistent recommendation content generation.
[0032] (4) This invention addresses the problems of insufficient semantic understanding, independent modal features, and inadequate utilization of collaborative signals in existing generative recommendation systems. It aligns and fuses multimodal semantic information from items with collaborative signals such as user-item interactions and contextual co-occurrence, jointly generating a unified identifier with rich semantics and structural awareness. This improves the performance of multimodal generative recommendation systems in terms of content relevance, personalization, and semantic consistency. This invention effectively enhances the ability of recommendation systems to fuse complex heterogeneous information, promotes the evolution of recommendation systems from a discriminative to a generative paradigm, and provides key technical support for intelligent content distribution, personalized services, and human-computer interaction. Attached Figure Description
[0033] Figure 1 This is a flowchart of the collaborative generative recommendation method of the present invention;
[0034] Figure 2 This is a schematic diagram illustrating the principle of the collaborative generative recommendation method in the embodiment.
[0035] Figure 3 This is a graph showing the hyperparameter variation results of Example 3 on the Instruments dataset;
[0036] Figure 4 This is a graph showing the hyperparameter variation results of Example 3 on the Arts dataset;
[0037] Figure 5 This is a graph showing the hyperparameter variation results of Example 3 on the Games dataset;
[0038] Figure 6 A comparison chart showing the impact of codebook size on model performance on the Arts dataset;
[0039] Figure 7 A comparison chart showing the impact of codebook contribution rate on model performance on the Arts dataset. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to embodiments:
[0041] Example
[0042] like Figure 1 As shown, a collaborative generative recommendation method based on multimodal semantic enhancement identifiers is proposed, the method comprising:
[0043] S1. Construct an entity database, where entities contain both text and image data. Extract text semantic embeddings from each entity in the entity database. Image semantic embedding and collaborative embedding The entity database can be an e-commerce product database, a movie / video database, etc. If the entity database is an e-commerce item and service database, then the entities in the entity database are e-commerce items or services, and the generated recommendation results are related recommended e-commerce items or services. Entities in the entity database are associated with and stored as sub-entities that are semantically identical or have a similarity greater than a threshold. All sub-entities with semantically identical or similar identities are stored in the entity database using the same entity name and their corresponding sub-entities are associated with the entity attributes. Collaborative embedding. This includes associated information and / or text-image collaborative attributes. Associated information refers to the information associated with the text and / or image (associated information includes user-item interaction information, such as user intent, comments, etc.; if used in a recommendation system, it includes a list of recommended items, user choices and comments for different tags, etc.; associated information contains a large amount of rich association information between users and items). The multimodal data of this invention's items consists of several multimodal data entries for the same item (for example, in an e-commerce marketplace, if there are multiple entries for the same item from different merchants, they are recorded as the same item; during recommendation, the item is recommended first, and then the various suppliers providing the item are recommended accordingly). Text-image collaborative attributes are the association signals or features between the extracted text semantics and image semantics of the item (capable of identifying and associating the association information between text and image, facilitating auxiliary prompts during item recommendation).
[0044] S2. Construct a semantically enhanced collaborative alignment and fusion representation method model. This model utilizes entity-related textual semantic embeddings. With image semantic embedding Semantic enhancement, co-alignment, and fusion processing are performed to obtain quantized text embeddings. and quantized image embedding And the corresponding data is stored in the entity database.
[0045] The preferred semantic enhancement collaborative alignment fusion representation method of this invention has the following model processing method:
[0046] S11. Embedding text semantics With image semantic embedding Compressed into latent semantic embeddings of the same dimension and In some embodiments, pre-trained models (such as LLaMA and ViT) are used to extract textual semantic embeddings from the project's multimodal data. With image semantic embedding Text semantics are embedded through a text encoder (a text encoder implemented using a multilayer perceptron). Compression into latent semantic embeddings Image semantics are embedded through an image encoder (an image encoder implemented using a multilayer perceptron). Compression into latent semantic embeddings The expression is as follows:
[0047] ,
[0048] Text encoder Image encoder These are implemented using multilayer perceptrons.
[0049] Embed latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. .
[0050] In some embodiments, latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. . Each level in the layer shared codebook Each is equipped with a shared codebook , , Let N represent the codeword with sequence number n, and N be the size of the codebook. The residual quantization expression for modal shared residual quantization is as follows:
[0051] ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals The selected codewords are used. Modal shared code sequences are obtained through modal shared residual quantization. .
[0052] S12. The last shared layer is processed by modal shared residual quantization (preferably, the residual embedding of the shared layer is selected). The input residual is embedded into a mode-specific quantizer for mode separation, and then processed separately. Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences .
[0053] In some embodiments, a modality-specific quantizer is used to embed the residuals of the last shared layer in modality-shared quantization. Modality separation is performed to extract text semantics for text modality quantization. and image semantics for image modality quantization Regarding text semantics pass The layered text's specific codebook is quantized into a sequence of codewords. Layer of a specific codebook The corresponding text modal codebook is obtained. The residual quantization expression for modality-specific residual quantization is as follows:
[0054] ,in Indicates the selection of the first code from a specific codebook of text. Level code words, It is the first Level of text semantic residuals The selected codewords are used to obtain the text-specific code sequence through modality-specific residual quantization. .
[0055] Image semantics pass The layer image-specific codebook is quantized into a codeword sequence, and the residual quantization processing expression for mode-specific residual quantization is as follows:
[0056] ,in Indicates the selection of the first image from a specific codebook. Level code words, It is the first Level-1 image semantic residuals The selected codewords are used to obtain the image-specific code sequence through modality-specific residual quantization. .
[0057] The residual quantization processing of methods S11 and S12 of this invention constructs the following loss function:
[0058] ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, A coefficient representing the balance between code embedding and encoder optimization intensity (set to 0.25 in this example).
[0059] S13, Modal sharing of code sequences With text-specific code sequences Jointly constructing quantized text embeddings Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings .
[0060] In some embodiments, during modality-specific residual quantization (including two process stages: obtaining text and image-specific code sequences), the present invention quantizes text embeddings... Reconstructing the original input features and encoding them yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The aggregated quantization embedding for each mode is reconstructed using a decoder based on a variational autoencoder. The reconstruction decoding process is expressed as follows: , , , These are the decoders.
[0061] S14. Semantic Enhancement Collaborative Alignment Fusion Representation Method Model Output Quantized Text Embedding of Entities and quantized image embedding Preferably, extract the quantized text embedding. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers. Extract quantized image embeddings. Shared codewords and image-specific codewords constitute an image lexical sequence and serve as the image identifier. Cooperative embedding representation. It also records quantified text embeddings With quantized image embedding Collaborative sensing alignment information (cooperative sensing and deep alignment are performed in methods S11 to S14 to obtain semantically enhanced collaborative information, collaborative embedding representation) It records the collaborative information of the pre-trained model and the collaborative information of the final text and image embedding representations, possessing rich and reliable comprehensive collaborative information. (For example...) Figure 2 As shown, to facilitate better structuring of text and image identifiers and to enable ordered machine recognition and judgment, text identifiers (i.e., text-specific modalities) use a set of lowercase letters {a, b, ...} to represent levels, while image identifiers (image-specific modalities) use a set of uppercase letters {A, B, ...} to represent levels. Then, complete semantically enhanced identifiers are constructed sequentially as new lexical sequences. For example, the text identifier for a project can be labeled as...<s1_1><s2_2><a_3><b_4> The corresponding image identifier can be marked as<S1_1><S2_2><A_5><B_6> .
[0062] S3. The user inputs entities of interest. Preferably, the user inputs multimodal data of interest, which includes text data and / or image data. The core entities in the multimodal data of interest are identified and extracted using a large language model. Text semantic embeddings and image semantic embeddings are extracted for each entity of interest and input into a semantically enhanced collaborative alignment fusion representation model to obtain quantized text embeddings, quantized image embeddings, and collaborative embedding representations of the entities of interest. Using the quantized text embeddings, quantized image embeddings, and / or collaborative embedding representations of the entities of interest as the search task objective, entities and / or combinations of entities associated with the entities of interest are searched from the entity database to generate recommendation results.
[0063] In some embodiments, the semantic loss of the semantic enhancement collaborative alignment fusion representation method model of the present invention is... The expression is as follows:
[0064] ,in Indicates semantic loss, This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.
[0065] This invention constructs a collaborative embedding representation Collaborative sensing alignment loss and collaborative relationship preservation loss are used to constrain the loss (improving the accuracy of the entire model processing stage) to bridge the gap between semantic features and collaborative signals. A pre-trained sequence recommendation model, SASRec, is used to obtain collaborative filtering embeddings of items. To better utilize collaborative signals to align semantic information, the InfoNCE loss function is employed to quantize the item embeddings. (Includes quantized text embedding) and quantized image embedding The collaborative filtering embedding is closer to the item while keeping it away from the embeddings of other irrelevant items in the training batch, resulting in a collaboratively perceptual alignment loss. The expression is as follows:
[0066] ,in Represents the cosine similarity function. This indicates temperature hyperparameters. Indicates the training batch size. This represents the quantized embedding representation of training batch m. This represents the collaborative filtering embedding of training batch m. Let represent the quantized embedding representation of training batch j. To preserve the cooperative relationship structure encoded in the cooperative embedding space, a cooperative relationship preservation loss is introduced to enforce structural consistency between the quantized embedding space and the cooperative embedding space; specifically, given a training batch of an item... Calculate the pairwise similarity matrix in two spaces, and calculate the cooperative relationship preservation loss. The expression is as follows:
[0067] ,in This represents the cosine similarity of the embeddings quantized in training batches m and j. Let represent the cosine similarity of the system embeddings in training batches m and j. Preferably, to maintain the stability of the semantic identifier training process, optimization is performed by jointly considering semantic loss, collaborative awareness alignment loss, and collaborative relation preservation loss. This invention constructs a collaborative embedding representation. semantic loss Collaborative perception alignment loss Losses in maintaining collaborative relationships The total loss is expressed as follows: ,in , Representing the control collaborative sensing alignment loss respectively Maintaining the relationship and loss Hyperparameters of capability strength.
[0068] Example 2
[0069] This embodiment uses all the techniques from Embodiment 1 to perform the recommendation generation task, such as... Figure 2 As shown, this invention designs a series of structured recommendation tasks in the execution of recommendation generation tasks, aiming to fully mine and utilize the multimodal semantic information and collaborative signals contained in the constructed semantically enhanced identifiers, thereby improving the generative recommendation capability. The recommendation generation tasks can introduce explicit cross-modal fusion and alignment mechanisms to achieve deep semantic fusion and generation guidance between modalities. In the execution of recommendation generation tasks, this invention can, based on the user's sequential behavior patterns in a sequential interaction environment, predict the user's next possible interaction item, and further subdivide it into the following two types of subtasks according to the composition of the input modality:
[0070] (1) Single-mode generation
[0071] In this task, the input is user item data in a certain modality (such as plain text description or pure image features) to generate single-modal or multimodal recommendations.
[0072] Project generation (text modal) example:
[0073]
[0074] Project generation (image modality) example:
[0075]
[0076] (2) Multimodal generation
[0077] To more realistically simulate the complexity of cross-modal interactions between users in real-world scenarios, this task enables the model to explicitly model the semantic complementarity and behavioral synergy between different modalities during the generation process, thereby improving the multimodal consistency of the generated results.
[0078] Project generation (text modal) example:
[0079]
[0080] Project generation (image modality) example:
[0081]
[0082] While large language models can implicitly learn cross-modal associations through context during sequence generation, this alignment is often loose and unstable, especially in scenarios with missing modalities or noise interference, which can easily lead to semantic drift. To address this, this invention introduces an explicit cross-modal alignment objective to force the model to establish strong semantic consistency constraints in the identifier space, ensuring high alignment of representations of the same item across different modalities. This alignment objective comprises two symmetrical subtasks.
[0083] (1) Text to Image Alignment: Given a text modal token representation of an item, the model needs to reconstruct or generate its corresponding image modal token representation.
[0084] (2) Image to text alignment: Given an image modal token representation of an item, the model needs to generate its corresponding text description token sequence.
[0085] Text-to-image alignment example:
[0086]
[0087] Image-to-text alignment example:
[0088]
[0089] This task not only enhances the model's understanding of cross-modal semantic equivalence, but also significantly improves the robustness of identifiers in scenarios with missing or partially observed modalities. Even when only a single modality is observed, the model can still infer the complete item semantics through the alignment structure in the identifier space, thereby supporting high-quality generative recommendations.
[0090] After obtaining the semantic identifiers for each modality of each item, the sequence recommendation task is modeled as a sequence-to-sequence generation problem. Specifically, each user's interaction history is uniformly transformed into an input sequence composed of item semantic identifiers, while the recommendation target is formalized as generating the complete semantic identifier of the next target item. Let the user's historical interaction sequence be... Each identifier All are made by A sequence of discrete semantic tags, generated by the aforementioned multimodal semantic and cooperative signal fusion mechanism, contains rich cross-modal semantic and structured behavioral information. This invention employs a T5 model based on an encoder-decoder architecture as the generation backbone. This model takes the input sequence... Encodes a context-aware latent state representation and progressively generates a semantic identifier for the next target item through an autoregressive decoder. , recorded as The entire process follows the standard sequence-to-sequence paradigm, but its input and output are both structured sequences of semantic identifiers, rather than raw text or item IDs, thus achieving full utilization of multimodal semantics and cooperative signals.
[0091] The training objective of the model is to maximize the performance of a given historical sequence. Correctly generate target identifier under the condition The probability. Optimization is achieved by minimizing the standard cross-entropy loss function; for each training sample, the model parameters... The optimization objective is defined as follows: , express The A token.
[0092] During the reasoning process, generative recommendation models are based on users' historical interaction sequences. Autoregressive generation of semantic identifiers for the next project To improve the accuracy and diversity of the generated results, each identifier code... A beam search strategy is used for selection to balance generation quality and search efficiency during decoding. However, since the semantic identifier space is discrete and structured, direct generation may correspond to a "fictitious" identifier that does not exist in the item library. To ensure the generation results... To consistently map to real-world items, a trie-based constraint decoding strategy is introduced: during decoding, each step only allows the selection of valid tokens that can form prefixes of existing item identifiers, thus forcing the model to generate within the valid item space. This strategy not only ensures the feasibility of the recommendation results but also significantly reduces invalid or semantically drifting generated samples.
[0093] To further improve recommendation quality and fully utilize the complementarity of multimodal information, a cross-modal re-ranking strategy is adopted. During the inference phase, textual and visual modal generation tasks are performed separately to obtain two preliminary recommendation lists. and Each list contains a relevance score. The final score for each item is calculated as follows: for items that appear in both lists, the score is calculated as follows: ,in and The extra 1 point is awarded to emphasize multimodal consistency, indicating that the item is highly relevant to the user's history at both the textual and visual semantic levels; for items that exist in only one list, their original scores are retained. or This re-ranking strategy effectively integrates the ability of different modalities to characterize user preferences: items whose preferences are significant across multiple modalities are given higher confidence; while items that stand out only in a single modality are still given a recommendation opportunity. This mechanism not only improves the accuracy and robustness of recommendations but also enhances the system's adaptability to heterogeneous user behavior patterns, thereby enabling more accurate and personalized generative recommendations in diverse real-world application scenarios.
[0094] like Figure 2 As shown, the generation process of multimodal semantically enhanced identifiers in this invention integrates multimodal item semantics through shared modality segmentation and specific modality segmentation, and employs a collaborative semantic alignment mechanism to preserve collaborative relationships, including collaborative perceptual alignment and a collaborative relationship preservation loss function component. In the recommendation result generation stage, this process includes three main generative sequence recommendation training tasks, with the encoder-decoder model autoregressively generating recommendation results. These are three dedicated tasks: next item generation (unimodal), next item generation (multimodal), and multimodal item alignment, to effectively utilize the generated identifiers during recommendation training. The synergistic optimization of these three tasks enables the model to fully utilize the semantic and structural information encoded in the identifiers during training, thereby achieving high-precision and high-consistency generative recommendations in the inference stage. Overall, this framework, through a two-stage paradigm of semantically enhanced identifier construction and multi-task generative learning, achieves an organic unity of multimodal content understanding, collaborative signal utilization, and sequence generation capabilities.
[0095] Example 3
[0096] This embodiment employs all the techniques from Embodiment 1 to perform the recommendation generation task. It systematically evaluates the data on three publicly available real-world benchmark datasets, all derived from the Amazon Product Review Dataset. This dataset is widely used in recommender system research, containing detailed metadata about products on the Amazon platform (such as titles, descriptions, categories, images, etc.) and large-scale user behavior records (including ratings, reviews, and interaction timestamps), covering multiple product category domains and providing a rich and representative experimental foundation for multimodal generative recommender tasks. This experiment selects three product category subsets with significant multimodal characteristics: "Instruments," "Arts," and "Games," for sequence recommendation tasks. To ensure data quality and the reliability of model training, a 5-core filtering strategy is adopted, retaining only users who have interacted with at least 5 different products, and products that have been interacted with by at least 5 different users. This preprocessing method is consistent with mainstream research in the recommender system field, effectively eliminating noise from sparse interactions and improving the fairness and stability of the evaluation. Subsequently, the interaction records of each user were strictly sorted by timestamp to construct their behavior sequence, thus realistically reflecting the dynamic evolution of user interests. To strike a balance between computational efficiency and contextual modeling capabilities, the maximum length of all user sequences was standardized to 20 items. Table 4 shows the statistical information for each dataset, where "average length" refers to the average number of items in the filtered user interaction sequence, reflecting the typical browsing depth of users within that category. For multimodal representation, each product item incorporates two types of heterogeneous features. Textual features are extracted from the product title and description text, encoded into fixed-dimensional semantic vectors using a pre-trained language model. Visual features are extracted from the product illustration, obtaining deep image embeddings through a pre-trained visual model.
[0097] Table 4 Statistical analysis of dataset composition
[0098]
[0099] To delve into the impact of key hyperparameters in the model on the utilization of multimodal semantics and collaborative signal alignment mechanisms, this section systematically evaluates their effect on generative recommendation accuracy from the perspective of hyperparameter sensitivity analysis. Experiments were conducted while keeping other components and training settings constant, adjusting two core hyperparameters and the collaborative sensing alignment weights respectively. Maintaining strength of collaborative relationships We observed the changing trends of model performance on three benchmark datasets, using click-through rate and normalized depreciation cumulative gain as the main evaluation metrics.
[0100] First, for the collaborative sensing alignment mechanism, the hyperparameters are... The value range is set to {1e-1, 1e-2, 2e-2, 1e-3}. For example... Figure 3 – Figure 5 As shown, the experimental results indicate that when When the coefficient of performance is 1e-2, the model achieves optimal performance on all three datasets. At this point, the supervisory role of the cooperative signal in multimodal semantic alignment reaches its optimal balance, effectively guiding semantic identifiers to learn structured behavioral patterns without suppressing the semantic expressive power of the modality itself due to excessive constraints. When the value is increased to 1e-1, the dominance of the cooperative signal becomes too strong, causing the model to overfit the interaction structure and ignore fine-grained multimodal semantic differences; while when When the value is too small, the alignment constraint is too weak to effectively establish the association between semantics and collaboration, thus weakening the discriminative ability of the identifier. This phenomenon verifies the sensitivity of the collaboration-aware alignment mechanism to α and highlights the crucial role of setting this hyperparameter appropriately in improving model performance.
[0101] Secondly, to evaluate the mechanism for maintaining collaborative relationships, hyperparameters will be used. The value range of is set to {1e-2, 1e-3, 1e-4, 1e-5}. Experimental results show that the model performance also exhibits a significant non-monotonic trend. When When the value is 1e-4, the evaluation metrics on all three datasets reach their peak. This result indicates that moderate relation-preserving regularization helps stabilize the embedding space of identifiers and enhances their ability to perceive higher-order cooperative relationships. However, if... If the size is too large, it will introduce an excessively strong smoothing effect, causing the identifiers of different users or items to become homogenized, reducing the ability to express personalization; conversely, if... If the value is too small, it will not be able to effectively constrain the identifier learning process and will be difficult to retain the topological information in the original interaction graph.
[0102] In summary, hyperparameters and Semantic-cooperative alignment strength and structural consistency preservation strength are controlled separately, each with a clearly defined optimal value range. Experiments not only verify the effectiveness of the proposed alignment and preservation mechanisms but also provide clear parameter tuning guidance for practical deployments, allowing for fine-tuning to maximize performance. This analysis further corroborates the necessity of deep fusion of multimodal semantics and cooperative signals and their decisive impact on generation quality.
[0103] This embodiment explores the mechanism by which the quantization codebook size affects the identifier representation ability of generative recommendation models by comparing the impact of different codebook sizes on model performance during shared-specific residual quantization. While keeping other hyperparameters and training configurations constant, this experiment systematically evaluates the performance of the MusicRec model under four different codebook sizes (64, 128, 256, and 512) to measure its impact on recommendation accuracy.
[0104] like Figure 6 As shown, experimental results indicate that codebook size has a significant and non-monotonic impact on model performance. As the codebook size gradually increases from 64 to 256, the model's evaluation metrics on all three datasets continuously improve, reaching peak performance at a codebook size of 256. However, when the codebook size is further increased to 512, performance declines significantly. This phenomenon can be attributed to the inherent trade-off between representational power and generalization robustness: when the codebook is too small, its discrete representation space capacity is limited, making it difficult to fully characterize the fine-grained differences in semantics and structure among multimodal items, leading to identifiers of different items being mapped to similar or identical codewords, thus weakening the model's discriminative ability; when the codebook is too large, although theoretical expressive power is enhanced, a large number of redundant codewords are not fully activated or are only activated by noisy samples, introducing unstable spurious patterns, making the model prone to overfitting sparse or anomalous interaction behaviors, thus impairing generalization performance. In contrast, a codebook of size 256 achieves an optimal balance between expressive power and robustness, providing sufficiently rich discrete semantic units to distinguish multimodal features and collaborative contexts of different items, while avoiding noise interference and training instability caused by codeword redundancy. This result also verifies the efficiency and practicality of the shared-specific residual quantization structure adopted in this invention at medium codebook sizes.
[0105] This embodiment employs a shared-specific segmentation quantization mechanism in the multimodal semantic enhancement stage, aiming to model cross-modal common semantics and modality-specific features separately. To deeply evaluate the relative contributions of shared codebooks and specific codebooks in identifier construction, this experiment, with a fixed total number of codebooks of 4, adjusts the allocation ratio of the two types of codebooks, including three configurations: [1:3] (1 shared codebook + 3 specific codebooks), [1:1] (2 shared codebooks + 2 specific codebooks), and [3:1] (3 shared codebooks + 1 specific codebook). The recommendation performance of the model is evaluated on three benchmark datasets.
[0106] like Figure 7As shown, experimental results indicate that the model achieves optimal results on evaluation metrics when the ratio of shared codebook to specific codebook is 1:1. This configuration optimally balances cross-modal semantic alignment and modality-specific preservation, ensuring that the generated semantic identifiers possess both good multimodal consistency and sufficient expression of unique information from each modality. In contrast, with a smaller shared codebook, model performance slightly decreases. This phenomenon can be attributed to the limited capacity of shared semantic representation; too few shared codebooks make it difficult to effectively model the deep semantic relationships between text and images, leading to insufficient cross-modal alignment and weakening the gains from multimodal fusion. With too few specific codebooks, model performance significantly decreases, fundamentally due to the overemphasis on shared semantics at the expense of modality-specificity. Insufficient specific codebook capacity prevents the model from fully preserving the semantic details of text or the visual characteristics of images, resulting in the loss of key modal information and consequently affecting the discriminative ability and generation quality of identifiers.
[0107] In summary, the experiments verified the complementarity of sharing and specific semantics in the construction of multimodal identifiers, and showed that the two need to be maintained in a reasonable ratio. The 1:1 codebook allocation ratio achieves the optimal balance between cross-modal collaboration and modal individual expression, and is a key design choice for the efficient operation of the multimodal semantic enhancement mechanism in this model. For detailed comparison results of identifier codebook contribution rates, please refer to... Figure 6 .
[0108] This invention constructs a generative recommendation method that fuses multimodal semantics and collaborative signals, effectively addressing key issues in current generative recommendation systems such as shallow semantic understanding, fragmented multimodal information, and insufficient utilization of collaborative signals. By constructing a semantic-collaborative signal joint alignment mechanism, deep semantic representations from multimodal content such as text and images are deeply fused with collaborative signals such as user-item interaction sequences and contextual co-occurrence patterns at the item identifier level, generating a joint semantically enhanced identifier. This identifier not only carries rich semantic information but also embeds structured behavioral priors, enabling precise driving of the generative recommendation model to output semantically consistent, distinctive, and context-relevant recommended content.
[0109] Experiments on multiple real-world e-commerce recommendation datasets demonstrate that our proposed method significantly outperforms existing traditional discriminative recommendation methods and mainstream generative recommendation baselines in terms of recommendation accuracy. Ablation studies further validate the necessity of the multimodal semantic fusion module and the collaborative signal alignment mechanism: using collaborative signals alone cannot reach the performance ceiling of joint modeling, indicating that the collaborative enhancement of semantics and interaction signals is irreplaceable. In scenarios rich in multimodal content, heterogeneous information such as images, text, and audio should be fully utilized to improve generation quality; simultaneously, the stability of joint training between the multimodal encoder and the large language model must be ensured to avoid modal noise interfering with identifier construction. In cold-start or sparse interaction scenarios, although collaborative signals are weak, multimodal semantics can provide strong priors; in this case, the cross-modal alignment loss should be strengthened to ensure the semantic integrity of identifiers. In high-interaction-density scenarios, the weights of collaborative signals can be appropriately increased to capture fine-grained preferences.
[0110] In summary, the generative recommendation model based on collaborative signal-supervised multimodal semantic enhancement identifiers proposed in this invention provides new ideas and technologies for exploring scalable multimodal intelligent representation frameworks, realizing cross-platform interest transfer and unified semantic identifier construction, and further improving recommendation performance in cold-start and sparse scenarios. Systematic experiments conducted on multiple real-world e-commerce product datasets (covering musical instruments, art, games, etc.) in this embodiment demonstrate that the MusicRec model significantly outperforms existing traditional discriminative recommendation models and mainstream generative recommendation models in terms of click-through rate (HR) and normalized discounted cumulative gain (NDCG), and the significance test is satisfactory. The values are all less than 0.05, verifying the high degree of consistency between the generated results and the user's intent and the semantics of the items, indicating the effective utilization of multimodal semantic information and collaborative-semantic alignment mechanisms in the identifier generation process. The experimental results also highlight the complementarity and irreplaceability of multimodal semantics and collaborative signals in generative recommendation systems, emphasizing the importance of their deep integration. Based on the above experimental results, this invention provides key technical support for constructing a semantically driven, structure-aware, and controllable generative intelligent recommendation system, and has broad application prospects and significant engineering promotion value in multimodal intensive recommendation scenarios such as e-commerce, short videos, and content platforms.
[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A collaborative generative recommendation method based on multimodal semantic enhancement tagging, characterized in that: The methods include: S1. Construct an entity database, where entities contain both text and image data; extract text semantic embeddings from each entity in the entity database. Image semantic embedding and collaborative embedding Collaborative embedding This includes associated information and / or text-image collaborative attributes, where associated information is the information associated with the text and / or image; S2. Construct a semantically enhanced collaborative alignment and fusion representation method model. This model utilizes entity-related textual semantic embeddings. With image semantic embedding Semantic enhancement, co-alignment, and fusion processing are performed to obtain quantized text embeddings. and quantized image embedding And the corresponding data is stored in the entity database; the semantic enhancement collaborative alignment fusion expression method model processing method is as follows: S11. Embedding text semantics With image semantic embedding Compressed into latent semantic embeddings of the same dimension and Embedding latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. ; S12. The last shared layer input to the modality-specific quantizer is processed by modal shared residual quantization for modal separation, and then... Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences ; S13, Modal sharing of code sequences With text-specific code sequences Jointly constructing quantized text embeddings ; Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings ; S14. Semantic Enhancement Collaborative Alignment Fusion Representation Method Model Output Quantized Text Embedding of Entities and quantized image embedding ; S3. The user inputs an entity of interest. Text semantic embedding and image semantic embedding are extracted for the entity of interest and input into the semantic enhancement collaborative alignment fusion representation model to obtain the quantized text embedding, quantized image embedding and collaborative embedding representation of the entity of interest. Taking the quantized text embedding, quantized image embedding and / or collaborative embedding representation of the entity of interest as the search task target, the entity database is searched for entities or / and combinations of entities associated with the entity of interest to generate recommendation results.
2. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 1, characterized in that: In method S1, if the entity database is an e-commerce item and service database, then the entities in the entity database are e-commerce items or services, and the generated recommendation result is the associated recommended e-commerce items or services; the entities in the entity database are associated with and stored as sub-entities with the same semantics or similarity greater than a threshold. All sub-entities with the same semantics or similarity greater than a threshold are stored in the entity database with the same entity and the corresponding sub-entities are associated with the entity attributes.
3. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 1, characterized in that: In method S3, the user inputs multimodal data of interest, which includes text data and / or image data. The core entities in the multimodal data of interest are extracted by large language model recognition.
4. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 1, characterized in that: In method S14, the quantized text embedding is extracted. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers; image embeddings are extracted. Shared codewords and image-specific codewords constitute an image lexical sequence and serve as an image identifier; the cooperative embedding representation It also records quantified text embeddings With quantized image embedding Collaboratively perceive alignment information.
5. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 1, characterized in that: In method S11, the latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. ; Each level in the layer-shared codebook Each is equipped with a shared codebook , , The codeword with sequence number n is represented by N, where N is the size of the codebook; the residual quantization expression for modal shared residual quantization is as follows: ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals For the selected codeword; modal shared code sequence .
6. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 5, characterized in that: In method S12, a modality-specific quantizer is used to perform modality-shared quantization on the residual embedding of the last shared layer. Modality separation is performed to extract text semantics for text modality quantization. and image semantics for image modality quantization ; on text semantics pass The layered text's specific codebook is quantized into a sequence of codewords. Layer of a specific codebook The corresponding text modal codebook is obtained. Text-specific code sequences ; Image semantics pass The layer image is quantized into a codeword sequence based on a specific codebook. Layer image specific codebook level The corresponding image modality codebook is obtained. Image-specific code sequence .
7. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 6, characterized in that: In quantified text embedding Reconstructing the original input features and encoding them yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The following loss function is constructed based on the residual quantization processing of methods S11 and S12: ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, This represents a coefficient that balances the strength of code embedding and encoder optimization.
8. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 7, characterized in that: The total loss function for quantization embedding is constructed, and its expression is as follows: ,in This represents the total loss of quantization embedding. This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.
9. The collaborative generative recommendation method based on multimodal semantic enhancement identifiers according to claim 8, characterized in that: The total loss function expression for the semantic enhancement, collaborative alignment, and fusion representation method model is as follows: , For the collaborative alignment loss of collaborative embedding representation, To preserve the loss for the cooperative relationships in the cooperative embedding representation, , They represent control , Hyperparameters of capability strength.