Semantic enhancement collaborative alignment and fusion expression method based on multi-modal data
By performing text and image semantic embedding and shared quantization on the multimodal data of the project, the problems of modal fragmentation and signal isolation in generative recommendation models were solved. Multimodal semantic enhancement identifiers were constructed, which improved the semantic richness and structure awareness of the generative recommendation system and realized the synergistic enhancement of semantic understanding and structure awareness of the generative recommendation system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- COMMUNICATION UNIVERSITY OF CHINA
- Filing Date
- 2025-11-19
- Publication Date
- 2026-05-12
AI Technical Summary
Existing generative recommendation models suffer from modal fragmentation and signal isolation when processing multimodal information. They lack cross-modal semantic deep fusion and alignment, and cannot effectively utilize collaborative signals from user interaction history, resulting in insufficient semantic understanding and a lack of guidance in the generation process.
By extracting text and image semantic embeddings from the project's multimodal data, generating joint latent representations using linear projection layers, and constructing modality-sharing and specific code sequences through modality sharing and specific quantization processing, outputting quantized text and image embeddings, recording collaborative perception alignment information, and reconstructing the original features using a pre-trained model and variational autoencoder to construct collaborative embedding representations.
It achieves cross-modal semantic alignment and interaction-aware modeling, constructs multimodal semantic enhancement identifiers, improves the semantic richness and structure awareness of generative recommendation systems, enhances the content relevance and personalization of recommendation systems, and promotes the evolution of recommendation systems from discriminative to generative paradigms.
Smart Images

Figure CN121502676B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic enhancement and collaborative processing of multimodal data, and in particular to a semantic enhancement, collaborative alignment, and fusion expression method based on multimodal data. Background Technology
[0002] Recommender systems play a crucial role in exploring personalized content across various scenarios, such as video platforms, e-commerce shopping, and movie recommendations. With the rapid development of large language models and recommender system technologies, generative recommendation, as a new generation of recommendation paradigm, is gradually evolving from the traditional "retrieval-ranking" model to an "understanding-generation" model. Compared to traditional methods that rely solely on user-item interaction history for matching, generative recommendation, through a constrained generation mechanism, can dynamically generate personalized recommendation results and improve content distribution quality in cold-start scenarios, significantly enhancing user experience and system flexibility. Encoding the multimodal information of items (including goods, video entities, objects, etc.) is the core of the "understanding-generation" model, effectively addressing the representational limitations of discriminative methods in the traditional "retrieval-ranking" model.
[0003] Currently, there are still some problems to be solved in the research on semantic understanding depth, multi-source information fusion capability, and utilization of collaborative signals. First, in real-world scenarios, multimodal information (such as text data, image data, user preferences, and item features) is often presented in multimodal forms such as text and images. However, the semantic identifier construction of current generative recommendation models relies on unimodal information. Some studies only process multimodal information at the stage of simple splicing or independent encoding, lacking deep fusion and alignment of cross-modal semantics. The isolation of different modalities hinders the model from learning a unified semantic representation that utilizes complementary information. Second, existing research only focuses on the semantic attributes of unimodality, ignoring collaborative signals derived from user interaction history, resulting in the inability to retain the interaction patterns that exist simultaneously between semantic understanding and collaborative behavior. Existing methods usually model collaborative signals and semantic content separately, failing to achieve their collaborative construction at the semantic enhancement level, resulting in a lack of deep joint guidance of multimodal semantics in the generation process.
[0004] In summary, against the backdrop of the explosive growth of multimodal content and the vigorous development of generative artificial intelligence, there is an urgent need for a semantically enhanced collaborative alignment fusion expression method that can deeply integrate multimodal semantics and collaborative signals to construct a new type of multimodal generative expression, thereby breaking through the technical bottlenecks of modal fragmentation and signal isolation. Summary of the Invention
[0005] The purpose of this invention is to provide a semantic enhancement collaborative alignment and fusion representation method based on multimodal data, which breaks through the technical bottlenecks of modality fragmentation and signal isolation, and realizes the collaborative enhancement of semantic understanding and structure awareness, thereby providing a unified representation foundation for generative recommendation that combines semantic richness, structure awareness and generative guidance capabilities.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] A semantically enhanced collaborative alignment and fusion representation method based on multimodal data, the method comprising:
[0008] S1. Extract text semantic embeddings from the project's multimodal data. Image semantic embedding and collaborative embedding representation Collaborative embedding This includes associated information and / or text-image co-attributes, where associated information is the information associated with the text and / or image, embedding text semantics. With image semantic embedding Compressed into latent semantic embeddings of the same dimension and ;
[0009] S2, embedding latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. ;
[0010] S3. The last shared layer is input into a mode-specific quantizer for modal separation after modal shared residual quantization. Then, the modes are separated separately... Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences ;
[0011] S4. Modal sharing of code sequences With text-specific code sequences Jointly constructing quantitative text embeddings ; Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings ;
[0012] S5, Quantitative text embedding of interrelated output items and quantized image embedding .
[0013] To better implement this invention, in method S4, the quantized text embedding is extracted. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers; extract quantized image embeddings. The shared codewords and image-specific codewords constitute an image lexical sequence and serve as the image identifier.
[0014] Preferably, the collaborative embedding representation It also records quantified text embeddings With quantized image embedding Collaboratively perceive alignment information.
[0015] Preferably, in method S1, a pre-trained model is used to extract text semantic embeddings from the multimodal data of the project. With image semantic embedding Text semantic embedding through a text encoder Compression into latent semantic embeddings Image semantic embedding is achieved through an image encoder. Compression into latent semantic embeddings .
[0016] Preferably, in method S2, the latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. ; Each level in the layer-shared codebook Each is equipped with a shared codebook , , The codeword with sequence number n is represented by N, where N is the size of the codebook; the residual quantization expression for modal shared residual quantization is as follows:
[0017] ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals For the selected codeword; modal shared code sequence .
[0018] Preferably, in method S3, a modality-specific quantizer is used to perform modality-shared quantization on the residual embedding of the last shared layer. Modality separation is performed to extract text semantics for text modality quantization. and image semantics for image modality quantization .
[0019] Preferably, in method S3, the text semantics are... pass The layered text's specific codebook is quantized into a sequence of codewords. Layer of a specific codebook The corresponding text modal codebook is obtained. The residual quantization expression for modality-specific residual quantization is as follows:
[0020] ,in Indicates the selection of the first code from a specific codebook of text. Level code words, It is the first Level of text semantic residuals For selected codewords; text-specific code sequences .
[0021] Preferably, in method S3, the image semantics are... pass The layer image-specific codebook is quantized into a codeword sequence, and the residual quantization processing expression for mode-specific residual quantization is as follows:
[0022] ,in Indicates the selection of the first image from a specific codebook. Level code words, It is the first Level-1 image semantic residuals For selected codewords; image-specific code sequences .
[0023] Preferably, in quantized text embedding Reconstructing and encoding the original input features yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The residual quantization processing of methods S2 and S3 constructs the following loss function:
[0024] ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, This represents a coefficient that balances the strength of code embedding and encoder optimization.
[0025] Preferably, the total loss function expression for methods S2 to S5 is as follows:
[0026] ,in Indicates the total loss. This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.
[0027] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0028] (1) This invention breaks through the technical bottleneck of modal fragmentation and signal isolation, and realizes the synergistic enhancement of semantic understanding and structural perception, thereby providing a unified representation basis for generative recommendation that combines semantic richness, structural perception and generative guidance capabilities.
[0029] (2) This invention constructs a multimodal semantic enhancement identifier explicitly supervised by collaborative signals through cross-modal semantic alignment, interaction-aware modeling and joint coding optimization. As the core input representation of the generative recommendation model, the identifier not only integrates deep semantic information from multimodal content such as text description and product images, but also embeds high-order collaborative signals such as user-item interaction graph and context co-occurrence pattern, thereby achieving synergistic enhancement of semantic richness and structural sensitivity in a unified semantic space.
[0030] (3) This invention addresses the problems of insufficient semantic understanding, independent modal features, and inadequate utilization of collaborative signals in existing generative recommendation systems. It aligns and fuses multimodal semantic information from items with collaborative signals such as user-item interactions and contextual co-occurrence, jointly generating a unified identifier with rich semantics and structural awareness. This improves the performance of multimodal generative recommendation systems in terms of content relevance, personalization, and semantic consistency. This invention effectively enhances the ability of recommendation systems to fuse complex heterogeneous information, promotes the evolution of recommendation systems from a discriminative to a generative paradigm, and provides key technical support for intelligent content distribution, personalized services, and human-computer interaction.
[0031] (4) The application of this invention in multiple application scenarios demonstrates its effective utilization of multimodal semantic information and collaborative-semantic alignment mechanisms in the identifier generation process. The experimental results also highlight the complementarity and irreplaceability of multimodal semantics and collaborative signals in generative recommendation systems, emphasizing the importance of their deep integration. This invention provides key technical support for building next-generation semantically driven, structure-aware, and generatively controllable intelligent recommendation systems, and has broad application prospects and significant engineering promotion value in multimodal intensive recommendation scenarios such as e-commerce, short videos, and content platforms.
[0032] (5) This invention can better integrate multimodal item semantics, extract shared quantization codes between different modalities, and learn specific quantization codes for each modality based on shared representation features, enabling the generated identifiers to capture richer cross-modal semantic representations. To effectively utilize collaborative information, a collaborative semantic alignment mechanism is introduced during the identifier construction learning process. This invention ensures that the semantic quantization representation and collaborative embedding maintain distributional alignment, thereby maintaining the proximity between items with similar interaction patterns in the quantization space while effectively driving the generative recommendation model to achieve more accurate and semantically consistent recommendation content generation. Attached Figure Description
[0033] Figure 1 This is a flowchart of the semantic enhancement collaborative alignment fusion expression method of the present invention;
[0034] Figure 2 This is a schematic diagram illustrating the principle of the semantic enhancement collaborative alignment fusion expression method in the embodiment;
[0035] Figure 3 This is a graph showing the hyperparameter variation results of Example 3 on the Instruments dataset;
[0036] Figure 4 This is a graph showing the hyperparameter variation results of Example 3 on the Arts dataset;
[0037] Figure 5 This is a graph showing the hyperparameter variation results of Example 3 on the Games dataset;
[0038] Figure 6 A comparison chart showing the impact of codebook size on model performance on the Arts dataset;
[0039] Figure 7 A comparison chart showing the impact of codebook contribution rate on model performance on the Arts dataset. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to embodiments:
[0041] Example 1
[0042] like Figure 1 As shown, a semantic enhancement collaborative alignment and fusion representation method based on multimodal data is proposed, the method comprising:
[0043] S1. Extract text semantic embeddings from the multimodal data of the project (in this invention, the project refers to multimodal entities such as commodities and objects). Image semantic embedding and collaborative embedding representation Collaborative embedding This includes associated information and / or text-image collaborative attributes. Associated information refers to the information associated with the text and / or image (associated information includes user-item interaction information, such as user intent, comments, etc.; if used in a recommendation system, it recommends a list of items to the user, including selections and comments from users with different tags; associated information contains a large amount of rich association information between users and items). The multimodal data of this invention's items consists of several multimodal data entries for the same item (for example, in an e-commerce marketplace, if there are multiple entries for the same item from different merchants, they are recorded as the same item; during recommendation, the item is recommended first, and then the various suppliers providing the item are recommended accordingly). Text-image collaborative attributes are the association signals or features between the extracted text semantics and image semantics of the item (capable of identifying and associating the association information between text and image, facilitating auxiliary prompts during item recommendation). Text semantics are embedded... With image semantic embedding Compressed into latent semantic embeddings of the same dimension and More preferably, pre-trained models (such as LLaMA and ViT) are used to extract textual semantic embeddings from the project's multimodal data. With image semantic embedding Text semantics are embedded through a text encoder (a text encoder implemented using a multilayer perceptron). Compression into latent semantic embeddings Image semantics are embedded through an image encoder (an image encoder implemented using a multilayer perceptron). Compression into latent semantic embeddings The expression is as follows:
[0044] ,
[0045] Text encoder Image encoder These are implemented using multilayer perceptrons.
[0046] S2, embedding latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. .
[0047] In some embodiments, latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. . Each level in the layer-shared codebook Each is equipped with a shared codebook , , Let N represent the codeword with sequence number n, and N be the size of the codebook. The residual quantization expression for modal shared residual quantization is as follows:
[0048] ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals The selected codewords are used to obtain the modal shared code sequence through modal shared residual quantization. .
[0049] S3. Quantize the modal shared residuals of the last shared layer (preferably by selecting the residual embedding of the shared layer). The input residual is embedded into a mode-specific quantizer for mode separation, and then processed separately. Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences .
[0050] In some embodiments, a modality-specific quantizer is used to embed the residuals of the last shared layer in modality-shared quantization. Modality separation is performed to extract text semantics for text modality quantization. and image semantics for image modality quantization Regarding text semantics pass The layered text's specific codebook is quantized into a sequence of codewords. Layer of a specific codebook The corresponding text modal codebook is obtained. The residual quantization expression for modality-specific residual quantization is as follows:
[0051] ,in Indicates the selection of the first code from a specific codebook of text. Level code words, It is the first Level of text semantic residuals The selected codewords are used to obtain the text-specific code sequence through modality-specific residual quantization. .
[0052] Image semantics pass The layer image-specific codebook is quantized into a codeword sequence, and the residual quantization processing expression for mode-specific residual quantization is as follows:
[0053] ,in Indicates the selection of the first image from a specific codebook. Level code words, It is the first Level-1 image semantic residuals The selected codewords are used to obtain the image-specific code sequence through modality-specific residual quantization. .
[0054] S4. Modal sharing of code sequences With text-specific code sequences Jointly constructing quantitative text embeddings Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings .
[0055] In some embodiments, during modality-specific residual quantization (including two process stages: obtaining text and image-specific code sequences), the present invention quantizes text embeddings... Reconstructing the original input features and encoding them yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The aggregated quantization embedding for each mode is reconstructed using a decoder based on a variational autoencoder. The reconstruction decoding process is expressed as follows: , , , These are the decoders.
[0056] The following loss function is constructed for the residual quantization processing of methods S2 and S3 of the present invention:
[0057] ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, A coefficient representing the balance between code embedding and encoder optimization intensity (set to 0.25 in this example).
[0058] S5, Quantitative text embedding of interrelated output items and quantized image embedding Preferably, in method S4, the quantized text embedding is extracted. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers. Extract quantized image embeddings. Shared codewords and image-specific codewords constitute an image lexical sequence and serve as the image identifier. Cooperative embedding representation. It also records quantified text embeddings With quantized image embedding Collaborative sensing alignment information (performing collaborative sensing and deep alignment in methods S2 to S4 to obtain semantically enhanced collaborative information, collaborative embedding representation) It records the collaborative information of the pre-trained model and the collaborative information of the final text and image embedding representations, possessing rich and reliable comprehensive collaborative information. (For example...) Figure 2 As shown, to facilitate better structuring of text and image identifiers and to enable ordered machine recognition and judgment, text identifiers (i.e., text-specific modalities) use a set of lowercase letters {a, b, ...} to represent levels, while image identifiers (image-specific modalities) use a set of uppercase letters {A, B, ...} to represent levels. Then, complete semantically enhanced identifiers are constructed sequentially as new lexical sequences. For example, the text identifier for a project can be labeled as...<s1_1><s2_2><a_3><b_4> The corresponding image identifier can be marked as<S1_1><S2_2><A_5><B_6> .
[0059] In some embodiments, the total loss function expression for methods S2 to S5 is as follows:
[0060] ,in Indicates the total loss. This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.
[0061] In some embodiments, the present invention also constructs a collaborative embedding representation. Collaborative sensing alignment loss and collaborative relationship preservation loss are used to constrain the loss (improving the accuracy of the entire model processing stage) to bridge the gap between semantic features and collaborative signals. A pre-trained sequence recommendation model, SASRec, is used to obtain collaborative filtering embeddings of items. To better utilize collaborative signals to align semantic information, the InfoNCE loss function is employed to quantize the item embeddings. (Includes quantized text embedding) and quantized image embedding The collaborative filtering embedding is closer to the item while keeping it away from the embeddings of other irrelevant items in the training batch, resulting in a collaboratively perceptual alignment loss. The expression is as follows:
[0062] ,in Represents the cosine similarity function. This indicates temperature hyperparameters. Indicates the training batch size. This represents the quantized embedding representation of training batch m. This represents the collaborative filtering embedding of training batch m. Let represent the quantized embedding representation of training batch j. To preserve the cooperative relationship structure encoded in the cooperative embedding space, a cooperative relationship preservation loss is introduced to enforce structural consistency between the quantized embedding space and the cooperative embedding space; specifically, given a training batch of an item... Calculate the pairwise similarity matrix in two spaces, and calculate the cooperative relationship preservation loss. The expression is as follows:
[0063] ,in This represents the cosine similarity of the embeddings quantized in training batches m and j. Let represent the cosine similarity of the system embeddings in training batches m and j. Preferably, to maintain the stability of the semantic identifier training process, optimization is performed by jointly considering semantic loss, collaborative awareness alignment loss, and collaborative relation preservation loss. This invention constructs a collaborative embedding representation. The total loss, consisting of collaborative sensing alignment loss, collaborative relationship preservation loss, and semantic loss, is expressed as follows:
[0064] ,in , They represent control and Hyperparameters of capability strength.
[0065] Example 2
[0066] This embodiment uses all the techniques from Embodiment 1 to perform the recommendation generation task, such as... Figure 2 As shown, this invention designs a series of structured recommendation tasks in the execution of recommendation generation tasks, aiming to fully mine and utilize the multimodal semantic information and collaborative signals contained in the constructed semantically enhanced identifiers, thereby improving the generative recommendation capability. The recommendation generation tasks can introduce explicit cross-modal fusion and alignment mechanisms to achieve deep semantic fusion and generation guidance between modalities. In the execution of recommendation generation tasks, this invention can, based on the user's sequential behavior patterns in a sequential interaction environment, predict the user's next possible interaction item, and further subdivide it into the following two types of subtasks according to the composition of the input modality:
[0067] (1) Single-mode generation
[0068] In this task, the input is user item data in a certain modality (such as plain text description or pure image features) to generate single-modal or multimodal recommendations.
[0069] Project generation (text modal) example:
[0070]
[0071] Project generation (image modality) example:
[0072]
[0073] (2) Multimodal generation
[0074] To more realistically simulate the complexity of cross-modal interactions between users in real-world scenarios, this task enables the model to explicitly model the semantic complementarity and behavioral synergy between different modalities during the generation process, thereby improving the multimodal consistency of the generated results.
[0075] Project generation (text modal) example:
[0076]
[0077] Project generation (image modality) example:
[0078]
[0079] While large language models can implicitly learn cross-modal associations through context during sequence generation, this alignment is often loose and unstable, especially in scenarios with missing modalities or noise interference, which can easily lead to semantic drift. To address this, this invention introduces an explicit cross-modal alignment objective to force the model to establish strong semantic consistency constraints in the identifier space, ensuring high alignment of representations of the same item across different modalities. This alignment objective comprises two symmetrical subtasks.
[0080] (1) Text to Image Alignment: Given a text modal token representation of an item, the model needs to reconstruct or generate its corresponding image modal token representation.
[0081] (2) Image to text alignment: Given an image modality token representation of an item, the model needs to generate its corresponding text description token sequence.
[0082] Text-to-image alignment example:
[0083]
[0084] Image-to-text alignment example:
[0085]
[0086] This task not only enhances the model's understanding of cross-modal semantic equivalence, but also significantly improves the robustness of identifiers in scenarios with missing or partially observed modalities. Even when only a single modality is observed, the model can still infer the complete item semantics through the alignment structure in the identifier space, thereby supporting high-quality generative recommendations.
[0087] After obtaining the semantic identifiers for each modality of each item, the sequence recommendation task is modeled as a sequence-to-sequence generation problem. Specifically, each user's interaction history is uniformly transformed into an input sequence composed of item semantic identifiers, while the recommendation target is formalized as generating the complete semantic identifier of the next target item. Let the user's historical interaction sequence be... Each identifier All are made by A sequence of discrete semantic tags, generated by the aforementioned multimodal semantic and cooperative signal fusion mechanism, contains rich cross-modal semantic and structured behavioral information. This invention employs a T5 model based on an encoder-decoder architecture as the generation backbone. This model takes the input sequence... Encodes a context-aware latent state representation and progressively generates a semantic identifier for the next target item through an autoregressive decoder. , recorded as The entire process follows the standard sequence-to-sequence paradigm, but its input and output are both structured sequences of semantic identifiers, rather than raw text or item IDs, thus achieving full utilization of multimodal semantics and cooperative signals.
[0088] The training objective of the model is to maximize the performance of a given historical sequence. Correctly generate target identifier under the condition The probability. Optimization is achieved by minimizing the standard cross-entropy loss function; for each training sample, the model parameters... The optimization objective is defined as follows: , express The A token.
[0089] During the reasoning process, generative recommendation models are based on users' historical interaction sequences. Autoregressive generation of semantic identifiers for the next project To improve the accuracy and diversity of the generated results, each identifier code... A beam search strategy is used for selection to balance generation quality and search efficiency during decoding. However, since the semantic identifier space is discrete and structured, direct generation may correspond to a "fictitious" identifier that does not exist in the item library. To ensure the generation results... To consistently map to real-world items, a trie-based constraint decoding strategy is introduced: during decoding, each step only allows the selection of valid tokens that can form prefixes of existing item identifiers, thus forcing the model to generate within the valid item space. This strategy not only ensures the feasibility of the recommendation results but also significantly reduces invalid or semantically drifting generated samples.
[0090] To further improve recommendation quality and fully utilize the complementarity of multimodal information, a cross-modal re-ranking strategy is adopted. During the inference phase, textual and visual modal generation tasks are performed separately to obtain two preliminary recommendation lists. and Each list contains a relevance score. The final score for each item is calculated as follows: for items that appear in both lists, the score is calculated as follows: ,in and The extra 1 point is awarded to emphasize multimodal consistency, indicating that the item is highly relevant to the user's history at both the textual and visual semantic levels; for items that exist in only one list, their original scores are retained. or This re-ranking strategy effectively integrates the ability of different modalities to characterize user preferences: items whose preferences are significant across multiple modalities are given higher confidence; while items that stand out only in a single modality are still given a recommendation opportunity. This mechanism not only improves the accuracy and robustness of recommendations but also enhances the system's adaptability to heterogeneous user behavior patterns, thereby enabling more accurate and personalized generative recommendations in diverse real-world application scenarios.
[0091] like Figure 2As shown, the generation process of multimodal semantically enhanced identifiers in this invention integrates multimodal item semantics through shared modality segmentation and specific modality segmentation, and employs a collaborative semantic alignment mechanism to preserve collaborative relationships, including collaborative perceptual alignment and a collaborative relationship preservation loss function component. In the recommendation result generation stage, this process includes three main generative sequence recommendation training tasks, with the encoder-decoder model autoregressively generating recommendation results. These are three dedicated tasks: next item generation (unimodal), next item generation (multimodal), and multimodal item alignment, to effectively utilize the generated identifiers during recommendation training. The synergistic optimization of these three tasks enables the model to fully utilize the semantic and structural information encoded in the identifiers during training, thereby achieving high-precision and high-consistency generative recommendations in the inference stage. Overall, this framework, through a two-stage paradigm of semantically enhanced identifier construction and multi-task generative learning, achieves an organic unity of multimodal content understanding, collaborative signal utilization, and sequence generation capabilities.
[0092] Example 3
[0093] This embodiment employs all the techniques from Embodiment 1 to perform the recommendation generation task. It systematically evaluates the data on three publicly available real-world benchmark datasets, all derived from the Amazon Product Review Dataset. This dataset is widely used in recommender system research, containing detailed metadata about products on the Amazon platform (such as titles, descriptions, categories, images, etc.) and large-scale user behavior records (including ratings, reviews, and interaction timestamps), covering multiple product category domains and providing a rich and representative experimental foundation for multimodal generative recommender tasks. This experiment selects three product category subsets with significant multimodal characteristics: "Instruments," "Arts," and "Games," for sequence recommendation tasks. To ensure data quality and the reliability of model training, a 5-core filtering strategy is adopted, retaining only users who have interacted with at least 5 different products, and products that have been interacted with by at least 5 different users. This preprocessing method is consistent with mainstream research in the recommender system field, effectively eliminating noise from sparse interactions and improving the fairness and stability of the evaluation. Subsequently, the interaction records of each user were strictly sorted by timestamp to construct their behavioral sequences, realistically reflecting the dynamic evolution of user interests. To strike a balance between computational efficiency and contextual modeling capabilities, the maximum length of all user sequences was standardized to 20 items. Table 1 shows the statistical information for each dataset, where "average length" refers to the average number of items in the filtered user interaction sequence, reflecting the typical browsing depth of users within that category. For multimodal representation, each product item incorporates two types of heterogeneous features. Textual features are extracted from the product title and description text, encoded into fixed-dimensional semantic vectors using a pre-trained language model. Visual features are extracted from the product illustration, obtaining deep image embeddings through a pre-trained visual model.
[0094] Table 1 Statistical analysis of dataset composition
[0095]
[0096] To delve into the impact of key hyperparameters in the model on the utilization of multimodal semantics and collaborative signal alignment mechanisms, this section systematically evaluates their effect on generative recommendation accuracy from the perspective of hyperparameter sensitivity analysis. Experiments were conducted while keeping other components and training settings constant, adjusting two core hyperparameters and the collaborative sensing alignment weights respectively. Maintaining strength of collaborative relationships We observed the changing trends of model performance on three benchmark datasets, using click-through rate and normalized depreciation cumulative gain as the main evaluation metrics.
[0097] First, for the collaborative sensing alignment mechanism, the hyperparameters are... The value range is set to {1e-1, 1e-2, 2e-2, 1e-3}. For example... Figure 3 – Figure 5 As shown, the experimental results indicate that when When the coefficient of performance is 1e-2, the model achieves optimal performance on all three datasets. At this point, the supervisory role of the cooperative signal in multimodal semantic alignment reaches its optimal balance, effectively guiding semantic identifiers to learn structured behavioral patterns without suppressing the semantic expressive power of the modality itself due to excessive constraints. When the value is increased to 1e-1, the dominance of the cooperative signal becomes too strong, causing the model to overfit the interaction structure and ignore fine-grained multimodal semantic differences; while when When the value is too small, the alignment constraint is too weak to effectively establish the association between semantics and collaboration, thus weakening the discriminative ability of the identifier. This phenomenon verifies the sensitivity of the collaboration-aware alignment mechanism to α and highlights the crucial role of setting this hyperparameter appropriately in improving model performance.
[0098] Secondly, to evaluate the mechanism for maintaining collaborative relationships, hyperparameters will be used. The value range of is set to {1e-2, 1e-3, 1e-4, 1e-5}. Experimental results show that the model performance also exhibits a significant non-monotonic trend. When When the value is 1e-4, the evaluation metrics on all three datasets reach their peak. This result indicates that moderate relation-preserving regularization helps stabilize the embedding space of identifiers and enhances their ability to perceive higher-order cooperative relationships. However, if... If the size is too large, it will introduce an excessively strong smoothing effect, causing the identifiers of different users or items to become homogenized, reducing the ability to express personalization; conversely, if... If the value is too small, it will not be able to effectively constrain the identifier learning process and will be difficult to retain the topological information in the original interaction graph.
[0099] In summary, hyperparameters and Semantic-cooperative alignment strength and structural consistency preservation strength are controlled separately, each with a clearly defined optimal value range. Experiments not only verify the effectiveness of the proposed alignment and preservation mechanisms but also provide clear parameter tuning guidance for practical deployments, allowing for fine-tuning to maximize performance. This analysis further corroborates the necessity of deep fusion of multimodal semantics and cooperative signals and their decisive impact on generation quality.
[0100] This embodiment explores the mechanism by which the quantization codebook size affects the identifier representation ability of generative recommendation models by comparing the impact of different codebook sizes on model performance during shared-specific residual quantization. While keeping other hyperparameters and training configurations constant, this experiment systematically evaluates the performance of the MusicRec model under four different codebook sizes (64, 128, 256, and 512) to measure its impact on recommendation accuracy.
[0101] like Figure 6 As shown, experimental results indicate that codebook size has a significant and non-monotonic impact on model performance. As the codebook size gradually increases from 64 to 256, the model's evaluation metrics on all three datasets continuously improve, reaching peak performance at a codebook size of 256. However, when the codebook size is further increased to 512, performance declines significantly. This phenomenon can be attributed to the inherent trade-off between representational power and generalization robustness: when the codebook is too small, its discrete representation space capacity is limited, making it difficult to fully characterize the fine-grained differences in semantics and structure among multimodal items, leading to identifiers of different items being mapped to similar or identical codewords, thus weakening the model's discriminative ability; when the codebook is too large, although theoretical expressive power is enhanced, a large number of redundant codewords are not fully activated or are only activated by noisy samples, introducing unstable spurious patterns, making the model prone to overfitting sparse or anomalous interaction behaviors, thus impairing generalization performance. In contrast, a codebook of size 256 achieves an optimal balance between expressive power and robustness, providing sufficiently rich discrete semantic units to distinguish multimodal features and collaborative contexts of different items, while avoiding noise interference and training instability caused by codeword redundancy. This result also verifies the efficiency and practicality of the shared-specific residual quantization structure adopted in this invention at medium codebook sizes.
[0102] This embodiment employs a shared-specific segmentation quantization mechanism in the multimodal semantic enhancement stage, aiming to model cross-modal common semantics and modality-specific features separately. To deeply evaluate the relative contributions of shared codebooks and specific codebooks in identifier construction, this experiment, with a fixed total number of codebooks of 4, adjusts the allocation ratio of the two types of codebooks, including three configurations: [1:3] (1 shared codebook + 3 specific codebooks), [1:1] (2 shared codebooks + 2 specific codebooks), and [3:1] (3 shared codebooks + 1 specific codebook). The recommendation performance of the model is evaluated on three benchmark datasets.
[0103] like Figure 7As shown, experimental results indicate that the model achieves optimal results on evaluation metrics when the ratio of shared codebook to specific codebook is 1:1. This configuration optimally balances cross-modal semantic alignment and modality-specific preservation, ensuring that the generated semantic identifiers possess both good multimodal consistency and sufficient expression of unique information from each modality. In contrast, with a smaller shared codebook, model performance slightly decreases. This phenomenon can be attributed to the limited capacity of shared semantic representation; too few shared codebooks make it difficult to effectively model the deep semantic relationships between text and images, leading to insufficient cross-modal alignment and weakening the gains from multimodal fusion. With too few specific codebooks, model performance significantly decreases, fundamentally due to the overemphasis on shared semantics at the expense of modality-specificity. Insufficient specific codebook capacity prevents the model from fully preserving the semantic details of text or the visual characteristics of images, resulting in the loss of key modal information and consequently affecting the discriminative ability and generation quality of identifiers.
[0104] In summary, the experiments verified the complementarity of sharing and specific semantics in the construction of multimodal identifiers, and showed that the two need to be maintained in a reasonable ratio. The 1:1 codebook allocation ratio achieves the optimal balance between cross-modal collaboration and modal individual expression, and is a key design choice for the efficient operation of the multimodal semantic enhancement mechanism in this model. For detailed comparison results of identifier codebook contribution rates, please refer to... Figure 6 .
[0105] This invention constructs a generative recommendation method that fuses multimodal semantics and collaborative signals, effectively addressing key issues in current generative recommendation systems such as shallow semantic understanding, fragmented multimodal information, and insufficient utilization of collaborative signals. By constructing a semantic-collaborative signal joint alignment mechanism, deep semantic representations from multimodal content such as text and images are deeply fused with collaborative signals such as user-item interaction sequences and contextual co-occurrence patterns at the item identifier level, generating a joint semantically enhanced identifier. This identifier not only carries rich semantic information but also embeds structured behavioral priors, enabling precise driving of the generative recommendation model to output semantically consistent, distinctive, and context-relevant recommended content.
[0106] Experiments on multiple real-world e-commerce recommendation datasets demonstrate that our proposed method significantly outperforms existing traditional discriminative recommendation methods and mainstream generative recommendation baselines in terms of recommendation accuracy. Ablation studies further validate the necessity of the multimodal semantic fusion module and the collaborative signal alignment mechanism: using collaborative signals alone cannot reach the performance ceiling of joint modeling, indicating that the collaborative enhancement of semantics and interaction signals is irreplaceable. In scenarios rich in multimodal content, heterogeneous information such as images, text, and audio should be fully utilized to improve generation quality; simultaneously, the stability of joint training between the multimodal encoder and the large language model must be ensured to avoid modal noise interfering with identifier construction. In cold-start or sparse interaction scenarios, although collaborative signals are weak, multimodal semantics can provide strong priors; in this case, the cross-modal alignment loss should be strengthened to ensure the semantic integrity of identifiers. In high-interaction-density scenarios, the weights of collaborative signals can be appropriately increased to capture fine-grained preferences.
[0107] In summary, the generative recommendation model based on cooperative signal supervision and multimodal semantic enhancement identifiers proposed in this invention provides new ideas and technologies for exploring scalable multimodal intelligent representation frameworks, realizing cross-platform interest transfer and unified semantic identifier construction, and further improving recommendation performance in cold start and sparse scenarios.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A semantically enhanced collaborative alignment and fusion representation method based on multimodal data, characterized in that: The methods include: S1. Extract text semantic embeddings from the project's multimodal data. Image semantic embedding and collaborative embedding Collaborative embedding This includes associated information and / or text-image collaborative attributes, where associated information includes user-item interaction information, embedding text semantics. With image semantic embedding Compressed into latent semantic embeddings of the same dimension and ; S2, embedding latent semantics and Generate joint latent representation through linear projection layer , will jointly represent potential pass The layer-shared codebook is used to generate codeword sequences through hierarchical quantization, and then modality-shared code sequences are obtained through modality-shared residual quantization. ; S3. The last shared layer is input into a mode-specific quantizer for modal separation after modal shared residual quantization. Then, the modes are separated separately... Layer-specific codebooks are used to generate codeword sequences through hierarchical quantization, and then modality-specific residual quantization is used to obtain text-specific code sequences. Image-specific code sequences ; S4. Modal sharing of code sequences With text-specific code sequences Jointly constructing quantitative text embeddings ; Modal sharing of code sequences Image-specific code sequences Jointly construct quantized image embeddings ; S5, Quantitative text embedding of interrelated output items and quantized image embedding .
2. The semantic enhancement, collaborative alignment, and fusion representation method based on multimodal data according to claim 1, characterized in that: In method S4, the quantized text embedding is extracted. Shared codewords and text-specific codewords constitute a text lexical sequence and serve as text identifiers; extract quantized image embeddings. The shared codewords and image-specific codewords constitute an image lexical sequence and serve as an image identifier.
3. The semantic enhancement, collaborative alignment, and fusion representation method based on multimodal data according to claim 1, characterized in that: The collaborative embedding It also records quantified text embeddings With quantized image embedding Collaboratively perceive alignment information.
4. The semantic enhancement, collaborative alignment, and fusion representation method based on multimodal data according to claim 1, characterized in that: In method S1, a pre-trained model is used to extract textual semantic embeddings from the multimodal data of the project. With image semantic embedding Text semantic embedding through a text encoder Compression into latent semantic embeddings Image semantic embedding is achieved through an image encoder. Compression into latent semantic embeddings .
5. The semantic enhancement collaborative alignment and fusion representation method based on multimodal data according to claim 1, characterized in that: In method S2, the latent semantic embedding is first performed. and The data is then stitched together, and a joint latent representation is generated through a linear projection layer. ; Each level in the layer shared codebook Each is equipped with a shared codebook , , The codeword with sequence number n is represented by N, where N is the size of the codebook; the residual quantization expression for modal shared residual quantization is as follows: ,in This indicates the minimum Euclidean distance. Indicates the first selected from the shared codebook Level code words, It is the first Level-shared semantic residuals For the selected codeword; modal shared code sequence .
6. The semantic enhancement collaborative alignment and fusion representation method based on multimodal data according to claim 5, characterized in that: In method S3, a modality-specific quantizer is used to perform modality-shared quantization on the residual embedding of the last shared layer. Modality separation is performed, and the resulting text semantics are used for text modality quantization. and image semantics for image modality quantization .
7. The semantic enhancement, collaborative alignment, and fusion representation method based on multimodal data according to claim 6, characterized in that: In method S3, text semantics are... pass The layer text-specific codebook is quantized into a codeword sequence. Layer of a specific codebook The corresponding text modal codebook is obtained. The residual quantization expression for modality-specific residual quantization is as follows: ,in Indicates the selection of the first code from a specific codebook of text. Level code words, It is the first Level of text semantic residuals For selected codewords; text-specific code sequences .
8. The semantic enhancement collaborative alignment and fusion representation method based on multimodal data according to claim 7, characterized in that: In method S3, image semantics are analyzed. pass The layer image-specific codebook is quantized into a codeword sequence, and the residual quantization processing expression for mode-specific residual quantization is as follows: ,in Indicates the selection of the first image from a specific codebook. Level code words, It is the first Level-1 image semantic residuals For selected codewords; image-specific code sequences .
9. The semantic enhancement collaborative alignment and fusion representation method based on multimodal data according to claim 8, characterized in that: In quantified text embedding Reconstructing the original input features and encoding them yields... In quantized image embedding Reconstructing the original input features and encoding them yields... The residual quantization processing of methods S2 and S3 constructs the following loss function: ,in This indicates that the gradient stopping operation is in progress. To quantify the loss in the residuals, The loss is processed by quantizing the text modal residuals. For image modal residual quantization loss, This represents a coefficient that balances the strength of code embedding and encoder optimization.
10. The semantic enhancement collaborative alignment and fusion representation method based on multimodal data according to claim 9, characterized in that: The total loss function expressions for methods S2 to S5 are as follows: ,in Indicates the total loss. This represents the reconstruction loss during the reconstruction process, where text and quantized image embeddings are used. This represents minimizing the loss between the residual vector and its corresponding codebook embedding.