A multimodal content representation and recommendation method, system, device, medium and product
Patent Information
- Application Number
- CN202611011578.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
然而,现有技术往往将语义表示生成与图结构优化割裂处理:部分技术仅侧重于对交互图结构进行扩散降噪,缺乏对多模态深层语义的显式建模;另一些技术利用扩散模型生成蕴含高阶语义的物品表示,却局限于潜在向量空间,难以直接修复交互图中显式存在的拓扑结构噪声
本申请提供了一种多模态内容表征及推荐方法、系统、设备、介质及产品,通过将统一特征向量作为先验信息,在离散空间引入无分类器引导机制,基于多模态语义锚点执行图扩散拓扑生成得到候选协同图谱的显示引导方式,能够在无连边或交互极少的长尾物品间建立有效的潜在拓扑关联,提高本申请在冷启动和极度稀疏环境下的特征连通性与召回率,进而能够实现多模态语义与图谱拓扑结构的紧密结合,有效缓解数据稀疏场景下的推荐冷启动问题,解决现有技术中针对语义与结构割裂的问题。通过在离散空间引入无分类器引导机制,基于多模态语义锚点执行图扩散拓扑生成,得到候选协同图谱,能够在拓扑构建阶段主动剔除热门趋势的干扰,进而能够有效增强目标用户的个性化偏好信号,提升对长尾特征的捕捉能力及最终推荐列表的精准排序性能。并且,针对多模态非对称噪声问题,本申请对候选协同图谱进行模态自适应的拓扑组装与修剪,构建纯净协同图谱骨架,并在纯净协同图谱骨架上进行信息传播与协同过滤,能够在保留高价值稀疏信号(如精准文本特征)的同时,有效滤除高噪模态(如热门背景音频)引入的假阳性连边,进而使得在处理结构噪声复杂或数据极度稀疏的业务场景时,均能保持稳定的推荐性能。
Smart Images

Figure CN122818239A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal information processing and personalized content recommendation technology, and in particular to a multimodal content representation and recommendation method, system, device, medium and product. Background Technology
[0002] As user needs become increasingly complex, multimodal recommender systems (MRS) are widely used in e-commerce, short videos, and other scenarios. Traditional multimodal recommender methods based on graph neural networks (GNNs) heavily rely on the quality of the user-item interaction graph. However, in practical applications, interaction graphs generally face the dual challenges of structural noise (such as false positive edges caused by popularity bias) and data sparsity (such as a lack of long-tail item interactions), severely compromising the accuracy and robustness of recommendations.
[0003] In recent years, diffusion probability models have been introduced into recommender systems, enabling a shift from a discriminative to a generative paradigm. However, existing technologies often treat semantic representation generation and graph structure optimization separately: some techniques focus only on diffusion denoising of the interaction graph structure, lacking explicit modeling of multimodal deep semantics; others utilize diffusion models to generate item representations containing high-order semantics, but are limited to the latent vector space, making it difficult to directly repair the explicit topological noise in the interaction graph.
[0004] In summary, existing technologies rarely utilize high-quality multimodal semantic representations as prior knowledge to explicitly guide the denoising and reconstruction of graph topology. Therefore, constructing a cascading mechanism that leverages strong semantic consistency features to constrain graph diffusion generation, achieving a closed loop where semantics guides structure and structure enhances semantics, while ensuring recommendation accuracy and improving recommendation performance stability, has become a pressing technical challenge in this field. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this application provides a method, system, device, medium, and product for multimodal content representation and recommendation.
[0006] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a multimodal content representation and recommendation method, including: Acquire multimodal interaction data, perform feature fusion and noise reduction in a continuous latent space, generate a unified feature vector, and use the unified feature vector as a multimodal semantic anchor point; A classifier-free guidance mechanism is introduced in the discrete space, and graph diffusion topology generation is performed based on the multimodal semantic anchors to obtain candidate cooperative graphs; The candidate cooperative graphs are subjected to modality-adaptive topology assembly and pruning to construct a clean cooperative graph skeleton; Information propagation and collaborative filtering are performed on the pure collaborative graph skeleton to generate a target recommendation list.
[0007] Secondly, this application provides a multimodal content representation and recommendation system, including: The multimodal latent semantic fusion module is used to acquire multimodal interaction data, perform feature fusion and noise reduction in a continuous latent space, generate a unified feature vector, and use the unified feature vector as a multimodal semantic anchor point. The intent-guided graph topology generation module is used to introduce a classifier-free guidance mechanism in the discrete space, and perform graph diffusion topology generation based on the multimodal semantic anchors to obtain candidate cooperative graphs; The modality-adaptive graph denoising module is used to perform modality-adaptive topology assembly and pruning on the candidate cooperative graphs to construct a clean cooperative graph skeleton. The dual-view collaborative aggregation and recommendation module is used to perform information dissemination and collaborative filtering on the pure collaborative graph skeleton to generate a target recommendation list.
[0008] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multimodal content representation and recommendation method provided above.
[0009] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal content representation and recommendation method provided above.
[0010] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multimodal content representation and recommendation method provided above.
[0011] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a multimodal content representation and recommendation method, system, device, medium, and product. By using a unified feature vector as prior information, a classifier-free guidance mechanism is introduced in the discrete space. Based on multimodal semantic anchors, a graph diffusion topology generation method is used to generate candidate collaborative graphs. This method can establish effective potential topological associations between long-tail items with no edges or minimal interactions, improving feature connectivity and recall in cold-start and extremely sparse environments. Furthermore, it achieves a close integration of multimodal semantics and graph topology, effectively alleviating the cold-start problem in data-sparse scenarios and solving the problem of semantic and structural separation in existing technologies. By introducing a classifier-free guidance mechanism in the discrete space and generating candidate collaborative graphs based on multimodal semantic anchors, interference from popular trends can be actively eliminated during the topology construction stage. This effectively enhances the personalized preference signals of target users, improves the ability to capture long-tail features, and enhances the accuracy of the final recommendation list's ranking performance. Furthermore, to address the multimodal asymmetric noise problem, this application performs modality-adaptive topology assembly and pruning on candidate collaborative graphs to construct a clean collaborative graph skeleton. Information propagation and collaborative filtering are then performed on the clean collaborative graph skeleton. This effectively filters out false positive connections introduced by high-noise modalities (such as popular background audio) while preserving high-value sparse signals (such as precise text features). As a result, stable recommendation performance can be maintained even when dealing with business scenarios with complex structural noise or extremely sparse data. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a multimodal content representation and recommendation method provided in an embodiment of this application; Figure 2 A schematic diagram of the functional modules of a multimodal content representation and recommendation system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] In one exemplary embodiment, this application provides a multimodal content representation and recommendation method. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is described using an application to a server as an example. Figure 1 As shown, the method includes: Step 100: Obtain multimodal interaction data, perform feature fusion and noise reduction in the continuous latent space, generate a unified feature vector, and use the unified feature vector as a multimodal semantic anchor point.
[0017] Step 101: Introduce a classifier-free guidance mechanism in the discrete space, perform graph diffusion topology generation based on multimodal semantic anchors, and obtain candidate collaborative graphs.
[0018] Step 102: Perform modality-adaptive topology assembly and pruning on the candidate co-cographies to construct a clean co-cograph skeleton.
[0019] Step 103: Perform information propagation and collaborative filtering on the clean collaborative graph skeleton to generate a target recommendation list.
[0020] By implementing steps 100-103 above, this application constructs a cascading mechanism that uses strong semantic consistency features to constrain graph diffusion generation, achieving a closed loop where semantics guides structure and structure enhances semantics, thereby improving the stability of recommendation performance while ensuring recommendation accuracy.
[0021] In an exemplary embodiment of this application, in order to achieve a tight integration of multimodal semantics and graph topology and effectively alleviate the cold start problem of recommendation in data-sparse scenarios, the implementation process of step 100 above may include: Step 100-1: Construct a continuous spatial diffusion network based on the U-Net architecture.
[0022] Step 100-2: A continuous spatial diffusion network is used to perform nonlinear fusion and noise reduction on the multimodal interaction data to obtain an aligned unified feature vector. The multimodal interaction data includes the target user's historical sequence of interactive items and the multimodal features of the interactive items. The multimodal features of the interactive items include visual features, auditory features, and textual features.
[0023] Based on the descriptions of steps 100-1 and 100-2 above, in practical applications, the historical sequence of interactive items by the target user within a preset time window, as well as the multimodal features (visual features, auditory features, and textual features, etc.) constituting the interactive items, can be obtained first. Then, the heterogeneous multimodal interaction data is input into a pre-constructed continuous spatial diffusion network based on the U-Net architecture. During the continuous Markov chain process of forward noise addition and reverse denoising in this network, the heterogeneous multimodal interaction data undergoes nonlinear spatial alignment and feature fusion in the latent space. The aligned unified feature vector output by the continuous spatial diffusion network can be configured as the basic condition input for the subsequent discrete map reconstruction process, i.e., the multimodal semantic anchor point. .
[0024] Based on the above description, to address the problem of semantic and structural separation, this application adopts a cascaded architecture of first fusing continuous spatial features and then reconstructing discrete spatial topology, utilizing multimodal joint representations (i.e., multimodal semantic anchors) extracted by a continuous spatial diffusion network. As prior information, it can explicitly guide the generation process of discrete spatial graph diffusion denoising network, and can establish effective potential topological associations between long-tailed items with no edges or very few interactions, thereby improving the feature connectivity and recall of discrete spatial graph diffusion denoising network in cold start and extremely sparse environments.
[0025] In an exemplary embodiment of this application, in order to reduce popularity bias and improve the personalized recommendation accuracy of long-tail content, this application introduces a classifier-free guidance (CFG) mechanism in step 101 above. Based on this, the implementation process of step 101 above may include: Step 101-1: Obtain the target user's historical interaction intent features .
[0026] Step 101-2: Construct a discrete spatial graph diffusion denoising network. In terms of network structure, the discrete spatial graph diffusion denoising network adopts a multi-layer perceptron (MLP) architecture, and is internally decoupled into a feature mixing backbone module, a time-step encoding branch, and a user intent mapping branch.
[0027] Step 101-3: Using multimodal semantic anchors Historical interaction intent characteristics of target users Using the current diffusion time step t as the conditional input, a discrete spatial graph diffusion denoising network is employed to perform both unconditional and conditional predictions, yielding a first probability distribution and a second probability distribution. Specifically, during the model inference phase, the discrete spatial graph diffusion denoising network dynamically controls the signal flow of the user intent mapping branch, introducing a classifier-free guidance mechanism to perform two forward predictions (both unconditional and conditional) for the same target graph, including: (1) Unconditional prediction: During the forward propagation of the discrete spatial graph diffusion denoising network, a physical masking operation is performed on the user intent mapping branch (e.g., setting it to an empty condition or zero vector) to forcibly mask the historical interaction intent features of the target user. This allows the feature fusion backbone module to perform calculations only based on multimodal semantic anchors, thereby predicting the basic graph edge probabilities that do not incorporate personalized preferences and reflect global popularity trends, and obtaining the first probability distribution. .
[0028] (2) Conditional prediction: During the forward propagation of the discrete spatial graph diffusion denoising network, the user intent mapping branch is activated to include the historical interaction intent features of the target user. With multimodal semantic anchors Joint computation is performed in the feature fusion backbone module to predict the edge probabilities of the target graph containing user personalized preferences, thereby obtaining the second probability distribution. .
[0029] Step 101-4: Based on the preset preference amplification factor ω (preferred interval [1.2, 2.0]), use the extrapolation formula to perform difference amplification calculation on the first probability distribution and the second probability distribution to generate the final candidate edge probability distribution. The final candidate edge probability distribution. The calculation formula is: .
[0030] Step 101-5: Based on the final candidate edge probability distribution The reconstruction generates a candidate collaborative graph containing implicit long-tail associations between target users and items.
[0031] Furthermore, during the model training phase of the discrete spatial graph diffusion denoising network (referred to as the model), a conditional masking mechanism with a preset probability (e.g., 0.1) can be introduced to set the user interaction intent features in some training batches to zero, so that the model can jointly learn the conditional topological distribution (i.e., the second probability distribution) and the unconditional topological distribution (i.e., the first probability distribution).
[0032] Based on the above description, to address the issue of susceptibility to popularity bias, this application introduces a CFG mechanism in the graph diffusion inference stage. By separately performing basic graph probability prediction without incorporating personalized preferences and target graph probability prediction with incorporating personalized preferences, and performing extrapolation calculations in the mathematical space, the interference of popular trends can be actively eliminated in the topology construction stage. This can effectively enhance the personalized preference signals of target users, improve the model's ability to capture long-tail features, and improve the accuracy of the final recommendation list's ranking performance.
[0033] In an exemplary embodiment of this application, to effectively isolate pseudo-association noise from heterogeneous modalities and improve the data robustness of the model, this application proposes a modality-adaptive graph topology pruning strategy. Based on this, the implementation process of step 102 above can be described as follows: Based on the signal-to-noise ratio differences exhibited by candidate cooperative graphs under different modal features, configure mutually independent modality-adaptive pruning strategies, perform modality-adaptive topology assembly and pruning, and obtain a clean cooperative graph skeleton. In practical applications, the implementation of this process may include: (1) The candidate edge probability distribution (i.e. the final candidate edge probability distribution) is obtained by extrapolating and correcting the output of the linear prediction head at the end of the discrete spatial graph diffusion denoising network after being processed by the classifier-free guidance mechanism (CFG). The confidence score of the graph diffusion edges is determined. The Top-K truncation algorithm is then applied to extract the probability distribution of candidate edges. The system uses the top K edges with the highest confidence scores for each node as the initial skeleton edges to achieve initial screening and filtering of low-confidence noisy edges.
[0034] (2) The extracted initial skeleton edges are mapped and aligned to the vector space of the pre-extracted visual, auditory, and textual multimodal features, thereby constructing a parallel and independent topological structure for each modality (such as visual modal subgraph, auditory modal subgraph and text modal subgraph).
[0035] (3) Set retention rate thresholds for each modality's parallel and independent topology (e.g., retention rate of visual subgraph, auditory subgraph, and text subgraph). Based on the set retention rate thresholds, perform asymmetric random edge dropping operations on the corresponding modality subgraphs. For example, for each skeleton connection in the corresponding modality subgraph, independently generate a random scalar that follows a uniform distribution of [0,1]. If the random scalar is greater than the retention rate threshold corresponding to the modality, the connection is filtered out; otherwise, it is retained. This operation filters out pseudo-associative noise connections specific to each modality, and finally assembles a clean cooperative graph skeleton for downstream graph neural network propagation.
[0036] For example, in a specific application scenario, if the current dataset belongs to a multimedia user-generated content (UGC) scenario, given that auditory modalities are prone to generating a large number of false positive edges due to the same background music, this application sets the retention rate threshold of the auditory modal subgraph to a first interval (e.g., 0.4 to 0.6). Given that text modal features are relatively sparse and accurate, the retention rate threshold of the text modal subgraph is set to a second interval (e.g., 0.90 to 0.98). Based on the above dynamically configured independent retention rate thresholds, random edge dropping operations are performed on the corresponding modal subgraphs to filter out pseudo-associative edges with probabilities lower than the threshold, thereby assembling to form the final pure basic collaborative graph skeleton (i.e., the pure collaborative graph skeleton).
[0037] Based on the above description, this application differs from existing global unified graph enhancement pruning strategies in addressing the multimodal asymmetric noise problem. It configures independent retention rate thresholds for different modal subgraphs during the graph assembly stage. This effectively filters out false positive connections introduced by high-noise modalities (such as popular background audio) while retaining high-value sparse signals (such as accurate text features). This ensures the purity of information input to the backbone graph network, enabling the model to maintain stable recommendation performance when dealing with business scenarios with complex structural noise or extremely sparse data.
[0038] In one exemplary embodiment of this application, in order to further improve the accuracy of recommendations, the implementation process of step 103 above includes: Step 103-1: Employ a graph convolutional network (GCN) to perform high-order information transfer on the clean collaborative graph skeleton, obtaining aggregated multimodal high-order representations and user features. For example, along the reconstructed topological connections, perform multi-layer (preferably 2 to 3 layers) message passing and node representation aggregation.
[0039] Step 103-2: Using a global static weight fusion strategy, the aggregated multimodal high-order representations are combined with user features to calculate the inner product, thereby obtaining the target user's preference prediction score for each candidate item.
[0040] Step 103-3: Sort and filter the preferences based on the predicted scores in descending order, and output the final personalized recommendation list to the target user.
[0041] Based on the above description, and in view of the defects and shortcomings of the existing technology, this application aims to solve the following three major technical problems: 1. This application addresses the problem of cold-start difficulties in sparse scenarios caused by the disconnect between semantic representation generation and graph structure optimization. Existing recommendation techniques based on diffusion models struggle to correct explicit topological noise in semantic representation, while graph structure optimization lacks sufficient depth for multimodal semantic mining. This application aims to eliminate the technological gap between semantic modeling and structural optimization, overcoming the weakness in feature connectivity caused by insufficient interaction in long-tail items under sparse data environments.
[0042] 2. Addressing the issue of models being susceptible to global popularity bias, leading to low accuracy in personalized recommendations for long-tail content. Existing multimodal recommendation methods based on graph neural networks heavily rely on the initial interaction graph, making them highly susceptible to false positive connections (structural noise) due to the influence of global popularity trends within the project. This application aims to overcome the problem of models overly favoring popular content, improving the ability to characterize fine-grained user preferences and ensuring fairness in the distribution of long-tail content.
[0043] 3. Addressing the issue of insufficient robustness caused by asymmetric noise interference in multimodal features when using globally uniform denoising. Existing methods typically employ a single global denoising strategy (such as uniformly dropping edges) for graph networks, ignoring the natural signal-to-noise ratio differences between different modalities (such as noisy audio and sparse, precise text), and failing to effectively isolate asymmetric interference. This application aims to achieve precise noise reduction intervention for different modalities, improving data robustness in complex business scenarios.
[0044] Based on the same inventive concept, this application also provides a multimodal content representation and recommendation system for implementing the multimodal content representation and recommendation method described above. The solution provided by this system is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more embodiments of the multimodal content representation and recommendation system provided below can be found in the limitations of the multimodal content representation and recommendation method described above, and will not be repeated here.
[0045] In one exemplary embodiment, such as Figure 2 As shown, a multimodal content representation and recommendation system is provided, including: a multimodal latent semantic fusion module, an intent-guided graph topology generation module, a modality-adaptive graph denoising module, and a dual-view collaborative aggregation and recommendation module.
[0046] The multimodal latent semantic fusion module is configured to acquire multimodal interaction data, perform feature fusion and noise reduction in a continuous latent space, generate a unified feature vector (i.e., perform multimodal joint representation), and use the unified feature vector as a multimodal semantic anchor point.
[0047] The intent-guided graph topology generation module is configured to introduce a classifier-free guidance mechanism in the discrete space, perform graph diffusion topology generation based on multimodal semantic anchors, and extract popularity trends through extrapolation formulas during the inference phase to obtain candidate collaborative graphs.
[0048] The modality-adaptive graph denoising module is configured to set independent topology retention rate thresholds based on the signal-to-noise ratio differences of heterogeneous modes, perform modality-adaptive topology assembly and pruning on candidate co-cographies, and construct a clean co-cograph skeleton (i.e., a basic co-cograph skeleton).
[0049] The dual-view collaborative aggregation and recommendation module is configured to perform graph convolution information propagation and collaborative filtering on a clean collaborative graph skeleton, calculate preference scores, and generate a target recommendation list.
[0050] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores multimodal content representation and recommendation data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a multimodal content representation and recommendation method.
[0051] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment to which the present application is applied. Specific computer equipment may include, for example, [the following is a list of possible additional structures]. Figure 3 The diagram shows more or fewer components, or combinations of certain components, or different component arrangements.
[0052] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0053] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0054] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0055] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0056] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (RRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0057] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0058] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0059] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A multimodal content representation and recommendation method, characterized in that, include: Multimodal interaction data is acquired, and feature fusion and noise reduction are performed in a continuous latent space to generate a unified feature vector, which is then used as a multimodal semantic anchor point. A classifier-free guidance mechanism is introduced in the discrete space, and graph diffusion topology generation is performed based on the multimodal semantic anchors to obtain candidate cooperative graphs; The candidate cooperative graphs are subjected to modality-adaptive topology assembly and pruning to construct a clean cooperative graph skeleton; Information propagation and collaborative filtering are performed on the pure collaborative graph skeleton to generate a target recommendation list.
2. The multimodal content representation and recommendation method according to claim 1, characterized in that, Acquire multimodal interaction data, perform feature fusion and noise reduction in a continuous latent space, and generate a unified feature vector, including: Constructing a continuous spatial diffusion network based on the U-Net architecture; The continuous spatial diffusion network is used to perform nonlinear fusion and noise reduction on the multimodal interaction data to obtain the aligned unified feature vector; the multimodal interaction data includes the target user's historical interaction item sequence and the multimodal features of the interaction items; the multimodal features of the interaction items include visual features, auditory features and text features.
3. The multimodal content representation and recommendation method according to claim 1, characterized in that, A classifier-free guidance mechanism is introduced in the discrete space, and graph diffusion topology generation is performed based on the multimodal semantic anchors to obtain candidate cooperative graphs, including: Obtain the historical interaction intent characteristics of the target user; The multilayer perceptron architecture is decoupled into a feature mixing backbone module, a time step encoding branch, and a user intent mapping branch, and a classifier-free guidance mechanism is introduced to construct a discrete spatial graph diffusion denoising network. Using the multimodal semantic anchor points, the target user's historical interaction intent features, and the current diffusion time step as conditional inputs, the discrete spatial graph diffusion denoising network is used to perform unconditional and conditional dual predictions to obtain the first probability distribution and the second probability distribution. Based on the preset preference amplification factor, the difference amplification calculation is performed on the first probability distribution and the second probability distribution using the extrapolation formula to generate the final candidate edge probability distribution. Based on the final candidate edge probability distribution, the candidate collaborative graph containing implicit long-tail associations between target users and items is reconstructed.
4. The multimodal content representation and recommendation method according to claim 3, characterized in that, Using the multimodal semantic anchor points, the target user's historical interaction intent features, and the current diffusion time step as conditional inputs, the discrete spatial graph diffusion denoising network performs unconditional and conditional dual predictions to obtain a first probability distribution and a second probability distribution, including: Using the multimodal semantic anchor points, the historical interaction intent features of the target user, and the current diffusion time step as conditional inputs, the signal flow of the user intent mapping branch is dynamically controlled, and a classifier-free guidance mechanism is introduced to perform unconditional and conditional dual prediction. The unconditional prediction includes: during the forward propagation of the discrete spatial graph diffusion denoising network, performing a physical masking operation on the user intent mapping branch to forcibly mask the historical interaction intent features of the target user, so that the feature mixing backbone module only calculates based on the multimodal semantic anchor point to predict the basic graph connection probability that does not incorporate personalized preferences and reflects the global popularity trend, and obtains the first probability distribution. Conditional prediction includes: during the forward propagation of the discrete spatial graph diffusion denoising network, activating the user intent mapping branch, jointly calculating the historical interaction intent features of the target user and the multimodal semantic anchor points in the feature mixing backbone module, predicting the target graph connection probability containing the user's personalized preferences, and obtaining the second probability distribution.
5. The multimodal content representation and recommendation method according to claim 1, characterized in that, The candidate cooperative graphs are subjected to modality-adaptive topology assembly and pruning to construct a clean cooperative graph skeleton, including: Based on the signal-to-noise ratio differences exhibited by the candidate co-location graph under different modal features, mutually independent modal adaptive pruning strategies are configured to perform modal adaptive topology assembly and pruning, thereby obtaining the pure co-location graph skeleton.
6. The multimodal content representation and recommendation method according to claim 1, characterized in that, Information propagation and collaborative filtering are performed on the pure collaborative graph skeleton to generate a target recommendation list, including: A backbone graph convolutional network is used to perform high-order information transfer on the pure collaborative graph skeleton to obtain the aggregated multimodal high-order representation and user features. A global static weight fusion strategy is adopted to perform inner product calculation between the aggregated multimodal high-order representation and the user features to obtain the target user's preference prediction score for each candidate item. The target recommendation list is generated by sorting and filtering the preferences based on the predicted scores in descending order.
7. A multimodal content representation and recommendation system, characterized in that, include: The multimodal latent semantic fusion module is used to acquire multimodal interaction data, perform feature fusion and noise reduction in a continuous latent space, generate a unified feature vector, and use the unified feature vector as a multimodal semantic anchor point. The intent-guided graph topology generation module is used to introduce a classifier-free guidance mechanism in the discrete space, and perform graph diffusion topology generation based on the multimodal semantic anchors to obtain candidate cooperative graphs; The modality-adaptive graph denoising module is used to perform modality-adaptive topology assembly and pruning on the candidate cooperative graphs to construct a clean cooperative graph skeleton. The dual-view collaborative aggregation and recommendation module is used to perform information dissemination and collaborative filtering on the pure collaborative graph skeleton to generate a target recommendation list.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the multimodal content representation and recommendation method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the multimodal content representation and recommendation method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the multimodal content representation and recommendation method as described in any one of claims 1-6.