Multi-modal recommendation method and system based on flow matching and causal perception negative sampling

By constructing a multimodal graph and combining conditional flow matching and causal awareness to remove biased weights, high-quality hard negative samples are generated, which solves the problems of negative sample quality and training stability in multimodal recommendation and improves the accuracy and robustness of recommendations.

CN121834069APending Publication Date: 2026-04-10CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing multimodal recommendation technologies struggle to balance negative sample quality, training stability, and preference authenticity, leading to model learning deviating from real user preferences and reducing recommendation reliability.

Method used

By constructing a multimodal graph, a conditional flow matching strategy is adopted to generate semantically enhanced item representations. Combined with causal perception bias removal weights, semantically similar difficult negative samples are selected to optimize the embedding representations of users and items to train the recommendation model.

Benefits of technology

It significantly improves the accuracy and robustness of personalized recommendations, alleviates data sparsity and popularity bias, and enhances the diversity and stability of recommendation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834069A_ABST
    Figure CN121834069A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal recommendation method and system based on flow matching and causal perception negative sampling, and relates to the technical field of intelligent recommendation. Based on ID embedding and multi-modal features of the article, constructing a multi-modal graph and obtaining multi-modal enhanced article representation; with the multi-modal features as conditions, combining prior distribution reflecting article popularity, and generating semantic enhanced flow enhanced article representation along a deterministic trajectory through a conditional flow matching strategy; based on a causal perception thought, constructing a depolarization weight by using an article frequency, performing embedded adaptive fusion on the stream enhanced article representation and the original ID to obtain a depolarized article representation, and selecting difficult negative samples with similar semantics from non-interactive articles based on the representation; according to the method, stream enhanced article representation and multi-modal enhanced article representation are fused, a difficult negative sample is combined, and final embedded representation of a user and an article is optimized through graph propagation, so that a recommendation model is trained to generate a personalized recommendation list, and the accuracy and robustness of personalized recommendation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent recommendation technology, and in particular to a multimodal recommendation method and system based on stream matching and causal perception negative sampling. Background Technology

[0002] Multimodal recommendation aims to integrate visual, textual, and other multimodal semantic information of items to accurately characterize user preferences and improve recommendation performance. High-quality negative sampling is a key prerequisite for enhancing the model's discriminative ability. Exposure bias refers to the phenomenon in recommendation systems where frequently exposed items are easily misidentified by the model as the user's true preferences. The resulting spurious relevance is a false association between the frequency of item exposure and the probability of user selection, which can cause the model to learn from a deviation from true preferences and reduce the reliability of recommendations.

[0003] Current performance improvements in multimodal recommendation heavily rely on the effectiveness of negative sampling strategies, necessitating the capture of negative samples that reveal fine-grained semantic differences across modalities to enhance model learning. However, existing technologies struggle to balance negative sample quality, training stability, and preference authenticity, creating significant technical bottlenecks that hinder the practical application of multimodal recommendation.

[0004] Traditional heuristic negative sampling methods rely solely on interactive statistical sampling and do not utilize multimodal content features, thus failing to capture cross-modal semantic differences. While recent diffusion model-based methods can generate modality-aware hard negative samples, they depend on curve trajectory generation, leading to training instability due to noise accumulation. Furthermore, they do not consider exposure bias and cannot eliminate spurious correlations, making it difficult to generate high-quality negative samples to support accurate recommendations. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a multimodal recommendation method and system based on flow matching and causal-aware negative sampling. Through the synergistic optimization of multimodal graph construction, conditional flow matching enhancement, and causal-aware debiased negative sampling, the method effectively improves the recommendation model's ability to capture item semantics and the quality of negative samples. While alleviating data sparsity and popularity bias, it significantly enhances the accuracy and robustness of personalized recommendations.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a multimodal recommendation method based on flow matching and causal-aware negative sampling, comprising: Based on the ID embedding and multimodal features of items, a multimodal graph that integrates visual and textual semantic relationships is constructed, and a multimodal enhanced item representation is obtained; Using multimodal features as conditions and combining prior distributions that reflect the popularity of items, semantically enhanced flow-enhanced item representations are generated along deterministic trajectories through a conditional flow matching strategy. Based on the idea of ​​causal perception, the bias removal weight is constructed using item frequency. The stream-enhanced item representation is adaptively fused with the original ID embedding to obtain the bias removal item representation. Based on the representation, semantically similar difficult negative samples are selected from non-interactive items. By fusing the stream-enhanced item representation with the multimodal-enhanced item representation and incorporating the hard negative samples, the final embedding representation of users and items is optimized through graph propagation to train the recommendation model and generate a personalized recommendation list.

[0007] Secondly, the present invention provides a multimodal recommendation system based on flow matching and causal-aware negative sampling, comprising: The multimodal graph construction module is used to construct a multimodal graph that integrates visual and textual semantic relationships based on the item's ID embedding and multimodal features, and obtain a multimodal enhanced item representation; The flow enhancement generation module is used to generate semantically enhanced flow-enhanced item representations along a deterministic trajectory by taking multimodal features as conditions and combining them with prior distributions that reflect the popularity of items through a conditional flow matching strategy. The causal negative sampling module is used to construct debiased weights based on the idea of ​​causal perception using item frequency, adaptively fuse the stream-enhanced item representation with the original ID embedding to obtain the debiased item representation, and select semantically similar difficult negative samples from non-interactive items based on the representation. The joint optimization recommendation module is used to fuse the flow-enhanced item representation and the multimodal-enhanced item representation, and combine the hard negative samples to optimize the final embedding representation of users and items through graph propagation, so as to train the recommendation model to generate a personalized recommendation list.

[0008] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal recommendation method based on stream matching and causal-aware negative sampling described in the first aspect.

[0009] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the multimodal recommendation method based on stream matching and causal-aware negative sampling described in the first aspect.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention addresses the issue of insufficient semantics in single-ID embedding by constructing a multimodal graph to integrate visual and textual semantic associations in item representation. Secondly, a conditional flow matching strategy, combined with prior item popularity, generates enhanced flow representations along deterministic trajectories, ensuring generation stability while enriching the semantic dimensions of items. Thirdly, causal-aware negative sampling, through popularity-based weighting and adaptive fusion, filters semantically similar but difficult negative samples, avoiding the inefficiency of training signals caused by random negative sampling and mitigating popularity bias. Finally, graph propagation further optimizes user and item embeddings, enabling the recommendation model to more accurately capture genuine preferences. The overall solution effectively addresses issues such as insufficient multimodal semantic fusion, low-quality negative samples, data sparsity, and popularity bias, significantly improving the accuracy, diversity, and robustness of personalized recommendations.

[0011] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0012] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.

[0013] Figure 1 This is a flowchart illustrating the main steps of a multimodal recommendation method based on flow matching and causal-aware negative sampling, as provided in an embodiment of the present invention. Figure 2 A flowchart illustrating a multimodal recommendation method based on flow matching and causal-aware negative sampling provided in an embodiment of the present invention; Figure 3 The following is an overall framework diagram of FMCNS provided in the embodiments of the present invention; wherein, (a) indicates that compared with the existing DDPM, flow matching makes the training process stable and has high sampling efficiency, and it uses a straight path instead of a curved diffusion path; (b) indicates that negative sample sampling for causal perception is achieved through a frequency-based debiasing method. Detailed Implementation

[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0015] With the surge in online multimedia content, multimodal recommendation has become the mainstream paradigm. Unlike traditional collaborative filtering, which relies solely on user-item interactions, this paradigm integrates multimodal semantic information such as visual and textual data of items with collaborative signals. This allows for a more comprehensive characterization of user preferences and item features, providing effective solutions to core issues such as cold start and data sparsity. However, the realization of these advantages is highly dependent on high-quality training signals. In sparse implicit feedback scenarios, negative sampling is a crucial step in constructing discriminative user-item representations, and its quality directly determines recommendation performance.

[0016] Exposure bias is a major interfering factor in the negative sampling process. It refers to the tendency for frequently exposed items to be misjudged by the model as genuine user preferences. For example, if an e-commerce platform repeatedly pushes a popular brand of hoodie, users may click on it multiple times due to frequent exposure, leading the model to mistakenly identify the hoodie as a user preference. This results in a false correlation between item exposure frequency and user selection probability, causing the model to learn from the true preferences and severely reducing recommendation reliability. Current negative sampling techniques struggle to balance the semantic quality of negative samples, training stability, and preference authenticity, creating a technical bottleneck. Traditional heuristic methods, in particular, rely solely on interactive statistical sampling, failing to utilize multimodal content features and thus unable to capture fine-grained semantic differences across modalities. Furthermore, they are prone to introducing popularity bias, such as selecting negative samples based solely on historical sales figures, excluding niche styles that align with the user's taste.

[0017] While recent methods based on diffusion models (DDPM) can generate modality-aware hard negative samples, they rely on curve trajectories with random noise injection, resulting in training instability due to noise accumulation. Furthermore, they lack explicit mechanisms to address exposure bias and cannot eliminate spurious correlations, making it difficult to generate high-quality negative samples.

[0018] In summary, current multimodal recommendation negative sampling faces two major challenges: the instability of multimodal semantic generation modeling and causal bias in sample selection. Stream matching technology, based on the deterministic linear trajectory of optimal transport theory, offers a potential direction for solving these problems, but its application in multimodal recommendation remains underexplored, and the collaborative challenges of cross-modal fusion and causal bias correction need to be overcome. Therefore, this invention proposes a multimodal recommendation method, system, medium, and device based on stream matching and causal-aware negative sampling. Through the synergistic effect of the conditional stream matching module and the causal-aware negative sampling module, semantically coherent and robust to exposure bias difficult negative samples are generated, overcoming the limitations of existing technologies.

[0019] Example 1 like Figure 1 As shown, this embodiment discloses a multimodal recommendation method based on flow matching and causal-aware negative sampling, including the following steps: S1: Based on the item ID embedding and multimodal features, construct a multimodal graph that integrates visual and textual semantic relationships, and obtain a multimodal enhanced item representation; S2: Using multimodal features as conditions and combining prior distributions that reflect the popularity of items, semantically enhanced flow-enhanced item representations are generated along a deterministic trajectory through a conditional flow matching strategy; S3: Based on the idea of ​​causal perception, the bias removal weight is constructed using the item frequency. The stream-enhanced item representation is adaptively fused with the original ID embedding to obtain the bias removal item representation. Based on the representation, semantically similar difficult negative samples are selected from non-interactive items. S4: Integrate the flow-enhanced item representation with the multimodal-enhanced item representation, and combine the hard negative samples to optimize the final embedding representation of users and items through graph propagation, so as to train the recommendation model to generate a personalized recommendation list.

[0020] Next, combined Figure 2 This embodiment provides a detailed description of a multimodal recommendation method based on flow matching and causal perception negative sampling.

[0021] (a) Problem Definition set up and These represent the user and item sets, respectively. Each user... With a learnable ID embedding Associated, embedded via ID As a digital identity carrier for users, the model will continuously optimize based on users' interactions with items. The vector value ultimately makes This indirectly corresponds to the user's personalized preferences.

[0022] And each item i Embedded by ID Together with multimodal content features, it is characterized, including visual representation. and text representation That is, embedding through ID Visual representation and text representation The combination of these elements is used to describe the attributes and content characteristics of the item.

[0023] Observed user-item interactions are encoded in a binary matrix. Among them This indicates that user u has interacted with item i. The multimodal recommendation task aims to learn a rating function. By combining from Collaborative filtering signals and features from multimodal sources The semantic information is used to effectively predict user-item preferences, and finally generate personalized top-k recommendations for each user.

[0024] (II) Overall Framework This embodiment proposes a multimodal recommendation framework based on flow matching with causal-aware negative sampling (FMCNS), which addresses the fundamental challenges of negative sampling and generative modeling in multimodal recommendation through deterministic flow trajectories and explicit bias mitigation. Figure 3 As shown, the proposed FMCNS framework includes two key collaborative modules: 1. with Figure 3 Unlike existing diffusion-based (DDPM) methods shown in (a), which employ curved probability paths with injected random noise and suffer from training instability leading to unfavorable signal-to-noise ratios, the proposed conditional flow matching enhancement generation constructs a semantic graph from heterogeneous visual and textual features. Then, it employs deterministic straight-line trajectory FlowMatching to generate high-quality enhanced embeddings in the latent space. By leveraging rich semantic information from multiple modalities, it learns stable flow dynamics to achieve efficient sampling with fewer steps, while incorporating frequency-based priors to reflect cooperative patterns.

[0025] 2. For example Figure 3 As shown in (b), causal-perceived negative sampling introduces a principled causal intervention method that explicitly breaks the spurious correlation between item frequency and selection probability through frequency-based debiasing. By combining flow-enhanced item representations with causal intervention weights, semantically meaningful hard negative samples are generated, while effectively mitigating exposure bias and ensuring that the model learns invariant user preferences rather than spurious environmental relevance.

[0026] (III) Constructing a multimodal graph Effective integration of heterogeneous multimodal features is key to capturing rich semantic relationships to support generative modeling and negative sampling. This problem is addressed by constructing a graph structure based on the original multimodal feature representations to model the relationships between features.

[0027] To capture semantic relationships between items, for each modality, the k most similar items are selected based on the cosine similarity of the original multimodal embeddings, thus constructing a k-nearest neighbor graph for that modality: ; in, and These represent the original visual and text feature embeddings, respectively.

[0028] For each item i, only the top-k similar items are retained to construct a sparse adjacency matrix. and The two matrices are collectively referred to as These matrices are symmetrically normalized to ,in It is a degree matrix. The joint multimodal adjacency matrix combines two modes: ; in, The relative contributions of visual and textual semantics are controlled.

[0029] To embed the item's ID Integrating multimodal semantic information, in Perform K-level graph propagation from start: ; Generate multimodal augmented representation It aggregates information on semantically similar items across visual and textual dimensions.

[0030] For the conditional flow matching process, the original multimodal features are transformed into a common d-dimensional space through modality-specific linear projection: ; in, and It is a learnable projection matrix. Alignment common features. and Used in the conditional flow matching generation process.

[0031] (iv) Enhanced generation of conditional flow matching In recommender systems, generative models are often used to optimize item features or generate recommendation candidates, such as predicting items that users might like. Traditional Diffusion Modeling (DDPM) is one of the commonly used methods for this type of generative task. Its core idea is to first add random noise to the real item features, and then let the model remove the noise step by step to restore the features. Through this process, the model learns the true distribution of item features and thus generates item representations that match user preferences.

[0032] However, DDPM has obvious limitations in recommendation scenarios: on the one hand, it relies on a large amount of random noise injection, which leads to large gradient fluctuations and unstable model parameter updates during training, especially when the item features have multimodal heterogeneous data, making it difficult to align different types of features; on the other hand, DDPM inference requires multiple steps of denoising, which is slow and not suitable for the real-time requirements of actual recommendation systems.

[0033] To address the aforementioned issues, this embodiment introduces conditional flow matching technology. By constructing deterministic and continuous transformation paths to simulate the data evolution process, more stable model training and more efficient single-step or few-step generation can be achieved, while better integrating multimodal information of user historical behavior and items.

[0034] 1. Basic Principles of Stream Matching FlowMatching is a continuous-time generative model based on deterministic ordinary differential equations (ODEs), providing a principled alternative to stochastic diffusion processes. Its core idea is to define a smooth, continuous "flow" of transformations from a simple initial distribution (source distribution) to a complex target distribution (target distribution). Specifically, given... 3D space Source probability distribution on and target distribution Define stream mapping The sample is smoothly transformed from the source distribution to the target distribution over time t, and its change is determined by the velocity field. Control, satisfy:

[0035] in, Representing time t The time-dependent velocity field at that location. Flow From the sample Switch to ; Represents a sample in the source distribution; As initial conditions, At that time, the stream mapping does not change the initial sample. The sample still belongs to the source distribution. The goal of the learning process is to fit a parameterized velocity field that can drive any sample from... Accurately flowing to .

[0036] Compared to the random and tortuous noisy paths of diffusion models, this embodiment employs a simple and effective straight-line path construction method inspired by optimal transport theory. For any pair of samples sampled from the source and target respectively... We can directly define the intermediate state at time t as the linear interpolation of the two:

[0037] Accordingly, the velocity field driving this linear motion is constant and deterministic: .

[0038] This embodiment employs a straight-line trajectory design, which offers significant advantages in gradient stability compared to traditional curved random paths. This advantage stems from the characteristics of straight-line trajectories: on the one hand, it reduces gradient fluctuations caused by path deviations, decreasing error accumulation during parameter iteration; on the other hand, it allows gradient changes to exhibit predictable linear characteristics, improving the system's adaptability to environmental changes and its response consistency.

[0039] Specifically, let's set and These represent the training objectives for flow matching with a straight path and diffusion models with a curved random trajectory, respectively. For a neural network with a parameterized velocity field... The gradient variance satisfies:

[0040] The equality holds only when the diffusion process degenerates into a deterministic path.

[0041] Proof: For flow matching, the gradient is , where given hour It is deterministic. For diffusion models, ,in This introduces additional randomness. According to the full variance formula:

[0042] The first item captures random noise. The variance of this term disappears in stream matching due to deterministic interpolation. Therefore, It is equal to the second term only, thus establishing an inequality.

[0043] This reduction in gradient variance directly translates into more stable optimization, which is particularly beneficial for multimodal recommendations that require consistent gradient signals for heterogeneous features to achieve effective cross-modal alignment.

[0044] 2. Enhanced Conditional Flow Matching Based on the above principles, this embodiment designs a conditional flow matching enhancement generation method specifically for recommendation systems, which includes the following steps: (1) Frequency-aware source distribution design Traditional methods typically start with simple distributions such as the standard Gaussian distribution. This embodiment, however, designs the source distribution to encode the item's popularity prior. Specifically, for item i, its source distribution samples... Each dimension is sampled based on the interaction frequency of the item: ; in, Indicates user; This indicates the interaction between user u and item i, and its value is either 0 or 1. This represents the historical interaction frequency of item i. The generation process incorporates the collaborative filtering prior knowledge that "popular items are more likely to be generated" from the initial moment.

[0045] (2) Conditional generation path construction In recommendation scenarios, the generation process needs to consider factors such as user history and multimodal information about items. Let the embedded true features of the target item be... (like ), the condition information is (May include user profiles, multimodal features, etc.)

[0046] Constructing samples from frequency sources Embedded to target Straight path: ; The model's task is to learn a conditional velocity field estimation network. This enables it to adapt to intermediate states. and conditions Accurately predict the direction of velocity pointing towards the target, that is Its training objective (conditional flow matching loss) is:

[0047] (3) Cross-modal attention fusion mechanism To effectively integrate heterogeneous modal information such as text and vision as generation conditions, this embodiment uses a velocity field estimation network... The paper employs a cross-modal attention mechanism, which adaptively weighs the importance of different modal information to the current generation task.

[0048] Specifically, the current intermediate state Projection visual characteristics Projected text features and time embedding spliced ​​into a sequence Calculated through self-attention: ; in It has The linear projection of the model. This mechanism enables cross-modal adaptive weighting, emphasizing unique visual features for aesthetic items and prioritizing textual descriptions for specification-oriented products. In other words, attention weights automatically learn modal preferences in different contexts. For example, for products where appearance is important, such as clothing, the model may focus more on visual features; for products where functional description is important, such as electronic products, it may focus more on textual features.

[0049] Through a cross-modal attention mechanism, based on multimodal context at t=1 Generative Stream Enhanced Embedding It can maintain the semantic similarity structure with the original multimodal features:

[0050] in, The attention weights represent the visual modality. The attention weights represent the text modality. It is to satisfy Attention-derived weights.

[0051] That is, items that are semantically similar in the original feature space also remain similar in their flow-enhanced embedding space, which provides a foundation for the subsequent construction of high-quality hard negative samples.

[0052] (v) Causal perception negative sampling 1. Causal bias-removing framework Drawing on the causal inference theory in the recommendations, this embodiment establishes the theoretical basis for solving exposure deviations in multimodal systems.

[0053] First, a confounding framework is proposed. In observed data, spurious correlations arise when confounding variables simultaneously influence treatment assignment and outcome. Consider a causal system comprising observable variables Y (observed interactions), latent true signals Z (true preferences), and confounding factors S (e.g., item frequencies). Causal relationships form a structural causal model: , S creates a backdoor path that obscures the true causal relationship between Z and Y.

[0054] Frequency debiasing principle. To eliminate spurious correlations caused by confounding factors, causal intervention theory manipulates the causal graph by intervening in weights. When the confounding factor S represents the frequency of an item, it is considered to have frequency. Item i is used to construct intervention weights:

[0055] in It is the sigmoid function. >0 controls intervention intensity, This represents the mean frequency of all items. The standard deviation of the frequency of all items.

[0056] 2. Causal perception negative sampling Therefore, although flow matching generates high-quality enhanced embeddings, directly using them for negative sampling may still perpetuate the exposure bias inherent in the training data. Based on the aforementioned causal debiasing framework, a principled causal intervention method is proposed, combining flow enhancement representations with explicit bias mitigation.

[0057] Frequency-based causal intervention. (This refers to the use of item frequencies.) Modeling is performed to simultaneously influence real user preferences Z_{u,i} and observed interactions. The confusion factor. Let... This represents the frequency of all items, with an empirical mean. and standard deviation For those with high interaction frequency For item i, the causal intervention weight is:

[0058] in It is a sigmoid function that ensures bounded weights. Control the intensity of intervention.

[0059] Debiased embedding construction. Debiased embeddings are constructed by adaptively combining the enhanced representation of the compositing stream with the original co-representation:

[0060] To demonstrate the effectiveness of this adaptive combination in removing bias: set up The exposure deviation in the embedding is measured. The adaptive combination satisfies:

[0061] in When stream enhancement embedding It exhibits better performance than the original embedding With lower frequency correlation, debiased embedding achieves reduced exposure bias.

[0062] As can be seen, the above formula confirms that the adaptive weighting strategy effectively reduces exposure bias by utilizing the frequency decorrelation characteristics of the multimodal guided flow representation.

[0063] The above formula achieves a principled trade-off, in which high-frequency items with reliable but potentially biased cooperative signals retain the original embedding (low-frequency items). Low-frequency items with sparse cooperative signals are represented using multimodal guided flow representations (high-frequency items with sparse cooperative signals). This compensates for data scarcity while avoiding popularity-driven artifacts.

[0064] Hard negative sample selection. Using debiased embeddings. The system performs hard negative sampling through the following steps: (1) Through random sampling Each item forms a candidate pool; (2) Calculate semantic similarity score At the same time, the user's interaction history is blocked; (3) Select from the top-K most similar candidates. This produces semantically challenging negative samples that respect multimodal relationships. Frequency-based interventions explicitly break down spurious frequency-selective correlations.

[0065] (vi) Joint optimization Embedding flow enhancement through adaptive fusion With multimodal enhancement representation Integration:

[0066] in The balance between control flow enhancement and graph propagation. This fusion representation combines the generative power of flow matching with the structural semantics of neighborhood aggregation capture.

[0067] Following LightGCN, in the case of a normalized adjacency matrix users Execution on the item's second part diagram Layer propagation. Initialization is performed to obtain the final embedding in the following way:

[0068] Optimize FMCNS using Bayesian Personalized Ranking (BPR) loss:

[0069] Where (u,p,n) represents a training triplet with user u, positive sample p, and randomly sampled negative sample n. To utilize the difficult negative samples from the causal perception mechanism, the following is introduced:

[0070] in This represents the difficult negative samples selected through causal perception. The overall objective combines two losses with flow matching regularization:

[0071] in Balanced standard and enhanced negative sampling, Weighted flow matching loss weights.

[0072] In this embodiment, we propose FMCNS, a novel framework that combines flow matching with causal-aware negative sampling for multimodal recommendation. By replacing the traditional curved diffusion path with a deterministic straight-line trajectory, we address key challenges in negative sampling and generative modeling, achieving stable and robust optimization as well as significantly fewer sampling steps. A frequency-based causal debiasing mechanism is introduced to explicitly mitigate exposure bias through multimodal semantic relationships. Furthermore, an integrated framework combining conditional flow matching and graph neural networks is designed to achieve effective alignment of multimodal semantics and collaborative filtering signals for comprehensive user preference modeling. Extensive experiments on three real-world datasets demonstrate that FMCNS consistently outperforms state-of-the-art baselines across various metrics, exhibiting particular advantages in training stability and computational efficiency.

[0073] Example 2 This embodiment provides a multimodal recommendation system based on stream matching and causal-aware negative sampling, including: The multimodal graph construction module is used to construct a multimodal graph that integrates visual and textual semantic relationships based on the item's ID embedding and multimodal features, and obtain a multimodal enhanced item representation; The flow enhancement generation module is used to generate semantically enhanced flow-enhanced item representations along a deterministic trajectory by taking multimodal features as conditions and combining them with prior distributions that reflect the popularity of items through a conditional flow matching strategy. The causal negative sampling module is used to construct debiased weights based on the idea of ​​causal perception using item frequency, adaptively fuse the stream-enhanced item representation with the original ID embedding to obtain the debiased item representation, and select semantically similar difficult negative samples from non-interactive items based on the representation. The joint optimization recommendation module is used to fuse the flow-enhanced item representation and the multimodal-enhanced item representation, and combine the hard negative samples to optimize the final embedding representation of users and items through graph propagation, so as to train the recommendation model to generate a personalized recommendation list.

[0074] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in a multimodal recommendation method based on stream matching and causal-aware negative sampling as described in Embodiment 1 above.

[0075] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the multimodal recommendation method based on stream matching and causal-aware negative sampling as described in Embodiment 1 above.

[0076] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal recommendation method based on flow matching and causal-aware negative sampling, characterized in that, include: Based on the ID embedding and multimodal features of items, a multimodal graph that integrates visual and textual semantic relationships is constructed, and a multimodal enhanced item representation is obtained; Using multimodal features as conditions and combining prior distributions that reflect the popularity of items, semantically enhanced flow-enhanced item representations are generated along deterministic trajectories through a conditional flow matching strategy. Based on the idea of ​​causal perception, the bias removal weight is constructed using item frequency. The stream-enhanced item representation is adaptively fused with the original ID embedding to obtain the bias removal item representation. Based on the representation, semantically similar difficult negative samples are selected from non-interactive items. By fusing the stream-enhanced item representation with the multimodal-enhanced item representation and incorporating the hard negative samples, the final embedding representation of users and items is optimized through graph propagation to train the recommendation model and generate a personalized recommendation list.

2. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The method involves constructing a multimodal graph that integrates visual and textual semantic relationships based on item ID embedding and multimodal features, and obtaining a multimodal enhanced item representation, specifically including: The similarity between items is calculated based on the visual features and text features in the multimodal features, and the most similar items are retained for each item to construct a sparse visual nearest neighbor map and text nearest neighbor map. After symmetric normalization of the two nearest neighbor graphs, a joint multimodal adjacency matrix is ​​obtained by weighted fusion; Based on the joint multimodal adjacency matrix, the item ID is embedded as the initial feature, and multi-round graph propagation is performed to aggregate multi-hop neighbor information, ultimately obtaining a multimodal enhanced item representation that integrates visual and textual semantic relationships.

3. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The prior distribution reflecting the popularity of items is specifically as follows: based on historical user-item interaction data, the frequency of each item being interacted with is counted; using the frequency as a parameter, a Bernoulli distribution source is constructed to initialize the flow matching generation process.

4. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The generation of semantically enhanced flow-enhanced item representations along a deterministic trajectory using a conditional flow matching strategy specifically involves: The projection-aligned representation of the multimodal features is used as the generation condition; Starting with samples drawn from the prior distribution of popularity, and ending with the ID embedding of the target item, a deterministic linear interpolation trajectory is constructed. By training a neural network to fit the velocity field along the trajectory, a mapping from the starting point to the ending point is learned; Using the aforementioned generation conditions as input, the neural network generates a flow-enhanced item representation that integrates multimodal semantics and popularity priors.

5. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The method based on causal perception, which utilizes item frequency to construct bias-free weights, specifically includes: Calculate the mean and standard deviation of the interaction frequency of all items in the dataset; For each item, the frequency is standardized, and the intervention weight is calculated through a monotonically decreasing mapping function, so that high-frequency items receive low weight and low-frequency items receive high weight.

6. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The adaptive fusion of the stream-enhanced item representation with the original ID embedding to obtain a bias-free item representation, and the selection of semantically similar difficult negative samples from non-interactive items based on the representation, specifically includes: The bias-reduction weights are used to perform a weighted summation of the stream-enhanced item representation and the original ID embedding of the item to obtain the final bias-reduction item representation. For each positive sample interaction, candidate items are randomly sampled from the set of items that have never interacted with that user; Calculate the similarity between the user embedding and the biased item representation of each candidate item, and select the items with the highest similarity as hard negative samples.

7. The multimodal recommendation method based on flow matching and causal-aware negative sampling as described in claim 1, characterized in that, The process of fusing the stream-enhanced item representation with the multimodal-enhanced item representation, and combining the hard negative samples, optimizes the final embedding representation of users and items through graph propagation to train the recommendation model and generate a personalized recommendation list. Specifically, this includes: The flow-enhanced item representation and the multimodal-enhanced item representation are merged according to a preset ratio to form an initial comprehensive item representation; The user ID embedding and the comprehensive representation are input into the user-item interaction graph. Multi-layer graph convolutional propagation is performed to capture high-order cooperative signals, and the outputs of each layer are averaged to obtain the final user and item embeddings. Construct a hybrid loss function that includes random negative samples and the hard negative samples, and train the model parameters by optimizing the loss function; After training, the final user and item embeddings are used to calculate preference scores and generate a personalized recommendation list.

8. A multimodal recommendation system based on flow matching and causal-aware negative sampling, characterized in that, include: The multimodal graph construction module is used to construct a multimodal graph that integrates visual and textual semantic relationships based on the item's ID embedding and multimodal features, and obtain a multimodal enhanced item representation; The flow enhancement generation module is used to generate semantically enhanced flow-enhanced item representations along a deterministic trajectory by taking multimodal features as conditions and combining them with prior distributions that reflect the popularity of items through a conditional flow matching strategy. The causal negative sampling module is used to construct debiased weights based on the idea of ​​causal perception using item frequency, adaptively fuse the stream-enhanced item representation with the original ID embedding to obtain the debiased item representation, and select semantically similar difficult negative samples from non-interactive items based on the representation. The joint optimization recommendation module is used to fuse the flow-enhanced item representation and the multimodal-enhanced item representation, and combine the hard negative samples to optimize the final embedding representation of users and items through graph propagation, so as to train the recommendation model to generate a personalized recommendation list.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the multimodal recommendation method based on flow matching and causal-aware negative sampling as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multimodal recommendation method based on flow matching and causal-aware negative sampling as described in any one of claims 1-7.