Transform-based multi-modal mask reinforcement learning recommendation state representation method

CN121434490BActive Publication Date: 2026-09-11NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511560038.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-09-11
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

[0005]1.融合机制简单,难以捕捉复杂非线性关系:由于该方案采用直接拼接或固定权重的线性加和方式进行融合,其本质是一种浅层的线性操作

Benefits of technology

[0045]This invention designs a unified feature encoding framework that maps heterogeneous multimodal features to the same semantic space. By introducing a learnable attention network, it can dynamically evaluate the importance of different modal features in the current recommendation scenario and adaptively adjust the contribution weights of each modality based on specific users, items, and environments, thereby more accurately capturing key signals of user interest. This invention applies a soft mask to the original feature representation, enabling it to automatically enhance the feature representation of important modalities while suppressing interference from secondary or noisy modalities. This achieves adaptive filtering and enhancement at the feature level, significantly improving the purity and robustness of the state representation. This invention employs a multi-layer Transformer encoder to deeply fuse the masked multimodal features, capturing complex nonlinear interactions between different modalities and learning fine-grained semantic associations across modalities, thus generating richer and more accurate state representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434490B_ABST
    Figure CN121434490B_ABST
Patent Text Reader

Abstract

The application provides a multi-modal mask reinforcement learning recommendation state representation method based on a Transformer, relates to the technical field of artificial intelligence, and can dynamically evaluate the importance of different modal characteristics in the current recommendation scene by introducing a learnable attention network, and can adaptively adjust the contribution weight of each mode according to specific users, items and environments, so that the key signals of user interest can be more accurately captured. The application calculates a soft mask and applies it to the original feature representation, can automatically enhance the feature representation of important modes, and at the same time suppress the interference of secondary or noise modes, realizing adaptive filtering and enhancement at the feature level. The application uses a multi-layer Transformer encoder to deeply fuse the masked multi-modal features, can capture complex nonlinear interaction relationships between different modes, learn cross-modal fine-grained semantic associations, and thus generate richer and more accurate state representations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a multimodal mask reinforcement learning recommendation state representation method based on Transformer. Background Technology

[0002] Reinforcement learning-based recommendation systems model the recommendation process as a Markov decision process, the core of which lies in generating a state that accurately represents the user's current interests and the environment. The quality of this state directly determines the accuracy of the recommendation agent's decisions.

[0003] Several schemes are proposed to construct effective state representations, specifically a state representation method based on simple fusion of multimodal features. This scheme aims to enrich the state representation by utilizing information from multiple sources (i.e., multimodal features) to overcome the problem of sparse information from a single feature source. Its implementation process is generally as follows: First, extract multimodal raw data from the system, including user-side features (such as user ID, historical click sequences), item-side features (such as item ID, category labels), and contextual features (such as time, location). Next, map these sparse or high-dimensional raw features into low-dimensional dense feature vectors through embedding layers or fully connected layers. Then, fuse these feature vectors from different modalities using simple operations such as direct concatenation or weighted averaging. For example, directly concatenate the user feature vector, item feature vector, and contextual feature vector into a longer joint vector, or assign fixed weights to different modalities and then sum them. Finally, use this fused vector as the state representation. The input is fed into the policy network of the reinforcement learning agent for recommendation decision-making.

[0004] The above-mentioned scheme based on simple fusion of multimodal features, while conceptually superior to single-feature-source methods, suffers from the following significant drawbacks in practical applications:

[0005] 1. The fusion mechanism is simple and struggles to capture complex nonlinear relationships: Because this scheme uses direct concatenation or fixed-weight linear summation for fusion, it is essentially a shallow linear operation. However, the interactions between user interests, item attributes, and contextual information are complex and nonlinear. Simple linear fusion models cannot effectively capture the deep intrinsic connections between these modalities, resulting in limited expressive power of the fused state representation, a "feature gap," and a significant amount of valuable information not being fully utilized.

[0006] 2. Lack of dynamic adaptability and inability to highlight key information: The fusion weights in this scheme (such as when using weighted averaging) are usually statically preset or generated by a simple network, and cannot be dynamically adjusted according to specific users, items, and contextual scenarios. For example, when recommending fashion items, the image features of the items may be crucial; while when recommending news articles, text features become key. Static fusion strategies cannot adaptively identify and strengthen key modalities and suppress secondary or noisy modalities in the current decision-making environment, resulting in generated state representations that are not targeted and contain redundant information.

[0007] 3. Sensitive to noise and poor robustness of state representation: Multimodal data often contains a large amount of noise irrelevant to the recommendation task (e.g., expired tags in user profiles, irrelevant words in item descriptions). Existing solutions lack effective noise filtering mechanisms and treat all feature information "equally" during the fusion process, causing noise information to be included in the state representation to the same extent. This reduces the purity and robustness of the state, thereby affecting the stability and accuracy of the downstream reinforcement learning agent's strategy.

[0008] Therefore, the state representations ultimately generated by existing technologies still fail to completely solve the problems of inaccurate and unrobust information, which restricts further improvement in the performance of reinforcement learning recommendation systems. Summary of the Invention

[0009] To address the shortcomings of existing technologies, the present invention aims to propose a Transformer-based multimodal mask reinforcement learning recommendation state representation method, including:

[0010] Receive the environmental characteristics currently accessed by the user, and obtain user characteristics, item characteristics, user behavior sequence characteristics, image characteristics, and text characteristics;

[0011] Multiple different encoders are used to encode user features, item features, user behavior sequence features, environmental features, image features, and text features respectively, to obtain a multimodal feature matrix. The multimodal feature matrix includes user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors.

[0012] The multimodal feature matrix is ​​processed by an attention network to obtain a multimodal feature weighting matrix;

[0013] Weighted matrix of multimodal features Add positional encoding to each weighted feature vector, and then... Inputting a multi-layer Transformer encoder yields the enhanced feature matrix. This will enhance the feature matrix. After average pooling, the fused state representation vector is obtained. ;

[0014] The fused state representation vector Input the policy network of the reinforcement learning agent to obtain recommended actions. The recommended actions include the click probability of each item in the candidate pool. The top N items with the highest click probability values ​​are displayed and output.

[0015] The reward is calculated based on the user's feedback on the output item. The Policy Gradient Algorithm (PPO) is used to update the parameters of the encoder, the attention network, and the reinforcement learning agent until the test metric does not improve within a preset number of rounds, thus obtaining the trained encoder, attention network, and reinforcement learning agent.

[0016] The trained encoder, attention network, and reinforcement learning agent process environmental features, user features, item features, user behavior sequence features, image features, and text features to display the N items with the highest click probability values.

[0017] Optionally, the user features include a static user profile and a dynamic interest vector, wherein the static user profile includes at least age and gender embedding vectors, and the dynamic interest vector represents the user's preferences;

[0018] The item features include the features of candidate items in the candidate pool, which is obtained by filtering all items based on user access requests within a preset time period using a collaborative filtering recall algorithm.

[0019] The user behavior sequence feature is a sequence of all items accessed by the user within a preset time period;

[0020] The environmental characteristics include at least the user's current access time, the device the user is currently using, and the network environment the user is currently accessing.

[0021] The image features include an image of the candidate item, and the text features include at least a brief description of the candidate item.

[0022] Optionally, multiple different encoders are used to encode user features, item features, user behavior sequence features, environmental features, image features, and text features respectively, to obtain a multimodal feature matrix, including:

[0023] User features are processed through an embedding layer and a fully connected layer to obtain a user feature vector u, and item features are processed through an embedding layer and an MLP to obtain an item feature vector v.

[0024] Embed all items in the user behavior sequence features to obtain a sequence matrix. Encode the sequence matrix using a bidirectional LSTM or self-attention network to obtain the behavior feature vector h.

[0025] The device and the current network environment are embedded in the environmental features, and the embedded representation is concatenated with the current access time. The concatenated vector is then processed by an MLP to obtain the environmental feature vector c.

[0026] The image features are processed through a pre-trained convolutional neural network or Vision Transformer, and then through an MLP to obtain the image feature vector i.

[0027] The text features are processed by a pre-trained language model or Bi-LSTM to extract the semantic vector representation of the text. The semantic vector representation is then processed by an MLP to obtain the text feature vector t.

[0028] The multimodal feature matrix is ​​composed of user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors. , where d represents the dimension of the feature vector;

[0029] Optionally, the multimodal feature matrix can be processed using an attention network to obtain a multimodal feature weighting matrix, including:

[0030] The importance score of each eigenvector in the multimodal feature matrix X is calculated using the following formula:

[0031] ;

[0032] Where, x j ∈X, For user feature vectors, For item feature vectors, For behavioral feature vectors, For environmental feature vectors, For image feature vectors, Let T be the text feature vector, and T denote the transpose. Represents the weight matrix. And k represents the bias, σ represents the activation function, for Importance score;

[0033] The importance scores of all feature vectors are Softmax normalized to obtain the final attention weights. Specifically, this is achieved through the following formula:

[0034] ;

[0035] Generate a soft mask α based on all attention weights;

[0036] The attention weights in the soft mask α are multiplied by the eigenvectors in their corresponding multimodal feature matrix X to obtain weighted eigenvectors. All weighted eigenvectors form the multimodal feature weighting matrix. .

[0037] Optionally, a soft mask α is generated based on all attention weights, including:

[0038] The vector composed of all attention weights is used as a soft mask α, or a threshold-based gating mask mechanism is used to numerically adjust the vector composed of all attention weights to obtain the soft mask α.

[0039] Optionally, the attention weights in the soft mask α are multiplied by the eigenvectors in their corresponding multimodal feature matrix X to obtain a weighted eigenvector, which is achieved using the following formula:

[0040] X masked =α⊙X;

[0041] Here, ⊙ represents multiplication row by row.

[0042] Optionally, the parameters of the encoder, the attention network, and the reinforcement learning agent can be updated, including:

[0043] The weights of the embedding layer and MLP, the weights of the bidirectional LSTM or the self-attention network, the weights of the projection layer of the convolutional neural network or the projection layer of the Vision Transformer, the weights of the projection layer of the pre-trained language model or the Bi-LSTM projection layer, and the weight matrix and bias of the attention network are updated. At the same time, the Q, K, V projection matrices in the self-attention layer of the Transformer encoder, the weights of the feedforward neural network, and the layer normalization parameters are updated. The weights of the MLP in the policy network and the weights of the MLP in the value network of the reinforcement learning agent are also updated.

[0044] The beneficial effects of adopting the above technical solution are as follows:

[0045] This invention designs a unified feature encoding framework that maps heterogeneous multimodal features to the same semantic space. By introducing a learnable attention network, it can dynamically evaluate the importance of different modal features in the current recommendation scenario and adaptively adjust the contribution weights of each modality based on specific users, items, and environments, thereby more accurately capturing key signals of user interest. This invention applies a soft mask to the original feature representation, enabling it to automatically enhance the feature representation of important modalities while suppressing interference from secondary or noisy modalities. This achieves adaptive filtering and enhancement at the feature level, significantly improving the purity and robustness of the state representation. This invention employs a multi-layer Transformer encoder to deeply fuse the masked multimodal features, capturing complex nonlinear interactions between different modalities and learning fine-grained semantic associations across modalities, thus generating richer and more accurate state representations. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the process of the Transformer-based multimodal mask reinforcement learning recommendation state representation method in an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the process for generating the fused state representation vector in an embodiment of the present invention. Detailed Implementation

[0048] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0049] To address the problems existing in the prior art, the purpose of this invention is to overcome the shortcomings of existing reinforcement learning recommendation systems, such as inaccurate user interest characterization and limited recommendation performance due to single state representation, insufficient multimodal feature fusion, and noise interference. This invention aims to solve the problems of simple and inefficient fusion methods, lack of adaptability, and sensitivity to noise in the prior art by introducing a dynamic multimodal feature fusion mechanism with noise suppression capabilities. This will generate a more accurate and robust state representation, improve the decision-making quality of the reinforcement learning recommendation agent, and ultimately optimize the long-term benefits of the recommendation system and the user experience.

[0050] Specifically, this invention provides a Transformer-based multimodal mask reinforcement learning recommendation state representation method, combined with... Figure 1 This may include the following steps:

[0051] It receives the environmental characteristics of the user's current visit and obtains user characteristics, item characteristics, user behavior sequence characteristics, image characteristics, and text characteristics, which can be obtained from the user's database or real-time logs.

[0052] Among them, the user features ( This includes static user profiles and dynamic interest vectors. The static user profile includes at least age and gender embedding vectors, and may also include user registration time and membership level / status. Registration time reflects the user's familiarity with the platform; older users may have more stable preferences. Membership level / status can be that of a regular user or a VIP user, which directly affects their privileges and spending power. The dynamic interest vector represents the user's preferences. The dynamic interest vector is constructed and trained based on the user's historical behavioral data. This vector can be seen as the user's "interest fingerprint" or "interest summary" at the current moment, condensing the preferences reflected in the user's recent behavior. For example, user characteristics could be that user 123 is a young male who frequently watches science fiction and action short films recently.

[0053] The characteristics of the item ( The candidate pool includes the characteristics of candidate items in the candidate pool, which is obtained by filtering all items based on user access requests within a preset time period using a collaborative filtering recall algorithm.

[0054] To illustrate the characteristics of an item, an item could be a candidate video (such as "Interstellar"), and its characteristics could include ID embedding, category tags, statistical popularity, etc.

[0055] This invention can be applied to multiple scenarios. In the video recommendation scenario, the item is a video. That is, when a user accesses the video recommendation platform through a mobile app, the item refers to a video. The user's access request within a preset time period refers to the video watched by the user within that preset time period. This invention can also be applied to the music recommendation scenario and the shopping scenario. The corresponding user access requests within a preset time period are the music listened to by the user and the goods purchased by the user within that preset time period.

[0056] The user behavior sequence features ( ) represents the sequence of all items accessed by a user within a preset time period. When the item is a video, the user behavior sequence feature is represented as [Video A, Video B, Video C].

[0057] The environmental characteristics include at least the user's current access time, the device the user is currently using, and the network environment the user is currently accessing, and may also include the user's geographical location; the device the user is currently using can be a mobile phone, computer, iPad, etc., and the network environment the user is currently accessing can be Wi-Fi, mobile data, etc.

[0058] The image features ( This includes images of candidate items. If the item is a video, the image feature can be the cover image or keyframe image of the candidate video. If the item is a product, the image feature can be an image of the product.

[0059] Text features ( The text features should include at least a brief description of the candidate item. If the item is a video, the text features can include the title, description, cast list, and other textual descriptions of the candidate video. If the item is a commodity, the text features can include the description of the commodity. If the item is music, the text features can include the textual description of the music, or information such as lyrics and singer.

[0060] Multiple different encoders are used to encode user features, item features, user behavior sequence features, environmental features, image features, and text features respectively, to obtain a multimodal feature matrix. The multimodal feature matrix includes user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors.

[0061] Specifically, user features are processed through an embedding layer and a fully connected layer to obtain a user feature vector u∈R. d The item features are processed through an embedding layer and an MLP to obtain the item feature vector v∈R. d ;

[0062] Embed all items from the user behavior sequence features to obtain the sequence matrix H∈R L×d Where L represents the number of items in the user behavior sequence, the sequence matrix is ​​encoded using a bidirectional LSTM or self-attention network to extract sequence information, and the behavior feature vector h∈R is obtained. d ;

[0063] The device and current network environment are embedded in the environmental features, and the embedded representation is concatenated with the current access time. The concatenated vector is then processed by an MLP to obtain the environmental feature vector c∈R. d ;

[0064] Image features are processed through a pre-trained convolutional neural network (CNN, such as ResNet) or Vision Transformer (ViT) to extract high-level feature representations of the image. These representations are then mapped to the same dimension as other modalities using an MLP to obtain the image feature vector i∈R. d ;

[0065] The text features are processed by a pre-trained language model (BERT, Sentence-BERT) or Bi-LSTM to extract semantic vector representations of the text. These semantic vector representations are then processed by an MLP to obtain the text feature vector t∈R. d ;

[0066] The multimodal feature matrix is ​​composed of user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors. , where d represents the feature vector dimension, and d can also be understood as the model's "ability" or "degree of freedom" to learn the abstract relationships within the features.

[0067] The multimodal feature matrix is ​​processed using an attention network to obtain a multimodal feature weighting matrix, specifically including:

[0068] Combination Figure 2 The importance score of each feature vector in the multimodal feature matrix X is calculated using the following formula:

[0069] ;

[0070] Where, x j ∈X, For user feature vectors, For item feature vectors, For behavioral feature vectors, For environmental feature vectors, For image feature vectors, Let T be the text feature vector, and T denote the transpose. Represents the weight matrix. And k represents the bias, σ represents the activation function, for Importance score;

[0071] The importance scores of all feature vectors are Softmax normalized to obtain the final attention weights. Specifically, this is achieved through the following formula:

[0072] ;

[0073] In the current scenario (weekend night, mobile device, user's historical preferences), the algorithm may calculate higher weights for user historical behavior (h) and item features (v) (for matching interests), medium weights for user profile (u), and lower weights for context features (c) (because the information about weekend nights is already partially included in historical behavior).

[0074] A soft mask α is generated based on all attention weights. Specifically, the vector composed of all attention weights is used as the soft mask α, or a threshold-based gating mask mechanism is used to adjust the values ​​of the vector composed of all attention weights to obtain the soft mask α. Specifically, a threshold range is set, increasing the values ​​of attention weights greater than the maximum value and decreasing the values ​​of attention weights less than the minimum value.

[0075] The attention weights in the soft mask α are multiplied by the eigenvectors in its corresponding multimodal feature matrix X to obtain the weighted eigenvectors, which is achieved through the following formula:

[0076] X masked =α⊙X;

[0077] Here, ⊙ represents row-by-row multiplication (broadcast mechanism). This means that features with high importance are enhanced, while features with low importance are suppressed.

[0078] All weighted eigenvectors form a multimodal feature weighting matrix. .

[0079] This process enables dynamic and adaptive feature selection and noise reduction. It generates a view of the most relevant features for the current decision context, effectively suppressing interference from irrelevant or noisy information.

[0080] Next, leveraging the powerful sequence modeling capabilities of Transformer, we will delve into the intrinsic connections between the masked and enhanced multimodal features.

[0081] Since Transformer itself does not have the ability to perceive sequence order, this invention uses a multimodal feature weighting matrix. Each weighted feature vector is encoded with a positional code to inject modality type information. The resulting vector with the added positional code... Inputting a multi-layer Transformer encoder yields an enhanced feature matrix. This will enhance the feature matrix. After average pooling, the fused state representation vector is obtained. Specifically, it is expressed by the following formula:

[0082]

[0083] in, This represents the fused state representation vector. This indicates average pooling.

[0084] The self-attention mechanism in the Transformer encoder further learns the interaction relationships between features of different modalities. For example, it can learn that "when the item feature contains 'Nolan,' it is highly correlated with the user's previous science fiction movie viewing history."

[0085] By stacking multiple layers, Transformer is able to capture the complex, non-linear dependencies between these modalities and perform deep fusion.

[0086] The fused state representation vector Input the policy network of the reinforcement learning agent The recommended action is obtained, which includes the click probability of each item in the candidate pool. The top N items with the highest click probability values ​​are then displayed.

[0087] Rewards are calculated based on user feedback regarding the output items. User feedback may include whether the user clicked on the item, whether they purchased it, and the duration of their interaction with the item.

[0088] Based on the Policy Gradient Algorithm (PPO) using rewards, the parameters of the encoder, the attention network, and the reinforcement learning agent are updated. Specifically, the weights of the embedding layer and MLP, the weights of the bidirectional LSTM or self-attention network, the weights of the projection layer of the convolutional neural network or the projection layer of the Vision Transformer, the weights of the projection layer of the pre-trained language model or the Bi-LSTM projection layer, and the weight matrix and bias of the attention network are updated. At the same time, the Q, K, and V projection matrices in the self-attention layer of the Transformer encoder, the weights of the feedforward neural network, and the layer normalization parameters are updated. Simultaneously, the weights of the MLP in the policy network and the weights of the MLP in the value network of the reinforcement learning agent are updated.

[0089] For the updated parameters, the feature encoding part is further explained in conjunction with the previous technical solutions. Specifically, it includes the weights of the embedding layers and MLPs for the user feature encoding part and the item feature encoding part, the weights of the bidirectional LSTM or self-attention network for the historical behavior sequence encoding part, the weights of the MLP for the environment feature encoding part, the weights of the projection layer of the convolutional neural network or the projection layer of the Vision Transformer for the image feature encoding part, and the weights of the projection layer of the pre-trained language model or the Bi-LSTM projection layer for the text feature encoding part.

[0090] Among them, the reinforcement learning recommendation agent can be a policy model based on the Actor-Critic network.

[0091] The training continues until the test metrics do not improve within a preset number of rounds. For example, if the core metrics such as click-through rate and dwell time do not improve for several consecutive rounds, the encoder, attention network, and reinforcement learning agent are obtained.

[0092] The trained encoder, attention network, and reinforcement learning agent process environmental features, user features, item features, user behavior sequence features, image features, and text features to display the N items with the highest click probability values.

[0093] This invention addresses the problems of inaccurate state representation, insufficient information fusion, and noise interference in reinforcement learning recommendation systems. Specifically, it is reflected in the following aspects:

[0094] 1. Dynamic Modal Importance Assessment and Weighting Mechanism: This invention introduces a learnable attention network to dynamically assess the importance of different modal features in the current recommendation scenario. Unlike traditional fixed-weight or simple concatenation fusion methods, this method adaptively adjusts the contribution weights of each modality based on the specific user, item, and context, thereby more accurately capturing key signals of user interest.

[0095] 2. Feature Masking Denoising Method Based on Attention Weights: This invention innovatively uses the calculated modal attention weights as soft masks applied to the original feature representation. This mechanism can automatically enhance the feature representation of important modalities while suppressing interference from secondary or noisy modalities, achieving adaptive filtering and enhancement at the feature level, and significantly improving the purity and robustness of the state representation.

[0096] 3. Deep Cross-Modal Feature Fusion Based on Transformer: This invention employs a multi-layer Transformer encoder to deeply fuse masked multi-modal features. Through a self-attention mechanism, the model can capture complex nonlinear interactions between different modalities and learn fine-grained semantic associations across modalities, thereby generating richer and more accurate state representations.

[0097] 4. End-to-End Reinforcement Learning State Representation Optimization Framework This invention tightly integrates multimodal feature extraction, attention masking mechanisms, and Transformer fusion modules with a reinforcement learning recommendation system, forming an end-to-end optimization framework. The entire state representation module is trained together with the recommendation policy network, enabling the generated state representation to directly optimize the final recommendation target, rather than merely reconstructing input features.

[0098] 5. Unified Encoding and Alignment of Multimodal Features This invention designs a unified feature encoding framework that maps heterogeneous multimodal data (including user features, item features, behavior sequences, contextual information, and image and text content features) to the same semantic space, laying the foundation for subsequent deep fusion and interactive computing.

[0099] Compared with existing technologies, the reinforcement learning recommendation state representation method based on multimodal mask Transformer provided by this invention has the following significant advantages and beneficial effects:

[0100] 1. Recommendation accuracy and user satisfaction have been significantly improved.

[0101] Results: By fusing dynamic attention masking and deep Transformer, the generated state representation can more comprehensively and accurately depict the user's complex and dynamically changing interests, thereby significantly improving recommendation relevance.

[0102] Theoretical Analysis: Traditional methods suffer from biased state representations due to insufficient feature fusion or the presence of noise. This invention effectively reduces this bias through modal-level weighting and feature-level fusion. Theoretical derivation and characteristic analysis show that the state representation generated by this invention exhibits a significantly improved cosine similarity to the user's true interest compared to baseline methods.

[0103] Advantages: Compared to simple splicing and fusion methods (which are prone to introducing noise) and static weighting methods (which cannot adapt), this invention can dynamically highlight important features and suppress noise, and is expected to significantly improve click-through rate (CTR) and average user dwell time.

[0104] 2. Enhanced system robustness and anti-interference capability

[0105] Results: It has a greater tolerance for noise and redundant information in multimodal data, and the recommendation quality is more stable.

[0106] Theoretical analysis: The masking mechanism is essentially a feature selection process. By reducing the weight of unimportant modalities, it is equivalent to reducing the dimensionality of noisy input.

[0107] Advantages: Traditional models show a significant drop in recommendation performance when encountering poor-quality image features (such as blurry covers) or exaggerated text descriptions. This invention can automatically reduce the weights of these noisy modalities, ensuring the stability of the system output, and is expected to significantly improve recommendation stability in noisy environments.

[0108] 3. Improved model convergence speed and training efficiency

[0109] Results: The end-to-end training framework and more accurate state representation accelerate the policy learning process of reinforcement learning agents.

[0110] Theoretical Analysis: More accurate state representation means that the environment state faced by the agent is more distinctive and deterministic, reducing the difficulty of policy learning. Analysis shows that, with the same amount of training data, this invention is expected to significantly shorten the convergence time of the policy network.

[0111] Advantages: Compared to learning strategies from raw, noisy multimodal data, the "cleaned" state representation provided by this invention offers clearer learning signals to the RL agent, improving the utilization of training data and learning efficiency.

[0112] 4. Optimization of computing and storage resources

[0113] Results: While improving performance, it also optimized computational and storage efficiency.

[0114] Theoretical Analysis: Although a Transformer structure is introduced, the computational overhead of directly processing high-dimensional raw features (such as image pixels and text sequences) is avoided through unified low-dimensional feature encoding and efficient masking operations. The feature-level fusion employed in this invention has a computational complexity far lower than models that perform fusion at the raw data level.

[0115] Advantages: Compared to earlier fusion solutions (such as directly concatenating high-dimensional features), this invention is expected to reduce the computational load of online inference. Furthermore, the unified feature encoding dimension facilitates model parallelization and hardware acceleration, making it more suitable for large-scale online recommendation systems.

[0116] 5. Scalability and versatility

[0117] Results: The framework design is flexible and can be easily integrated into new modalities or adapted to different recommendation scenarios.

[0118] Theoretical analysis: The mask attention mechanism and Transformer fusion module of this invention are not sensitive to the number and type of input modalities. New modalities can be introduced simply by adding modal encoders and position encoders accordingly, without changing the core architecture.

[0119] Advantages: Traditional fusion models designed for specific modalities have poor scalability. This invention provides a general multimodal state representation solution that can be quickly applied to various recommendation scenarios such as e-commerce, news, and video, significantly reducing development and maintenance costs.

[0120] 6. End-user experience and business value

[0121] Results: Ultimately, it brings users a more personalized and accurate recommendation experience, creating greater value for the business.

[0122] Theoretical analysis: More accurate recommendations directly translate into higher user engagement (click-through rate, dwell time, conversion rate) and long-term satisfaction (retention rate).

[0123] Comparative advantages: It is expected to significantly improve core business metrics. For example, in video recommendation scenarios, user viewing time and video completion rate are expected to see considerable improvement; in e-commerce scenarios, click-through rate and order value are also expected to see significant growth.

[0124] In summary, this invention, through its innovative multimodal mask Transformer architecture, not only achieves superior state representation at the technical level, but also brings comprehensive improvements in recommendation performance, system robustness, computational efficiency, and adaptability to multiple scenarios at the application level, demonstrating significant technological advancement and broad application prospects.

[0125] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A Transformer-based multimodal mask reinforcement learning recommendation state representation method, characterized in that, include: Receive the environmental characteristics currently accessed by the user, and obtain user characteristics, item characteristics, user behavior sequence characteristics, image characteristics, and text characteristics; Multiple different encoders are used to encode user features, item features, user behavior sequence features, environmental features, image features, and text features respectively, to obtain a multimodal feature matrix. The multimodal feature matrix includes user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors. The multimodal feature matrix is ​​processed using an attention network to obtain a multimodal feature weighting matrix, including: Calculate the importance score of each eigenvector in the multimodal feature matrix X. Perform Softmax normalization on all eigenvector importance scores to obtain the final attention weights. Generate a soft mask α based on all attention weights. Multiply the attention weights in the soft mask α by the corresponding eigenvectors in the multimodal feature matrix X to obtain weighted eigenvectors. All weighted eigenvectors form the multimodal feature weighting matrix. ; Weighted matrix of multimodal features Add positional encoding to each weighted feature vector, and then... Inputting a multi-layer Transformer encoder yields the enhanced feature matrix. This will enhance the feature matrix. After average pooling, the fused state representation vector is obtained. ; The fused state representation vector Input the policy network of the reinforcement learning agent to obtain recommended actions. The recommended actions include the click probability of each item in the candidate pool. The top N items with the highest click probability values ​​are displayed and output. The reward is calculated based on the user's feedback on the output item. The Policy Gradient Algorithm (PPO) is used to update the parameters of the encoder, the attention network, and the reinforcement learning agent until the test metric does not improve within a preset number of rounds, thus obtaining the trained encoder, attention network, and reinforcement learning agent. The trained encoder, attention network, and reinforcement learning agent process environmental features, user features, item features, user behavior sequence features, image features, and text features to display the N items with the highest click probability values.

2. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 1, characterized in that, The user features include static user profiles and dynamic interest vectors. The static user profiles include at least age and gender embedding vectors, and the dynamic interest vectors represent the user's preferences. The item features include the features of candidate items in the candidate pool, which is obtained by filtering all items based on user access requests within a preset time period using a collaborative filtering recall algorithm. The user behavior sequence feature is a sequence of all items accessed by the user within a preset time period; The environmental characteristics include at least the user's current access time, the device the user is currently using, and the network environment the user is currently accessing. The image features include an image of the candidate item, and the text features include at least a brief description of the candidate item.

3. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 1, characterized in that, Multiple encoders are used to encode user features, item features, user behavior sequence features, environmental features, image features, and text features respectively, resulting in a multimodal feature matrix, including: User features are processed through an embedding layer and a fully connected layer to obtain a user feature vector u, and item features are processed through an embedding layer and an MLP to obtain an item feature vector v. Embed all items in the user behavior sequence features to obtain a sequence matrix. Encode the sequence matrix using a bidirectional LSTM or self-attention network to obtain the behavior feature vector h. The device and the current network environment are embedded in the environmental features, and the embedded representation is concatenated with the current access time. The concatenated vector is then processed by an MLP to obtain the environmental feature vector c. The image features are processed through a pre-trained convolutional neural network or Vision Transformer, and then through an MLP to obtain the image feature vector i. The text features are processed by a pre-trained language model or Bi-LSTM to extract the semantic vector representation of the text. The semantic vector representation is then processed by an MLP to obtain the text feature vector t. The multimodal feature matrix is ​​composed of user feature vectors, item feature vectors, behavior feature vectors, environmental feature vectors, image feature vectors, and text feature vectors. , where d represents the dimension of the feature vector.

4. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 3, characterized in that, The importance score of each eigenvector in the multimodal feature matrix X is calculated using the following formula: ; Where, x j ∈X, For user feature vectors, For item feature vectors, For behavioral feature vectors, For environmental feature vectors, For image feature vectors, Let T be the text feature vector, and T denote the transpose. Represents the weight matrix. And k represents the bias, σ represents the activation function, for Importance score; Specifically, the importance scores of all feature vectors are Softmax normalized to obtain the final attention weights. Specifically, this is achieved through the following formula: 。 5. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 1, characterized in that, Based on all attention weights, generate a soft mask α, including: The vector composed of all attention weights is used as a soft mask α, or a threshold-based gating mask mechanism is used to numerically adjust the vector composed of all attention weights to obtain the soft mask α.

6. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 1, characterized in that, The attention weights in the soft mask α are multiplied by the eigenvectors in its corresponding multimodal feature matrix X to obtain the weighted eigenvectors, which is achieved through the following formula: X masked =α⊙X; Here, ⊙ represents multiplication row by row.

7. The Transformer-based multimodal mask reinforcement learning recommendation state representation method according to claim 3, characterized in that, The parameters of the encoder, the attention network, and the reinforcement learning agent are updated, including: The weights of the embedding layer and MLP, the weights of the bidirectional LSTM or the self-attention network, the weights of the projection layer of the convolutional neural network or the projection layer of the Vision Transformer, the weights of the projection layer of the pre-trained language model or the Bi-LSTM projection layer, and the weight matrix and bias of the attention network are updated. At the same time, the Q, K, V projection matrices in the self-attention layer of the Transformer encoder, the weights of the feedforward neural network, and the layer normalization parameters are updated. The weights of the MLP in the policy network and the weights of the MLP in the value network of the reinforcement learning agent are also updated.

Citation Information

Patent Citations

  • Advertisement marketing recommendation method based on deep reinforcement learning

    CN118396685A

  • Aligning Sequence Processing Models with Recommendation Knowledge

    US20250200440A1