Personalized recommendation method based on multi-modal behavior sequence modeling

Through the Transformer-based multimodal self-attention mechanism and long-term and short-term interest modeling, combined with online learning of real-time user feedback, the problems of low efficiency and insufficient accuracy of multimodal data fusion and time series modeling in the recommendation system are solved, and the accurate capture of user interests and real-time updating of recommendation results are achieved.

CN120670665APending Publication Date: 2025-09-19HUBEI UNIV

Patent Information

Application Number
CN202510737440.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing recommendation systems suffer from inefficiency and lack of accuracy in multimodal data fusion, time series modeling, and real-time feedback driving. They find it difficult to effectively capture the dynamic changes and long-term dependencies of user interests, and fail to optimize the model in real time to adapt to the dynamic changes in user behavior.

Method used

The Transformer-based multimodal self-attention mechanism (MMSA) is used to model multimodal behavior sequences. Combined with the long-term and short-term interest divide-and-conquer modeling and the online learning method driven by real-time user feedback, the dynamic fusion and real-time optimization of multimodal data are achieved by dynamically adjusting weights and model parameters.

Benefits of technology

It improves the information sharing and expression efficiency of multimodal data, can accurately capture the multi-dimensional dynamic changes of user interests, improve the personalization and real-time nature of recommendations, and enhance the performance of recommendation systems in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670665A_ABST
    Figure CN120670665A_ABST
Patent Text Reader

Abstract

The invention relates to the field of recommendation systems, and particularly discloses a personalized recommendation method based on multi-modal behavior sequence modeling, which comprises the following steps of: acquiring multi-modal behavior data such as user text, image, time and place, preprocessing, and realizing dynamic fusion of the data by utilizing a multi-modal self-attention mechanism (MMSA) to obtain a multi-modal behavior sequence model; and the interest evolution of the user is accurately captured. An independent RNN module is adopted to model long-term and short-term interests of a user, and the long-term and short-term interests are combined through a self-learning weight coefficient, so that the change of the user interests is reflected more accurately. In addition, by introducing an online learning and incremental learning mechanism, model parameters are dynamically adjusted according to real-time feedback of the user, and it is ensured that a recommendation result can respond to user interest changes in time. According to the method, the defects of an existing recommendation system in the aspects of data fusion, time sequence modeling and real-time adaptability are effectively overcome, recommendation individuation and accuracy are improved, and the real-time updating capacity and scene adaptability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of recommendation systems, and in particular to a personalized recommendation method based on multimodal behavior sequence modeling. Background Art

[0002] Recommendation systems are a core technology for internet services. Their goal is to provide users with accurate content or product recommendations by analyzing user behavior, preferences, and contextual information. With the multimodal and dynamic nature of user behavior data, traditional recommendation methods face significant challenges in data fusion, time series modeling, and real-time adaptability. Existing technologies struggle to balance recommendation accuracy and efficiency in scenarios involving the dynamic correlations of multimodal data, differentiated modeling of users' long-term and short-term interests, and real-time feedback-driven model optimization. More efficient solutions are urgently needed.

[0003] Currently, research on recommendation systems focuses on the following areas: (1) Collaborative filtering methods make recommendations by mining similarities in the user-item interaction matrix; (2) Content-based recommendation methods generate recommendations by matching user preferences with item features; and (3) Sequence modeling methods attempt to capture temporal dependencies by modeling user behavior sequences. The following are related research and applications of two existing technologies:

[0004] (1) Patent CN109885756B proposes a hybrid neural network recommendation scheme, whose technical solution mainly includes the following steps: first, obtain user historical behavior sequence data (such as click and browsing records), and perform data cleaning and standardization; second, use convolutional neural network (CNN) to extract local features of user recent behavior data to capture the correlation between adjacent behaviors; then, input the local features extracted by CNN into recurrent neural network (RNN), and learn the long-term dependencies and short-term preferences of user behaviors through RNN's sequence modeling capability; finally, fuse the features output by CNN and RNN, and use multi-layer perceptron (MLP) to predict the items that users may interact with in the future and generate a recommendation list. However, the CNN used in this method focuses on the local correlation of recent behaviors and may ignore the global temporal patterns in long-term behavior sequences, resulting in incomplete interest modeling; in addition, RNN has the problem of vanishing gradient when processing long sequences, making it difficult to effectively capture long-distance behavior dependencies.

[0005] Patent CN116089701A proposes a keyword expansion recommendation method based on a semantic tree. Its technical solution mainly includes the following steps: first, extract core keywords from the user's text browsing history, such as high-frequency words in the user's reading articles; second, map the first keyword to a preset semantic tree, which is composed of multiple nodes and semantic relationship links between nodes; then, according to the position of the first keyword in the semantic tree, expand along the semantic relationship link to obtain the second keyword to form a semantically enhanced interest tag set; finally, based on the combination of the first keyword and the second keyword, match the candidate text content and generate a recommendation list. However, the construction of the semantic tree preset by this method requires manual annotation or a fixed knowledge base, which is difficult to dynamically update to adapt to emerging vocabulary or domain-specific semantic relationships, resulting in limited coverage of semantic expansion; in addition, it only relies on text keywords for recommendation and does not integrate the user's multimodal behavior data, making it difficult to fully capture user interests.

[0006] Disadvantages of existing technology:

[0007] (1) Low efficiency of multimodal data fusion: Existing methods use static feature splicing or fixed weight fusion strategies, which do not consider the dynamic correlation and weight changes of different modalities (such as text, images, and spatiotemporal data) in time series. As a result, the temporal expression ability of multimodal features is limited, making it difficult to accurately capture the dynamic evolution of user interests.

[0008] (2) Insufficient time series modeling capabilities: The sequence modeling method based on RNN / LSTM has the gradient vanishing problem, which makes it difficult to effectively model long-distance dependencies in long sequences. It also does not incorporate temporal context information (such as timestamps and user activity), resulting in insufficient differentiation between users' long-term interests and short-term behaviors.

[0009] (3) Lack of real-time feedback and adaptability: Existing technologies mostly use offline batch training mode, and do not combine users' real-time behavior feedback (such as clicks, skips) to dynamically optimize the model. As a result, the recommendation results lag behind the changes in user interests and are difficult to cope with the needs of dynamic scenarios.

[0010] Based on this, this study proposed a personalized recommendation method based on multimodal behavior sequence modeling, which can effectively capture the multidimensional dynamic changes of user interests. Summary of the Invention

[0011] In response to the above-mentioned problems in the prior art, the present invention provides a personalized recommendation method based on multimodal behavior sequence modeling, which effectively improves the efficiency of information sharing and expression between different modalities and can effectively capture the multi-dimensional dynamic changes of user interests.

[0012] To achieve the above objectives, the present invention proposes a personalized recommendation method based on multimodal behavior sequence modeling, comprising:

[0013] S1, multimodal behavioral data collection and preprocessing;

[0014] S2, multimodal behavior sequence modeling;

[0015] S3, interest representation generation and recommendation generation;

[0016] S4, user feedback and model update;

[0017] S5. Model optimization and expansion.

[0018] Preferably, in S1, multimodal data acquisition and preprocessing includes:

[0019] S11. Collect multimodal behavioral data: collect text data entered by users on e-commerce platforms, search engines, and social media; collect image data of pictures viewed or uploaded by users on e-commerce platforms and social platforms; collect time data and location data consisting of time tags and location information generated by user behavior;

[0020] Among them, text data includes search history, comments, and social media interactions; image data includes product images and social platform images that users have browsed or interacted with; time data includes the timestamps of user actions; location data includes the user's geographic location information;

[0021] S12. Data preprocessing: Segment the collected text data, remove stop words, and convert it into a low-dimensional vector using the pre-trained BERT model;

[0022] Among them, the vector of each text data is represented as d t is the dimension of text embedding; the image data is extracted using a pre-trained convolutional neural network (CNN) and converted into a fixed-dimensional feature vector d i It is the dimension of graphic features; it converts the time data into periodic features and generates digital codes for location data through positioning information. d i Encoded features for time and place.

[0023] Preferably, in S2, the multimodal behavior sequence modeling includes:

[0024] S21. Construct a multimodal behavior data sequence: Arrange the user's behavior data in chronological order to form a multimodal behavior sequence; wherein the behavior data expression of user u at time t is:

[0025]

[0026] Where, Represents text features, Represents image features, For time and location characteristics;

[0027] S22. Multimodal sequence modeling and temporal context modeling: We use the Transformer-based multimodal self-attention mechanism (MMSA) to model multimodal behavior sequences. We weight the data of each modality and capture the dynamic changes and correlations of each modality in the time series. The formula for calculating the state of interest at each time step t is:

[0028]

[0029] Where, Represents text features, Represents image features, Represents characteristics of time and place;

[0030] S23. Cross-modal dynamic feature fusion: After obtaining the interest state at each time step, a fusion layer is used to dynamically fuse the features of all modalities. The features of each modality are combined through weighted averaging. The weight coefficient automatically adjusts the importance of each modality at a specific time point to form a fused user interest state. The calculation formula for the fused user interest state is:

[0031]

[0032] Where, Represent the self-attention output of text, image and time / place respectively, is the weight coefficient of self-learning,

[0033] Preferably, in S22, the multimodal self-attention mechanism based on the Transformer structure consists of an input mapping layer, a self-attention mechanism layer, a multi-head attention layer and an output layer.

[0034] Preferably, in S22, the specific processing steps of the multimodal self-attention mechanism based on the Transformer structure at each level are:

[0035] Input mapping layer: data for each modality Mapped to the same dimensional space through a linear transformation, we get q t ,v t ,k t , where q t ,v t ,k t They are query, key and value respectively;

[0036] Self-attention mechanism layer: Calculates the similarity between different modalities and calculates the weighted representation of each modality based on the similarity:

[0037]

[0038] Where Q is the query, K is the key, V is the value, and d k is the dimension of the key;

[0039] Multi-head attention layer: Multiple attention heads are used to simultaneously calculate the attention of different modalities and finally stitch them together to obtain the final weighted fusion representation;

[0040] Output layer: Map the output of the multi-head attention to the final interest representation h through a fully connected layer t .

[0041] Preferably, in S3, the steps of interest representation generation and recommendation generation adopt a dynamic context-aware long-term and short-term interest divide-and-conquer modeling method, combined with adaptive time window division and behavior type weighting mechanism, to refine user interest modeling, specifically:

[0042] S31. Long-term and short-term interest modeling:

[0043] S311, long-term interest modeling: The input is the user's long-term behavior sequence, the time window defaults to 30 days, and the time-aware bidirectional LSTM model is used for modeling. The input is the fusion feature h output in step S23. u (t), obtain the BiLSTM hidden state at time step t, introduce the time decay factor based on the BiLSTM hidden state to enhance the contribution of recent behavior and generate the final long-term interest vector h long , the calculation formula is:

[0044]

[0045] w t =exp(-γ·(T current -t));

[0046]

[0047] Where h BiLSTM (t) is the BiLSTM hidden state at time step t, γ is the learnable decay coefficient that controls the time decay rate, T current is the current timestamp;

[0048] S312, short-term interest modeling: The input is the user's recent behavior sequence, the time window defaults to 3 days, and the gated recurrent unit GRU is used for short-term interest modeling. The input is the fusion feature h output in step S23. u(t), obtain the GRU hidden state at time step t, capture the short-term changes in user interest, and generate a short-term interest vector h by weighting the hidden state of the behavior sequence in the aggregation window through the attention mechanism. short , the calculation formula is:

[0049]

[0050] α j =softmax(W a ·h G RU(t));

[0051]

[0052] Where h GRU (t) is the GRU output state at time step t, T short is the set of all time steps within the short time window;

[0053] S313, dynamic weight fusion: long-term interest and short-term interest vectors are weighted and aggregated through the attention mechanism to obtain the final interest vector representation:

[0054] u interesting =α·h long +β·h short ;

[0055] Where α and β are the learned weight coefficients;

[0056] S32, personalized recommendation generation;

[0057] S321, based on the generated user interest representation u interesting , a complex recommendation model is used to generate item recommendations, and the matching degree between the user interest vector and the item feature vector is calculated to predict the rating. The calculation formula is:

[0058]

[0059] In the formula, j is the candidate item, v j is a feature vector of fixed dimension, σ is the sigmoid activation function, which is used to normalize the score to the interval [0,1];

[0060] S322. Improve the mapping accuracy between user interests and item features through multiple fully connected layers or other complex nonlinear mappings;

[0061] S33, Recommended sorting and output: Based on the calculated score Sort all candidate items and output the user's top k recommended items, the recommended item set R u The calculation formula is:

[0062] R u ={r1,r2,...,r k};

[0063] Where r1, r2, ..., r k are the top k recommended items for user u.

[0064] Preferably, in S4, user feedback and model update include:

[0065] S41. Collect user feedback: Collect user feedback data on recommendation results in real time, including user click, skip, purchase, and favorite behaviors;

[0066] S42, Model Adaptive Optimization and Online Learning: Using online learning and incremental learning methods, the LSTM, MMSA, and Fusion Layer modules are trained online using current feedback data. The gradient of the loss function is calculated based on the back-propagation algorithm to update the model parameters. The loss function is minimized to optimize the model. The update formula is:

[0067]

[0068] Where θ t are the parameters of the model, η is the learning rate, L(θ t ) is the loss function;

[0069] S43. Dynamically adjust model parameters: Dynamically adjust model parameters based on behavioral differences among different user groups; among them, the weights of long-term and short-term interest models and multimodal self-attention mechanisms are retrained by retraining the module weights based on user behavior feedback.

[0070] Preferably, in S5, the model optimization and extension is specifically cross-modal alignment optimization, and the cross-modal alignment optimization includes modal space alignment and shared representation learning.

[0071] Preferably, the modal space alignment is to align the feature space of each modality, compare data of different modalities in a unified vector space, and minimize the similarity difference between the modalities.

[0072] Preferably, the shared representation learning is to map the feature vectors of text, image and time / place modalities into the same shared space through a shared representation learning framework, and perform multimodal data fusion in a unified representation space.

[0073] Therefore, the present invention proposes a personalized recommendation method based on multimodal behavior sequence modeling, which has the following beneficial effects:

[0074] The present invention's multimodal self-attention mechanism (MMSA), cross-modal sequence modeling, and long-term / short-term interest modeling strategies effectively capture the multi-dimensional dynamics of user interests. These innovations not only enhance the personalization and accuracy of recommendations, but also ensure the real-time updating and diversity of recommendation results, enabling the recommendation system to perform better in complex multimodal data environments.

[0075] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is a complete flow chart of a personalized recommendation method based on multimodal behavior sequence modeling of the present invention;

[0077] Figure 2 This is a diagram of the overall model structure of a personalized recommendation method based on multimodal behavior sequence modeling of the present invention;

[0078] Figure 3 It is a structural diagram of MMSA of a personalized recommendation method based on multimodal behavior sequence modeling of the present invention;

[0079] Figure 4 This is a framework diagram of a personalized recommendation method based on multimodal behavior sequence modeling in the present invention. DETAILED DESCRIPTION

[0080] To make the technical solutions, advantages, and purposes of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are part of the embodiments of the present invention, not all of them. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0081] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0082] like Figures 1-4 As shown, a personalized recommendation method based on multimodal behavior sequence modeling according to an embodiment of the present invention specifically includes the following steps:

[0083] S1, multimodal behavioral data collection and preprocessing;

[0084] In S1, multimodal data acquisition and preprocessing include:

[0085] S11. Collect multimodal behavioral data: collect text data entered by users on e-commerce platforms, search engines, and social media; collect image data of pictures viewed or uploaded by users on e-commerce platforms and social platforms; collect time data and location data consisting of time tags and location information generated by user behavior;

[0086] Among them, text data includes search history, comments, and social media interactions; image data includes product images and social platform images that users have browsed or interacted with; time data includes the timestamps of user actions; location data includes the user's geographic location information;

[0087] Text data reflects the interests and emotions expressed by users in information exchange; image data reflects users' preferences for visual content;

[0088] S12. Data preprocessing: Segment the collected text data, remove stop words, and convert it into a low-dimensional vector using the pre-trained BERT model;

[0089] Among them, the vector of each text data is represented as d t is the dimension of text embedding; the image data is extracted using a pre-trained convolutional neural network (CNN) and converted into a fixed-dimensional feature vector d i It is the dimension of graphic features; it converts the time data into periodic features and generates digital codes for location data through positioning information. d i Encoding characteristics for time and place;

[0090] Among them, time data includes hours, weeks, and months;

[0091] S2, multimodal behavior sequence modeling;

[0092] Multimodal behavior sequence modeling includes:

[0093] S21. Construct a multimodal behavior data sequence: Arrange the user's behavior data in chronological order to form a multimodal behavior sequence; wherein the behavior data of user u at time t is expressed as:

[0094]

[0095] Where, Represents text features, represents image features, For time and location characteristics;

[0096] S22. Multimodal sequence modeling and temporal context modeling: We use the Transformer-based multimodal self-attention mechanism (MMSA) to model multimodal behavior sequences. We weight the data of each modality and capture the dynamic changes and correlations of each modality in the time series. The formula for calculating the state of interest at each time step t is:

[0097]

[0098] Where, Represents text features, Represents image features, Represents characteristics of time and place;

[0099] S23. Cross-modal dynamic feature fusion: After obtaining the interest state at each time step, a fusion layer is used to dynamically fuse the features of all modalities. The features of each modality are combined through weighted averaging. The weight coefficient automatically adjusts the importance of each modality at a specific time point to form a fused user interest state. The calculation formula for the fused user interest state is:

[0100]

[0101] Where, Represent the self-attention output of text, image and time / place respectively, is the weight coefficient of self-learning,

[0102] In S22, the multimodal self-attention mechanism based on the Transformer structure consists of an input mapping layer, a self-attention mechanism layer, a multi-head attention layer, and an output layer.

[0103] In S22, the specific processing steps of the Transformer-based multimodal self-attention mechanism at each level are as follows:

[0104] Input mapping layer: data for each modality Mapped to the same dimensional space through a linear transformation, we get q t ,v t ,k t , where q t ,v t ,k t They are query, key and value respectively;

[0105] Self-attention mechanism layer: Calculates the similarity between different modalities and calculates the weighted representation of each modality based on the similarity:

[0106]

[0107] Where Q is the query, K is the key, V is the value, and d k is the dimension of the key;

[0108] Multi-head attention layer: Multiple attention heads are used to simultaneously calculate the attention of different modalities and finally stitch them together to obtain the final weighted fusion representation;

[0109] Output layer: Map the output of the multi-head attention to the final interest representation h through a fully connected layer t .

[0110] S3, interest representation generation and recommendation generation;

[0111] In S3, the steps for interest representation generation and recommendation generation adopt a dynamic context-aware divide-and-conquer modeling approach for long- and short-term interests, combined with an adaptive time window partitioning and behavior type weighting mechanism to refine user interest modeling. Specifically:

[0112] S31. Long-term and short-term interest modeling:

[0113] S311, long-term interest modeling: The input is the user's long-term behavior sequence, the time window defaults to 30 days, and the time-aware bidirectional LSTM model is used for modeling. The input is the fusion feature h output in step S23. u (t), obtain the BiLSTM hidden state at time step t, introduce the time decay factor based on the BiLSTM hidden state to enhance the contribution of recent behavior and generate the final long-term interest vector h long , the calculation formula is:

[0114]

[0115] w t =exp(-γ·(T current -t));

[0116]

[0117] Where h BiLSTM (t) is the BiLSTM hidden state at time step t, γ is the learnable decay coefficient that controls the time decay rate, T current is the current timestamp;

[0118] S312, short-term interest modeling: The input is the user's recent behavior sequence, the time window defaults to 3 days, and the gated recurrent unit GRU is used for short-term interest modeling. The input is the fusion feature h output in step S23. u (t), obtain the GRU hidden state at time step t, capture the short-term changes in user interest, and generate a short-term interest vector h by weighting the hidden state of the behavior sequence in the aggregation window through the attention mechanism. short, the calculation formula is:

[0119]

[0120] α j =softmax(W a ·h G RU(t));

[0121]

[0122] Where h GRU (t) is the GRU output state at time step t, T short is the set of all time steps within the short time window;

[0123] S313, dynamic weight fusion: long-term interest and short-term interest vectors are weighted and aggregated through the attention mechanism to obtain the final interest vector representation:

[0124] u interesting =α·h long +β·h short ;

[0125] Where α and β are the learned weight coefficients;

[0126] S32, personalized recommendation generation;

[0127] S321, based on the generated user interest representation u interesting , a complex recommendation model is used to generate item recommendations, and the matching degree between the user interest vector and the item feature vector is calculated to predict the rating. The calculation formula is:

[0128]

[0129] In the formula, j is the candidate item, v j is a feature vector of fixed dimension, σ is the sigmoid activation function, which is used to normalize the score to the interval [0,1];

[0130] S322. Improve the mapping accuracy between user interests and item features through multiple fully connected layers or other complex nonlinear mappings;

[0131] S33, Recommended sorting and output: Based on the calculated score Sort all candidate items and output the user's top k recommended items, the recommended item set R u The calculation formula is:

[0132] R u ={r1,r2,...,r k};

[0133] Where r1, r2, ..., r k are the top k recommended items for user u.

[0134] S4, user feedback and model update;

[0135] User feedback and model updates include:

[0136] S41. Collect user feedback: Collect user feedback data on recommendation results in real time, including user click, skip, purchase, and favorite behaviors;

[0137] S42, Model Adaptive Optimization and Online Learning: Using online learning and incremental learning methods, the LSTM, MMSA, and Fusion Layer modules are trained online using current feedback data. The gradient of the loss function is calculated based on the back-propagation algorithm to update the model parameters. The loss function is minimized to optimize the model. The update formula is:

[0138]

[0139] Where θ t are the parameters of the model, η is the learning rate, L(θ t ) is the loss function;

[0140] S43. Dynamically adjust model parameters: Dynamically adjust model parameters based on behavioral differences among different user groups; among them, the weights of long-term and short-term interest models and multimodal self-attention mechanisms are retrained by retraining the module weights based on user behavior feedback.

[0141] S5. Model optimization and expansion.

[0142] Model optimization and expansion specifically include cross-modal alignment optimization, which includes modal space alignment and shared representation learning.

[0143] Modal space alignment is to align the feature space of each modality, compare the data of different modalities in a unified vector space, and minimize the similarity difference between the modalities;

[0144] Shared representation learning is a learning framework that uses shared representations to map the feature vectors of text, image, and time / place modalities into the same shared space, and to perform multimodal data fusion in a unified representation space.

[0145] Therefore, the present invention provides a personalized recommendation method based on multimodal behavior sequence modeling, which accurately captures the evolution of user interests through dynamic multimodal time series fusion, improves personalized recommendations by dividing long-term and short-term interests into separate modeling, and enhances the real-time and scenario adaptability of recommendations through real-time feedback-driven online optimization, significantly improving the accuracy of the recommendation system and user experience.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A personalized recommendation method based on multimodal behavior sequence modeling, characterized in that: include: S1, multimodal behavioral data collection and preprocessing; S2, multimodal behavior sequence modeling; S3, interest representation generation and recommendation generation; S4, user feedback and model update; S5. Model optimization and expansion.

2. A personalized recommendation method based on multimodal behavior sequence modeling according to claim 1, characterized in that: In S1, multimodal data acquisition and preprocessing include: S11. Collect multimodal behavioral data: collect text data entered by users on e-commerce platforms, search engines, and social media; collect image data of pictures viewed or uploaded by users on e-commerce platforms and social platforms; collect time data and location data consisting of time tags and location information generated by user behavior; Among them, text data includes search history, comments, and social media interactions; image data includes product images and social platform images that users have browsed or interacted with; time data includes the timestamps of user actions; location data includes the user's geographic location information; S12. Data preprocessing: Segment the collected text data, remove stop words, and convert it into a low-dimensional vector using the pre-trained BERT model; Among them, the vector of each text data is represented as d t is the dimension of text embedding; the image data is extracted using a pre-trained convolutional neural network (CNN) and converted into a fixed-dimensional feature vector d i It is the dimension of graphic features; it converts the time data into periodic features and generates digital codes for location data through positioning information. d i Encoded features for time and place.

3. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 1, characterized in that: In S2, multimodal behavior sequence modeling includes: S21. Construct a multimodal behavior data sequence: Arrange the user's behavior data in chronological order to form a multimodal behavior sequence; wherein the behavior data expression of user u at time t is: Where, Represents text features, Represents image features, For time and location characteristics; S22. Multimodal sequence modeling and temporal context modeling: We use the Transformer-based multimodal self-attention mechanism (MMSA) to model multimodal behavior sequences. We weight the data of each modality and capture the dynamic changes and correlations of each modality in the time series. The formula for calculating the state of interest at each time step t is: Where, Represents text features, Represents image features, Represents characteristics of time and place; S23. Cross-modal dynamic feature fusion: After obtaining the interest state at each time step, a fusion layer is used to dynamically fuse the features of all modalities. The features of each modality are combined through weighted averaging. The weight coefficient automatically adjusts the importance of each modality at a specific time point to form a fused user interest state. The calculation formula for the fused user interest state is: Where, Represent the self-attention output of text, image and time / place respectively, is the weight coefficient of self-learning, 4. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 3, characterized in that: In S22, the multimodal self-attention mechanism based on the Transformer structure consists of an input mapping layer, a self-attention mechanism layer, a multi-head attention layer, and an output layer.

5. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 4, characterized in that: In S22, the specific processing steps of the Transformer-based multimodal self-attention mechanism at each level are as follows: Input mapping layer: data for each modality Mapped to the same dimensional space through a linear transformation, we get q t ,v t ,k t , where q t ,v t ,k t They are query, key and value respectively; Self-attention mechanism layer: Calculates the similarity between different modalities and calculates the weighted representation of each modality based on the similarity: Where Q is the query, K is the key, V is the value, and d k is the dimension of the key; Multi-head attention layer: Multiple attention heads are used to simultaneously calculate the attention of different modalities and finally stitch them together to obtain the final weighted fusion representation; Output layer: Map the output of the multi-head attention to the final interest representation h through a fully connected layer t .

6. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 1, characterized in that: In S3, the steps for interest representation generation and recommendation generation adopt a dynamic context-aware divide-and-conquer modeling approach for long- and short-term interests, combined with an adaptive time window partitioning and behavior type weighting mechanism to refine user interest modeling. Specifically: S31. Long-term and short-term interest modeling: S311, long-term interest modeling: The input is the user's long-term behavior sequence, the time window defaults to 30 days, and the time-aware bidirectional LSTM model is used for modeling. The input is the fusion feature h output in step S23. u (t), obtain the BiLSTM hidden state at time step t, introduce the time decay factor based on the BiLSTM hidden state to enhance the contribution of recent behavior and generate the final long-term interest vector h long , the calculation formula is: w t =exp(-γ·(T current -t)); Where h BiLSTM (t) is the BiLSTM hidden state at time step t, γ is the learnable decay coefficient that controls the time decay rate, T current is the current timestamp; S312, short-term interest modeling: The input is the user's recent behavior sequence, the time window defaults to 3 days, and the gated recurrent unit GRU is used for short-term interest modeling. The input is the fusion feature h output in step S23. u (t), obtain the GRU hidden state at time step t, capture the short-term changes in user interest, and generate a short-term interest vector h by weighting the hidden state of the behavior sequence in the aggregation window through the attention mechanism. short , the calculation formula is: a j =softmax(W a ·h G RU(t); Where h GRU (t) is the GRU output state at time step t, W a is the attention weight matrix, used to calculate the attention score, α j Attention weight, indicating the importance of the j-th time step; S313, dynamic weight fusion: long-term interest and short-term interest vectors are weighted and aggregated through the attention mechanism to obtain the final interest vector representation: you interesting =a·h long +β·h short ; Where α and β are the learned weight coefficients; S32, personalized recommendation generation; S321, based on the generated user interest representation u interesting , a complex recommendation model is used to generate item recommendations, and the matching degree between the user interest vector and the item feature vector is calculated to predict the rating. The calculation formula is: In the formula, j is the candidate item, v j is a feature vector of fixed dimension, σ is the sigmoid activation function, which is used to normalize the score to the interval [0,1]; S322. Improve the mapping accuracy between user interests and item features through multiple fully connected layers or other complex nonlinear mappings; S33, Recommended sorting and output: Based on the calculated score Sort all candidate items and output the user's top k recommended items, the recommended item set R u The calculation formula is: R u ={r1,r2,...,r k}; Where r1, r2, ..., r k are the top k recommended items for user u.

7. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 1, characterized in that: In S4, user feedback and model updates include: S41. Collect user feedback: Collect user feedback data on recommendation results in real time, including user click, skip, purchase, and favorite behaviors; S42, Model Adaptive Optimization and Online Learning: Using online learning and incremental learning methods, the LSTM, MMSA, and Fusion Layer modules are trained online using current feedback data. The gradient of the loss function is calculated based on the back-propagation algorithm to update the model parameters. The loss function is minimized to optimize the model. The update formula is: i t =θ t-1 -η▽ θ L(θ t ); Where θ t are the parameters of the model, η is the learning rate, L(θ t ) is the loss function; S43. Dynamically adjust model parameters: Dynamically adjust model parameters based on behavioral differences among different user groups; among them, the weights of long-term and short-term interest models and multimodal self-attention mechanisms are retrained by retraining the module weights based on user behavior feedback.

8. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 1, characterized in that: In S5, model optimization and expansion are specifically cross-modal alignment optimization, which includes modal space alignment and shared representation learning.

9. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 8, characterized in that: The modal space alignment is to align the feature space of each modality, compare data of different modalities in a unified vector space, and minimize the similarity difference between the modalities.

10. The personalized recommendation method based on multimodal behavior sequence modeling according to claim 8, characterized in that: The shared representation learning is to map the feature vectors of text, image and time / place modalities into the same shared space through a shared representation learning framework, and perform multimodal data fusion in a unified representation space.

Citation Information

Patent Citations

  • Serialization recommendation methods based on CNN and RNN

    CN109885756B

Cited By

  • Personalized content recommendation method and system based on multi-modal perception

    CN120974015A

  • A personalized content recommendation method and system based on multi-modal perception

    CN120974015B

  • E-commerce platform commodity recommendation method based on AI intelligence

    CN121052904A

  • Multi-task interest point recommendation method based on multiple modes and time perception

    CN121301672A

  • Multi-round dialogue logic optimization method in intelligent question-answering system based on knowledge graph

    CN121478989A