A Multimodal Sequence Recommendation Method Based on Large Multimodal Models
Through the combination of the visual understanding and generation capabilities of LMM, IVQA, IMSG and USBC components, a multimodal interaction sequence of users is built, and the supervised fine-tuning SFT optimization model with quantitative low-rank adaptation is used to solve the limitations of multimodal information processing in the existing technology, and an accurate understanding and personalized recommendation of user dynamic preferences are achieved.
Patent Information
- Application Number
- CN202411747012.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing recommendation system based on large language models lacks the learning of user-project sequence interaction specific knowledge when processing multimodal information. Especially in the context of continuous multiple, noisy ordered text and image interactions, it is difficult to accurately understand and generate user preferences that change dynamically over time, and ignores the fine-grained content learning and joint fine-tuning of multimodal information, resulting in limitations in accuracy and robustness of recommendation systems.
IVQA, IMSG and USBC components are adopted, combined with the visual understanding and generation capabilities of LMM, and the user's multimodal interaction sequence is constructed through sliding windows and step sizes. The supervised fine-tuning SFT is used to optimize parameters, and the four-step thinking chain prompt learning optimization model is achieved to achieve efficient integration and personalized recommendation of multimodal information.
It improves the accuracy and robustness of the multimodal sequence recommendation system, can better capture the dynamic changes of user interests, enhances the ability to understand and integrate multimodal information, overcomes the preference forgetting or distortion caused by the length of user interaction sequences, and improves the model's response performance and personalized recommendation capabilities.
Smart Images

Figure CN119862318B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a multimodal sequence recommendation method based on a large multimodal model. Background Art
[0002] In recent years, the great success of large language models (LLMs) in the field of natural language processing and other fields has demonstrated excellent semantic understanding and generation capabilities, promoting the rapid development of generative recommendation systems. Due to their rich world knowledge and strong context understanding capabilities, LLMs have shown significant improvements in capturing the dynamic changes in user interest preferences and processing complex item content, which is crucial for enhancing the accuracy and interpretability of sequence recommendations. However, most current recommendation systems based on LLMs rely on extracting deep semantic information from single-modal information generated by users (such as users' historical behaviors, chat records, text comments, etc.), while ignoring the fusion and utilization of other potential modal information. At the same time, with the widespread application of personalized intelligent products, the generation of more sequence multimodal information such as comments and vision, and the increase in data diversity, people's attention to multimodal information has become higher and higher. Technologies such as the fusion and alignment of multimodal information have been adopted to integrate auxiliary data from various sources (text, image, audio, and video), and this information is used to represent and discover the hidden relationships between different modalities and users' preferences in different modalities, and may recover complementary information that cannot be captured by single-modal methods and implicit interactions, which more precisely characterizes user-item information, intentions and goals, and item consumption trends, provides a deeper understanding of the recommendation context, thereby alleviating the data sparsity and cold start problems, resulting in a more accurate, diverse, and personalized multimodal recommendation system (MRS).
[0003] In addition, large multimodal models (LMMs) have demonstrated excellent performance in understanding various types of data such as text, image, audio, and video, and cross-modal generation. The latest progress of LMMs has paved the way for fusing and utilizing multiple modal information and constructing large multimodal generative recommendation models. By jointly learning the data representations of different modalities, LMMs align multimodal information into a unified semantic space, which can not only capture more comprehensive and rich user preference information, but also achieve specific modal generation or cross-modal conversion. This not only improves the overall performance of the recommendation system, but also shows better adaptability and robustness in complex scenarios, promoting the innovation and development of recommendation technologies. Although there are currently individual works that use large visual models or large vision-language models as modal encoders to assist recommendation systems, there is still a considerable amount of unstudied space in directly using LMMs as multimodal sequence recommendation systems.
[0004] However, there are many significant challenges and obstacles when applying LMM to sequential recommendation tasks. First of all, existing LMMs lack the learning of specific knowledge about user-item sequential interactions in the recommendation field. Especially in the context of consecutive, noisy, ordered text and image interactions, they often show limitations in understanding and generating user preferences that change dynamically over time. This key limitation weakens the system's ability to accurately capture and reflect the dynamic personalization of user interests over time. Secondly, LMM mainly understands visual information through image captioning, and rarely learns and accurately understands the finer-grained content of multimodal information, which hinders the accuracy of recommendation results. In addition, different from the fine-tuning strategy of LMM in NLP tasks, the joint fine-tuning of multimodal content such as text and pictures in specific recommendation tasks is ignored, and at the same time, ensuring the model's inference and generalization ability is a serious challenge. How to design specific multimodal optimization and prompting strategies based on LMM to integrate visual knowledge into recommendation tasks to enhance the understanding ability and robustness of multimodal sequential recommendation. Summary of the Invention
[0005] In view of the above deficiencies, the present invention proposes a multimodal sequential recommendation method based on a large multimodal model. In order to learn user dynamic multimodal sequential interaction data, the IVQA, IMSG, and USBC components are proposed based on LMM, and through the efficient parameter fine-tuning technology of SFT, the system has the ability to accurately capture and reflect the dynamic personalization of user interests over time.
[0006] To achieve the above object, the present invention provides the following technical solutions: A multimodal sequential recommendation method based on a large multimodal model, comprising the following steps:
[0007] S1: Construct a dynamic multimodal interaction sequence according to the user's historical records, perform picture understanding IVQA based on visual question answering, and perform individual summarization on each item's multimodal information through item multimodal summarization IMSG;
[0008] S2: Understand the user's dynamic preferences through the user sequence behavior understanding USBC assisted by LLM, and form the user's multimodal interaction sequence preference by understanding each item's multimodal summary in the order of interaction time;
[0009] In step S2, in the order of user interaction time, use the prompt p3 to iteratively model the constructed user dynamic multimodal behavior sequence, where p3 = "Summarize the user's multimodal summary sequence in the order of time to refine the user's interest preference";
[0010] The user's dynamic multimodal behavior sequence is constructed by a sliding window and a step size, and the user's historical interactions are {i1, i2,..., i n}, where any i is composed of (image, text), n is the maximum number of historical interactions. When 2 ≤ n ≤ 5, the user's multimodal summary sequence is:
[0011] [(image1, text1), (image2, text2), …, (image n―1 , text n―1 )]. When n ≥ 6, the user's multimodal summary sequence is:
[0012] [(image n―5 , text n―5 ), (image n―4 , text n―4 ), …, (image n―1 ,
[0013] text n―1 )];
[0014] S3: Combine the obtained user multimodal interaction sequence preference, user configuration information, and multimodal information of the next item to construct instruction fine-tuning data, and optimize the model parameters of the LMM through supervised fine-tuning SFT of quantized low-rank adaptation;
[0015] S4: On this basis, further align the LMM with the sequence recommendation task through chain-of-thought prompting learning to train a multimodal sequence recommendation model based on the LMM.
[0016] As an improvement, in the image understanding IVQA based on visual question answering in step S1, utilize the visual understanding and generation ability of the LMM, and define the IVQA visual summary υs i of any picture image i with the following formal definition:
[0017] υs i = IVQA(p1, image i )
[0018] where p1 = "Describe the content contained in the picture and its required relevant features".
[0019] As an improvement, in step S1, by utilizing the semantic understanding ability of the LLM and designing a joint personalized prompt p2 for the multimodal information of the LLM, combine the visual summary υs i obtained by IVQA and the content text i of the corresponding text modality as the input to achieve the multimodal summary IMSG of the item. The formal definition of the multimodal summary IMSG is:
[0020] ms i= IMSG(p2, υs i , text i )
[0021] where p2 = "Summarize the item content by combining visual summary and text information".
[0022] As an improvement, the instruction data composed of the {instruction, input, response} structure generated according to each user's multimodal interaction history in step S3 is used to form instruction fine-tuning data by simulating the scenario of instruction fine-tuning. Among them, instruction provides a description of the task that the model needs to execute, and the task description is "Predict whether the user likes the next item according to the user's interaction history". Input is the input data that needs to be processed, and the input data is to provide partial multimodal interaction preferences of the user and multimodal information of the next item to simulate actual user behavior data. Response is the recommended result expected by the generation model, and the recommended result is whether the user likes the next item.
[0023] As an improvement, in step S3, the following SFT training objective is used to train the LMM-based recommender, and efficient parameter fine-tuning is performed through quantized low-rank adaptation:
[0024]
[0025] where y i is the i-th word in the prompt text, L is the length of the prompt text, and the probability P(y i |y <i , M) is calculated by the LMM model according to the causal language model next token prediction paradigm.
[0026] As an improvement, if the recommended result is that the user likes the next item, then "yes" is output. If the recommended result is that the user does not like the next item, then "no" is output. The softmax function is used to calculate the probability of item interaction:
[0027]
[0028] As an improvement, in step S4, on the basis of performing the efficient parameter fine-tuning, the multimodal sequence recommendation model is further optimized through four-step thought chain prompt learning. The prompts and the returned answers of each step of IVQA, IMSG, and USBC are used as the prefix of the next step's prompt in turn, and together with the information of the current step as the input, clearly instructing the LMM to generate whether the user likes or dislikes the next target item from the given user configuration, user preferences summarized from historical interactions, and multimodal information of the target item.
[0029] Compared with the prior art, the advantages of the present invention are as follows:
[0030] (1) Through picture understanding based on IVQA, the present invention uses more targeted question prompts at p1 to obtain key information such as the name of the item, its category, type, color, style, brand, specifications, and features, and summarizes them into a unified text description. By leveraging the powerful visual understanding and generation capabilities of the LMM, noise and useless information are avoided, key features and effective semantics in each image are extracted, and the personalized content of the image is deeply understood to provide a detailed response. Precise guidance is crucial for filtering information and capturing elements that are directly consistent with the user's interests, thus laying a foundation for subsequent integration with the Text modality and understanding of user preferences;
[0031] (2) The present invention realizes the multi-modal fusion summary of items that combines the item text modality and the visual modality by leveraging the semantic understanding ability of the LLM, thereby exploring the hidden relationships between different modalities and the user's preference patterns in different modalities. It can learn the deeper complementary semantic relationships between the user and the item in different modalities, understand the user's preferences and item characteristics more comprehensively, and better integrate multi-modal information. For multi-modal item information including image and text content, a large model backbone suitable for different modalities is selected to ensure more detailed fine-grained features are captured and a more thorough understanding is achieved from each modality. Then, specific fusion prompts are used to integrate the independently processed visual summary and text modality together to achieve an all-round and multi-faceted summary and understanding of the item;
[0032] (3) The present invention summarizes and refines the user's historical preferences by adding the LLM-assisted User Sequence Behavior Comprehension USBC. This method takes the user interaction time as the order and uses prompts to iteratively model the constructed dynamic multi-modal sequence behavior, thereby improving cross-sequence context awareness to optimize the current user interest preference. It effectively overcomes the limitations of the LMM in processing multi-modal data sequences and dynamic interest preferences, and thus can more effectively handle multi-modal interaction sequences. By analyzing each interaction in detail to understand the user's preferences, it realizes effective control and learning of multi-modal interaction sequences, avoiding problems such as preference forgetting or distortion due to the long user interaction sequence;
[0033] (4) In the present invention, the open-source large multi-modal model is supervised and fine-tuned SFT to optimize the model parameters to build a multi-modal sequence recommender based on the LMM, which can minimize the difference between the prediction and the actual user interaction and improve the response performance of the model;
[0034] (5) In the present invention, instruction data composed of the {instruction, input, response} structure can be used for instruction fine-tuning to enable the LMM to have the ability to handle multi-modal sequence recommendation tasks;
[0035] (6) During the training process of the present invention, quantization low-rank adaptation is utilized for efficient parameter fine-tuning, thereby significantly reducing the number of trainable parameters and accelerating the training process;
[0036] (7) Based on the efficient parameter fine-tuning, we further optimize the multi-modal sequence recommendation model through the chain-of-thought prompting method in four steps, thereby being able to further enhance the multi-modal understanding and personalized recommendation capabilities of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0038] Figure 1 is a flowchart of a multi-modal sequence recommendation method based on a large multi-modal model;
[0039] Figure 2 is a schematic diagram of the process of the IVQA method for image understanding based on visual question answering;
[0040] Figure 3 is a schematic diagram of the process of the IMSG method for generating item multi-modal summaries;
[0041] Figure 4 is an example diagram of the process of the USBC method for summarizing dynamic user multi-modal interaction sequences and refining preferences;
[0042] Figure 5 is an example diagram of the overall framework diagram of a multi-modal sequence recommendation method based on a large multi-modal model. SPECIFIC EMBODIMENTS
[0043] As Figures 1 to 5 shown, a multi-modal sequence recommendation method based on a large multi-modal model includes the following steps:
[0044] S1.1: Construct and input user configuration information, a user dynamic multi-modal interaction sequence of length N, multi-modal information of the next target item, and positive and negative sample label data generated by negative sampling according to the user's historical records;
[0045] S1.2: In the IVQA for image understanding based on visual question answering, using the visual understanding and generation capabilities of the LMM, define the visual summary υs i of any image image i in the form of:
[0046] υs i = IVQA(p1, image i )
[0047] Among them, p1 = "What is in the picture, and describe its classification, type, color, style, brand, specifications, and main features?";
[0048] S1.3: By leveraging the semantic understanding ability of the LLM and designing a joint personalized prompt p2 for the multimodal information of the LLM, the visual summary υs obtained from IVQA i and the content text of the corresponding text modality i are combined as inputs to achieve the multimodal summary IMSG of the item. The formal definition of the multimodal summary IMSG is:
[0049] ms i = IMSG(p2, υs i , text i )
[0050] Among them, p2 = "Please summarize the item content by combining the visual summary and text information";
[0051] S2: Through the LLM-assisted user sequence behavior understanding USBC, the dynamic user preferences are understood for each item multimodal summary in the order of interaction time. The prompt p3 is used to iteratively model the constructed user dynamic multimodal behavior sequence, where p3 = "Please summarize the user's multimodal summary sequence in the order of time and refine the user's interest preferences". The user dynamic multimodal behavior sequence is constructed by a sliding window and a step size. The user's historical interactions are {i1, i2,..., i n}, and any i is composed of (image, text). n is the maximum number of historical interactions. When 2 ≤ n ≤ 5, the user's multimodal summary sequence is
[0052] [(image1, text1), (image2, text2),..., (image n-1 , text n-1 )],
[0053] When n ≥ 6, the user's multimodal summary sequence is
[0054] [(image n-5 , text n-5 ), (image n-4 , text n-4 ),..., (image n-1 , text n-1 )];
[0055] S3.1: Instruction data consisting of a structure of {instruction, input, response} generated based on each user's multimodal interaction history, where instruction provides a description of the task that the model needs to perform. The task description is "Based on the user's interaction history, please predict whether the user likes the next item." Input is the input data that needs to be processed. The input data provides part of the user's multimodal interaction preferences and the multimodal information of the next item to simulate actual user behavior data. Response is the recommendation result expected by the generated model. The recommendation result is whether the user likes the next item, that is, the label data. This structure is used to simulate the scenario of instruction fine-tuning, and the instruction fine-tuning data is constructed by combining the obtained user multimodal interaction sequence preferences, user configuration information, and the multimodal information and label data of the next item.
[0056] S3.2: Efficient parameter fine-tuning of LMM via supervised fine-tuning of quantized low-rank adaptive SFT:
[0057]
[0058] where y i is the i words in the prompt text, L is the length of the prompt text, and the probability P(y i |y <i , M) is calculated by the LMM model according to the next token prediction paradigm of the causal language model;
[0059] S3.3: If the recommendation result is that the user likes the next item, then output "yes", if the recommendation result is that the user does not like the next item, then output "no", use the softmax function to calculate the probability of item interaction:
[0060]
[0061] S4: On the basis of efficient parameter fine-tuning, the multimodal sequence recommendation model is further optimized through four-step thought chain prompt learning. The prompts and returned answers of each step of IVQA, IMSG, and USBC are used as the prefix of the next step prompt, and together with the information of the current step, it is used as input to explicitly instruct LMM to generate whether the next target item is liked or not from the given user configuration, the user preferences summarized by historical interactions, and the multimodal information of the target item, where:
[0062] Set prompt 1 to "What is in the picture? Please describe its category, type, color, style, brand, specifications and main features?" LMM returns the picture content describing the category, type, color, style, brand, specifications, etc. through visual question answering;
[0063] Set Prompt 2 as "Please summarize the key features of the item by combining visual summaries and text information", and summarize the features that best reflect multiple modalities of the item by combining the answers obtained from the visual understanding of Prompt 1 and the item text;
[0064] Set Prompt 3 as "Please summarize the user's multi-modal summary sequence in chronological order and refine the user's interest preferences", and ensure that the most accurate dynamic interest preferences of the user can be refined and summarized by combining the answers obtained from Prompt 2 and the text;
[0065] Set Prompt 4 as "If you are an expert in multi-modal sequence recommendation, based on the user's profile, the aggregated user preferences, and the picture and text information of the next item, predict whether the user will interact with the next item by simply replying 'Yes' or 'No'", and clarify that LMM generates a like or dislike for the next target item from the given user configuration, the user preferences summarized from historical interactions, and the multi-modal information of the target item by combining the answers obtained from Prompt 3 and the text.
[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0067] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor, and at least part of the functions described above can be performed by one or more hardware logic components.
[0068] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made using the technical solutions of the present invention, or the concepts and technical solutions of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A multimodal sequence recommendation method based on a large multimodal model, characterized in that, The following steps are involved: S1: Build dynamic multimodal interaction sequences based on user history records, image understanding based on visual question answering IVQA, and summarize the multimodal information of each item one by one through item multimodal summary IMSG; S2: Understand USBC through LLM-assisted user sequence behavior, summarize each item multimodally in order of interaction time to understand the user's dynamic preferences, and form the user's multimodal interaction sequence preferences; In step S2, the user interaction time is used as the order, and prompt p3 is used to iteratively model the constructed user dynamic multimodal behavior sequence, where p3 = "summarize the user's multimodal summary sequence in time order and refine the user's interest preference"; The user's dynamic multi-modal behavior sequence is constructed by a sliding window and a step size. The user's historical interactions are {i1, i2, …, i n}, where any i is composed of (image, text), and n is the maximum number of historical interactions. When 2 ≤ n ≤ 5, the user's multi-modal summary sequence is: [(image1,text1),(image2,text2),…,(image n-1 ,text n-1 )], When n≥6, the multimodal summary sequence of the user is: [(image n-5 , text n-5 ), (image n-4 , text n-4 ), …, (image n-1 , text n-1 )]; S3: Combine the obtained user multimodal interaction sequence preferences, user configuration information and multimodal information of the next item to construct instruction fine-tuning data, and optimize the model parameters of LMM through quantized low-rank adaptive supervised fine-tuning SFT; S4: On this basis, LMM is further aligned with the sequence recommendation task through thought chain-based prompt learning to train a LMM-based multimodal sequence recommendation model.
2. The multimodal sequence recommendation method based on a large multimodal model according to claim 1, wherein: In the image understanding based on visual question answering (IVQA) in step S1, the visual understanding and generation capabilities of the LMM are utilized to define any image i The IVQA visual summary vs i The formal definition is as follows: vs i = IVQA(p1, image i ) Where p1 = "describes the content contained in the image and its required related features".
3. The multimodal sequence recommendation method based on a large multimodal model according to claim 2, wherein: In the step S1, by utilizing the semantic understanding ability of the LLM and designing a joint personalized prompt p2 for the multimodal information of the LLM, the visual summary vs obtained by the IVQA i and the content text of the corresponding text modality i are combined as the input to implement the multimodal summary IMSG of the item. The formal definition of the multimodal summary IMSG is as follows: ms i = IMSG(p2, vs i , text i ) Wherein p2 = "combine visual summary and text information to summarize the content of the item".
4. The multimodal sequence recommendation method based on a large-scale multimodal model according to claim 1, wherein: In step S3, instruction data consisting of a structure of {instruction, input, response} is generated according to each user's multimodal interaction history, and instruction fine-tuning data is formed by structurally simulating the instruction fine-tuning scenario, wherein the instruction provides a description of the task that the model needs to perform, and the task description is "predict whether the user likes the next item based on the user's interaction history", and the input is input data that needs to be processed, and the input data provides part of the user's multimodal interaction preference and the multimodal information of the next item to simulate actual user behavior data, and the response is the recommendation result expected by the generated model, and the recommendation result is whether the user likes the next item.
5. The multimodal sequence recommendation method based on a large multimodal model according to claim 4, characterized in that: In step S3, the following SFT training objective is used to train the LMM-based recommender, and efficient parameter fine-tuning is performed through quantized low-rank adaptation: where y i is the i-th word in the prompt text, L is the length of the prompt text, and the probability P(y i |y <i , M) is calculated by the LMM model according to the next-token prediction paradigm of the causal language model.
6. The multimodal sequence recommendation method based on a large multimodal model according to claim 4, wherein: If the recommendation result is that the user likes the next item, then "yes" is output; if the recommendation result is that the user does not like the next item, then "no" is output. The softmax function is used to calculate the probability of item interaction:
7. A multimodal sequence recommendation method based on a large multimodal model according to claim 4, characterized in that: In step S4, on the basis of efficient parameter fine-tuning, the multimodal sequence recommendation model is further optimized through four-step thinking chain prompt learning, and the prompts and returned answers of each step of the IVQA, IMSG, and USBC are used as prefixes of the next step prompts in turn, and together with the information of the current step, are used as inputs, explicitly instructing LMM to generate whether the next target item is liked or not based on the given user configuration, user preferences summarized from historical interactions, and multimodal information of the target item.
Citation Information
Patent Citations
Recommendation method and system based on multi-level comparative learning and multi-modal knowledge graph
CN116091152A
Video hot comment generation method and device oriented to multi-modal application
CN117998156A