An intelligent collaborative recommendation method and device for multi-modal immersive content

By collecting and analyzing multimodal data and using multi-agent system to optimize recommendation strategies, the problem of insufficient personalized content recommendation in the existing technology is solved, and more accurate and efficient reflection of user interests and emotional status is achieved, and the effect of immersive content recommendation is improved.

CN120086453BActive Publication Date: 2025-07-08JIMEI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541128.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-08
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art personalized content recommendation systems in the fields of virtual reality and augmented reality fail to make full use of multimodal data, resulting in the recommendation results being far from users' expectations and cannot meet the actual needs of users.

Method used

Multimodal data of users in the virtual environment, including user behavior, physiological reactions and display feedback, and by calculating attention weights and fusion characteristics between modals, multi-agents are used to determine recommended actions, including content recommendation, experience adjustment, exploration guidance and social connection agents, combining dynamic weight communication networks, interdependence strategy evolution method and hierarchical differentiated reward architecture to optimize recommendation strategies.

Benefits of technology

It improves the accuracy and adaptability of personalized content recommendations, enhances the collaboration efficiency and stability of multi-agent systems in recommended scenarios, and ensures that the recommended content is in line with the user's dynamic interests and emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086453B_ABST
    Figure CN120086453B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent collaborative recommendation method and device for multi-modal immersive content, which relates to the field of immersive interaction technologies. The method includes: collecting multi-modal data and calculating the attention weights of the interaction between itself and other modalities for each modality, obtaining the corresponding fusion features by weighting the representation after internal attention processing of its own modality with the calculated attention weights, and determining the recommendation actions using pre-defined multi-agents according to the obtained fusion features. The present invention solves the problem that the prior art does not comprehensively consider multiple modal data for understanding the user state in a virtual environment and their interaction effects, resulting in the recommended immersive content not meeting the actual needs of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of immersive interaction technologies, and particularly to an intelligent collaborative recommendation method and device for multi-modal immersive content. Background Art

[0002] With the rapid development of technology, virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies have brought revolutionary changes to the game and entertainment industries, and immersive experiences have gradually become the industry standard. However, in terms of personalized content recommendation, for example, in the field of cultural tourism, the application process of these technologies is relatively slow, and there are still many problems to be solved.

[0003] Currently, when different industries such as the cultural tourism industry provide users with immersive experiences, the degree of personalization is seriously insufficient. Content recommendation systems such as cultural tourism recommendation systems usually rely too much on simple user portraits and historical behavior data, and it is difficult to deeply explore users' real interests and needs. This situation not only makes the integration of immersive content and real cultural knowledge remain on the surface, but also causes the system to respond severely lagged to the dynamic interest changes of users. In immersive scenarios such as immersive cultural tourism scenarios, rich multi-modal interaction data, such as users' voice commands, gesture actions, line-of-sight focus, physiological reactions, etc., contains a large amount of information that can reflect users' real-time interests, but traditional recommendation systems cannot make full use of multi-modal data resources, resulting in a large gap between the recommended results and users' expectations. Summary of the Invention

[0004] Embodiments of the present invention provide an intelligent collaborative recommendation method and device for multi-modal immersive content, aiming to solve the problem that the existing technology does not comprehensively consider multiple modal data used to understand the user state in a virtual environment and their interaction effects, resulting in the recommended immersive content not meeting the actual needs of users.

[0005] To achieve the above object, in a first aspect, the present invention provides an intelligent collaborative recommendation method for multi-modal immersive content, including the following steps:

[0006] Collect multi-modal data associated with multiple modalities of a user in a virtual environment, where the multiple modalities include: user behavior, physiological reaction, and display feedback;

[0007] For each modality, calculate the attention weight of the interaction between the modality itself and other modalities, and obtain the corresponding fusion feature by weighting the representation processed by the internal attention of the modality itself with the calculated attention weight, where the fusion feature corresponding to the user behavior is the user interest feature, the fusion feature corresponding to the physiological reaction is the emotional state feature, and the fusion feature corresponding to the display feedback is the group interaction tendency feature;

[0008] Based on the obtained fused features, predefined multi - agents are used to determine the recommended actions, where the multi - agents include:

[0009] A content recommendation agent, which is used to calculate the first matching degree between the user interest feature and the candidate cultural and tourism content, and when the first matching degree meets a predetermined first threshold, recommend personalized content based on the user interest;

[0010] An experience adjustment agent, which is used to calculate the second matching degree between the emotional state feature and the preset ideal emotional state, and adjust the visual effect and / or volume of the recommended content display according to the relationship between the second matching degree and a predetermined second threshold;

[0011] An exploration guidance agent, which is used to calculate the third matching degree between the user interest feature and the un - recommended content when the number of times of the current recommended content exceeds a preset threshold, and when the third matching degree meets a third threshold, recommend the un - recommended content;

[0012] A social connection agent, which is used to recommend groups with the same interests according to the group interaction tendency feature.

[0013] Furthermore, the data of the user behavior includes: the movement trajectory represented by three - dimensional space coordinates, movement speed, and staying time; the line - of - sight focus represented by the fixation point coordinates, fixation duration, and saccade path; and the operation record represented by the click position and operation frequency;

[0014] The data of the physiological reaction includes the heart rate signal represented by the time - series heart rate value and heart rate variability, the electroencephalogram signal represented by the power spectral density of the α wave, β wave, and θ the wave power spectral density, and the electromyogram signal represented by the muscle activation degree;

[0015] The data of the display feedback includes voice commands, satisfaction ratings, and text evaluations.

[0016] Furthermore, after collecting the multi - modal data, there are steps of feature extraction and standardization processing for the modal data, including:

[0017] For the non - time - series data in the multi - modal data, feature extraction is performed through the following formula:

[0018] ,

[0019] In the formula, represents the mapped feature of the th modal data ; tanh is the activation function; is the feature extraction weight matrix, which is used to transform the original modal data Linear transformation to a high-dimensional feature space; Is the feature bias term; Used to generate the input of the sigmoid activation function; Is the gating bias term;

[0020] For the time-series data in the multi-modal data, feature extraction is performed through the following formula:

[0021] ,

[0022] In the formula ,X time Is the original time-series data of the modality X raw The time-series feature of; FFT is the fast Fourier transform; Wavelet is the wavelet transform; α Is the fusion weight coefficient;

[0023] Among them, it also includes noise reduction and filtering of the original time-series data using the following formula:

[0024] ,

[0025] In the formula, X filtered Is the original time-series data after noise reduction; Mask is the noise recognition function based on the threshold ; Represents the Hadamard product.

[0026] Furthermore, the attention weights are calculated through the following formula:

[0027] ,

[0028] In the formula, Represents the th modality and the th modality interaction attention weight; Is the th modality and the th modality specific Query projection matrix between; Is the th modality and the th modality specific Key projection matrix between; Represents the th modality and the th modality correlation matrix between; And Are the representations after internal attention processing of the th modality and the th modality respectively, and is the The representation of the input data of one modality after encoding is understood as abstracting the input data into a high-dimensional space; where K d K is the dimension of the Key projection matrix.

[0029] Furthermore, the fused feature corresponding to one modality is calculated by the following formula:

[0030] ,

[0031] ,

[0032] In the formula, H fused is the fused feature vector representation of all modalities; v interest is the user interest feature; v emotion is the emotional state feature; v social is the group interaction tendency feature; represents the fused feature of the th modality. When the th modality is user behavior, is the user interest feature; when the th modality is physiological reaction, is the emotional state feature; when the th modality is display feedback, is the group interaction tendency feature; is the dynamically calculated modality weight, which is activated by the sigmoid activation function for the cosine similarity between the th modality and the th modality; is the projection parameter.

[0033] Furthermore, the multi-agents also optimize the recommended content through a cooperation mechanism, and the cooperation mechanism includes: a dynamic weight communication network, an interdependent strategy evolution method, and / or a hierarchical differential reward architecture, where,

[0034] The dynamic weight communication network calculates the communication dynamic weights between agents according to the task correlation degree and determines the communication path with the highest communication score;

[0035] The interdependent strategy evolution method optimizes the recommendation strategy by evaluating the influence of the recommended content updates among the multi-agents;

[0036] The hierarchical differential reward architecture evaluates the performance of the multi-agents in the virtual environment through a multi-level reward mechanism.

[0037] Furthermore, the communication score is calculated by the following formula:

[0038] ,

[0039] ,

[0040] ,

[0041] ,

[0042] ,

[0043] ,

[0044] wherein, M j is the communication score of the j rd agent; is the communication weight; is the non-linearly transformed message sent by the th agent to the j th agent; represents the recommended action of the th agent; is the activation function; is the weight matrix for adjusting the original message u ; is the bias term; is the feature extracted by the agent according to the fused feature and context content , which is the state representation of the agent; represents connecting the state and recommended action of the th agent, so that the transmitted information contains the complete decision context; is the agent-specific non-linear mapping function, including M attention heads; is the Kronecker product; is the Hadamard product; GELU and Swish are activation functions; is the output projection matrix; is the Query projection matrix of the t th attention head; is the Key projection matrix of the t th attention head; is the Value projection matrix of the t th attention head; and are the weight and bias respectively; dK is the dimension of the Key projection matrix; Indicates Agent and j The task correlation between agents; ELU is the exponential linear unit; Indicates Agent and j The cosine similarity of the state representations between agents; v and is a learnable parameter, v Used to measure the importance of state features between agents to the communication weight, is the state projection matrix; T is the transpose.

[0045] It should be noted that Express The value of is not j All the results are accumulated.

[0046] Furthermore, the recommendation strategy objective function is:

[0047] ,

[0048] In the formula, Indicates that all agents follow the joint strategy After running, the global state of the agent s and joint actions a In following Under distribution, the agent The average reward expected to be obtained; represents the joint strategy of all agents; Indicates that except The joint strategy of all agents except the agent; s is the global environment state of the agent; Indicates that except Recommended actions for agents other than the agent; is the state-action distribution; β is a parameter that adjusts the constraint strength; To limit the current policy With reference strategy KL divergence of differences; Indicates Agents in the global state s Take action , while other agents take joint actions Instant rewards received;

[0049] The gradient formula used for strategy update is:

[0050] ,

[0051] ,

[0052] ,

[0053] ,

[0054] In the formula, is the gradient function; is the communication score of the th agent; is the policy parameter of the th agent; is the interdependent advantage function of the th agent considering other agents except itself; represents the representation state - recommended action value function of the th agent; represents the representation state - recommended action value function of the j th agent; is the state value function of the th agent; is the state value function of the j th agent; is the interdependent weight between the th agent and the j th agent; MLP is a multi - layer perceptron; measures the average gradient influence of the policy change of the th agent on the value function of the j th agent; is the correlation between the value functions of the th agent and the j th agent.

[0055] It should be noted that, represents the sum of all results where the value of j is not .

[0056] Furthermore, the multi - level reward mechanism includes: professional task reward, interactive collaboration reward and / or overall experience reward, where,

[0057] The professional task reward is used to evaluate the performance score of the agent in its professional field, and the calculation formula is as follows:

[0058] ,

[0059] ,

[0060] ,

[0061] wherein, represents the professional task reward; is the recommended content and the user vector the correlation between them, is the corresponding balance parameter; the recommended content relative to the user's historical vector diversity, is the corresponding balance parameter; represents the user vector, including the user's interests and historical behavior information; represents the user vector of past history, h represents each record in the set;

[0062] The interaction and collaboration reward is used to evaluate the performance score of an agent under the influence of the recommended actions of other agents. The calculation formula is as follows:

[0063] ,

[0064] wherein, is the interaction and collaboration reward; is the th agent and the j th agent's collaboration coefficient; represents the value function under the global state s where the j th agent takes the recommended action i a i , and at the same time, the other agents except the th agent take the joint action a -i function; represents the value function under the global state s where the j th agent takes the benchmark action , and at the same time, the other agents except the th agent take the joint action function; is the benchmark action; is the recommended action of the agents except the th agent;

[0065] The overall experience reward is used to evaluate the overall performance score of the multi-agent recommendation action, and the calculation formula is as follows:

[0066] ,

[0067] In the formula, is the overall experience reward, f global is a non-linear mapping function that integrates multiple metrics into a single global score; satisfaction is the user satisfaction; duration is the residence time of the user in the current recommended content; revisit_intention represents the intention to revisit;

[0068] The comprehensive reward of the agent is:

[0069] ,

[0070] In the formula, SUM i is the comprehensive reward function of the th agent; , and are adaptive weights, which are dynamically adjusted according to the system state and learning stage. The adjustment calculation formula is:

[0071] ,

[0072] In the formula, is the learning rate; is the gradient function; is the weight optimization objective function; t represents time, t -1 is the previous time state.

[0073] In a second aspect, the present invention provides an electronic device, which is characterized by including a memory and a processor. The memory stores at least one program, and the at least one program is executed by the processor to implement the intelligent collaborative recommendation method for multi-modal immersive content as described above.

[0074] The above technical solution has the following technical effects:

[0075] The technical solution of the embodiment of the present invention collects multi-modal data and calculates the interaction attention between itself and other modalities for each modality. The corresponding fusion features are obtained by weighting the representation processed by the internal attention of its own modality with the calculated attention weights. According to the obtained fusion features, predefined multi-agents are used to determine the recommended actions. The present invention solves the problem that the prior art does not comprehensively consider multiple modal data for understanding the user state and their interaction effects in the virtual environment, resulting in the immersive content recommended not meeting the actual needs of users.

[0076] In a further embodiment, by constructing a dynamic weight communication network between multi-agents, calculating the communication dynamic weights between agents according to the task correlation degree, and determining the communication path with the highest communication score, the agents can preferentially receive the most relevant information, improving the adaptability and cooperation efficiency of the multi-agent system in the recommendation scenario.

[0077] In a further embodiment, by evaluating the influence of the recommended content updates between multi-agents on each other, and updating the strategies for each agent with the strategies of other agents as constraint conditions, it is ensured that the strategy update can not only improve the performance but also will not deviate too far and cause system instability.

[0078] In a further embodiment, by comprehensively measuring the performance of agents in a multi-objective environment from multiple levels of professional field, interactive collaboration and overall experience, accurately reflecting the collaborative contributions between agents, and realizing the accurate evaluation and guidance of the performance of agents. Brief Description of the Drawings

[0079] Figure 1 It is a schematic flow chart of an intelligent collaborative recommendation method for multi-modal immersive content according to an embodiment of the present invention;

[0080] Figure 2 It is a schematic flow chart of immersive content recommendation after establishing a preset collaboration mechanism when using the intelligent collaborative recommendation method for multi-modal immersive content according to an embodiment of the present invention;

[0081] Figure 3 It is a schematic structural diagram of an electronic device for immersive content recommendation based on multi-modal data interaction according to an embodiment of the present invention. Detailed Embodiment

[0082] To further illustrate each embodiment, the present invention provides drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used to explain the operating principles of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.

[0083] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments.

[0084] Embodiment 1:

[0085] Figure 1 The following is a schematic flowchart of an intelligent collaborative recommendation method for multi-modal immersive content according to an embodiment of the present invention. The method of this embodiment includes the following steps:

[0086] Collect multi-modal data associated with multiple modalities of a user in a virtual environment. The multiple modalities include: user behavior, physiological response, and display feedback. In a specific implementation, the data of user behavior includes: a movement trajectory represented by three-dimensional spatial coordinates, movement speed, and residence time; a line-of-sight focus represented by a fixation point coordinate, fixation duration, and saccade path; and, an operation record represented by a click position and operation frequency. The data of physiological response includes a heart rate signal represented by a sequential heart rate value and heart rate variability, and an α wave, β wave, and θ electroencephalogram signal represented by the power spectral density of the

[0087] For each modality, calculate the attention weights of the interaction between the modality itself and other modalities, and obtain the corresponding fusion features by weighting the representation after internal attention processing of the modality itself with the calculated attention weights. Among them, the fusion feature corresponding to user behavior is the user interest feature, the fusion feature corresponding to physiological response is the emotional state feature, and the fusion feature corresponding to display feedback is the group interaction tendency feature. In a specific implementation, the user interest feature is mainly obtained through the synergistic effect of user behavior data and physiological response data. The core lies in the temporal correspondence and complementarity of the two modalities of data: user behavior data records the specific actions of users in the experience, reflecting which content or activities users show attention to, while physiological response data captures the emotional reactions of users when performing these actions. Since the physiological response data corresponds one by one with the user behavior data in time, the system can infer the degree of preference of users for relevant scenarios or recommended content by analyzing the physiological changes of users when performing a certain action at a certain moment. The emotional state feature is mainly based on physiological response data, but is cross-fused with user behavior data and explicit feedback data. User behavior data can provide context for physiological response data. For example, if a user shows an accelerated heart rate in a certain horror scene, and the user behavior data also shows that they quickly leave the area, it may be inferred that they are in a state of fear. Explicit feedback data provides users' subjective expressions of emotions. For example, a user says "This is so interesting" or gives a five-star rating, indicating a positive emotion. Combining it with physiological response data can verify and strengthen the inference of the emotional state. The group interaction tendency feature is mainly based on explicit feedback data, and is cross-fused with user behavior data and physiological response data. Explicit feedback data directly reflects users' attitudes towards social activities. User behavior data provides the actual performance of users in a social environment. For example, whether a user actively approaches crowded areas or participates in multi-person interaction tasks reflects objective evidence of their social behavior. Physiological response data reveals the emotional reactions of users in social situations. By clustering users based on the comprehensive three types of data, the system can judge the degree of gregariousness or isolation of users among the same group of people.

[0088] According to the obtained fusion features, use a pre-defined multi-agent to determine the recommended actions. Among them, the multi-agent includes:

[0089] A content recommendation agent, which is used to calculate the first matching degree between the user interest feature and the candidate cultural and tourism content, and when the first matching degree meets a predetermined first threshold, recommend personalized content based on the user's interest;

[0090] An experience adjustment agent, which is used to calculate the second matching degree between the emotional state feature and the preset ideal emotional state, and adjust the visual effect and / or volume of the recommended content display according to the relationship between the second matching degree and a predetermined second threshold;

[0091] An exploration and guidance agent is used to calculate the third matching degree between the user interest feature and the un-recommended content when the number of times of the current recommended content exceeds a preset threshold, and recommend the un-recommended content when the third matching degree meets the third threshold.

[0092] A social connection agent is used to recommend groups with the same interests according to the group interaction tendency feature.

[0093] Furthermore, after collecting the multi-modal data, it also includes the steps of feature extraction and standardization processing of the modal data, including:

[0094] For the non-temporal data in the multi-modal data, feature extraction is performed through the following formula:

[0095] ,

[0096] In the formula, represents the mapping feature of the th modal data ; tanh is the activation function; is the feature extraction weight matrix, which is used to linearly transform the original modal data to the high-dimensional feature space; is the feature bias term; is used to generate the input of the sigmoid activation function; is the gating bias term;

[0097] For the temporal data in the multi-modal data, feature extraction is performed through the following formula:

[0098] ,

[0099] In the formula ,X time is the temporal feature of the modal original temporal data X raw ; FFT is the fast Fourier transform; Wavelet is the wavelet transform; is the fusion weight coefficient;

[0100] Among them, it also includes noise reduction filtering of the original temporal data using the following formula:

[0101] ,

[0102] In the formula, X filtered is the original temporal data after noise reduction; Mask is the noise recognition function based on the threshold ; represents the Hadamard product.

[0103] Further, the attention weights are calculated by the following formula:

[0104] ,

[0105] In the formula, represents the attention weight of the interaction between the -th modality and the -th modality; is the specific Query projection matrix between the -th modality and the -th modality; is the specific Key projection matrix between the -th modality and the -th modality; represents the correlation matrix between the -th modality and the -th modality; and are the representations after internal attention processing of the -th modality and the -th modality respectively; d K is the dimension of the Key projection matrix.

[0106] Further, the fused feature corresponding to a modality is calculated by the following formula:

[0107] ,

[0108] ,

[0109] In the formula, H fused is the fused feature vector representation of all modalities; v interest is the user interest feature; v emotion is the emotional state feature; v social is the group interaction tendency feature; represents the fused feature of the -th modality. When the -th modality is user behavior, is the user interest feature; when the -th modality is physiological reaction, is the emotional state feature; when the -th modality is display feedback, is the group interaction tendency feature; is the dynamically calculated modality weight, which is obtained by applying the sigmoid activation function to the -th modality and the obtained by activating the cosine similarity between modalities; is the projection parameter.

[0110] It should be noted that denotes the value of which is not accumulates all the results of.

[0111] Furthermore, the multi - agents also optimize the recommended content through a collaboration mechanism, which includes: a dynamic weight communication network, an interdependent strategy evolution method, and / or a hierarchical differential reward architecture. Among them,

[0112] The dynamic weight communication network attempts to solve the existing problems, that is, traditional multi - agent systems often adopt fixed - structure or simple - average communication methods, which cannot adapt to the dynamically changing relationship importance among agents in different scenarios. Agents in the dynamic weight communication network not only need to decide what information to transmit, but also need to clarify to whom to transmit the information and how to weigh information from different sources. The message sent from the th agent to the j th agent is generated through a special non - linear transformation. The dynamic weight communication network calculates the communication dynamic weights between agents according to the task correlation degree and determines the communication path with the highest communication score;

[0113] Furthermore, the communication score is calculated through the following formula:

[0114] ,

[0115] ,

[0116] ,

[0117] ,

[0118] ,

[0119] ,

[0120] In the formula, M j is the communication score of the j th agent; is the communication weight; is the th agent's message sent to the j th agent after non - linear transformation; denotes the th agent's recommended action; is the activation function; is to adjust the original message The weight matrix; is the bias term; is the feature extracted by the agent according to the fused feature and context content and is the state representation of the agent; denotes connecting the state and recommended action of the -th agent, so that the transmitted information contains the complete decision context; is the agent-specific non-linear mapping function, which contains M attention heads; is the Kronecker product; is the Hadamard product; GELU and Swish are activation functions; is the output projection matrix; is the Query projection matrix of the t -th attention head; is the Key projection matrix of the t -th attention head; is the Value projection matrix of the t -th attention head; and are the weight and bias respectively; d K is the dimension of the Key projection matrix; denotes the task correlation degree between the -th agent and the j -th agent; ELU is the exponential linear unit; denotes the cosine similarity of the state representations between the -th agent and the j -th agent; and are learnable parameters, which are used to measure the importance of the state features between agents to the communication weight, is the state projection matrix; T is the transpose.

[0121] The unique value of the dynamic weight communication network lies in its ability to automatically form an optimal communication path according to the task progress. For example, in the content exploration stage, the information of the exploration guiding agent may obtain a higher weight, while when the user engagement decreases, the information of the experience adjustment agent becomes more important. This dynamic adjustment mechanism greatly improves the adaptability and cooperation efficiency of the multi-agent system in the cultural and tourism recommendation scenario.

[0122] The interdependent strategy evolution method optimizes the recommendation strategy by evaluating the influence of the recommended content updates among multi-agents on each other;

[0123] Furthermore, traditional multi-agent reinforcement learning methods often update agent policies in a completely independent or simply additive manner, ignoring the complex interdependent relationships among agent decisions. In a recommendation scenario, the decision of one agent may significantly affect the performance of other agents. For example, content selection affects user engagement, and engagement in turn affects content preference. To address this challenge, this embodiment proposes the Interdependent Policy Evolution Method, which is an innovative optimization method that takes into account the mutual influence of policies among agents.

[0124] The core idea of the Interdependent Policy Evolution Method is to regard the policy update of each agent as a conditional optimization problem, that is, to optimize its own policy under the condition that the policies of other agents are fixed, and at the same time, by introducing a policy stability constraint, to prevent drastic fluctuations during the optimization process.

[0125] Among them, the policy objective function is:

[0126] ,

[0127] This objective function consists of two key parts: the first part is the expected reward obtained under the joint policy of all agents where denotes the joint policy of all agents except the -th agent, and is the state-action distribution; the second part is the policy stability constraint, which restricts the difference between the current policy and the reference policy through the KL divergence , and β is a parameter that adjusts the constraint strength. This design ensures that the policy update can not only improve performance but also not deviate too far to cause system instability; denotes the immediate reward obtained when the -th agent takes action s in the global state , while other agents take joint action .

[0128] The gradient formula used for policy update is:

[0129] ,

[0130] ,

[0131] ,

[0132] ,

[0133] In the formula, Indicates that all agents follow the joint strategy After running, the global state of the agents s And the joint action a Under the following Distribution, the agents Expected average reward to be obtained; Is the communication score of the th agent; Is the th agent's policy parameters; Is the th agent's interdependent advantage function considering other agents except itself, which not only considers the agent's own advantages but also incorporates the impact of its decisions on other agents; Represents the th agent's representation state - recommended action value function; Represents the j th agent's representation state - recommended action value function; Is the th agent's state value function; Is the j th agent's state value function; Is the interdependent weight between the th agent and the j th agent, indicating the degree of influence of the th agent's decision on the j th agent's goal; MLP is a multi - layer perceptron; Measures the average gradient impact of the policy change of the th agent on the value function of the j th agent; Is the th agent and the j th agent's value function correlation. This method of calculating interdependent weights can automatically identify the key dependency relationships between agents and dynamically adjust them during the training process.

[0134] Furthermore, in the immersive recommendation scenario, the criteria for evaluating agent performance are diverse, including both the quality of professional task completion and the contribution to the overall experience. Traditional single - reward functions are difficult to comprehensively measure agent performance in a multi - objective environment, especially unable to accurately reflect the collaborative contributions between agents. To address this challenge, the hierarchical differential reward architecture in this embodiment evaluates the performance of multi - agents in a virtual environment through a multi - level reward mechanism.

[0135] Furthermore, the multi - level reward mechanism includes: professional task reward, interaction and collaboration reward, and / or overall experience reward, where

[0136] For professional task rewards, "professional" means that different agents focus on different professional fields. For example, content agents focus on the recommended content, and experience agents focus on the user's emotional state and engagement, etc. Professional task rewards are used to evaluate the performance scores of agents in their professional fields, and the calculation formula is as follows:

[0137] ,

[0138] ,

[0139] ,

[0140] In the formula, represents the professional task reward; is the recommended content c and the user vector the correlation between them, is the corresponding balance parameter; the recommended content c relative to the user's historical vector the diversity of, is the corresponding balance parameter; represents the user vector, which contains the user's interests and historical behavior information; represents the user vector of past history, h represents each record in the set;

[0141] The interaction and cooperation reward specifically evaluates how the decision of one agent enhances or weakens the effects of other agents. Different from the traditional method of simply adding individual rewards, the interaction and cooperation reward precisely quantifies the mutual influence between agents through counterfactual analysis, and is used to evaluate the performance score of an agent under the influence of the recommended actions of other agents. The calculation formula is as follows:

[0142] ,

[0143] In the formula, is the interaction and cooperation reward; is the th agent and the j th agent's cooperation coefficient; represents the value function under the global state s when the j th agent takes the recommended action while the other agents except the th agent take the joint action and the value function under the action of; represents the value function under the global state represents the global states Under the condition that the j -th agent takes the benchmark action at the same time, the other agents except the -th agent take the joint action under the action of the value function; is the benchmark action; is the recommended action for agents other than the -th agent; In a specific implementation, if the content recommendation agent recommends an exhibit that highly attracts the user's attention, enabling the experience adjustment agent to more effectively adjust the immersion parameters, the content recommendation agent will obtain a positive interaction and collaboration reward. This design encourages agents to not only optimize their own goals but also consider how to create favorable conditions for other agents.

[0144] The overall experience reward is a global reward signal shared by all agents, evaluating the quality of the overall recommendation experience, including indicators such as user satisfaction, stay time, and revisit intention, and is used to evaluate the overall performance score of the multi-agent recommendation action. The calculation formula is as follows:

[0145] ,

[0146] In the formula, is the overall experience reward, f global is a non-linear mapping function that integrates multiple indicators into a single global score; satisfaction is user satisfaction; duration is the stay time of the user in the current recommended content; revisit_intention represents the revisit intention;

[0147] The comprehensive reward of the agent is:

[0148] ,

[0149] In the formula, SUM i is the comprehensive reward function of the -th agent; , and are adaptive weights, dynamically adjusted according to the system state and learning stage. The adjustment calculation formula is:

[0150] ,

[0151] In the formula, is the learning rate; t represents time, t -1 is the previous time state; is the weight optimization objective function to ensure that the system can automatically find the optimal reward balance point. For example, at the initial stage of learning, more attention may be paid to professional task rewards to build basic capabilities, while as learning progresses, the weight of collaborative rewards will gradually increase to promote agent cooperation.

[0152] Embodiment 2:

[0153] Figure 2 It is a schematic flowchart of immersive content recommendation after establishing a preset collaboration mechanism when using the intelligent collaborative recommendation method for multi-modal immersive content in an embodiment of the present invention for recommendation. In this embodiment, obtaining the multi-modal data fusion feature and establishing the multi-agent collaboration mechanism includes:

[0154] Collect multi-modal data associated with multiple modalities of the user in the virtual environment, and the multiple modalities include: user behavior, physiological reaction, and display feedback;

[0155] For each modality, calculate the attention weight of the interaction between this modality and other modalities, and obtain the corresponding fusion feature by weighting the representation after internal attention processing of the self-modality with the calculated attention weight;

[0156] After obtaining the fusion feature, the multi-agents determine the recommendation actions of the immersive content according to the fusion feature. However, the recommendation actions between different agents affect and restrict each other, which often leads to a reduction in the recommendation effect. To ensure the accuracy and efficiency of the recommendation actions, it is necessary to continuously improve the collaboration ability between the agents. Therefore, the multi-agents in this embodiment perform immersive content recommendation through a preset collaboration mechanism, and the preset collaboration mechanism includes: a dynamic weight communication network, an interdependent strategy evolution method, and / or a hierarchical differential reward architecture, where,

[0157] The dynamic weight communication network calculates the communication dynamic weight between the agents according to the task relevance, and determines the communication path with the highest communication score;

[0158] The interdependent strategy evolution method optimizes the recommendation strategy by evaluating the influence of the update of the recommended content between the multi-agents on each other;

[0159] The hierarchical differential reward architecture evaluates the performance of the multi-agents in the virtual environment through a multi-level reward mechanism.

[0160] Such as Figure 2 , after the preset collaboration mechanism, the agents of each specialty first calculate the communication dynamic weight between the agents according to the task relevance through the dynamic weight communication network based on the fusion feature and their own states, determine the communication path with the highest communication score, and according to the communication path, each agent exchanges information to establish a comprehensive understanding of the current virtual situation;

[0161] Based on the understanding of the current virtual scenario, multiple agents make collaborative decision-making to recommend immersive content according to the policy objective function of the interdependent strategy evolution method, considering the expected rewards under the joint strategy of multiple agents and the mutual constraints between agents.

[0162] The system integrates the decisions of each agent into a unified recommendation plan and presents it to the user through an immersive interface.

[0163] During the interaction between the user and the recommended content, the system continuously collects feedback data and evaluates the professional task rewards, interaction collaboration rewards, and / or overall experience rewards of each agent through a hierarchical differential reward architecture, and then obtains a comprehensive reward by weighting the adaptive weights of the three rewards.

[0164] Based on the collected feedback data and evaluation results, the system then updates the recommendation strategies of each agent through the interdependent strategy evolution method using the interdependent advantage function while considering the influence of other agents.

[0165] Embodiment 3:

[0166] Figure 3 This is a schematic structural diagram of an electronic device for immersive content recommendation based on multimodal data interaction in an embodiment of the present invention. As Figure 3 shown, the device includes a processor 301, a memory 302, a bus 303, and a computer program stored in the memory 302 and executable on the processor 301. The processor 301 includes one or more processing cores. The memory 302 is connected to the processor 301 through the bus 303. The memory 302 is used to store program instructions. When the processor executes the computer program, it implements the steps in the above method embodiment of Embodiment 1 of the present invention.

[0167] Furthermore, as an executable solution, the electronic device for immersive content recommendation based on multimodal data interaction can be a computer unit, which can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer unit may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above composition structure of the computer unit is only an example of the computer unit and does not constitute a limitation on the computer unit. It may include more or fewer components than the above, or combine some components, or different components. For example, the computer unit may further include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make limitations in this regard.

[0168] Further, as an executable solution, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer unit and connects all parts of the computer unit using various interfaces and lines.

[0169] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer unit by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0170] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in form and detail without departing from the spirit and scope of the present invention defined by the appended claims, and all such changes are within the protection scope of the present invention.

Claims

1. An intelligent collaborative recommendation method for multi-modal immersive content, characterized in that, Including the following steps: Collect multi-modal data associated with multiple modalities of the user in a virtual environment, where the multiple modalities include: user behavior, physiological reaction, and display feedback; For each modality, calculate the attention weight of the interaction between this modality and other modalities, and obtain the corresponding fusion feature by weighting the representation after internal attention processing of the self-modality with the calculated attention weight, where the fusion feature corresponding to the user behavior is the user interest feature, the fusion feature corresponding to the physiological reaction is the emotional state feature, and the fusion feature corresponding to the display feedback is the group interaction tendency feature; According to the obtained fusion features, use a predefined multi-agent to determine the recommended action, where the multi-agent includes: A content recommendation agent, which is used to calculate the first matching degree between the user interest feature and the candidate cultural and tourism content, and when the first matching degree meets a predetermined first threshold, recommend personalized content based on the user's interest; An experience adjustment agent, which is used to calculate the second matching degree between the emotional state feature and a preset ideal emotional state, and adjust the visual effect and / or volume of the recommended content display according to the relationship between the second matching degree and a predetermined second threshold; An exploration guidance agent, which is used to calculate the third matching degree between the user interest feature and the un-recommended content when the number of times of the current recommended content exceeds a preset threshold, and when the third matching degree meets the third threshold, recommend the un-recommended content; A social connection agent, which is used to recommend groups with the same interests according to the group interaction tendency feature; The attention weight is calculated by the following formula: , In the formula, represents the attention weight for the interaction between the -th modality and the -th modality; is the specific Query projection matrix between the -th modality and the -th modality; is the specific Key projection matrix between the -th modality and the -th modality; represents the correlation matrix between the m -th modality and the -th modality; and are respectively the representations after internal attention processing of the -th modality and the -th modality; d K is the dimension of the Key projection matrix; where, the -th modality and the -th modality are any of the multiple modalities, and the -th modality and the -th modality are different modalities; The multi-agent also optimizes the recommended content through a collaboration mechanism, and the collaboration mechanism includes: a dynamic weight communication network, an interdependent strategy evolution method, and / or a hierarchical differential reward architecture, where, The dynamic weight communication network calculates the communication dynamic weight between agents according to the task correlation degree, and determines the communication path with the highest communication score; The interdependent strategy evolution method optimizes the recommendation strategy by evaluating the influence of the recommended content updates among the multi-agents; The hierarchical differential reward architecture evaluates the performance of the multi-agents in the virtual environment through a multi-level reward mechanism.

2. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 1, characterized in that, The data of the user behavior includes: a movement trajectory represented by three-dimensional space coordinates, movement speed, and stay time; a line-of-sight focus represented by fixation point coordinates, fixation duration, and saccade path; and an operation record represented by click position and operation frequency; The data of the physiological responses includes a heart rate signal represented by sequential heart rate values and heart rate variability, an electroencephalogram signal represented by the power spectral density of α wave, β wave and θ wave, and an electromyogram signal represented by the degree of muscle activation; α wave, β wave and θ wave, and an electromyogram signal represented by the degree of muscle activation; The data of the display feedback includes voice instructions, satisfaction ratings, and text evaluations.

3. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 1, wherein After collecting the multi-modal data, it further includes the steps of feature extraction and normalization processing of the modal data, including: For the non-sequential data in the multi-modal data, feature extraction is performed by the following formula: , In the formula, represents the mapping feature of the ith modal data; tanh is the activation function; is the feature extraction weight matrix, which is used to linearly transform the original modal data into a high-dimensional feature space; is the feature bias term; is used to generate the input of the sigmoid activation function; is the gating bias term; For the sequential data in the multi-modal data, feature extraction is performed by the following formula: , wherein ,X time is the modal original time series data X raw is the time series feature; FFT is the fast Fourier transform; Wavelet is the wavelet transform; α is the fusion weight coefficient; Wherein, it also includes noise reduction filtering of the original sequential data using the following formula: , In the formula, X filtered is the original time series data after noise reduction; Mask is the noise recognition function based on the threshold . represents the Hadamard product.

4. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 1, wherein The fused feature vector representations of all modalities are calculated through the following formula: , , Wherein, H fused represents the fused feature vector representation of all modalities; v interest is the user interest feature; v emotion is the emotional state feature; v social is the group interaction tendency feature; represents the fused feature of the -th modality. When the -th modality is user behavior, is the user interest feature; when the -th modality is physiological reaction, is the emotional state feature; when the -th modality is display feedback, is the group interaction tendency feature; is the dynamically calculated modality weight, which is activated by the sigmoid activation function for the cosine similarity between the -th modality and the -th modality; is the projection parameter.

5. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 1, wherein The communication score is calculated through the following formula: , , , , , , Wherein, M j is the communication score of the j th agent; is the communication weight; is the non-linearly transformed message sent by the th agent to the j th agent; represents the recommended action of the th agent; is the activation function; is the weight matrix for adjusting the original message ; is the bias term; is the feature extracted by the agent based on the fused feature and context content C i , which is the state representation of the agent; represents connecting the state and recommended action of the th agent, so that the transmitted information contains the complete decision context; is the agent-specific non-linear mapping function, which contains M attention heads; is the Kronecker product; is the Hadamard product; GELU and Swish are activation functions; is the output projection matrix; is the Query projection matrix of the t th attention head; is the Key projection matrix of the t th attention head; is the Value projection matrix of the t th attention head; and are the weight and bias respectively; d K is the dimension of the Key projection matrix; represents the task correlation degree between the th agent and the j th agent; ELU is the exponential linear unit; denotes the cosine similarity of the state representations between the j th agent and the v and are learnable parameters, v used to measure the importance of the state feature pair between agents for the communication weight, is the state projection matrix; T is the transpose.

6. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 5, characterized in that The following policy objective function is adopted by the recommendation strategy: , wherein, represents the average reward that an agent expects to obtain under the distribution after all agents operate according to the joint policy s and the joint action a while following the distribution; is the average reward that an agent expects to obtain; is the policy objective function; represents the joint policy of all agents; represents the joint policy of all agents except the -th agent; s is the global environmental state of the agent; represents the recommended action of the agent except the -th agent; is the state - action distribution; β is a parameter for adjusting the constraint strength; is the KL divergence that restricts the difference between the current policy and the reference policy ; R i represents the immediate reward obtained when the -th agent takes the action s in the global state while other agents take the joint action ; The gradient formula adopted for policy update is: , , , , In the formula, is the gradient function; is the communication score of the -th agent; is the policy parameter of the -th agent; is the interdependent advantage function of the -th agent considering the state representation and recommended action effects of other agents except itself; represents the representation state-recommended action value function of the -th agent; represents the representation state-recommended action value function of the j -th agent; is the state value function of the -th agent; is the state value function of the j -th agent; is the interdependent weight between the -th agent and the j -th agent; MLP is a multi-layer perceptron; Measure the impact of the policy change of the th agent on the average gradient of the value function of the j th agent; For the th agent and the j th agent, it is the correlation between their value functions.

7. The intelligent collaborative recommendation method for multi-modal immersive content according to claim 5, wherein The multi-level reward mechanism includes: professional task rewards, interaction and collaboration rewards, and / or overall experience rewards, where The professional task rewards are used to evaluate the performance scores of the agent in its professional field, and the calculation formula is as follows: , , , In the formula, represents the professional task reward; is the recommended content c and the user vector u The correlation between is the corresponding balance parameter; is the recommended content c relative to the user's historical vector diversity, is the corresponding balance parameter; u represents the user vector, including the user's interests and historical behavior information; represents the set of user vectors of past history, h represents each record in the set; The interaction and collaboration rewards are used to evaluate the performance scores of an agent under the influence of the recommendation actions of other agents, and the calculation formula is as follows: , wherein, is the interactive collaboration reward; is the th agent and the j th collaborative coefficient between agents; represents the value function under the global state s wherein the j th agent takes the recommended action while the other agents except the a i th agent take the joint action ; is the value function under the global state wherein the s th agent takes the benchmark action j while the other agents except the th agent take the joint action ; is the benchmark action; is the recommended action of the agents except the th agent; is the recommended action of the agents except the th agent; The overall experience rewards are used to evaluate the overall performance scores of the recommendation actions of multiple agents, and the calculation formula is as follows: , In the formula, is the overall experience reward, f global is a non-linear mapping function that integrates multiple metrics into a single global score; satisfaction is the user satisfaction; duration is the time the user stays on the current recommended content; revisit_intention represents the intention to revisit; The comprehensive reward of the agent is: , wherein, SUM i is the comprehensive reward function of the th agent; , and are adaptive weights, dynamically adjusted according to the system state and learning stage, and the adjustment calculation formula is: , In the formula, is the learning rate; is the gradient function; is the weight optimization objective function; t represents time, t -1 is the previous time state.

8. An electronic device, characterized in that, It includes a memory and a processor. The memory stores at least one segment of program, and the at least one segment of program is executed by the processor to implement the intelligent collaborative recommendation method for multi-modal immersive content according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal adaptive digital human system based on deep learning and intelligent emotion optimization method thereof

    CN119443291A

  • A quality adaptive multimodal affect recognition system for user-centric multimedia indexing

    WO2017136938A1