Intelligent digital human recommendation method based on multimodal emotional computing and cross-domain behavior analysis
By collecting multi-domain behavioral data streams in real time for multimodal emotion calculation and generating emotion vectors and behavior vectors, the problems of single data and missing emotions in the digital human recommendation system are solved, the accuracy and emotional adaptability of personalized recommendations are achieved, and the conversion rate and user experience are improved.
Patent Information
- Application Number
- CN202511021002.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-24
AI Technical Summary
In existing technologies, digital human recommendation systems rely on a single data source, resulting in one-sided user portraits and an inability to fully capture the evolution of user interests. They also lack sentiment analysis, leading to a lack of emotional adaptability in recommendations, which affects recommendation accuracy and conversion rates.
By collecting cross-domain behavioral data streams from financial applications, e-commerce platforms, search engines, and social media in real time, multimodal sentiment calculation is performed to generate sentiment vectors and behavior vectors. By combining timestamps and domain credibility weights, the weights of conflicting behavioral events are dynamically adjusted to achieve personalized recommendations.
It achieves the precise construction of global user portraits, quantifies user emotional states, enhances the adaptability of recommendation scenarios, reduces recommendation lags, and improves conversion rates and user trust.
Smart Images

Figure CN120524042B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of personalized digital human recommendation, and specifically to an intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis. Background Art
[0002] Currently, large-scale multi-agent recommendation systems collect customer behavior data, product information, and historical conversation data to construct agent roles for product recommendations. This system leverages multi-agent collaboration to optimize the recommendation process. For example, patent application number CN202411535657.8 proposes a large-scale multi-agent financial management recommendation system and method. This system uses internal bank financial management data to derive structured user characteristics (such as risk preferences and investment amounts) and then match personalized product recommendations.
[0003] Although this method uses multi-agent collaboration to optimize the recommendation process, it still has some defects:
[0004] Single data source: Relying solely on internal bank wealth management data (such as customer holdings and transaction records) results in a one-sided user profile and fails to fully capture the evolution of user interests. For example, a user's recent high-frequency consumption on e-commerce platforms or investment discussions on social media are not included in the analysis, affecting the accuracy of recommendations.
[0005] Lack of sentiment analysis: Matching products based solely on structured features results in recommendations that lack emotional adaptability. For example, recommending high-risk products to users with negative emotions can exacerbate resistance and reduce conversion rates.
[0006] These defects lead to poor accuracy in personalized matching of traditional digital human recommendation methods, which in turn leads to low product conversion rates. Summary of the Invention
[0007] In view of the above-mentioned defects or deficiencies in the existing technology, this application aims to provide an intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis to improve the accuracy of matching users' personalized needs;
[0008] The recommended approach includes the following steps:
[0009] Real-time collection of target users' cross-domain behavioral data streams, including behavioral events and their occurrence timestamps from different application domains; application domains include financial applications, e-commerce platforms, search engines, and social media platforms;
[0010] Performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors; the emotional vectors include emotional polarity categories and emotional intensity values, and the emotional polarity categories include positive, neutral, and negative;
[0011] Generate a behavior vector of the target user based on the timestamp of the behavior event and the corresponding emotion vector, wherein the behavior vector is used to represent the key interests and demand states of the target user in the current time window;
[0012] According to the behavior vector of the target user, a corresponding recommended action is matched so that the digital human can perform personalized recommendation.
[0013] According to the technical solution provided by this application, generating the target user's behavior vector based on the timestamp of the behavior event and the corresponding emotion vector includes the following steps:
[0014] Calculate the time difference between the timestamp of each behavior event and the current system time to obtain the timeliness weight of each behavior event; wherein, the greater the time difference, the lower the timeliness weight of the behavior event;
[0015] Mapping the emotion intensity value in the emotion vector corresponding to each behavior event into an emotion weight; wherein the behavior event with a higher emotion intensity value has a greater emotion weight;
[0016] Assigning a domain credibility weight to each behavioral event based on the application domain of the behavioral event source; wherein the domain credibility weights of financial applications and search engines are higher than the domain credibility weights of e-commerce platforms and social media platforms;
[0017] Obtaining a first comprehensive weight of each of the behavioral events based on the timeliness weight, the emotion weight, and the domain credibility weight of each of the behavioral events;
[0018] The emotion vectors corresponding to all the behavior events are weightedly summed according to the corresponding first comprehensive weights to obtain the behavior vector of the target user.
[0019] According to the technical solution provided by the present application, before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included:
[0020] Determine whether there are conflicting behavioral events among all the behavioral events, wherein the conflicting behavioral events are behavioral events with opposite sentiment polarity categories that occur within the same application field and target the same or highly related specific objects within a first preset time period, wherein the specific objects include products, services, and topics;
[0021] The step of obtaining a first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event comprises the following steps:
[0022] If not, a first comprehensive weight of each behavioral event is obtained based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event.
[0023] According to the technical solution provided by the present application, before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included:
[0024] If so, performing dynamic weight adjustment of conflict perception for each group of said conflict behavior events;
[0025] The step of performing dynamic weight adjustment of conflict perception for each group of conflict behavior events includes the following steps:
[0026] Calculating the timeliness weight ratio and the emotion weight ratio of the two behavioral events in the group of conflicting behavioral events;
[0027] Obtaining a conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio;
[0028] Based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event, an initial comprehensive weight of each behavioral event is obtained;
[0029] Increasing the initial comprehensive weight of the behavior event with the smaller time difference and the larger emotion intensity value in the group of the conflicting behavior events by A times to obtain a corresponding first comprehensive weight;
[0030] The initial comprehensive weight of the behavior event with the larger time difference or the smaller emotion intensity value in the group of the conflicting behavior events is reduced to 1 / A times to obtain the corresponding first comprehensive weight;
[0031] The initial comprehensive weight of the non-conflicting behavior event is used as the corresponding first comprehensive weight, and the emotion vectors corresponding to all the behavior events are weighted and summed according to the corresponding first comprehensive weight to obtain the behavior vector of the target user.
[0032] According to the technical solution provided by this application, before performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors, the following steps are included:
[0033] Determining whether there is an application field in which the number of behavioral events is lower than a first preset threshold;
[0034] The multimodal analysis of each behavioral event, extraction of corresponding emotional features, and generation of corresponding emotional vectors includes the following steps:
[0035] If not, a multimodal analysis is performed on each of the behavioral events to extract corresponding emotional features and generate corresponding emotional vectors.
[0036] According to the technical solution provided by the present application, after determining whether there is an application field in which the number of behavioral events is lower than the first preset threshold, the following steps are further included:
[0037] If yes, then obtain the actual behavior vector of the first application field; the application field in which the number of behavior events is lower than the preset threshold is the second application field, and the application fields other than the second application field are the first application field;
[0038] Acquire a historical cross-domain behavior data stream of the target user, and obtain a predicted behavior vector for the second application domain based on the actual behavior vector of the target user in the first application domain;
[0039] The predicted behavior vector is combined with the actual behavior vector to form a complete behavior vector.
[0040] According to the technical solution provided by this application, the multimodal analysis of each behavioral event, extraction of corresponding emotional features, and generation of corresponding emotional vectors include the following steps:
[0041] Extracting corresponding primary emotional features from the text modality, visual modality, and audio modality of the behavioral event respectively;
[0042] Obtaining conflict coefficients between all of the primary emotional features;
[0043] If the conflict coefficient is less than or equal to a second preset threshold, a plurality of the primary emotion features are fused to generate a corresponding emotion vector.
[0044] According to the technical solution provided by the present application, after obtaining the conflict coefficients between all the primary emotional features, the following steps are also included:
[0045] If the conflict coefficient is greater than a second preset threshold, the corresponding behavior event is taken as the first behavior event, and the application field from which the first behavior event originates is taken as the first application field;
[0046] Obtaining a preset modality priority rule for the first application field; the preset modality priority rule defines relative importance weights of different modalities in the first application field in sentiment analysis;
[0047] According to the preset modality priority rule, the primary emotion features corresponding to different modalities are weightedly fused to generate an emotion vector.
[0048] According to the technical solution provided by this application, after matching the behavior vector of the target user to obtain the corresponding recommended action, the following steps are also included:
[0049] Record the target user's operational feedback after the digital human performs the recommended action to obtain feedback indicators, including click-through rate, dwell time, and secondary interaction depth;
[0050] When the feedback indicator is lower than a third preset threshold, extracting all the primary emotional features of the first behavioral event to generate multiple groups of weighted candidate emotional vectors;
[0051] generating a simulated recommended action based on the candidate emotion vector;
[0052] A weight combination that minimizes the difference between the simulated recommended action and the operation feedback of the target user is selected, and the preset modality priority rule of the first application field is updated.
[0053] According to the technical solution provided by the present application, obtaining the conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio includes the following steps:
[0054] Counting the proportion of the target user's conflicting behavior events within the first preset time period that are ultimately adopted, with the behavior events having a smaller time difference and a larger emotional intensity value, and using the proportion as the target user's historical behavior stability coefficient;
[0055] According to the formula , get the conflict balance factor;
[0056] Among them, A represents the conflict balance factor, S represents the timeliness weight ratio, and Q represents the emotion weight ratio. 、 They are respectively the preset time importance parameter and the preset emotion importance parameter, ; B represents the historical behavior stability coefficient.
[0057] Compared to existing technologies, this application offers the following advantages: By collecting behavioral events and timestamps from multiple domains, including financial applications, e-commerce platforms, search engines, and social media, this application constructs a global user profile. This cross-domain behavioral data integration overcomes the limitations of single-domain data, accurately identifying user interests across scenarios (e.g., how e-commerce consumption preferences influence financial management needs), and improving recommendation coverage and relevance. Furthermore, multimodal analysis is performed on each behavioral event to generate sentiment vectors (including sentiment polarity categories and intensity values) to quantify user emotional states. Emotional intelligence drives recommendations (e.g., recommending low-risk, reliable products to users with negative emotions), enhancing contextual adaptability of recommendations, reducing user resistance, and improving conversion rates. Furthermore, by fusing timestamps with sentiment vectors to generate behavior vectors, this real-time representation of user interests and needs within the current time window can be captured, enabling timely capture of dynamic interest shifts (e.g., short-term promotional responses or long-term demand changes), enabling real-time, accurate recommendations and reducing recommendation lag. Recommended actions are generated by matching behavior vectors, and personalized recommendations are executed by digital humans. Interactive content is adapted to the sentiment vectors, enhancing the naturalness of interactions and user trust, optimizing the user experience, and promoting high-frequency customer conversions. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A flowchart of the steps of the intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis provided in this application. DETAILED DESCRIPTION
[0059] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0060] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0061] Example 1
[0062] As mentioned in the background technology, in order to solve the problems in the existing technology, this application proposes a method for intelligent digital human recommendation based on multimodal emotion computing and cross-domain behavior analysis, such as Figure 1 As shown, the following steps are included:
[0063] S1. Real-time collection of cross-domain behavioral data streams of target users, including behavioral events and their occurrence timestamps from different application domains; the application domains include financial applications, e-commerce platforms, search engines, and social media platforms;
[0064] Specifically, cross-domain behavioral data streams consist of user behavior logs generated across various digital platforms (including at least financial apps, e-commerce platforms, search engines, and social media). Each log entry includes behavioral events, such as "stock purchase (financial apps)," "negative product reviews (e-commerce platforms)," "'new energy vehicle' keyword search (search engines)," and "liking negative news (social media)." Timestamps indicate the millisecond-accurate time of the event. Data is collected in real time through each platform's open APIs (such as e-commerce order interfaces and social media like event webhooks) or tracking SDKs, and then transmitted to a Kafka message queue to ensure time consistency.
[0065] S2. Performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors; the emotional vectors include emotional polarity categories and emotional intensity values, and the emotional polarity categories include positive, neutral, and negative;
[0066] Specifically, multimodal analysis of each of these behavioral events involves sentiment analysis across text, image (visual), and voice (audio) data. Text includes user comments, search terms, and chat logs; images include user-uploaded images, emoticons, and livestream screenshots; and voice includes e-commerce customer service call recordings and voice searches. Sentiment analysis of text uses the BERT model to output sentiment polarity (positive / neutral / negative) and intensity (0-1, e.g., "very satisfied" corresponds to an intensity of 0.9). Image sentiment analysis uses ResNet to identify expression tags (e.g., "angry") and combines it with scene object detection (e.g., "damaged goods" corresponds to a negative sentiment polarity).
[0067] S3. Based on the timestamp of the behavior event and the corresponding emotion vector, generate a behavior vector of the target user, where the behavior vector is used to represent the key interests and demand states of the target user in the current time window;
[0068] Specifically, behavioral events are aggregated using a sliding time window (e.g., 30 minutes). By weightedly merging sentiment vectors, the output vector dimensions are aligned with sentiment features, such as [financial interest, consumer propensity, social sentiment]. The behavioral vector is essentially a spatiotemporal weighted fusion of sentiment features, with recent strong sentiment behaviors receiving higher weights.
[0069] Furthermore, generating the target user's behavior vector based on the timestamp of the behavior event and the corresponding emotion vector includes the following steps:
[0070] Calculate the time difference between the timestamp of each behavior event and the current system time to obtain the timeliness weight of each behavior event; wherein, the greater the time difference, the lower the timeliness weight of the behavior event;
[0071] Mapping the emotion intensity value in the emotion vector corresponding to each behavior event into an emotion weight; wherein the behavior event with a higher emotion intensity value has a greater emotion weight;
[0072] Assigning a domain credibility weight to each behavioral event based on the application domain of the behavioral event source; wherein the domain credibility weights of financial applications and search engines are higher than the domain credibility weights of e-commerce platforms and social media platforms;
[0073] Obtaining a first comprehensive weight of each of the behavioral events based on the timeliness weight, the emotion weight, and the domain credibility weight of each of the behavioral events;
[0074] The emotion vectors corresponding to all the behavior events are weightedly summed according to the corresponding first comprehensive weights to obtain the behavior vector of the target user.
[0075] Alternatively, human memory decays over time (Ebbinghaus curve), and new events reflect immediate needs more, so the timeliness weight = e (-λ·Δt) ;in, Indicates the time difference, Indicates the attenuation coefficient, the default value is 0.5. Sentiment weight = ;in, =1.2, It represents the sentiment intensity value, and the domain credibility weight adopts a preset value. Optionally, the domain credibility weight of the search engine is 0.9, the domain credibility weight of social media is 0.5, the domain credibility weight of financial applications is 0.9, and the domain credibility weight of e-commerce platforms is 0.7.
[0076] Specifically, through the triple mechanisms of time decay index, emotion intensity power law amplification, and domain credibility filtering, the interference of outdated events, low emotion events, and low credibility domain data is reduced, so that the behavior vector focuses on the user's real-time strong needs.
[0077] Furthermore, before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included:
[0078] Determine whether there are conflicting behavioral events among all the behavioral events, wherein the conflicting behavioral events are behavioral events with opposite sentiment polarity categories that occur within the same application field and target the same or highly related specific objects within a first preset time period, wherein the specific objects include products, services, and topics;
[0079] The step of obtaining a first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event comprises the following steps:
[0080] If not, a first comprehensive weight of each behavioral event is obtained based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event.
[0081] S4. According to the behavior vector of the target user, a corresponding recommended action is matched so that the digital human can perform personalized recommendation.
[0082] In a preferred embodiment, before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included:
[0083] If so, performing dynamic weight adjustment of conflict perception for each group of said conflict behavior events;
[0084] The step of performing dynamic weight adjustment of conflict perception for each group of conflict behavior events includes the following steps:
[0085] Calculating the timeliness weight ratio and the emotion weight ratio of the two behavioral events in the group of conflicting behavioral events;
[0086] Obtaining a conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio;
[0087] Furthermore, obtaining the conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio includes the following steps:
[0088] Counting the proportion of the target user's conflicting behavior events within the first preset time period that are ultimately adopted, with the behavior events having a smaller time difference and a larger emotional intensity value, and using the proportion as the target user's historical behavior stability coefficient;
[0089] According to the formula , get the conflict balance factor;
[0090] Among them, A represents the conflict balance factor, S represents the timeliness weight ratio, and Q represents the emotion weight ratio. 、 They are respectively the preset time importance parameter and the preset emotion importance parameter, B represents the historical behavior stability coefficient. Generally, if A is greater than 1, it means that recent or strong emotional events dominate and their weight should be increased.
[0091] Specifically, B acts as a multiplicative factor, directly scaling the magnitude of the conflict balance factor A. High B values (stable user behavior) amplify A, significantly increasing the weight of recent / emotionally charged events (users tend to trust new behaviors). Low B values (fluctuating user behavior) reduce A and conservatively adjust weights (avoiding overreliance on a single event). The first half of the formula calculates a weighted balance between timeliness and emotion, while the historical behavior stability factor B incorporates user historical behavior habits, making conflict resolution more personalized.
[0092] Specifically, the conflicting behavior event group is a pair of behavior events with opposite emotional polarities that occur within the same application field and for the same / highly associated object (a certain financial product code) within a first preset time period (1 hour for finance and 6 hours for e-commerce).
[0093] Based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event, an initial comprehensive weight of each behavioral event is obtained;
[0094] Increasing the initial comprehensive weight of the behavior event with the smaller time difference and the larger emotion intensity value in the group of the conflicting behavior events by A times to obtain a corresponding first comprehensive weight;
[0095] The initial comprehensive weight of the behavior event with the larger time difference or the smaller emotion intensity value in the group of the conflicting behavior events is reduced to 1 / A times to obtain the corresponding first comprehensive weight;
[0096] The initial comprehensive weight of the non-conflicting behavior event is used as the corresponding first comprehensive weight, and the emotion vectors corresponding to all the behavior events are weighted and summed according to the corresponding first comprehensive weight to obtain the behavior vector of the target user.
[0097] Specifically, when there are conflicting behavioral events, the first comprehensive weight is calculated for non-conflicting behavioral events using the same weight calculation method. For newer and more emotional behavioral events among the conflicting behavioral events, the dominance of credible behaviors is strengthened. For older or less emotional behavioral events among the conflicting behavioral events, noise or outdated contradictory signals are suppressed. Finally, the first comprehensive weight of all behavioral events is obtained, and then the behavioral vector is obtained according to the above method.
[0098] This implementation considers that users' short-term positive and negative behaviors towards the same item (e.g., in a home-buying scenario, immediately complaining to the real estate agent after adding a listing to their favorites) can reflect indecisive decisions or incomplete information, necessitating a distinction between core demands and temporary emotions. Relying solely on the time ratio might overlook persistent and intense negative experiences (e.g., repeated complaints); relying solely on the emotion ratio might amplify occasional, brief emotional outbursts (e.g., a slip of the wrist). Therefore, a conflict balance factor is employed to balance temporal, spatial, and psychological factors, avoiding bias from a single indicator.
[0099] In a preferred embodiment, before performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors, the following steps are included:
[0100] Determining whether there is an application field in which the number of behavioral events is lower than a first preset threshold;
[0101] The multimodal analysis of each behavioral event, extraction of corresponding emotional features, and generation of corresponding emotional vectors includes the following steps:
[0102] If not, a multimodal analysis is performed on each of the behavioral events to extract corresponding emotional features and generate corresponding emotional vectors.
[0103] Furthermore, after determining whether there is an application field where the number of behavioral events is lower than a first preset threshold, the following steps are further included:
[0104] If yes, then obtain the actual behavior vector of the first application field; the application field in which the number of behavior events is lower than the preset threshold is the second application field, and the application fields other than the second application field are the first application field;
[0105] Acquire a historical cross-domain behavior data stream of the target user, and obtain a predicted behavior vector for the second application domain based on the actual behavior vector of the target user in the first application domain;
[0106] The predicted behavior vector is combined with the actual behavior vector to form a complete behavior vector.
[0107] Specifically, to detect sparse domains: Threshold setting: A preset lower limit for the number of behavioral events (e.g., <5 events within 24 hours in a specific domain). Domain classification: Second application domain (sparse domain): For example, users' recent behavior in "financial applications" is sparse. First application domain (non-sparse domain): For example, data from "e-commerce platforms" and "social media" is sufficient. Multimodal analysis is performed on behavioral events in the first application domain to generate actual behavior vectors, which are then used to predict behavior vectors in the second application domain. Historical data retrieval: The user's behavior sequence in the second application domain over the past 30 days is extracted to generate a historical cross-domain behavior data stream. Based on this historical cross-domain behavior data stream, cross-domain correlation modeling is performed: A domain correlation matrix is established (e.g., the correlation weight between purchase behavior in financial applications and e-commerce is 0.7). The target user's actual behavior vector in the first application domain is input into the time series prediction model, which outputs a predicted behavior vector. The predicted behavior vector and the actual behavior vector are concatenated to generate a complete behavior vector.
[0108] In a preferred embodiment, performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors includes the following steps:
[0109] Extracting corresponding primary emotional features from the text modality, visual modality, and audio modality of the behavioral event respectively;
[0110] Specifically, the text modality uses the RoBERTa model to parse user comments / search terms and output a sentiment probability distribution vector (e.g., [positive: 0.8, neutral: 0.15, negative: 0.05]). The visual modality uses the Facial Action Coding System (FACS) to analyze user-uploaded images / videos and output a 17-dimensional facial action unit vector (e.g., AU12_lip corner raised = 0.72). The audio modality uses OpenSMILE to extract an acoustic feature set, which is then fed into an LSTM network to output a sentiment intensity value (ranging from 0 to 1).
[0111] Obtaining conflict coefficients between all of the primary emotional features;
[0112] Specifically, the three-modal feature vectors are normalized to the interval [0, 1], and the cosine similarity between the two modalities is calculated using the following formula: , where S tv Represents the cosine similarity between the text modality and the visual modality, T represents the normalized primary emotional features of the text modality, and V represents the normalized primary emotional features of the visual modality; similarly, the cosine similarity S between the text modality and the audio modality is obtained. TA , and the cosine similarity S between the visual modality and the audio modality VA ; The conflict coefficient C is calculated by the following formula: .
[0113] If the conflict coefficient is less than or equal to a second preset threshold, a plurality of the primary emotion features are fused to generate a corresponding emotion vector.
[0114] Optionally, the second preset threshold is 0.75, and the conflict coefficient is less than or equal to 0.75, indicating low conflict. When in low conflict, domain priority fusion is triggered (the emotion vectors corresponding to all behavioral events are weighted and summed according to the corresponding first comprehensive weight to obtain the behavior vector of the target user).
[0115] Furthermore, after obtaining the conflict coefficients between all the primary emotional features, the method further includes the following steps:
[0116] If the conflict coefficient is greater than a second preset threshold, the corresponding behavior event is taken as the first behavior event, and the application field from which the first behavior event originates is taken as the first application field;
[0117] Obtaining a preset modality priority rule for the first application field; the preset modality priority rule defines relative importance weights of different modalities in the first application field in sentiment analysis;
[0118] Specifically, for high-conflict events, a preset modal weight rule is determined according to a preset modal priority rule to perform weighted fusion, and the preset modal priority rule is a predefined domain-modal weight mapping table.
[0119] For example, the preset modal priority rules are as follows:
[0120]
[0121] According to the preset modality priority rule, the primary emotion features corresponding to different modalities are weightedly fused to generate an emotion vector.
[0122] Specifically, the application domain is matched through event metadata, and the domain rule library is queried to obtain the relative importance weights of the text modality, visual modality and audio modality corresponding to the matched application domain. The emotion vector is obtained based on the relative importance weight corresponding to each modality and the normalized primary emotion features.
[0123] This implementation improves domain adaptability. For example, in financial applications, text is prioritized to avoid distracting expressions (e.g., a frown might be caused by lighting, not emotion). In e-commerce platforms, visuals are prioritized, where authenticity of product images outweighs textual descriptions. By quantifying intermodal differences (the conflict coefficient), intelligent fusion strategies are selected: low-conflict models are directly integrated, while high-conflict models are weighted according to domain rules.
[0124] In a preferred embodiment, after matching the target user's behavior vector to obtain a corresponding recommended action, the method further includes the following steps:
[0125] Record the target user's operational feedback after the digital human performs the recommended action to obtain feedback indicators, including click-through rate, dwell time, and secondary interaction depth;
[0126] Specifically, the user behavior log after the digital human performs the recommended action is recorded. The user behavior log represents the target user's operational feedback. The click-through rate represents the click ratio after the recommended content is exposed. It is calculated as the ratio of the number of clicks to the number of exposures. The dwell time is the number of seconds from clicking to leaving the interface. The secondary interaction depth represents the number of subsequent user interaction behaviors. The feedback index is calculated by the following formula ,in, They represent the click-through rate weight coefficient, the dwell time weight coefficient, and the interaction depth weight coefficient, which can be selected as 0.4, 0.3, and 0.3 respectively; T0 and D0 represent the reference benchmark value of dwell time (preset constant, typical value 20 seconds) and the reference benchmark value of interaction depth (preset constant, typical value 3 times).
[0127] When the feedback indicator is lower than a third preset threshold, extracting all the primary emotional features of the first behavioral event to generate multiple groups of weighted candidate emotional vectors;
[0128] Specifically, when the feedback index is higher than or equal to the third preset threshold, it means that the recommendation is very accurate and the process can be ended. When the feedback index is less than the third preset threshold, it indicates that the recommendation is biased and multiple groups of weighted candidate emotion vectors are generated. Optionally, the third preset threshold is 0.6.
[0129] For example, all normalized primary sentiment features of the first behavioral event are extracted: the primary sentiment features of the text modality are [0.8, 0.1, 0.1], the visual features are [0.2, 0.1, 0.7], and the audio features are [0.3, 0.4, 0.3]. Three candidate solutions are generated in the weight space: Weight combination 1: text weight 0.3, visual weight 0.5, audio weight 0.2; Weight combination 2: text weight 0.4, visual weight 0.4, audio weight 0.2; Weight combination 3: text weight 0.5, visual weight 0.3, audio weight 0.2. The candidate sentiment vector E1 for weight combination 1 is [0.8*0.3, 0.1*0.5, 0.1*0.2] = [0.24, 0.05, 0.02]. The candidate sentiment vectors corresponding to the other weight combinations are obtained similarly.
[0130] generating a simulated recommended action based on the candidate emotion vector;
[0131] Specifically, the corresponding simulated recommended actions are obtained according to the comparison table of candidate emotion vectors and recommended actions. For example, the simulated recommended action corresponding to the candidate emotion vector E1 = [0.24, 0.11, 0.38] of weight combination 1 is to recommend high-end mobile phones, the simulated recommended action corresponding to the candidate emotion vector of weight combination 2 is to recommend mid-range mobile phones, and the simulated recommended action corresponding to the candidate emotion vector of weight combination 3 is to recommend mobile phone accessories.
[0132] A weight combination that minimizes the difference between the simulated recommended action and the operation feedback of the target user is selected, and the preset modality priority rule of the first application field is updated.
[0133] For example, the digital human's previous recommendation action was to recommend a game product to the target user. The target user's feedback index was low, and the target user subsequently clicked and purchased a high-end mobile phone on the e-commerce platform. This indicates that the weight combination with the smallest difference in operation feedback with the target user is: the simulated recommendation action is to recommend a high-end mobile phone, which corresponds to weight combination 1. Therefore, the preset modal priority rule of the first application field (this is the e-commerce platform) is updated with weight combination 1. For example, the weight combination of the previous e-commerce platform was text weight 0.25, visual weight 0.65, and audio weight 0.10, which was updated to text weight 0.3, visual weight 0.5, and audio weight 0.2.
[0134] This implementation method can make the dynamic adjustment of multimodal weights explicit, so that the digital human can achieve closed-loop optimization of accurate consumer intention insights and reasonable product classification recommendations.
[0135] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of the present invention, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.
Claims
1. An intelligent digital human recommendation method based on multimodal emotional computing and cross-domain behavior analysis, characterized in that: The following steps are involved: Collect cross-domain behavior data streams of target users in real time, including behavior events from different application domains and their occurrence timestamps; Said application areas include financial applications, e-commerce platforms, search engines, and social media platforms; Performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors; The emotion vector includes an emotion polarity category and an emotion intensity value, wherein the emotion polarity category includes positive, neutral and negative; Generate a behavior vector of the target user based on the timestamp of the behavior event and the corresponding emotion vector, wherein the behavior vector is used to represent the key interests and demand states of the target user in the current time window; According to the behavior vector of the target user, a corresponding recommended action is matched so that the digital human can perform personalized recommendations; Generating the target user's behavior vector based on the timestamp of the behavior event and the corresponding emotion vector includes the following steps: Calculate the time difference between the timestamp of each behavior event and the current system time to obtain the timeliness weight of each behavior event; wherein, the greater the time difference, the lower the timeliness weight of the behavior event; Mapping the emotion intensity value in the emotion vector corresponding to each behavior event into an emotion weight; wherein the behavior event with a higher emotion intensity value has a greater emotion weight; Assigning a domain credibility weight to each behavioral event based on the application domain of the behavioral event source; wherein the domain credibility weights of financial applications and search engines are higher than the domain credibility weights of e-commerce platforms and social media platforms; Obtaining a first comprehensive weight of each of the behavioral events based on the timeliness weight, the emotion weight, and the domain credibility weight of each of the behavioral events; Performing weighted summation on the emotion vectors corresponding to all the behavior events according to the corresponding first comprehensive weights to obtain the behavior vector of the target user; The multimodal analysis of each behavioral event, extraction of corresponding emotional features, and generation of corresponding emotional vectors includes the following steps: Extracting corresponding primary emotional features from the text modality, visual modality, and audio modality of the behavioral event respectively; Obtaining conflict coefficients between all of the primary emotional features; If the conflict coefficient is less than or equal to a second preset threshold, a plurality of the primary emotion features are fused to generate a corresponding emotion vector.
2. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 1 is characterized by: Before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included: Determine whether there are conflicting behavioral events among all the behavioral events, wherein the conflicting behavioral events are behavioral events with opposite sentiment polarity categories that occur within the same application field and target the same or highly related specific objects within a first preset time period, wherein the specific objects include products, services, and topics; The step of obtaining a first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event comprises the following steps: If not, a first comprehensive weight of each behavioral event is obtained based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event.
3. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 2 is characterized by: Before obtaining the first comprehensive weight of each behavioral event based on the timeliness weight, the emotion weight, and the domain credibility weight of each behavioral event, the following steps are included: If so, performing dynamic weight adjustment of conflict perception for each group of said conflict behavior events; The step of performing dynamic weight adjustment of conflict perception for each group of conflict behavior events includes the following steps: Calculating the timeliness weight ratio and the emotion weight ratio of the two behavioral events in the group of conflicting behavioral events; Obtaining a conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio; Based on the timeliness weight, the emotion weight and the domain credibility weight of each behavioral event, an initial comprehensive weight of each behavioral event is obtained; Increasing the initial comprehensive weight of the behavior event with the smaller time difference and the larger emotion intensity value in the group of the conflicting behavior events by A times to obtain a corresponding first comprehensive weight; The initial comprehensive weight of the behavior event with the larger time difference or the smaller emotion intensity value in the group of the conflicting behavior events is reduced to 1 / A times to obtain the corresponding first comprehensive weight; The initial comprehensive weight of the non-conflicting behavior event is used as the corresponding first comprehensive weight, and the emotion vectors corresponding to all the behavior events are weighted and summed according to the corresponding first comprehensive weight to obtain the behavior vector of the target user.
4. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 1 is characterized by: Before performing multimodal analysis on each of the behavioral events, extracting corresponding emotional features, and generating corresponding emotional vectors, the following steps are included: Determining whether there is an application field in which the number of behavioral events is lower than a first preset threshold; The multimodal analysis of each behavioral event, extraction of corresponding emotional features, and generation of corresponding emotional vectors includes the following steps: If not, a multimodal analysis is performed on each of the behavioral events to extract corresponding emotional features and generate corresponding emotional vectors.
5. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 4 is characterized by: After determining whether there is an application field in which the number of behavioral events is lower than a first preset threshold, the following steps are further included: If yes, then obtain the actual behavior vector of the first application field; the application field in which the number of behavior events is lower than the preset threshold is the second application field, and the application fields other than the second application field are the first application field; Acquire a historical cross-domain behavior data stream of the target user, and obtain a predicted behavior vector for the second application domain based on the actual behavior vector of the target user in the first application domain; The predicted behavior vector is combined with the actual behavior vector to form a complete behavior vector.
6. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 1 is characterized by: After obtaining the conflict coefficients between all the primary emotional features, the following steps are also included: If the conflict coefficient is greater than a second preset threshold, the corresponding behavior event is taken as the first behavior event, and the application field from which the first behavior event originates is taken as the first application field; Obtaining a preset modality priority rule for the first application field; the preset modality priority rule defines relative importance weights of different modalities in the first application field in sentiment analysis; According to the preset modality priority rule, the primary emotion features corresponding to different modalities are weightedly fused to generate an emotion vector.
7. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 6 is characterized by: After obtaining the corresponding recommended action based on the behavior vector of the target user, the following steps are further included: Record the target user's operational feedback after the digital human performs the recommended action to obtain feedback indicators, including click-through rate, dwell time, and secondary interaction depth; When the feedback indicator is lower than a third preset threshold, extracting all the primary emotional features of the first behavioral event to generate multiple groups of weighted candidate emotional vectors; generating a simulated recommended action based on the candidate emotion vector; A weight combination that minimizes the difference between the simulated recommended action and the operation feedback of the target user is selected, and the preset modality priority rule of the first application field is updated.
8. The intelligent digital human recommendation method based on multimodal emotion computing and cross-domain behavior analysis according to claim 3 is characterized by: Obtaining the conflict balance factor A according to the timeliness weight ratio and the emotion weight ratio includes the following steps: Counting the proportion of the target user's conflicting behavior events within the first preset time period that are ultimately adopted, with the behavior events having a smaller time difference and a larger emotional intensity value, and using the proportion as the target user's historical behavior stability coefficient; According to the formula , get the conflict balance factor; Among them, A represents the conflict balance factor, S represents the timeliness weight ratio, and Q represents the emotion weight ratio. 、 They are respectively the preset time importance parameter and the preset emotion importance parameter, ; B represents the historical behavior stability coefficient.
Citation Information
Patent Citations
A financial recommendation system and method based on large model multi-agent
CN119048244B
Commodity recommendation method based on consumer behaviors
CN119151643A
Recommended forecast sheet generation system and method based on multi-modal distillation sentiment analysis
CN119484939A