A virtual human-computer interaction system and method for brand promotion

By collecting multimodal data to analyze user status, generating personalized interactive contextual commands, and dynamically adjusting the virtual human image and content, the problem of rigid interaction in existing virtual human systems is solved, and the accuracy and continuous optimization of brand promotion are achieved.

CN122086244AInactive Publication Date: 2026-05-26GUANGZHOU ZONGHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-05-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing virtual human advertising systems struggle to understand user emotions and interests in real time, resulting in rigid interactions, homogenized content, a lack of effective evaluation and optimization, and an inability to achieve personalized interaction and in-depth, precise brand advertising.

Method used

By collecting and analyzing multimodal data in real time, user status information is generated. Combined with historical interaction data, comprehensive interactive contextual instructions are generated, and the virtual human image and brand content are dynamically adjusted to build a data closed-loop optimization system.

Benefits of technology

It enables real-time personalized adaptation of virtual human images and promotional content, enhancing the naturalness and appeal of interactions, and achieving precision and continuous iterative optimization in brand promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086244A_ABST
    Figure CN122086244A_ABST
Patent Text Reader

Abstract

This invention discloses a virtual human-computer interaction system and method for brand promotion, relating to the field of human-computer interaction technology. The system includes: S1: collecting multimodal data from users during the interaction process via a user terminal and performing real-time analysis to generate real-time user status information; S2: generating comprehensive interaction context instructions; S3: based on the comprehensive interaction context instructions, executing a virtual human adaptive generation step and a brand content dynamic generation step in parallel; S4: forming and outputting an interaction response for the current user; S5: collecting user behavior data related to the interaction effect and updating and optimizing rules based on the user behavior data. The advantages of this invention are: through multimodal perception and context analysis, it achieves real-time personalized dynamic generation of virtual human image, interaction style, and promotional content, enabling brand promotion to accurately adapt to the real-time status and historical preferences of different users, greatly improving the naturalness, attractiveness, and user resonance of the interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, specifically to a virtual human-computer interaction system and method for brand promotion. Background Technology

[0002] In recent years, virtual human technology has been widely used in brand promotion, interacting with users through digital avatars to enhance brand awareness and user experience. However, existing virtual human promotional systems mostly rely on pre-set dialogue libraries and fixed avatars, resulting in rigid interaction processes. They struggle to perceive user emotions and interests, leading to highly homogenized promotional content and failing to achieve personalized interaction based on deep understanding. Furthermore, existing systems generally lack effective evaluation and optimization mechanisms for the interaction process. The performance of virtual humans and promotional content cannot be dynamically adjusted based on actual results, making the iteration of promotional strategies dependent on human experience—inefficient and difficult to quantify.

[0003] These technological limitations often result in virtual human advertising remaining at a superficial level, lacking emotional connection with users, and failing to deliver brand information with sufficient depth and precision. To overcome these shortcomings, the industry urgently needs an intelligent virtual human solution that can understand user status in real time, dynamically generate personalized interactive content, and drive system self-optimization based on a closed-loop data loop of interaction effects, thereby achieving more natural, attractive, and scientifically measurable brand advertising results. Summary of the Invention

[0004] To address the aforementioned technical problems, a virtual human-computer interaction system and method for brand promotion are provided. This technical solution resolves at least one of the technical problems mentioned in the background section.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A virtual human-computer interaction method for brand promotion, comprising: S1: Collect multimodal data of the user during the interaction process through the user terminal, and perform real-time analysis on the multimodal data to generate real-time status information of the user; S2: Generate a comprehensive interactive scenario instruction based on the real-time status information and the pre-stored user historical interaction data; S3: Based on the comprehensive interactive context instructions, execute the virtual human adaptive generation step and the brand content dynamic generation step in parallel; The virtual human adaptive generation step is used to dynamically determine at least one image presentation parameter and interaction style parameter of the virtual human; The aforementioned brand content dynamic generation step is used to dynamically select and assemble personalized brand promotional content from a structured brand knowledge base. S4: The virtual human instance generated based on the image presentation parameters and interaction style parameters is merged and rendered with the personalized brand promotion content to form and output an interactive response for the current user; S5: During the interaction and after each interaction, collect user behavior data related to the interaction effect, and update and optimize the rules based on the user behavior data; The optimization rules are used to adjust the logic for generating comprehensive interactive context instructions in step S2 and / or the strategies for adaptive virtual human generation and dynamic brand content generation in step S3 during subsequent interactions.

[0006] Preferably, in step S1, the real-time analysis of the multimodal data includes: S11: Perform image analysis on the collected visual data to identify at least one of the user's facial expression, posture, and direction of attention, and infer the user's primary emotional state and level of engagement in the interaction. S12: Perform speech recognition and speech emotion analysis on the collected audio data to obtain the text information input by the user and infer the user's second emotional state; S13: Perform natural language understanding on the text information to identify the user's interaction intent and key entities in the text; The real-time status information includes at least the emotional state obtained by fusing the first emotional state and / or the second emotional state, the interaction engagement, the interaction intent, and the key entities.

[0007] Preferably, in step S2, the generation of comprehensive interactive context instructions specifically includes: S21: Based on the real-time status information, generate a real-time user profile including real-time sentiment tags, real-time interest tags, and real-time interaction stage tags; S22: The real-time user profile is associated and fused with the user's historical interaction data to update and form a long-term user profile, which includes at least historical interest preference tags and historical behavior pattern tags. S23: Combining the real-time user profile, the long-term user profile, and the context information of the current interaction, generate the comprehensive interaction context instruction, which includes the target user feature identifier and the recommended interaction strategy identifier.

[0008] Preferably, in step S3, the virtual human adaptive generation step specifically includes: S31: Based on the target user feature identifier in the comprehensive interactive context instruction, select suitable virtual human basic image, clothing and accessory resources from the virtual human resource library; S32: Determine the corresponding facial expressions and body movement sequence of the virtual human based on the emotional state in the real-time status information; S33: Determine the virtual human's voice parameters and dialogue style template based on the recommended interaction strategy identifier in the comprehensive interactive context instruction; The image presentation parameters include at least the selected basic image, clothing, accessories, facial expressions, and action sequences, and the interaction style parameters include at least the voice parameters and dialogue style templates.

[0009] Preferably, in step S3, the dynamic generation step of brand content specifically includes: S34: Determine the core theme and level of detail of the brand promotion content based on the target user feature identifier and recommendation interaction strategy identifier in the comprehensive interactive context instruction; S35: Based on the core theme, retrieve matching product information tags, brand story tags, and marketing script tags from the structured brand knowledge base; S36: Based on the level of detail and the real-time interest tags in the real-time user profile, prioritize and crop the retrieved content materials to generate coherent personalized brand promotion content.

[0010] Preferably, in step S5, the user behavior data related to the interaction effect includes at least three of the following: total interaction time, number of rounds of dialogue, user emotional state change value, type and depth of user-initiated questions, number of positive feedbacks on preset brand keywords, and completion rate of preset guiding actions.

[0011] Preferably, in step S5, the rule update optimization based on user behavior data includes: S51: Based on machine learning models, train and analyze multi-round historical interaction data to establish a correlation model between "user profile features - virtual human and content strategy combination - performance indicators"; S52: Based on the association model, generate or update the strategy mapping rules, which define that when a specific user profile feature is identified, virtual human generation strategy and brand content generation strategy that have been verified by historical data to improve target performance indicators should be given priority.

[0012] Furthermore, this solution also proposes a virtual human-computer interaction system for brand promotion, which implements the aforementioned virtual human-computer interaction method for brand promotion, including: The multimodal perception module is configured to collect and analyze the user's multimodal data and generate the user's real-time status information. The user profiling and context analysis engine is connected to the multimodal perception module and configured to generate comprehensive interactive context instructions based on the real-time status information and historical interaction data. The virtual human adaptive generation engine is connected to the user profile and context analysis engine and is configured to determine the virtual human's image presentation parameters and interaction style parameters based on the comprehensive interactive context instructions. The brand knowledge base and content dynamic generation module are connected to the user profile and context analysis engine. It stores a structured brand knowledge base and is configured to dynamically generate personalized brand promotion content based on the comprehensive interactive context instructions. The rendering output module is connected to the virtual human adaptive generation engine and the brand knowledge base and content dynamic generation module, respectively, and is configured to integrate virtual human instances with brand promotional content to form an interactive response and output it. The interaction effect evaluation and optimization module is connected to the rendering output module and the multimodal perception module. It is configured to collect user behavior data and update optimization rules based on this data. The optimization rules are fed back to at least one of the user profile and context analysis engine, the virtual human adaptive generation engine, and the brand knowledge base and content dynamic generation module to adjust their subsequent processing strategies.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention, through multimodal perception and contextual analysis, enables real-time personalized dynamic generation of virtual avatars, interaction styles, and promotional content. This allows brand promotion to accurately adapt to the real-time status and historical preferences of different users, greatly enhancing the naturalness, attractiveness, and user resonance of the interaction. Simultaneously, the system constructs a complete data loop from interaction data collection and quantitative effect analysis to automatic strategy optimization. This allows promotional strategies to continuously iterate based on objective results, thereby achieving a fundamental transformation in brand promotion from one-way instruction to intelligent interaction, from experience-driven to data-driven, and from a one-size-fits-all approach to a personalized one. This significantly improves the accuracy, adaptability, and long-term effectiveness of brand promotion. Attached Figure Description

[0014] Figure 1 This is a flowchart of the virtual human-computer interaction method for brand promotion proposed in this solution. Detailed Implementation

[0015] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0016] Reference Figure 1 As shown, a virtual human-computer interaction method for brand promotion includes: S1: Collect multimodal data of users during the interaction process through user terminals, and perform real-time analysis on the multimodal data to generate real-time status information of users. By integrating visual, voice, text and other multimodal data for real-time analysis, it can transcend the limitations of single text analysis and more comprehensively and accurately understand the user's current emotional state, attention focus and true intention, providing a reliable and rich real-time data foundation for subsequent precise personalized interaction. S2: Based on the real-time status information and the pre-stored user historical interaction data, a comprehensive interactive context command is generated. By combining real-time perception with historical memory, the generated comprehensive interactive context command is no longer an isolated response to a single trigger, but a comprehensive judgment that integrates the user's real-time emotions, immediate interests and long-term preferences. This enables the system to have "memory" and "understanding" capabilities, providing core driving commands for achieving continuous and evolving personalized brand interaction throughout the user's entire life cycle. S3: Based on the comprehensive interactive context instructions, execute the virtual human adaptive generation step and the brand content dynamic generation step in parallel; The virtual human adaptive generation step is used to dynamically determine at least one image presentation parameter and interaction style parameter of the virtual human; The aforementioned brand content dynamic generation step is used to dynamically select and assemble personalized brand promotional content from a structured brand knowledge base. By employing a parallel generation strategy, it is ensured that the virtual avatar's style and promotional content can be collaboratively adapted and dynamically created based on unified interactive context commands. This solves the problem of the disconnect and rigid combination of virtual avatar image, behavior, and push content in existing technologies, achieving a high degree of consistency and synchronous personalization between "character settings" and "communication scripts," greatly enhancing the overall sense of brand information delivery and immersion. S4: The virtual human instance generated based on the image presentation parameters and interaction style parameters is integrated and rendered with the personalized brand promotional content to form and output an interactive response for the current user. Through technical rendering, the dynamically generated virtual human image, actions, and voice are seamlessly integrated with dynamically organized content text and visual elements, outputting a complete, natural, and consistent interactive response. This ensures the smoothness and realism of the front-end user experience, transforming the complex intelligent calculations in the back-end into a highly customized and user-friendly interaction that is perceptible to the front-end user. S5: During the interaction and after each interaction, collect user behavior data related to the interaction effect, and update and optimize the rules based on the user behavior data; The optimization rules are used to adjust the logic for generating comprehensive interactive context instructions in step S2 and / or the strategies for adaptive generation of virtual humans and dynamic generation of brand content in step S3 in subsequent interactions. A complete data loop of "perception-decision-execution-evaluation-optimization" has been constructed. By collecting performance data and updating optimization rules, the system no longer relies on fixed, pre-written scripts, but instead possesses the ability to learn and evolve based on actual interaction results. This allows the system's personalized strategies to continuously iterate and optimize with the accumulation of interaction data, solving the key problem that static systems cannot adapt and continuously improve their effectiveness, and realizing the intelligent and automated upgrade of brand promotion strategies.

[0017] Specifically, in step S1, the real-time analysis of multimodal data includes: S11: Perform image analysis on the collected visual data to identify at least one of the user's facial expression, posture, and direction of attention, and infer the user's primary emotional state and level of engagement in the interaction. S12: Perform speech recognition and speech emotion analysis on the collected audio data to obtain the text information input by the user and infer the user's second emotional state; S13: Perform natural language understanding on the text information to identify the user's interaction intent and key entities in the text; The real-time status information includes at least the emotional state obtained by fusing the first emotional state and / or the second emotional state, the interaction engagement, the interaction intent, and the key entities.

[0018] Specifically, in practical implementation, real-time analysis of multimodal data can be achieved through analysis service modules deployed in the cloud or edge servers. For visual data, real-time processing is performed using computer vision models deployed behind the user terminal's camera: face detection and key point localization algorithms are used to identify the user's facial regions, and then classification models trained with a large amount of sentiment-annotated data, such as convolutional neural networks, are used to analyze facial muscle movement units to identify basic expressions such as happiness, surprise, and confusion and infer the primary emotional state; simultaneously, human pose estimation algorithms such as OpenPose are used to analyze the user's body orientation, gestures, and head turning angles, combined with eye-tracking technology (if supported by hardware) or estimation methods based on attention models, to comprehensively determine whether the user's gaze is focused on the virtual human interface, thereby quantifying their level of engagement. Secondly, for audio data, the speech stream is transmitted in real time to the speech processing engine: Automatic Speech Recognition (ASR) services, such as an end-to-end deep learning model-based recognition system, convert the speech into text. In parallel, a speech sentiment analysis module extracts acoustic features from the audio, such as fundamental frequency, Mel-frequency cepstral coefficients, and speech rate, and inputs them into a pre-trained sentiment classifier, outputting a second sentiment state corresponding to excitement, calmness, frustration, etc. Finally, for the text information output by ASR, a Natural Language Understanding (NLU) service is invoked. This service, based on a pre-trained intent recognition model such as BERT or a fine-tuned model with a similar architecture, and a Named Entity Recognition (NER) model, identifies the user's interaction intent from the text, such as "querying product information," "complaining," or "general conversation," as well as key entities such as product name, technical parameters, and location. Ultimately, an information fusion center weights or makes a decision-level fusion of the first sentiment state from vision and the second sentiment state from speech. For example, in cases of conflict, visual sentiment is prioritized or selected based on confidence level, forming a unified sentiment state label. This label, along with interaction engagement, interaction intent, and key entities, is encapsulated into structured real-time state information for subsequent steps.

[0019] In step S2, the generation of comprehensive interactive context instructions specifically includes: S21: Based on the real-time status information, generate a real-time user profile including real-time sentiment tags, real-time interest tags, and real-time interaction stage tags; S22: The real-time user profile is associated and fused with the user's historical interaction data to update and form a long-term user profile, which includes at least historical interest preference tags and historical behavior pattern tags. S23: Combining the real-time user profile, the long-term user profile, and the context information of the current interaction, generate the comprehensive interaction context instruction, which includes the target user feature identifier and the recommended interaction strategy identifier.

[0020] In practical implementation, the process of generating comprehensive interactive contextual instructions in step S2 is completed by a dedicated contextual analysis engine. This engine logically comprises three core units: profile building, data fusion, and decision generation. In the profile building unit, the system automatically generates real-time user profiles based on the real-time status information output in step S1, using predefined rules and models: real-time sentiment tags can be generated by mapping sentiment state values ​​(e.g., continuous values ​​from positive to negative) to discrete tags such as "pleasant," "neutral," and "frustrated"; real-time interest point tags are generated by extracting and analyzing key entities in the text information, such as mentions of product characteristics like "battery life" and "photography," combined with interactive intents such as "query parameters"; and real-time interaction stage tags are automatically labeled by a conversation state machine based on historical dialogue rounds, current intent, and preset process nodes such as "opening greeting," "needs inquiry," "product explanation," "action facilitation," and "end." In the data fusion unit, the system maintains a long-term profile database indexed by anonymous user IDs. When a new real-time user profile is generated, the fusion algorithm updates the corresponding long-term user profile accordingly. Specifically, historical interest preference tags are not simply replaced, but rather updated using a weighted update mechanism based on time decay. For example, key entities that appear frequently recently are given higher weights, thus forming a dynamically evolving preference vector. Historical behavior pattern tags are derived by analyzing statistical data from users' past interactions, such as average session duration, preferred content types, and typical question depth, and using clustering or classification methods to categorize them into terms like "technology-oriented," "price-sensitive," and "browsing-oriented." In the decision generation unit, the contextual decision module receives real-time user profiles, long-term user profiles, and contextual information including the current session ID, timestamp, and number of interaction rounds. Internally, this module runs a mapper based on a rule engine or lightweight machine learning model such as a decision tree or gradient boosting tree. This mapper comprehensively analyzes the above diverse inputs, such as real-time emotion = pleasure, real-time interest = taking photos, historical preferences = cutting-edge technology, behavior patterns = in-depth exploration, and interaction stage = product explanation, and outputs a structured "comprehensive interaction context instruction." In practice, this instruction can be expressed as a JSON object or feature vector containing multiple dimensions. Its core elements include at least: 1. Target user characteristic identifiers, such as technology enthusiasts; 2. Recommendation interaction strategy identifiers, such as actively guided in-depth technical explanations. This instruction will serve as a unified and operable strategic framework for subsequent virtual human generation and content generation.

[0021] In step S3, the virtual human adaptive generation step specifically includes: S31: Based on the target user feature identifier in the comprehensive interactive context instruction, select suitable virtual human basic image, clothing and accessory resources from the virtual human resource library; S32: Determine the corresponding facial expressions and body movement sequence of the virtual human based on the emotional state in the real-time status information; S33: Determine the virtual human's voice parameters and dialogue style template based on the recommended interaction strategy identifier in the comprehensive interactive context instruction; The image presentation parameters include at least the selected basic image, clothing, accessories, facial expressions, and action sequences, and the interaction style parameters include at least the voice parameters and dialogue style templates.

[0022] In specific implementation, step S3, the virtual human adaptive generation step, is executed by a virtual human driving engine. This engine internally maintains a structured virtual human resource library, storing multiple basic character models, different styles of clothing and accessories such as glasses and jewelry, and can be associated with brand product resources. Each resource is accompanied by metadata tags for matching, such as "fashionable casual," "business formal," "technological," and "approachable." When the engine receives the comprehensive interactive context instruction from step S2, the resource scheduling unit parses the "target user feature identifier" in the instruction and, according to preset mapping rules or collaborative filtering recommendation algorithms, retrieves and selects the basic image, clothing, and accessory combination that best matches the metadata tags from the resource library, forming the current image base of the virtual human. The expression and action driving unit receives emotional states such as "pleasure" from real-time status information. This unit has a built-in "emotion-expression" mapping model, which can be a predefined lookup table or a lightweight generative neural network. Based on the input emotion tags, the model outputs corresponding parameterized facial expression codes, such as the degree of mouth twitching, eyelid opening and closing, and predefined or programmatically generated body movement sequence identifiers, such as "waving hello," "nodding in agreement," and "resting chin in thought." These parameters and sequences drive the skeleton and skin of the virtual human model, achieving real-time, emotionally consistent presentation of expressions and movements. The interaction style control unit parses the "recommended interaction strategy identifiers" in the comprehensive interaction context instructions. This unit is associated with a strategy configuration file library, with each strategy identifier corresponding to a configuration file, which defines in detail the speech parameters, such as speech rate, pitch, and timbre selection controlled by the TTS engine, and dialogue style templates, such as sentence structure, commonly used vocabulary, and level of honorifics. For example, for the "friendly and persuasive" strategy, the system will load a higher speech rate and higher pitch to appear more energetic, and select a dialogue template containing encouraging and hypothetical closed sentences. Finally, the virtual human adaptive generation engine packages all the above image presentation parameters and interaction style parameters, outputting a structured virtual human driving instruction set, which is sent to the rendering module for integration with the content.

[0023] In step S3, the dynamic generation step of brand content specifically includes: S34: Determine the core theme and level of detail of the brand promotion content based on the target user feature identifier and recommendation interaction strategy identifier in the comprehensive interactive context instruction; S35: Based on the core theme, retrieve matching product information tags, brand story tags, and marketing script tags from the structured brand knowledge base; S36: Based on the level of detail and the real-time interest tags in the real-time user profile, prioritize and crop the retrieved content materials to generate coherent personalized brand promotion content.

[0024] In specific implementation, step S3, the dynamic generation of brand content, is executed by a content generation engine. This engine first parses the comprehensive interactive context instructions: based on the "target user characteristic identifier" and "recommended interaction strategy identifier," it determines the core theme of the promotion, such as "high cost-effectiveness and durability of the product," and the level of detail, such as "medium," through a pre-configured "strategy-content" mapping rule table, including core parameters, 1-2 usage scenarios, and 1 comparative data point. The content retrieval unit, based on the determined core theme, initiates a query to the structured brand knowledge base. This knowledge base tags all brand materials, such as product manuals, promotional materials, and user cases, with multi-dimensional tags such as "power saving," "long battery life," "five-year warranty," and "outdoor scenarios," and stores them in vector form. The retrieval process calculates the similarity between the semantic vector of the core theme and the vector similarity of content fragments in the knowledge base, returning results with tags such as "cost-effectiveness" and "durability." Multiple highly relevant candidate content fragments, linked to specific product information, brand stories, and marketing messages, are received by the content assembly and rendering unit. It sets length and depth constraints for the output content based on a level of detail, such as "medium," and prioritizes candidate fragments according to "real-time interest tags" in the real-time user profile, such as the repeated mention of "battery" in the current conversation, significantly increasing the weight of content with "battery"-related tags. Following a preset narrative logic template such as "pain point-solution-evidence," this unit intelligently trims and strings together high-weight content fragments that meet the detail requirements, filling in natural conjunctions to generate a grammatically coherent, focused, and personalized brand promotional text that aligns with the user's immediate interests. This text is then output to the downstream speech synthesis and image rendering module.

[0025] In step S5, the user behavior data related to the interaction effect includes at least three of the following: total interaction time, number of rounds of dialogue, user emotional state change value, type and depth of user-initiated questions, number of positive feedbacks on preset brand keywords, and completion rate of preset guiding actions.

[0026] In step S5, the rule update optimization based on user behavior data includes: S51: Based on machine learning models, train and analyze multi-round historical interaction data to establish a correlation model between "user profile features - virtual human and content strategy combination - performance indicators"; S52: Based on the association model, generate or update the strategy mapping rules, which define that when a specific user profile feature is identified, virtual human generation strategy and brand content generation strategy that have been verified by historical data to improve target performance indicators should be given priority.

[0027] In practical implementation, the process described in step S5 is implemented by an interaction effect evaluation and optimization module. This module first collects user behavior data through probes deployed throughout the system: the total interaction duration is calculated from the timestamps of the start and end of the session; the number of multi-turn dialogue rounds is counted by recording the number of complete question-and-answer sessions completed between the user and the virtual human; the change value of the user's emotional state is calculated by comparing the emotional state values, such as the difference in positive emotional scores, obtained from the analysis at the start and end of the session in step S1; the type and depth of the user's proactive questions are identified by the natural language processing module through intent classification and semantic depth analysis of the statements raised by the user that are not direct responses to the virtual human's questions; the number of positive feedbacks for preset brand keywords such as "good", "like", and "awesome" is counted by real-time monitoring of the dialogue text and matching it with the keyword table; the completion rate of preset guiding actions such as "click to view details" and "claim coupon" is counted by reporting the user's actual click behavior through the event listener of the front-end interface.

[0028] The process of updating and optimizing rules based on user behavior data is completed by an offline model training and strategy management service. Specifically, in step S51, the system periodically (e.g., daily) collects massive amounts of historical interaction data, including user profile features from S2 (e.g., "young male tech enthusiast"), virtual human and content strategy combinations from S3 (e.g., "professional image + in-depth technical explanation strategy"), and multiple performance indicators as mentioned above. This data is then comprehensively calculated into an "interaction performance score" as training samples. Machine learning algorithms such as random forests, gradient boosting decision trees, or deep neural networks are used to train these samples to establish a predictive association model. This model can learn the potential impact and contribution of different user profile features under different strategy combinations on various performance indicators, such as the "interaction performance score." In step S52, the system automatically generates or updates strategy mapping rules based on the trained association model. This process can use model interpretation techniques such as SHAP value analysis to identify the strategy combinations that have the most positive impact on the target performance indicator, such as improving the "guided action completion rate," and then update an online service's "strategy recommendation query table" or rule engine. For example, the correlation model might discover that for customers profiled as "price-sensitive hesitant customers," a strategy combination of "friendly image + emphasis on discounts + customer case studies" has historically significantly increased "coupon redemption rates." The system would then generate a new rule to prioritize this strategy combination when the real-time user profile analysis matches "price-sensitive hesitant customers." This mechanism allows the system's interaction strategies to continuously evolve based on real data feedback, achieving intelligent, data-driven personalized optimization.

[0029] Furthermore, this solution also proposes a virtual human-computer interaction system for brand promotion, which implements the aforementioned virtual human-computer interaction method for brand promotion, including: The multimodal perception module is configured to collect and analyze the user's multimodal data and generate the user's real-time status information. The user profiling and context analysis engine is connected to the multimodal perception module and configured to generate comprehensive interactive context instructions based on the real-time status information and historical interaction data. The virtual human adaptive generation engine is connected to the user profile and context analysis engine and is configured to determine the virtual human's image presentation parameters and interaction style parameters based on the comprehensive interactive context instructions. The brand knowledge base and content dynamic generation module are connected to the user profile and context analysis engine. It stores a structured brand knowledge base and is configured to dynamically generate personalized brand promotion content based on the comprehensive interactive context instructions. The rendering output module is connected to the virtual human adaptive generation engine and the brand knowledge base and content dynamic generation module, respectively, and is configured to integrate virtual human instances with brand promotional content to form an interactive response and output it. The interaction effect evaluation and optimization module is connected to the rendering output module and the multimodal perception module. It is configured to collect user behavior data and update optimization rules based on this data. The optimization rules are fed back to at least one of the user profile and context analysis engine, the virtual human adaptive generation engine, and the brand knowledge base and content dynamic generation module to adjust their subsequent processing strategies.

[0030] In summary, the advantages of this invention are: through multimodal perception and contextual analysis, it enables real-time personalized dynamic generation of virtual human images, interaction styles, and promotional content, allowing brand promotion to accurately adapt to the real-time status and historical preferences of different users, greatly enhancing the naturalness, attractiveness, and user resonance of the interaction.

[0031] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A virtual human-computer interaction method for brand promotion, characterized in that, include: S1: Collect multimodal data of the user during the interaction process through the user terminal, and perform real-time analysis on the multimodal data to generate real-time status information of the user; S2: Generate a comprehensive interactive scenario instruction based on the real-time status information and the pre-stored user historical interaction data; S3: Based on the comprehensive interactive context instructions, execute the virtual human adaptive generation step and the brand content dynamic generation step in parallel; The virtual human adaptive generation step is used to dynamically determine at least one image presentation parameter and interaction style parameter of the virtual human; The aforementioned brand content dynamic generation step is used to dynamically select and assemble personalized brand promotional content from a structured brand knowledge base. S4: The virtual human instance generated based on the image presentation parameters and interaction style parameters is merged and rendered with the personalized brand promotion content to form and output an interactive response for the current user; S5: During the interaction and after each interaction, collect user behavior data related to the interaction effect, and update and optimize the rules based on the user behavior data; The optimization rules are used to adjust the logic for generating comprehensive interactive context instructions in step S2 and / or the strategies for adaptive virtual human generation and dynamic brand content generation in step S3 during subsequent interactions.

2. The virtual human-computer interaction method for brand promotion according to claim 1, characterized in that, In step S1, the real-time analysis of the multimodal data includes: S11: Perform image analysis on the collected visual data to identify at least one of the user's facial expression, posture, and direction of attention, and infer the user's primary emotional state and level of engagement in the interaction. S12: Perform speech recognition and speech emotion analysis on the collected audio data to obtain the text information input by the user and infer the user's second emotional state; S13: Perform natural language understanding on the text information to identify the user's interaction intent and key entities in the text; The real-time status information includes at least the emotional state obtained by fusing the first emotional state and / or the second emotional state, the interaction engagement, the interaction intent, and the key entities.

3. The virtual human-computer interaction method for brand promotion according to claim 2, characterized in that, In step S2, the generation of comprehensive interactive context instructions specifically includes: S21: Based on the real-time status information, generate a real-time user profile including real-time sentiment tags, real-time interest tags, and real-time interaction stage tags; S22: The real-time user profile is associated and fused with the user's historical interaction data to update and form a long-term user profile, which includes at least historical interest preference tags and historical behavior pattern tags. S23: Combining the real-time user profile, the long-term user profile, and the context information of the current interaction, generate the comprehensive interaction context instruction, which includes the target user feature identifier and the recommended interaction strategy identifier.

4. The virtual human-computer interaction method for brand promotion according to claim 3, characterized in that, In step S3, the virtual human adaptive generation step specifically includes: S31: Based on the target user feature identifier in the comprehensive interactive context instruction, select suitable virtual human basic image, clothing and accessory resources from the virtual human resource library; S32: Determine the corresponding facial expressions and body movement sequence of the virtual human based on the emotional state in the real-time status information; S33: Determine the virtual human's voice parameters and dialogue style template based on the recommended interaction strategy identifier in the comprehensive interactive context instruction; The image presentation parameters include at least the selected basic image, clothing, accessories, facial expressions, and action sequences, and the interaction style parameters include at least the voice parameters and dialogue style templates.

5. A virtual human-computer interaction method for brand promotion according to claim 4, characterized in that, In step S3, the dynamic generation step of brand content specifically includes: S34: Determine the core theme and level of detail of the brand promotion content based on the target user feature identifier and recommendation interaction strategy identifier in the comprehensive interactive context instruction; S35: Based on the core theme, retrieve matching product information tags, brand story tags, and marketing script tags from the structured brand knowledge base; S36: Based on the level of detail and the real-time interest tags in the real-time user profile, prioritize and crop the retrieved content materials to generate coherent personalized brand promotion content.

6. The virtual human-computer interaction method for brand promotion according to claim 5, characterized in that, In step S5, the user behavior data related to the interaction effect includes at least three of the following: total interaction time, number of rounds of dialogue, user emotional state change value, type and depth of user-initiated questions, number of positive feedbacks on preset brand keywords, and completion rate of preset guiding actions.

7. A virtual human-computer interaction method for brand promotion according to claim 6, characterized in that, In step S5, the rule update optimization based on user behavior data includes: S51: Based on machine learning models, train and analyze multi-round historical interaction data to establish a correlation model between "user profile features - virtual human and content strategy combination - performance indicators"; S52: Based on the association model, generate or update the strategy mapping rules, which define that when a specific user profile feature is identified, virtual human generation strategy and brand content generation strategy that have been verified by historical data to improve target performance indicators should be given priority.

8. A virtual human-computer interaction system for brand promotion, characterized in that, A virtual human-computer interaction method for brand promotion as described in any one of claims 1-7 includes: The multimodal perception module is configured to collect and analyze the user's multimodal data and generate the user's real-time status information. The user profiling and context analysis engine is connected to the multimodal perception module and configured to generate comprehensive interactive context instructions based on the real-time status information and historical interaction data. The virtual human adaptive generation engine is connected to the user profile and context analysis engine and is configured to determine the virtual human's image presentation parameters and interaction style parameters based on the comprehensive interactive context instructions. The brand knowledge base and content dynamic generation module are connected to the user profile and context analysis engine. It stores a structured brand knowledge base and is configured to dynamically generate personalized brand promotion content based on the comprehensive interactive context instructions. The rendering output module is connected to the virtual human adaptive generation engine and the brand knowledge base and content dynamic generation module, respectively, and is configured to integrate virtual human instances with brand promotional content to form an interactive response and output it. The interaction effect evaluation and optimization module is connected to the rendering output module and the multimodal perception module. It is configured to collect user behavior data and update optimization rules based on this data. The optimization rules are fed back to at least one of the user profile and context analysis engine, the virtual human adaptive generation engine, and the brand knowledge base and content dynamic generation module to adjust their subsequent processing strategies.