Multi-scene digital human interaction management method based on artificial intelligence
By adopting an AI-based multi-scenario digital human interaction management method, combined with reinforcement learning algorithms and multi-dimensional evaluation models, the problems of scenario adaptability and multimodal data processing in traditional digital human interaction technologies have been solved. This enables the generation and real-time optimization of personalized interaction strategies, thereby improving user experience and interaction efficiency.
Patent Information
- Application Number
- CN202610042288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional digital human interaction technologies suffer from weak scene adaptability, imperfect multimodal data processing mechanisms, and an inability to detect interaction deviations and deficiencies in a timely manner, thus failing to meet users' needs for high-quality, personalized, and long-lasting interaction.
An AI-based multi-scenario digital human interaction management method is adopted, including deep scene semantic analysis, multimodal interaction intent prediction, dynamic interaction strategy generation, real-time digital human behavior driving, and intelligent interaction quality evaluation modules. Combined with reinforcement learning algorithms and multi-dimensional evaluation models, personalized interaction strategies are generated and optimized in real time.
It significantly enhances the intelligence and personalization of digital human interaction in multiple scenarios, ensures the continuous stability of interaction quality, and improves user experience and interaction efficiency through real-time evaluation and early warning mechanisms.
Smart Images

Figure CN121501151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive management technology, specifically to a multi-scenario digital human interactive management method based on artificial intelligence. Background Technology
[0002] Multi-scenario digital humans represent a deep integration of virtual digital technology and artificial intelligence. Digital humans have been widely applied in diverse scenarios such as customer service, entertainment interaction, and intelligent navigation. The scene adaptability of the interaction, the accuracy of intent recognition, and the personalized experience have become core factors affecting users' willingness to use and commercialization. However, with the continuous growth of demand for multi-scenario digital human interaction, various practical shortcomings have gradually emerged.
[0003] Traditional digital human interaction technologies have weak scene adaptability and imperfect multimodal data processing mechanisms. The generation of interaction strategies has obvious rigidity and limitations. At the same time, they have not built an evaluation model that covers multiple dimensions such as user feedback, system operation data, and behavior matching degree. They cannot detect interaction deviations and deficiencies in a timely manner, so they cannot reasonably judge the interaction performance based on the interaction effect and make timely adjustments and optimizations. They cannot meet users' needs for high-quality, personalized, and long-lasting interaction, which is not conducive to improving user experience and interaction efficiency.
[0004] To address the aforementioned technical shortcomings, a solution is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-scenario digital human interaction management method based on artificial intelligence, so as to solve the technical defects mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-scenario digital human interaction management method based on artificial intelligence, comprising the following steps: Step 1: The scene semantic deep analysis module performs deep semantic analysis based on multi-source scene information to extract core scene requirements, interaction constraints, and environmental feature parameters. Step 2: The multimodal interaction intent prediction module integrates user multimodal interaction data to predict the user's core interaction intent and potential needs, and outputs the intent confidence score. Step 3: The dynamic interaction strategy generation module generates personalized interaction strategies that adapt to the current scene, user intent, and potential needs based on the scene semantic parsing results and multimodal intent prediction results, combined with reinforcement learning algorithms, thus clarifying the specific execution plan of the digital human. Step 4: The real-time behavior driving module of the digital human parses the standardized strategy instruction set and drives the digital human to complete multi-dimensional interactive behaviors, including voice playback, action execution and facial expression changes. Step 5: The intelligent interaction quality assessment module collects multi-source feedback data in real time during the interaction process and evaluates the interaction quality through a multi-dimensional comprehensive scoring model.
[0007] Furthermore, in step one, the scene semantic deep parsing module acquires multi-source scene information, including scene type information, environmental parameter data, scene context text, and user scene behavior data. The collected multi-source scene data is then input into a pre-trained scene semantic parsing model. First, the text-based scene context is segmented, entity-recognized, and semantically labeled to extract core scene keywords and map them into scene requirement vectors. Then, the environmental parameter data is normalized to generate scene constraint features. Next, the user scene behavior data is subjected to temporal feature extraction to determine the user's interaction priority in the scene. Finally, the core scene requirement vector, scene interaction constraint parameter set, and scene environment feature matrix are output.
[0008] Furthermore, in step two, the multimodal interaction intent prediction module collects user interaction data and preprocesses each modal data. After preprocessing, the features of each modality are input into the multimodal fusion intent prediction model, which outputs the user's current core interaction intent and intent confidence. At the same time, based on the preset intent-potential need association rule library, the module predicts the user's potential needs in combination with the core needs of the scenario, and finally outputs the core interaction intent tag, potential need description, and intent confidence.
[0009] Furthermore, during the user interaction data collection process, voice data is collected through a microphone and converted into a text sequence, facial expression data is collected through a camera to capture facial images and extract facial key point features, motion data is extracted through motion sensing devices or image recognition technology to extract limb joint motion parameters, and text data is collected through the user's manually input text information in the interactive interface input box. During the preprocessing of each modality of data, for the speech-text sequence, a denoising algorithm is used to remove environmental noise interference, for the facial expression key point features, for the body movement parameters, time sequence alignment is performed, and for the text data, word segmentation and stop word removal are performed.
[0010] Furthermore, in step three, the dynamic interaction strategy generation module calculates the matching degree between the core requirements of the scenario and the core interaction intent. If the matching degree is ≥0.7, the basic interaction strategy is directly generated based on the core interaction intent; if the matching degree is <0.7, the basic strategy is adjusted in combination with the description of potential requirements. The boundary conditions of the interaction strategy are determined based on the set of scene interaction constraint parameters. The matched intent information, potential needs and boundary conditions are input into the strategy generation model of fusion reinforcement learning to generate personalized interaction strategies that include voice content, action instructions, facial expression state and response timing, and then encapsulated into a standardized strategy instruction set.
[0011] Furthermore, in step four, the digital human behavior real-time driving module receives a standardized strategy instruction set, decodes various codes, and obtains directly executable speech text sequences, joint motion parameters, facial muscle control parameters, and timing control timestamps. The decoded parameters are optimized. For the speech-text sequence, TTS technology is used to synthesize speech audio, and the speech rate and tone are adjusted by combining tone style parameters. For joint motion parameters, a smooth interpolation algorithm is used to avoid abrupt movements. Facial muscle control parameters are matched and verified with the emotional tendency of the speech. The optimized parameters are transmitted to the digital human rendering engine to drive the digital human to synchronously perform speech playback, action display and facial expression changes. In addition, during the execution of interactive behavior, the actual interaction data of the digital human is collected in real time and fed back to the interaction quality intelligent evaluation module.
[0012] Furthermore, the specific operation process of the interaction quality intelligent assessment module includes: A multi-source feedback data collection system is constructed, including user feedback data, system operation data, and behavior matching data. A multi-dimensional comprehensive scoring model is used to calculate the comprehensive interaction quality score Q. Preset upper and lower limits of the comprehensive score Qmax and Qmin are retrieved. The comprehensive interaction quality score Q is compared with the preset upper and lower limits of the comprehensive score Qmax and Qmin respectively. If Q≥Qmax, the "Excellent Interaction Quality" label is assigned; if Qmax>Q>Qmin, the "Good Interaction Quality" label is assigned; if Q≤Qmin, the "Unsatisfactory Interaction Quality" label is assigned.
[0013] Furthermore, the interaction quality intelligent assessment module is connected to the interaction adjustment management and early warning module, which in turn is connected to the intelligent management terminal. The interaction quality intelligent assessment module sends the interaction quality assessment tag to the interaction adjustment management and early warning module. The interaction adjustment management and early warning module analyzes the tag to determine whether to generate an adjustment early warning signal. If an adjustment early warning signal is generated, it is sent to the intelligent management terminal. When the intelligent management terminal receives the adjustment early warning signal, it issues a corresponding early warning.
[0014] Furthermore, the specific analysis process of the interactive adjustment management early warning module is as follows: Set a management cycle with a duration of X1. When the duration reaches X1, obtain all interaction quality assessment tags received within the management cycle. Calculate the ratio of the number of times the "interaction quality is unqualified" tag is assigned to the total number of interaction quality assessment tags, and subtract the ratio result from the value 1 to obtain the interaction qualification coefficient. The interaction pass coefficient is compared with the preset interaction pass coefficient threshold. If the interaction pass coefficient does not exceed the preset interaction pass coefficient threshold, an adjustment warning signal is generated.
[0015] Furthermore, if the interaction pass coefficient exceeds the preset interaction pass coefficient threshold, the comprehensive interaction quality score Q corresponding to all interaction quality evaluation tags is obtained. The average of all comprehensive interaction quality scores Q is calculated to obtain the interaction quality feature value. The interaction pass coefficient and the interaction quality feature value are weighted and summed to obtain the adjustment warning coefficient. The adjustment warning coefficient is compared with the preset adjustment warning coefficient threshold. If the adjustment warning coefficient does not exceed the preset adjustment warning coefficient threshold, an adjustment warning signal is generated.
[0016] Compared with the prior art, the beneficial effects of the present invention are: In this invention, by combining reinforcement learning algorithms, strategies are dynamically adjusted according to scenario requirements and user intent. The real-time digital human behavior driving module drives the digital human to complete multi-dimensional interactive behaviors based on a standardized policy instruction set. The intelligent interaction quality evaluation module comprehensively and objectively evaluates the interaction effect, greatly improving the intelligence and personalization of digital human interaction in multiple scenarios.
[0017] In this invention, the interaction quality tags and scores within the management cycle are analyzed by the interaction adjustment management early warning module. Based on this, it is determined whether to generate an adjustment early warning signal. When an adjustment early warning signal is generated, a corresponding early warning is issued through the intelligent management terminal to promptly remind optimization and improvement, thereby ensuring the continuous stability of the digital human interaction quality. Attached Figure Description
[0018] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 2 This is a system block diagram of Embodiment 1 of the present invention; Figure 3 This is a system block diagram of Embodiment 2 of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1: As Figure 1-2 As shown, the multi-scenario digital human interaction management method based on artificial intelligence proposed in this invention includes the following steps: Step 1: The scene semantic deep analysis module performs deep semantic analysis based on multi-source scene information, accurately extracts the core requirements of the scene, interaction constraints and environmental feature parameters, and sends them to the multimodal interaction intent prediction module and the dynamic interaction strategy generation module respectively. The former integrates scene context constraints into the intent prediction process, and the latter uses scene data as the basis for strategy generation to solve the problem of poor multi-scene adaptability of the system and improve the scene adaptability of the interaction. Specifically, the scene semantic deep analysis module first acquires multi-source scene information, including scene type information (such as scene types like office, entertainment, and customer service obtained through scene identification sensors or user preset input), environmental parameter data (collecting ambient temperature and humidity, light intensity, and noise decibel values through temperature and humidity sensors, light sensors, and noise sensors), scene context text (such as meeting topics in office scenes, interactive task descriptions in entertainment scenes, and business type descriptions in service scenes), and user scene behavior data (collecting user dwell time, operation actions, and movement trajectories in the scene through cameras and motion capture devices). Subsequently, the collected multi-source scene data is input into a pre-trained scene semantic parsing model. This model is based on the Transformer architecture and incorporates a scene feature attention mechanism. It first performs word segmentation, entity recognition and semantic role labeling on the text-based scene context, extracts the core keywords of the scene (such as "report" and "approval" in office scenarios, and "consultation" and "complaint" in service scenarios) and maps them into scene demand vectors. The environmental parameter data is then normalized to generate scene constraint features (e.g., when the noise level is ≥60 dB, constraint parameters such as "voice volume increase" and "interaction statement simplification" are generated). Then, the user scene behavior data is subjected to temporal feature extraction to determine the user's interaction priority in the scene (e.g., if the dwell time is ≥5 minutes and the actions are concentrated, it is judged as a high-priority interaction). Finally, the core requirement vector of the scene, the set of scene interaction constraint parameters, and the scene environment feature matrix are output to provide accurate basic scene data support for subsequent modules.
[0021] Step 2: The multimodal interaction intent prediction module integrates user multimodal interaction data to predict the user's core interaction intent and potential needs, and outputs intent confidence to improve the accuracy and reliability of interaction intent prediction, avoid misjudgment of intent caused by single-modal data, and together with the output data of the scene semantic deep analysis module, it determines the core direction and specific content of the interaction strategy. Specifically, firstly, the multimodal interaction intent prediction module collects user interaction data. Among them, voice data is collected through a microphone and converted into a text sequence, facial expression data is collected through a camera to capture facial images and extract facial key point features, motion data is extracted through motion sensing devices or image recognition technology to extract limb joint motion parameters (such as arm swing angle and body posture), and text data is collected through the input box of the interactive interface to collect text information manually entered by the user. Subsequently, the data of each modality were preprocessed: for the speech-text sequence, a denoising algorithm was used to remove environmental noise interference; for the facial expression key point features, standardization was performed to eliminate the influence of lighting and angle; for the body movement parameters, time sequence alignment was performed to ensure data synchronization; and for the text data, word segmentation and stop word removal were performed. After preprocessing, the modal features are input into the multimodal fusion intent prediction model. This model fuses the modal features through a dynamic weight allocation mechanism (prioritizing the weight of modal features with complete information and high clarity), and then performs intent recognition through a classification network to output the user's current core interaction intent (such as consultation, command issuance, emotional exchange, task collaboration, etc.) and intent confidence (value range 0-1, confidence ≥0.8 is considered a highly reliable intent). Meanwhile, based on the preset intent-potential demand association rule library, and combined with the core needs of the scenario, the system predicts the user's potential needs (e.g., when the core intent is "inquire about product price", the potential needs are determined to be "learn about promotional activities" and "compare competitor prices"). Finally, the system outputs the core interaction intent tag, potential demand description and intent confidence, and transmits them to the dynamic interaction strategy generation module.
[0022] Step 3: The dynamic interaction strategy generation module generates personalized interaction strategies that adapt to the current scene, user intent, and potential needs based on the scene semantic parsing results and multimodal intent prediction results, combined with reinforcement learning algorithms. It clarifies the specific execution plan of the digital human and outputs it to the real-time behavior driving module of the digital human to guide the execution of interaction behavior. This solves the problem of fixed and rigid system interaction strategies and helps to improve the personalization and accuracy of interaction. Specifically, firstly, the dynamic interaction strategy generation module receives the core requirement vector of the scene and the set of scene interaction constraint parameters output by the scene semantic deep parsing module, as well as the core interaction intent label, potential requirement description and intent confidence output by the multimodal interaction intent prediction module. Subsequently, the matching degree between the core needs of the scenario and the core interaction intent is calculated. The matching degree is achieved by the cosine similarity algorithm. If the matching degree is ≥0.7, the basic interaction strategy is directly generated based on the core interaction intent. If the matching degree is <0.7, the basic strategy is adjusted in combination with the description of potential needs (e.g., when the scenario need is "efficient communication" and the intent is "emotional exchange", the strategy is adjusted to "simplified emotional response + rapid transmission of core information"). Next, the boundary conditions of the interaction strategy are determined based on the set of scene interaction constraint parameters. For example, when the noise is high, the voice response volume is increased and the interaction sentences are simplified; in public scenes, the action amplitude is reduced and the facial expressions are moderately restrained. For example, when the constraint parameter is "noise ≥ 60 decibels", the voice response volume is set to ≥ 75 decibels and the interaction sentence length is set to ≤ 15 words. When the constraint parameter is "public scene", the action amplitude is set to ≤ 30° and the degree of facial expression exaggeration is set to ≤ 0.5 (value range 0-1). Subsequently, the matched intent information, potential needs, and boundary conditions are input into a policy generation model that integrates reinforcement learning. This model uses "maximizing user satisfaction" and "optimizing interaction efficiency" as dual reward functions to generate personalized interaction strategies that include speech content (speech text, tone style, speech rate and intonation), action instructions (joint motion parameters, action duration), facial expression state (facial key point control parameters, expression duration), and response timing (synchronous triggering time difference between speech and action, matching timing between expression and semantics). These strategies are then encapsulated into a standardized policy instruction set (including speech text encoding, action parameter encoding, facial expression parameter encoding, and timing control parameters).
[0023] Furthermore, if the intent confidence level is less than 0.8, the dynamic interaction strategy generation module will also generate supplementary confirmation instructions (such as "Do you want to inquire about product prices?"), which will be integrated into the strategy instruction set to reduce the risk of intent misjudgment. Finally, the strategy instruction set will be output to the digital human behavior real-time driving module.
[0024] Step 4: The real-time driving module for digital human behavior parses the standardized strategy instruction set and drives the digital human to complete multi-dimensional interactive behaviors such as voice playback, action execution, and facial expression changes, ensuring the naturalness, synchronization, and real-time nature of the interactive behavior, and enhancing the user's immersion and experience. Specifically, firstly, the real-time driving module for digital human behavior receives a standardized set of policy instructions output by the dynamic interaction policy generation module and decodes various types of encoding: converting speech text encoding into a synthesizable text sequence, converting action parameters into motion parameters of the digital human's skeletal joints (such as shoulder joint angle, elbow joint flexion, and movement speed), converting facial expression parameters into facial muscle control parameters (such as eye muscle contraction, corner of the mouth upward angle, and eyebrow relaxation), and converting timing control parameters into trigger timestamps for each interactive behavior (such as speech playback start time t1, action trigger time t2, and facial expression change time t3, ensuring t2-t1≤0.2 seconds and t3-t1≤0.1 seconds). In other words, the decoded data yields directly executable speech text sequences, joint motion parameters, facial muscle control parameters, and timing control timestamps. Subsequently, the decoded parameters were optimized: For the speech-text sequence, TTS technology was used to synthesize speech audio, and the speech rate (150 words / minute for office scenarios and 200 words / minute for entertainment scenarios) and intonation were adjusted in combination with tone style parameters (the intonation rises by 5% when the intention is affirmative and falls by 3% when the intention is negative). For joint motion parameters, a smooth interpolation algorithm was used to avoid abrupt movements (such as from standing to raising an arm, the angle of the intermediate joints was calculated by interpolation to ensure the continuity of the movement trajectory). Facial muscle control parameters were matched and verified with the emotional tendency of the speech (such as when the emotional tendency of the speech is "pleasant", ensuring that the upward angle of the corners of the mouth is ≥15° and the degree of eye muscle contraction is ≥0.3). The optimized parameters are then transmitted to the digital human rendering engine to drive the digital human to synchronously perform voice playback, action display and facial expression changes. That is, synchronously play voice audio, control the skeletal joints to perform actions according to motion parameters, adjust facial muscles to achieve facial expression changes, and at the same time ensure the synchronization of voice, action and expression through the timing control module (synchronization error ≤ 0.1 seconds). Furthermore, during the execution of interactive behaviors, the actual interaction data of the digital human (such as voice playback duration, action completion rate, and facial expression synchronization accuracy) is collected in real time and fed back to the intelligent interaction quality evaluation module to ensure the accuracy of subsequent interaction effect evaluation.
[0025] Step 5: The intelligent interaction quality assessment module collects multi-source feedback data in real time during the interaction process. It evaluates the interaction quality through a multi-dimensional comprehensive scoring model, providing a comprehensive and objective assessment of the interaction effect and accurately identifying areas for improvement. Compared to traditional single-indicator assessment methods, this approach is more scientific and comprehensive, facilitating targeted optimization and improvement measures. The specific operation process is as follows: A multi-source feedback data collection system is constructed, including user feedback data (collecting user subjective satisfaction ratings through interactive interface pop-ups, user facial expressions through camera, and user voice feedback through microphone), system operation data (obtaining voice response latency, action completion accuracy, and facial expression synchronization accuracy from the digital human behavior real-time driving module), and behavior matching data (comparing the consistency between the digital human's actual interactive behavior and the strategy instruction set to obtain the speech matching degree and action matching degree). A multi-dimensional comprehensive scoring model is used to calculate the comprehensive interaction quality score Q, as shown in the following formula: Q=0.4×S+0.2×(1-T / Tmax)+0.1×A+0.1×E+0.1×M+0.1×N; Where S is the user's subjective satisfaction rating, with a value range of 1-5 points (1 point = very dissatisfied, 2 points = dissatisfied, 3 points = neutral, 4 points = satisfied, 5 points = very satisfied). The method of obtaining it is as follows: after a single interaction, the system will automatically pop up the rating option through the interactive interface pop-up window. The user clicks the corresponding score to complete the submission, and the module reads the submission result in real time. T is the voice response delay time, which represents the time interval from the output of the policy instruction set by the dynamic interaction policy generation module to the start of voice playback by the digital human. It is obtained by using the high-precision timer built into the module (timer precision 0.001s). The timer is triggered to start when the policy instruction set is received and to end when the digital human voice playback interface is detected to start. The time difference recorded by the timer is T. Tmax is the preset maximum acceptable response delay time, fixed at 1.5s. It is obtained by testing the interaction experience of at least 3,000 users of different ages in various scenarios such as office, entertainment, and service. The results show that when T > 1.5s, user satisfaction drops by more than 40%. Therefore, T_max is set to 1.5s to ensure that the real-time interaction meets user needs. A represents the accuracy of action completion, ranging from 0 to 1. It indicates the degree of consistency between the actual action performed by the digital human and the standard action in the strategy instruction set. The method for obtaining A is as follows: standard joint motion parameters (such as shoulder joint angle, elbow joint flexion, etc.) in the strategy instruction set are constructed into a standard vector, and the actual joint motion parameters collected by the real-time behavior driving module of the digital human are constructed into an actual vector. The results are obtained by calculating the cosine similarity algorithm. E represents the facial expression synchronization accuracy, ranging from 0 to 1. It indicates the degree of consistency between the actual facial expression presented by the digital human and the standard facial expression in the policy instruction set. The method for obtaining it is as follows: standard facial muscle control parameters (such as the angle of mouth corners and the degree of eye muscle contraction) in the policy instruction set are constructed into a standard vector, and the actual facial muscle control parameters collected by the real-time behavior driving module of the digital human are constructed into an actual vector. The result is obtained by calculating the cosine similarity algorithm. M represents the speech matching degree, ranging from 0 to 1. It indicates the degree of consistency between the actual speech text played by the digital human and the preset speech text in the policy instruction set. The method of obtaining it is as follows: convert the actual speech played by the digital human into text through speech-to-text (ASR) technology, extract the preset speech text in the policy instruction set, and use the edit distance algorithm to calculate the similarity between the two. N is the action matching degree, which ranges from 0 to 1. It represents the degree of consistency between the timing and parameters of the actual actions performed by the digital human and the standard actions in the policy instruction set. The method of obtaining it is to align the timing sequence of the standard actions (the curve of the change of each joint motion parameter over time) with the timing sequence of the actual actions, calculate the absolute error of the parameters at each time node, take the average value of the errors of all nodes and normalize it. Retrieve the preset upper limit threshold Qmax and the preset lower limit threshold Qmin of the comprehensive score. Compare the comprehensive score Q of the interaction quality with the preset upper limit threshold Qmax and the preset lower limit threshold Qmin respectively. If Q≥Qmax, the "Excellent Interaction Quality" label is assigned. If Qmax>Q>Qmin, the "Good Interaction Quality" label is assigned. If Q≤Qmin, the "Unsatisfactory Interaction Quality" label is assigned.
[0026] Example 2: Figure 3 As shown, the difference between this embodiment and Embodiment 1 is that the interaction quality intelligent assessment module is communicatively connected to the interaction adjustment management early warning module, the interaction adjustment management early warning module is communicatively connected to the intelligent management terminal, the interaction quality intelligent assessment module sends the interaction quality assessment tag to the interaction adjustment management early warning module, and the interaction adjustment management early warning module analyzes the data to determine whether to generate an adjustment early warning signal. When an adjustment warning signal is generated, it is sent to the intelligent management terminal. Upon receiving the adjustment warning signal, the intelligent management terminal issues a corresponding warning to promptly remind users to optimize and improve, ensuring the continuous and stable quality of digital human interaction. The specific analysis process is as follows: Set a management cycle with a duration of X1. When the duration reaches X1, obtain all interaction quality assessment tags received within the management cycle. Calculate the ratio of the number of times the "interaction quality is unqualified" tag is assigned to the total number of interaction quality assessment tags, and subtract the ratio result from the value 1 to obtain the interaction qualification coefficient. The interaction pass coefficient is compared with the preset interaction pass coefficient threshold. If the interaction pass coefficient does not exceed the preset interaction pass coefficient threshold, it indicates that the digital human interaction management performance is poor during the management period and corresponding adjustment and optimization measures need to be taken in time. Then, an adjustment warning signal is generated.
[0027] Furthermore, if the interaction pass coefficient exceeds the preset interaction pass coefficient threshold, the comprehensive interaction quality score Q corresponding to all interaction quality evaluation tags is obtained, and the average of all comprehensive interaction quality scores Q is calculated to obtain the interaction quality feature value. Furthermore, the interaction pass coefficient and the interaction quality feature value are weighted and summed to obtain the adjustment warning coefficient. That is, the interaction pass coefficient and the interaction quality feature value are respectively assigned corresponding preset weight coefficients, and the interaction pass coefficient and the interaction quality feature value are respectively multiplied by the corresponding preset weight coefficients, and the sum of the two sets of product results is marked as the adjustment warning coefficient. It should be noted that the larger the adjustment warning coefficient is, the worse the overall performance of digital human interaction management during the management period. The adjustment warning coefficient is compared with the preset adjustment warning coefficient threshold. If the adjustment warning coefficient does not exceed the preset adjustment warning coefficient threshold, it indicates that the overall performance of digital human interaction management during the management period is poor and corresponding adjustment and optimization measures need to be taken in a timely manner, and an adjustment warning signal is generated.
[0028] The working principle of this invention is as follows: In use, the scene semantic deep parsing module accurately extracts multi-source scene information; the multimodal interaction intent prediction module integrates multi-dimensional data and predicts interaction intent; the dynamic interaction strategy generation module combines reinforcement learning algorithms to dynamically adjust strategies according to scene requirements and user intent; the digital human behavior real-time driving module drives the digital human to complete multi-dimensional interaction behaviors based on a standardized strategy instruction set; the interaction quality intelligent evaluation module comprehensively and objectively evaluates the interaction effect; and the interaction adjustment management and early warning module determines whether to generate an adjustment early warning signal based on interaction quality data within a period. This achieves intelligent control of the entire chain from scene adaptation to interaction optimization. Relying on artificial intelligence technology, it significantly improves the intelligence, personalization, and reliability of digital human interaction in multiple scenarios, effectively meeting the interaction needs of multiple scenarios such as office, entertainment, and service, and significantly improving user experience and interaction efficiency.
[0029] In this invention, the threshold, preset value, or preset range settings are for result comparison and analysis to determine whether the result is good or bad. The magnitude of these values is determined by a combination of large-scale model analysis of sample data and human experience, and can also be appropriately adjusted based on seasonal or common-sense influence conditions. Similarly, the preset weight coefficients and influence factors are assigned specific values based on the magnitude of each parameter's influence on the result, ultimately reflecting the impact on the result. These settings are also determined by a combination of large-scale model analysis of sample data and human experience, and can also be appropriately adjusted based on seasonal or common-sense influence conditions.
[0030] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize it. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-scenario digital human interaction management method based on artificial intelligence, characterized in that: Includes the following steps: Step 1: The scene semantic deep analysis module performs deep semantic analysis based on multi-source scene information to extract core scene requirements, interaction constraints, and environmental feature parameters. Step 2: The multimodal interaction intent prediction module integrates user multimodal interaction data to predict the user's core interaction intent and potential needs, and outputs the intent confidence score. Step 3: The dynamic interaction strategy generation module generates personalized interaction strategies that adapt to the current scenario, user intent, and potential needs, clarifying the specific execution plan for the digital human; Step 4: The real-time behavior driving module of the digital human parses the standardized strategy instruction set and drives the digital human to complete multi-dimensional interactive behaviors. Step 5: The intelligent interaction quality assessment module collects multi-source feedback data in real time during the interaction process and evaluates the interaction quality through a multi-dimensional comprehensive scoring model.
2. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 1, characterized in that, In step one, the scene semantic deep parsing module acquires multi-source scene information, inputs the collected multi-source scene data into the pre-trained scene semantic parsing model, and finally outputs the scene core requirement vector, the scene interaction constraint parameter set, and the scene environment feature matrix.
3. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 1, characterized in that, In step two, the multimodal interaction intent prediction module collects user interaction data and preprocesses each modal data. After preprocessing, the features of each modality are input into the multimodal fusion intent prediction model, which outputs the user's current core interaction intent and intent confidence. At the same time, based on the preset intent-potential need association rule library, the module predicts the user's potential needs in combination with the core needs of the scenario.
4. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 3, characterized in that, During the user interaction data collection process, voice data is collected through a microphone and converted into a text sequence, facial expression data is collected through a camera to capture facial images and extract facial key point features, motion data is collected through motion sensing devices or image recognition technology to extract limb joint motion parameters, and text data is collected through the user's manually input text information in the interactive interface input box. During the preprocessing of each modality of data, for the speech-text sequence, a denoising algorithm is used to remove environmental noise interference, for the facial expression key point features, for the body movement parameters, time sequence alignment is performed, and for the text data, word segmentation and stop word removal are performed.
5. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 1, characterized in that, In step three, the dynamic interaction strategy generation module calculates the matching degree between the core requirements of the scenario and the core interaction intent. If the matching degree is ≥0.7, the basic interaction strategy is directly generated based on the core interaction intent. If the matching degree is less than 0.7, the basic strategy should be adjusted based on the description of potential needs. The boundary conditions of the interaction strategy are determined based on the set of scene interaction constraint parameters. The matched intent information, potential needs and boundary conditions are input into the strategy generation model that integrates reinforcement learning to generate personalized interaction strategies, which are then encapsulated into a standardized strategy instruction set.
6. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 1, characterized in that, In step four, the digital human behavior real-time driving module receives a standardized strategy instruction set, decodes various codes, and obtains directly executable speech text sequences, joint motion parameters, facial muscle control parameters, and timing control timestamps. The decoded parameters are optimized and then transmitted to the digital human rendering engine to drive the digital human to synchronously perform voice playback, action display, and facial expression changes. During the execution of interactive behaviors, the actual interaction data of the digital human is collected in real time and fed back to the interaction quality intelligent evaluation module.
7. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 6, characterized in that, The specific operation process of the interaction quality intelligent assessment module includes: A multi-source feedback data collection system is constructed, including user feedback data, system operation data, and behavior matching data. A multi-dimensional comprehensive scoring model is used to calculate the comprehensive interaction quality score Q. If Q≥Qmax, the interaction quality is labeled as "excellent"; if Qmax>Q>Qmin, the interaction quality is labeled as "good"; if Q≤Qmin, the interaction quality is labeled as "unsatisfactory".
8. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 7, characterized in that, The interactive quality intelligent assessment module communicates with the interactive adjustment management early warning module, which in turn communicates with the intelligent management terminal. The interactive adjustment management early warning module analyzes the data to determine whether to generate an adjustment early warning signal, and sends the signal to the intelligent management terminal when it is generated.
9. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 8, characterized in that, The specific analysis process of the interaction adjustment management early warning module is as follows: obtain all interaction quality assessment tags received within the management cycle, calculate the ratio of the number of times the "interaction quality unqualified" tag is assigned to the total number of interaction quality assessment tags, and subtract the ratio result from the value 1 to obtain the interaction qualification coefficient; if the interaction qualification coefficient does not exceed the preset interaction qualification coefficient threshold, an adjustment early warning signal is generated.
10. The multi-scenario digital human interaction management method based on artificial intelligence according to claim 9, characterized in that, If the interaction pass coefficient exceeds the preset interaction pass coefficient threshold, the interaction pass coefficient and the interaction quality feature value are weighted and summed to obtain the adjustment warning coefficient. If the adjustment warning coefficient does not exceed the preset adjustment warning coefficient threshold, an adjustment warning signal is generated.
Citation Information
Patent Citations
Digital interaction method and system based on artificial intelligence, and medium
CN117348736A
Digital human dynamic interaction method and system based on deep learning
CN120688535A
Digital human interaction method and system based on large model
CN121118962A