Virtual digital human interaction control method, device and equipment, storage medium and program product
By using multimodal fusion to obtain the voice activity, visual focus, and lip movement synchronization rate of interactive objects, and combining it with the dominance calculation logic, the problem of low recognition accuracy of traditional single-modal solutions in complex environments is solved, achieving higher recognition accuracy and interaction continuity for interactive objects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional single-modal solutions have low accuracy in identifying the main interactive object in complex environments, especially when noise exceeds limits or people obstruct the view, resulting in a high error rate. Furthermore, the lack of decision-making logic in multi-object interaction scenarios leads to insufficient interaction continuity.
A multimodal fusion method is adopted to obtain the voice activity, visual attention and lip movement synchronization rate of the interaction object, and combine the dominance calculation logic to dynamically adjust the weights, comprehensively judge the target interaction object, and combine historical interaction information to ensure the continuity of interaction.
It effectively reduces the impact of noise and occlusion on recognition results, improves the accuracy of judging the main interactive object, ensures the naturalness and immersion of the interaction, and adapts to the continuity of interaction in scenarios where multiple people take turns speaking.
Smart Images

Figure CN121833110A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an interactive control method, device, equipment, storage medium, and program product for a virtual digital human. Background Technology
[0002] In digital human interaction technology scenarios, the identification of the main interaction object in complex multi-person environments is always the primary interaction scenario. These scenarios are often accompanied by complex factors such as environmental noise, personnel movement, and limb occlusion, which places high demands on the digital human's ability to capture interaction targets.
[0003] Current traditional solutions to this problem are monomodal approaches, including purely audio-based solutions (such as sound source localization based on microphone arrays) and purely visual solutions (such as eye-tracking based on cameras). Monomodal solutions have inherent limitations: purely audio-based solutions have a high error rate (over 80%) in environments with excessive noise (e.g., noise exceeding 60dB); purely visual solutions have a high rate of misjudging eye-tracking when people are occluded or in scenes with side profiles.
[0004] Therefore, the traditional approach suffers from low accuracy in identifying the interaction objects. Summary of the Invention
[0005] This application provides an interactive control method, device, equipment, storage medium, and program product for virtual digital humans, which can improve the accuracy of identifying the interactive object.
[0006] To achieve the above objectives, this application adopts the following technical solution: Firstly, this application provides an interactive control method for a virtual digital human, including: Obtain the voice activity, visual focus, and lip-sync rate of each interactive object in the set of interactive objects; The dominance of each interactive object is obtained based on the voice activity, visual focus, and lip movement synchronization rate. The target interaction object is determined based on the dominance of each interaction object; Control the virtual digital human to interact with the target interactive object.
[0007] Optionally, the voice activity level of the interactive object is obtained in the following way: Acquire the voice data of the interactive object. The voice data includes the number of audio frames, the voice presence of each audio frame, and the voiceprint similarity. The voiceprint similarity is the similarity between the current voiceprint data of the interactive object and the historical voiceprint data obtained in the previous recognition. The voice activity level of the interactive object is determined based on the number of audio frames, the presence of speech in each audio frame, and the voiceprint similarity.
[0008] Optionally, the visual focus of the interactive object is obtained in the following ways: Acquire the first video data of the interactive object, the first video data including the frame number of the first video frame, instantaneous gaze focus, instantaneous facial orientation focus, and image quality factor; The visual attention of the interactive object is determined based on the frame number of the first video frame, instantaneous gaze attention, instantaneous facial orientation attention, and image quality factor.
[0009] Optionally, the lip-sync rate of the interactive object is obtained in the following way: Obtain the second video data of the interactive object, the second video data including the frame number of the second video frame and the lip movement activity; The lip movement synchronization rate of the interactive object is determined based on the frame number of the second video frame and the lip movement activity.
[0010] Optionally, determining the target interaction object based on the dominance of each interaction object includes: If there is only one interaction object whose dominance is greater than or equal to the first dominance threshold, then the interaction object with a dominance greater than the first dominance threshold is identified as the target interaction object. If there are at least two interactive objects whose dominance is greater than the second dominance threshold and less than the first dominance threshold, and there are no interactive objects whose dominance is greater than the first dominance threshold, then the voiceprint similarity and semantic coherence of each interactive object are determined respectively. The semantic coherence is the coherence between the current semantic data of the interactive object and the historical semantic data obtained from the previous identification. Based on the voiceprint similarity and semantic coherence, the interaction score of each interactive object is determined; The interaction object with the highest interaction score is identified as the target interaction object.
[0011] Optionally, the method further includes: If the identified speech activity belongs to the first interaction object, visual attention belongs to the second interaction object, and lip movement synchronization rate belongs to the third interaction object, then the weight of visual attention and the weight of lip movement synchronization rate in the process of calculating the dominance of each interaction object are increased, and the dominance of each interaction object is recalculated. The target interaction object is determined based on the recalculated dominance of each interaction object.
[0012] Optionally, determining the target interaction object based on the dominance of each interaction object includes: If any interaction object has a dominance less than the second dominance threshold, then the target interaction object is determined to be empty. The control of the virtual digital human to interact with the target interactive object includes: When the target interaction object is not empty, control the virtual digital human to interact with the target interaction object; When the target interaction object is empty, the virtual digital human is controlled to remain silent.
[0013] Secondly, this application provides an interactive control device for a virtual digital human, comprising: The acquisition module is used to acquire the voice activity, visual focus, and lip-movement synchronization rate of each interactive object in the interactive object set; The data processing module is used to obtain the dominance of each interactive object based on the speech activity, visual focus, and lip movement synchronization rate; and to determine the target interactive object based on the dominance of each interactive object. The control module is used to control the interaction between the virtual digital human and the target interactive object.
[0014] Thirdly, this application provides a computing device, including a memory and a processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of the first aspects.
[0015] Fourthly, this application provides a computer-readable storage medium for storing a computer program for performing the method as described in any one of the first aspects.
[0016] As can be seen from the above technical solution, this application has at least the following beneficial effects: This method integrates three major features—vocal activity, visual attention, and lip-movement synchronization rate—to calculate dominance. Vocal activity is verified through multi-dimensional checks using the number of audio frames, speech presence, and voiceprint similarity, avoiding interference from noise on a single audio signal. Visual attention combines instantaneous gaze, facial orientation, and image quality factors to reduce visual misjudgments caused by occlusion or side profiles. Lip-movement synchronization rate is verified by the temporal alignment of lip movements and speech, further filtering invalid interaction signals. These three features work together to form a multimodal cross-validation mechanism, effectively reducing the impact of complex factors such as noise and occlusion on the recognition results and improving the accuracy of identifying the target interaction object.
[0017] Furthermore, for complex scenarios such as simultaneous interaction among multiple users and modal information conflicts, this method designs a hierarchical decision-making logic and a dynamic weight adjustment method. When only one interaction object meets the dominance standard, the target can be directly locked to avoid redundant judgments. When there are multiple objects with high dominance, the interaction score is calculated by voiceprint similarity and semantic coherence, and objects with higher consistency with historical interactions and stronger dialogue coherence are prioritized to ensure the continuity of interaction in scenarios where multiple people take turns speaking. When voice, visual, and lip movement features belong to different objects, the dominance is recalculated by increasing the weight of visual and lip movement features to avoid misjudging the dominance result of a single modality. This method adapts to special interaction states such as the speaker not being a gazer, and greatly improves the adaptability of the method to diverse scenarios.
[0018] Furthermore, this method ensures the continuity of interaction through a multi-historical information association mechanism. The calculation of voice activity incorporates the similarity verification of previous voiceprint data, visual attention is judged based on the comprehensive judgment of multi-frame video data, and semantic coherence is directly linked to the historical dialogue content. This makes the recognition result not only dependent on real-time signals, but also deeply bound to the historical interaction state, reducing identity jumps caused by instantaneous interference. At the same time, the hierarchical decision-making logic avoids the problem of blind switching when there is no clear goal, ensuring that the digital human only initiates interaction when there is a clear goal, or selects the best option based on historical information when multiple objects are ambiguous, reducing the probability of interaction interruption and enhancing the naturalness and immersion of the user's interaction with the digital human.
[0019] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0020] Figure 1 A flowchart illustrating an interactive control method for a virtual digital human provided in this application embodiment; Figure 2 A schematic diagram of an interactive control device for a virtual digital human provided in an embodiment of this application; Figure 3 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation
[0021] The terms "first," "second," and "third," etc., used in this application specification and accompanying drawings are used to distinguish different objects, not to limit a specific order.
[0022] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0023] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first: Virtual digital humans refer to digital entities constructed using technologies such as computer graphics and artificial intelligence. They possess visual imagery and interactive capabilities, enabling them to engage in multimodal interactions with humans, including voice and vision, in virtual or real-world scenarios. They are widely used in scenarios such as virtual customer service, service robots, and metaverse social networking.
[0024] The primary interaction object is the target human being who engages in the main interactive behavior with the virtual digital human in an interactive scenario where multiple people are present simultaneously (such as actively speaking or continuously staring at the digital human). It is the subject that the digital human needs to respond to and pay attention to first.
[0025] In complex multi-person interaction scenarios of virtual digital humans, the primary technical problem to be solved is the insufficient recognition accuracy of traditional single-modal solutions in complex environments. Current pure audio solutions (such as microphone array sound source localization) rely on a single voice signal to determine the main interaction object. When the ambient noise exceeds 60dB, such as in public places like shopping malls and exhibition halls, the voice signal is easily submerged by noise, which not only leads to false positives for voice presence but also causes deviations in voiceprint recognition, ultimately resulting in an error rate of over 80% for the main interaction object. Pure vision solutions (such as camera eye tracking) lock onto the target only through visual signals. In scenarios where people are obstructed, faces are turned to the side, or in low light, it is difficult to guarantee the accuracy of extracting visual features such as instantaneous gaze focus and facial orientation, leading to an increased false positive rate for eye contact. Neither of these solutions can cope with the recognition failure caused by environmental interference, making it difficult to meet the basic interaction needs of digital humans.
[0026] Meanwhile, traditional solutions also suffer from a lack of decision-making logic and insufficient interaction continuity in multi-object interaction scenarios. On the one hand, they do not consider complex states such as multiple people speaking simultaneously or the speaker not being the observer. When multiple objects speak at the same time, it is impossible to effectively distinguish between the active interactor and the background speaker. When voice and visual signals point to different objects, there is a lack of an effective conflict arbitration mechanism, which easily leads to accidental switching of the main interaction object. On the other hand, traditional solutions rely solely on real-time signals to determine the target without associating historical interaction information, such as the voiceprint of the previous interacting object and the semantics of historical dialogue. When the instantaneous signal is interfered with, such as a sudden noise briefly masking the voice, it is easy to frequently switch the interaction object, causing the dialogue flow to be interrupted and seriously damaging the naturalness and immersion of the interaction.
[0027] In view of this, embodiments of this application provide an interactive control method for a virtual digital human, which can be executed by a processing device. The processing device can be a terminal or a server. Terminals include, but are not limited to, smartphones, tablets, laptops, personal digital assistants, or smart wearable devices. The server can be a cloud server, such as a central server in a central cloud computing cluster or an edge server in an edge cloud computing cluster. Alternatively, the server can be a server in a local data center. A local data center refers to a data center directly controlled by the user.
[0028] To address the issues of low accuracy, lack of multi-object decision-making, and insufficient interaction continuity in complex multi-person interaction scenarios of virtual digital humans, this application firstly overcomes the limitations of single-modal information by extracting key features from three interaction dimensions: speech, vision, and lip movement. The speech dimension uses speech activity to quantify vocal effectiveness and identity consistency; the vision dimension uses visual focus to quantify gaze state and visual signal reliability; and the lip movement dimension uses lip movement synchronization rate to verify the correlation between speech and lip movements. These three dimensions form a cross-validation mechanism to filter environmental interference. Secondly, dominance is introduced as an evaluation index, comprehensively quantifying the priority of interaction objects based on the three features to avoid misjudgment based on a single feature. Finally, dynamic decision-making logic is designed for different interaction scenarios, combining historical interaction information (voiceprint, semantics) to solve multi-object competition and modal conflict problems while ensuring interaction continuity. Ultimately, this achieves accurate identification of the main interaction object and natural interaction of the digital human in complex environments.
[0029] To make the technical solution of this application clearer and easier to understand, the interactive control method for a virtual digital human provided by the embodiments of this application will be described below with reference to the accompanying drawings. Figure 1 As shown, this figure is a flowchart of an interactive control method for a virtual digital human provided in an embodiment of this application. The method includes: S201, The processing device acquires the voice activity, visual focus, and lip-movement synchronization rate of each interactive object in the set of interactive objects.
[0030] The set of interactive objects refers to the entirety of all human objects that may interact with the virtual digital human in the virtual digital human interaction scenario. For example, multiple customers consulting with the digital human customer service in a shopping mall, or multiple users surrounding the digital human in a metaverse social scenario. It is not a single object, but a group within the scenario that meets the definition of potential interactors.
[0031] An interactive object is a single individual within a set of interactive objects, that is, a specific human being within a scene who may interact with the virtual digital human.
[0032] Voice activity level is an indicator used to quantify the validity and identity relevance of the voice signal of an interactive object. It is calculated by analyzing the voice data of the interactive object and reflects whether the object is effectively speaking and the consistency between the speaking identity and historical interactions. For example, an interactive object that speaks clearly and continuously with a voiceprint that matches the previous interaction will have a higher voice activity level. Voice activity level is obtained through the following methods: First, the processing device acquires the voice data of the interactive object. The voice data includes the number of audio frames, the presence of voice in each audio frame, and voiceprint similarity. Voiceprint similarity is the similarity between the current voiceprint data of the interactive object and the historical voiceprint data obtained from the previous recognition.
[0033] The frame count of an audio signal refers to the number of independent audio segments formed after a processing device divides a continuous raw audio signal into segments of fixed time length (e.g., 20ms / frame). For example, if a processing device acquires a 1-second audio signal, dividing it into 20ms / frame segments will result in 50 frames. The frame count is essentially a quantification of the time dimension of the audio signal.
[0034] The speech presence of each audio frame refers to the binary result obtained by the processing device after judging each audio frame whether the frame contains valid human speech. 1 indicates that it contains valid speech, and 0 indicates that it does not contain valid speech, such as only environmental noise or silence.
[0035] Voiceprint similarity refers to the quantitative value of similarity obtained by the processing device after comparing the current voiceprint data of the interactive object with the historical voiceprint data of the previous recognition through a voiceprint recognition algorithm. It is represented by a value of 0-1 or 0-100, with a higher value indicating stronger identity consistency. Among them, the historical voiceprint data is the voiceprint features stored by the processing device when the interactive object interacted with the digital human in the previous time.
[0036] The processing device first receives the original audio signal of the interactive object through a front-end microphone or other device. Then, it splits the continuous signal into discrete audio frames through frame-segmentation processing, that is, determines the number of audio frames. Next, it performs a speech existence judgment on each frame, filters out frames containing valid speech, and extracts the voiceprint features of the current audio frame. It compares these features with the historical voiceprint features of the object stored in the device to calculate the voiceprint similarity, and finally forms structured speech data containing these three types of information.
[0037] Then, the processing device determines the voice activity level of the interactive object based on the number of audio frames, the presence of speech in each audio frame, and the voiceprint similarity. The expression for calculating voice activity level is:
[0038]
[0039] in, Indicates voice activity level. Indicates the first The existence of speech in frames. Indicates the first Voiceprint similarity of frames Represents a sliding time window The number of audio frames within the content. Indicates the first Root mean square energy of the frame Indicates the energy threshold, the energy boundary that distinguishes speech from noise. This indicates the signal-to-noise ratio after speech enhancement. This indicates the preset signal-to-noise ratio threshold.
[0040] Visual attention is an indicator used to quantify the visual attention state and reliability of visual signals of an interactive object. It is calculated based on the video data of the interactive object and reflects whether the object is paying attention to the digital mannequin and whether the visual data is clear and usable. For example, if an interactive object is facing the digital mannequin, its gaze point is on the digital mannequin's area, and the screen brightness is appropriate, its visual attention will be higher. Visual attention is obtained through the following methods: First, the processing device acquires first video data of the interactive object, which includes the frame number of the first video frame, instantaneous gaze focus, instantaneous facial orientation focus, and image quality factor.
[0041] The first video data refers to the structured information set extracted by the processing device from the video signal of the interactive object for calculating visual attention, including four types of information: the frame number of the first video frame, instantaneous gaze attention, instantaneous facial orientation attention, and image quality factor.
[0042] The frame number of the first video frame is the number of independent video frames formed after the processing device divides the continuous video signal into fixed time intervals (such as 16.7ms / frame, corresponding to 60 frames / second). It serves as the time dimension basis for subsequent multi-frame comprehensive analysis, such as calculating visual attention over a period of time based on 30 frames of data.
[0043] Instantaneous gaze focus is a quantitative value derived by analyzing the gaze direction of an interactive object in a single frame of video. It reflects whether the interactive object is looking at the digital man in that frame. The closer the gaze is to the digital man, the higher the value; conversely, the further away the gaze is, the lower the value.
[0044] Instantaneous facial orientation focus is a quantitative value obtained by analyzing the facial angle of the interactive object in a single frame of video. It reflects whether the interactive object's face is facing the digital man in that frame. The more directly the face is facing the digital man, the higher the value; the value is lower when the face is turned to the side or away from the digital man.
[0045] The image quality factor is a quantitative value based on the brightness of a single frame of video, reflecting the reliability of the visual signal in that frame. The value is 1 when the brightness is within a suitable range for recognition (such as neither too dark nor too bright), and the value decreases when it is too bright or too dark, so as to avoid misjudgment of visual features due to poor image quality.
[0046] The processing device first acquires video signals from the interactive object using a camera, splitting the continuous video stream into discrete first video frames and simultaneously counting the frame count of each first video frame. Next, for each first video frame, the processing device calculates two core metrics: instantaneous gaze focus, used to determine if the interactive object's gaze is aligned with the digital persona (the closer the gaze is to the digital persona, the higher the metric); and instantaneous facial orientation focus, used to determine if the interactive object's face is directly facing the digital persona (the closer the facial orientation is to the digital persona, the higher the metric). Simultaneously, the processing device determines an image quality factor based on the brightness of the current frame to assess the reliability of the visual signal in that frame. If the brightness is within a suitable range for recognition, the image quality factor is higher; if the image is too bright or too dark, the value is lower. Finally, the processing device integrates these four types of information—frame count, instantaneous gaze focus, instantaneous facial orientation focus, and image quality factor—to form structured first video data.
[0047] Then, the processing device determines the visual attention of the interactive object based on the frame number of the first video frame, instantaneous gaze attention, instantaneous facial orientation attention, and image quality factor. The expression for calculating visual attention is:
[0048]
[0049]
[0050]
[0051] in, Indicates visual focus. Represents a sliding time window The frame number of the first video frame within the video. Indicates instantaneous focus of attention. Indicates the level of concentration of facial orientation at any given moment. Represents the image quality factor. Indicates the angle of gaze. This indicates the preset maximum tolerance value for the gaze angle. Indicates the angle of facial orientation. This indicates the preset maximum tolerance value for facial orientation. Indicates image brightness. This represents the minimum preset ideal brightness. This represents the maximum value of the preset ideal brightness. This represents a logical function used to smoothly adjust the image quality factor.
[0052] Lip-movement synchronization rate (LPSR) is an indicator used to verify the correlation between the lip movements of an interactive object and the speech signal. It is determined by analyzing the video data of the interactive object and combining it with concurrent speech data. It reflects whether the object's lip movements are synchronized with its speech in time. For example, if an interactive object's lip movements are consistent with the speech rhythm, its LPSR will be higher, eliminating invalid scenarios where only the lips move without uttering a sound or only uttering a sound without moving the lips. LPSR is obtained through the following methods: First, the processing device acquires second video data of the interactive object, which includes the frame number of the second video frame and the lip movement activity.
[0053] The second video data refers to the set of structured information extracted by the processing device from the video signal of the interactive object, which is specifically used to calculate the lip movement synchronization rate. It includes two types of data: the number of frames of the second video frame and the lip movement activity, focusing on the lip movement feature analysis of the interactive object.
[0054] The frame count of the second video frame refers to the number of independent video frames formed by the processing device after dividing the continuous video signal into fixed time intervals. The frame count is the basis for the time dimension of subsequent multi-frame comprehensive analysis of lip movements, ensuring that the continuity of lip movements over a period of time can be captured, rather than relying solely on a single static frame.
[0055] Lip movement activity is a quantified value calculated based on the lip features of the interactive object in the second frame of a single video frame. It reflects the activity level of lip movements in that frame. The processing device first identifies key lip points, such as the corners of the mouth and the cupid's bow, and then calculates the magnitude of positional changes of these key points. The greater the magnitude of change, the more obvious the lip movements, and the higher the lip movement activity value. If there are no obvious lip movements, the value is close to 0.
[0056] The processing device first acquires real-time video signals of the interactive object through a camera, splits the continuous video stream into discrete second video frames, and counts the number of frames to ensure that it can cover lip movements within a certain time range. Then, for each second video frame, the processing device uses a lip key point detection algorithm to identify the lip contour and key points, calculates the dynamic change amplitude of the points, and thus obtains the lip movement activity corresponding to each frame. Finally, the frame number of the second video frame and the lip movement activity of each frame are integrated to form structured second video data.
[0057] Then, the processing device determines the lip-sync rate of the interacting object based on the frame number of the second video frame and the lip movement activity. The expression for calculating the lip-sync rate is:
[0058]
[0059]
[0060] in, Indicates lip movement synchronization rate. Represents a sliding time window The number of frames in the second video frame within the frame. The corresponding sliding time window duration is uniform; for example, the sliding time window duration is 1 second, and the frame count differs only due to differences in frame rate. This indicates frame-level lip-sound synchronization determination. Indicates the level of lip movement activity. Indicates the first The set of key points for the lips in a frame. This indicates that the maximum variation of key lip points has been normalized, mapping the variation range to the 0-1 interval for easier quantification. This indicates the preset threshold for lip movement amplitude. Indicates the first The existence of speech in frames. Represents the lip movement amplitude sequence With speech energy sequence Cross-correlation, after considering physiological delay, reflects the temporal synchronicity between lip movement and speech. This represents the cross-correlation threshold, with a value ranging from 0.6 to 0.8.
[0061] The processing device first identifies the set of interactive objects within the current interactive scenario, that is, it identifies all human groups that may interact with the digital human. Next, for each interactive object in the set, that is, each viewer, the processing device collects data from two dimensions: acquiring the object's voice data through an audio acquisition device (such as a microphone) and acquiring the object's video data through a video acquisition device (such as a camera). Then, based on the above calculation expressions, the processing device analyzes and calculates the voice activity level from the voice data, and performs gaze / facial orientation analysis and lip movement analysis on the video data to obtain visual focus and lip movement synchronization rate. Finally, the processing device completes the calculation of the three major features of all individuals in the set, providing basic data support for subsequent quantification of dominance through features and identification of target interactive objects. This is a step that connects data collection and decision-making.
[0062] S202. The processing device determines the dominance of each interactive object based on voice activity, visual focus, and lip movement synchronization rate.
[0063] Dominance is a quantifiable value calculated by combining three indicators: speech activity, visual attention, and lip-sync rate. Its purpose is to prioritize each interaction object as the primary interaction target. A higher dominance indicates a stronger willingness and more reliable signal from that object to actively interact with the digital human, making it more likely to be identified as the digital human's interaction target. The expression for dominance is:
[0064] in, Indicates dominance. The first weight representing voice activity The second weight representing visual focus, The third weight represents the lip movement synchronization rate.
[0065] Traditional single-modal solutions rely on only one type of signal, such as judging whether a voice is emitted. These solutions are easily affected by environmental interference and may lead to misjudgments. However, by fusing speech, vision, and lip movement signals through dominance, it can not only comprehensively reflect the true intention of the interactive object (where the simultaneous occurrence of voice, gaze, and synchronized lip movement indicates a strong interactive intention), but also adapt to different complex scenarios by adjusting the weights of various signals. Furthermore, it can transform abstract interactive signals into quantifiable and comparable dominance, providing a clear judgment standard for selecting the target interactive object from multiple interactive objects, thereby significantly improving the accuracy of main interactive object recognition.
[0066] S203. The processing device determines the target interaction object based on the dominance of each interaction object.
[0067] The specific determination process is as follows: If there is only one interaction object whose dominance is greater than or equal to the first dominance threshold, then the interaction object with a dominance greater than the first dominance threshold is identified as the target interaction object.
[0068] The first dominance threshold is a quantitative standard preset by the processing device to determine whether an interactive object has the qualification to be the main interactive object. The value range is from 0 to 1. For example, the first dominance threshold is 0.8. This first dominance threshold is calibrated based on a large amount of interactive scenario data and represents the minimum dominance level that can be identified as the only interactive object.
[0069] The target interaction object refers to the interaction object that the virtual digital human ultimately identifies as the priority to respond to and focus on. It is the direct target of the digital human's subsequent interaction behaviors (such as voice response and eye focus).
[0070] In multi-user interaction scenarios, after calculating the dominance of all interactive objects, the processing device first performs a threshold judgment: if statistical analysis shows that only one interactive object's dominance reaches or exceeds the first dominance threshold, it indicates that this object's overall performance in the three dimensions of speech activity, visual focus, and lip-movement synchronization rate is superior to other objects. Its willingness to actively interact and signal reliability have reached a level of clarity that requires no further filtering. For example, in a quiet environment, if a user continuously gazes at the digital human and speaks clearly, with lip movements perfectly synchronized with speech, their dominance far exceeds that of other silent or low-focused users. In this case, the processing device will directly identify this object as the target interactive object, ensuring that the digital human quickly locks onto the core interactive subject and avoids response delays caused by redundant judgments.
[0071] If there are at least two interaction objects whose dominance is greater than the second dominance threshold and less than the first dominance threshold, and there are no interaction objects whose dominance is greater than the first dominance threshold, then the voiceprint similarity and semantic coherence of each interaction object are determined respectively. The semantic coherence is the coherence between the current semantic data of the interaction object and the historical semantic data obtained from the previous identification. Based on the voiceprint similarity and semantic coherence, the interaction score of each interaction object is determined. The interaction object with the largest interaction score is determined as the target interaction object.
[0072] The second dominance threshold is a quantitative standard preset by the processing device that is lower than the first dominance threshold. For example, the first dominance threshold is 0.8 and the second dominance threshold is 0.5, both ranging from 0 to 1. It represents the minimum dominance level of having a certain willingness to interact but not reaching the qualification of clear main interaction, and is used to filter out objects with potential interaction priority in multi-person scenarios.
[0073] Semantic coherence is a quantitative value derived by analyzing the semantic data of the current dialogue between interactive objects and the historical semantic data of the previous interaction. It reflects the correlation between the current dialogue content and the logic of the historical dialogue. The higher the value, the stronger the continuity of the dialogue content.
[0074] The interaction score is a comprehensive score calculated by combining voiceprint similarity and semantic coherence with preset weights. It is used to quantify the interaction priority of each object in multi-object competition scenarios. The calculation expression for the interaction score is:
[0075] in, Indicates interactive score, Indicates voiceprint similarity. The first weight representing voiceprint similarity. The second weight representing semantic coherence Indicates semantic coherence.
[0076] The processing device first determines the dominance of all interactive objects. If it finds that the dominance of at least two objects is in the range of greater than the second dominance threshold and less than the first dominance threshold, and no object reaches the first dominance threshold, it indicates that these objects all have certain interactive signals (such as speaking or gazing), but the signal strength has not formed an absolute advantage. This belongs to a fuzzy scenario of multi-object competition, such as multiple people taking turns to speak, and each person's dominance is at a medium level.
[0077] At this point, the processing device will introduce historical interaction information for further filtering: on the one hand, it will calculate the voiceprint similarity of each object to confirm whether it is the object that interacted with the digital human in the previous time, so as to ensure the continuity of identity; on the other hand, it will calculate the semantic coherence to determine whether the current dialogue is logically connected with the historical dialogue, so as to avoid the digital human from interrupting the ongoing dialogue process.
[0078] Subsequently, voiceprint similarity and semantic coherence are integrated into an interaction score based on preset weights. The higher the interaction score, the stronger the correlation between the object and historical interactions, and the better the continuity of the dialogue. Finally, the processing device determines the object with the highest interaction score as the target interaction object, ensuring that the digital human prioritizes objects with matching identities and coherent dialogue, avoiding blindly switching targets in multi-object competition, and guaranteeing the naturalness and continuity of the interaction.
[0079] If any interaction object has a dominance less than the second dominance threshold, then the target interaction object is determined to be empty.
[0080] If it is found that the dominance of any interactive object is lower than the second dominance threshold, and after considering the scenario, it is confirmed that the dominance of all interactive objects has not reached the second dominance threshold, it means that no object in the current scenario has shown a valid interaction signal. This may be because all objects are in a silent state, not looking at the digital human, or their vocalizations, gazes, or other behaviors are extremely weak, and no identifiable interaction intention has been formed.
[0081] At this point, the processing device will determine that the target interaction object is empty, that is, it will not set any object as the interaction target of the digital human.
[0082] S204. The processing device controls the virtual digital human to interact with the target interactive object.
[0083] When the target interaction object is not empty, the processing device controls the virtual digital human to interact with the target interaction object.
[0084] When the target interaction object is not empty, it means that there is a clear object in the scene that the user wants to interact with the digital human. For example, the user actively speaks and looks at the digital human. At this time, the processing device will trigger the digital human's interaction response mode: on the one hand, the voice module outputs appropriate response content, such as answering questions and taking over the conversation; on the other hand, the visual module adjusts the digital human's posture, such as aligning the gaze with the target object and making nodding or other actions, to ensure that the digital human's interaction behavior is directly directed at the target object, so that the user feels that they are being paid attention to, thereby enhancing the naturalness and immersion of the interaction.
[0085] When the target interaction object is empty, the processing device controls the virtual digital human to remain silent.
[0086] When the target interaction object is empty, it means that no user in the scene has shown a valid willingness to interact. For example, people around the digital human may just pass by normally without making a sound or looking at them. At this time, the processing device will switch the digital human to a silent standby mode: the digital human will not actively initiate a dialogue or perform aimless actions to avoid scene interference caused by behaviors such as not responding to conversations or wandering eyes. At the same time, it will reduce the system's computing power consumption, such as pausing unnecessary speech synthesis and motion rendering processes.
[0087] This method also includes the following cases: If the identified speech activity belongs to the first interaction object, visual attention belongs to the second interaction object, and lip movement synchronization rate belongs to the third interaction object, then the weight of visual attention and lip movement synchronization rate in the process of calculating the dominance of each interaction object is increased, and the dominance of each interaction object is recalculated. When the processing device identifies voice activity, visual focus, and lip-sync rate as belonging to three different interactive objects (first, second, and third respectively), it indicates a signal inconsistency in the current scene. For example, the first object only speaks but does not look at the digital human (possibly background dialogue), the second object only looks but does not speak (possibly waiting for interaction), and the third object only moves its lips but does not produce any effective speech (possibly silent imitation). If the dominance is still calculated using conventional weights, the first object may be mistakenly selected due to interference from a single voice signal, ignoring the object that truly intends to interact.
[0088] At this point, the processing device proactively increases the weight of visual attention and lip-sync rate: on the one hand, visual attention directly reflects whether the subject is actively paying attention to the digital human, which is a visual manifestation of the willingness to interact; on the other hand, lip-sync rate needs to be combined with voice verification of the correlation between actions and speech, which can better eliminate interference from invalid actions. The combination of the two is closer to the judgment logic of real interaction intentions. After recalculating the dominance of all interactive subjects through weight adjustment, the subject with better visual and lip-sync signal performance will obtain higher dominance, thus being accurately identified as a potential target.
[0089] Then, the processing device determines the target interaction object based on the recalculated dominance of each interaction object.
[0090] After recalculating the dominance, the processing device uses the same dominance threshold judgment logic as in regular scenarios to analyze the recalculated dominance of each object: First, it checks whether there is an object whose recalculated dominance is greater than or equal to the first dominance threshold. If there is only one such object, it means that after adjusting the weights, the object's advantages in visual focus and lip-movement synchronization have been amplified, and it has become a clear interaction target. The processing device directly identifies it as the target interaction object. If multiple objects have recalculated dominance between the second and first dominance thresholds, the interaction score is further calculated by combining voiceprint similarity and semantic coherence, and the object with the highest interaction score is selected as the target. If the recalculated dominance of all objects is still lower than the second threshold, the target interaction object is determined to be empty.
[0091] Based on the above description, this application has the following beneficial effects: This method integrates three major features—vocal activity, visual attention, and lip-movement synchronization rate—to calculate dominance. Vocal activity is verified through multi-dimensional checks using the number of audio frames, speech presence, and voiceprint similarity, avoiding interference from noise on a single audio signal. Visual attention combines instantaneous gaze, facial orientation, and image quality factors to reduce visual misjudgments caused by occlusion or side profiles. Lip-movement synchronization rate is verified by the temporal alignment of lip movements and speech, further filtering invalid interaction signals. These three features work together to form a multimodal cross-validation mechanism, effectively reducing the impact of complex factors such as noise and occlusion on the recognition results and improving the accuracy of identifying the target interaction object.
[0092] Furthermore, for complex scenarios such as simultaneous interaction among multiple users and modal information conflicts, this method designs a hierarchical decision-making logic and a dynamic weight adjustment method. When only one interaction object meets the dominance standard, the target can be directly locked to avoid redundant judgments. When there are multiple objects with high dominance, the interaction score is calculated by voiceprint similarity and semantic coherence, and objects with higher consistency with historical interactions and stronger dialogue coherence are prioritized to ensure the continuity of interaction in scenarios where multiple people take turns speaking. When voice, visual, and lip movement features belong to different objects, the dominance is recalculated by increasing the weight of visual and lip movement features to avoid misjudging the dominance result of a single modality. This method adapts to special interaction states such as the speaker not being a gazer, and greatly improves the adaptability of the method to diverse scenarios.
[0093] Furthermore, this method ensures the continuity of interaction through a multi-historical information association mechanism. The calculation of voice activity incorporates the similarity verification of previous voiceprint data, visual attention is judged based on the comprehensive judgment of multi-frame video data, and semantic coherence is directly linked to the historical dialogue content. This makes the recognition result not only dependent on real-time signals, but also deeply bound to the historical interaction state, reducing identity jumps caused by instantaneous interference. At the same time, the hierarchical decision-making logic avoids the problem of blind switching when there is no clear goal, ensuring that the digital human only initiates interaction when there is a clear goal, or selects the best option based on historical information when multiple objects are ambiguous, reducing the probability of interaction interruption and enhancing the naturalness and immersion of the user's interaction with the digital human.
[0094] The above text combined Figure 1 The interactive control method for virtual digital humans provided in the embodiments of this application has been described in detail. The apparatus and devices provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0095] like Figure 2 As shown in the figure, this is a schematic diagram of an interactive control device for a virtual digital human provided in an embodiment of this application. The device includes: The acquisition module 301 is used to acquire the voice activity, visual focus, and lip movement synchronization rate of each interactive object in the interactive object set; The data processing module 302 is used to obtain the dominance of each interactive object based on the voice activity, the visual focus and the lip movement synchronization rate; and to determine the target interactive object based on the dominance of each interactive object. The control module 303 is used to control the interaction between the virtual digital human and the target interactive object.
[0096] Optionally, the acquisition module 301 is specifically used to acquire the voice data of the interactive object. The voice data includes the number of audio frames, the voice presence of each audio frame, and the voiceprint similarity. The voiceprint similarity is the similarity between the current voiceprint data of the interactive object and the historical voiceprint data obtained in the previous recognition. The data processing module 302 is specifically used to determine the voice activity of the interactive object based on the number of audio frames, the voice presence of each audio frame, and the voiceprint similarity.
[0097] Optionally, the acquisition module 301 is specifically used to acquire the first video data of the interactive object, the first video data including the frame number of the first video frame, instantaneous gaze focus, instantaneous facial orientation focus, and image quality factor. The data processing module 302 is specifically used to determine the visual attention of the interactive object based on the frame number of the first video frame, instantaneous gaze attention, instantaneous facial orientation attention, and image quality factor.
[0098] Optionally, the acquisition module 301 is specifically used to acquire the second video data of the interactive object, the second video data including the frame number of the second video frame and the lip movement activity; The data processing module 302 is specifically used to determine the lip movement synchronization rate of the interactive object based on the frame number of the second video frame and the lip movement activity.
[0099] Optionally, the data processing module 302 is specifically used to determine the interaction object whose dominance is greater than the first dominance threshold as the target interaction object if there is only one interaction object whose dominance is greater than or equal to the first dominance threshold. If there are at least two interactive objects whose dominance is greater than the second dominance threshold and less than the first dominance threshold, and there are no interactive objects whose dominance is greater than the first dominance threshold, then the voiceprint similarity and semantic coherence of each interactive object are determined respectively. The semantic coherence is the coherence between the current semantic data of the interactive object and the historical semantic data obtained from the previous identification. Based on the voiceprint similarity and semantic coherence, the interaction score of each interactive object is determined; The interaction object with the highest interaction score is identified as the target interaction object.
[0100] Optionally, the data processing module 302 is further configured to, if the identified speech activity belongs to the first interactive object, the visual attention belongs to the second interactive object, and the lip movement synchronization rate belongs to the third interactive object, increase the weight of the visual attention and the lip movement synchronization rate in the process of calculating the dominance of each interactive object, and recalculate the dominance of each interactive object. The target interaction object is determined based on the recalculated dominance of each interaction object.
[0101] Optionally, the data processing module 302 is specifically used to determine that the target interaction object is empty if the dominance of any interaction object is less than the second dominance threshold. The control module is specifically used to control the virtual digital human to interact with the target interaction object when the target interaction object is not empty; and to control the virtual digital human to remain silent when the target interaction object is empty.
[0102] The interactive control device for a virtual digital human according to the embodiments of this application can correspond to the execution of the methods described in the embodiments of this application, and the other operations and / or functions of each module / unit of the interactive control device for a virtual digital human are respectively for implementing Figure 1 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.
[0103] This application also provides a computing device. For example... Figure 3 As shown in the figure, this is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 700 includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, the memory 704, and the communication interface 703 communicate with each other via the bus 701.
[0104] The 701 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0105] The processor 702 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0106] The communication interface 703 is used for communication with external devices.
[0107] Memory 704 may include volatile memory, such as random access memory (RAM). Memory 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0108] The memory 704 stores executable code, and the processor 702 executes the executable code to perform the aforementioned interactive control method for the virtual digital human.
[0109] Specifically, in achieving Figure 2 In the case of the illustrated embodiment, and Figure 2 When the modules or units of the virtual digital human interactive control device described in the embodiment are implemented through software, the following functions are executed: Figure 2 The software or program code required for the functions of each module / unit can be partially or entirely stored in the memory 704. The processor 702 executes the program code corresponding to each unit stored in the memory 704 to execute the aforementioned interactive control method for the virtual digital human.
[0110] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned interactive control method for a virtual digital human.
[0111] This application also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.
[0112] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0113] When the computer program product is executed by a computer, the computer executes any of the aforementioned interactive control methods for the virtual digital human. The computer program product can be a software installation package; when any of the aforementioned interactive control methods for the virtual digital human is required, the computer program product can be downloaded and executed on the computer.
[0114] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0115] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.
Claims
1. An interactive control method for a virtual digital human, characterized in that, The method includes: Obtain the voice activity, visual focus, and lip-sync rate of each interactive object in the set of interactive objects; The dominance of each interactive object is obtained based on the voice activity, visual focus, and lip movement synchronization rate. The target interaction object is determined based on the dominance of each interaction object; Control the virtual digital human to interact with the target interactive object.
2. The method according to claim 1, characterized in that, The voice activity level of the interactive object is obtained through the following methods: Acquire the voice data of the interactive object. The voice data includes the number of audio frames, the voice presence of each audio frame, and the voiceprint similarity. The voiceprint similarity is the similarity between the current voiceprint data of the interactive object and the historical voiceprint data obtained in the previous recognition. The voice activity level of the interactive object is determined based on the number of audio frames, the presence of speech in each audio frame, and the voiceprint similarity.
3. The method according to claim 1, characterized in that, The visual attention of the interactive object is obtained through the following methods: Acquire the first video data of the interactive object, the first video data including the frame number of the first video frame, instantaneous gaze focus, instantaneous facial orientation focus, and image quality factor; The visual attention of the interactive object is determined based on the frame number of the first video frame, instantaneous gaze attention, instantaneous facial orientation attention, and image quality factor.
4. The method according to claim 1, characterized in that, The lip-sync rate of the interactive object is obtained in the following way: Obtain the second video data of the interactive object, the second video data including the frame number of the second video frame and the lip movement activity; The lip movement synchronization rate of the interactive object is determined based on the frame number of the second video frame and the lip movement activity.
5. The method according to any one of claims 1-4, characterized in that, The step of determining the target interaction object based on the dominance of each interaction object includes: If there is only one interaction object whose dominance is greater than or equal to the first dominance threshold, then the interaction object with a dominance greater than the first dominance threshold is identified as the target interaction object. If there are at least two interactive objects whose dominance is greater than the second dominance threshold and less than the first dominance threshold, and there are no interactive objects whose dominance is greater than the first dominance threshold, then the voiceprint similarity and semantic coherence of each interactive object are determined respectively. The semantic coherence is the coherence between the current semantic data of the interactive object and the historical semantic data obtained from the previous identification. Based on the voiceprint similarity and semantic coherence, the interaction score of each interactive object is determined; The interaction object with the highest interaction score is identified as the target interaction object.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: If the identified speech activity belongs to the first interaction object, visual attention belongs to the second interaction object, and lip movement synchronization rate belongs to the third interaction object, then the weight of visual attention and the weight of lip movement synchronization rate in the process of calculating the dominance of each interaction object are increased, and the dominance of each interaction object is recalculated. The target interaction object is determined based on the recalculated dominance of each interaction object.
7. An interactive control device for a virtual digital human, characterized in that, The device includes: The acquisition module is used to acquire the voice activity, visual focus, and lip-movement synchronization rate of each interactive object in the interactive object set; The data processing module is used to obtain the dominance of each interactive object based on the speech activity, visual focus, and lip movement synchronization rate; and to determine the target interactive object based on the dominance of each interactive object. The control module is used to control the interaction between the virtual digital human and the target interactive object.
8. A computing device, characterized in that, Including memory and processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes one or more computer instructions, which, when executed by a computer, perform the method as described in any one of claims 1 to 6.