Virtual human interaction control method and system in immersion interaction space

By using multimodal signal acquisition and synchronous quantization technology, scores and confidence levels are calculated to generate an interactive control parameter set, which solves the problems of mismatched interaction rhythm and attention loss in virtual human interaction systems and achieves efficient immersive interactive control.

CN121900628APending Publication Date: 2026-04-21XIAMEN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN UNIV
Filing Date
2026-03-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing virtual human interaction systems rely on a single modality to determine user engagement and focus, lacking closed-loop evaluation and linkage control driven by stage nodes, resulting in mismatched interaction rhythms and easy loss of user attention.

Method used

By employing multimodal signal acquisition and synchronous quantization technology, scores and confidence levels are calculated within the same sliding window using posture, gaze, and speech prosody information. Combined with stage comprehensive scores and weight updates, an interactive control parameter set is generated to achieve dynamic adjustment of environmental linkage and virtual human expression.

Benefits of technology

It improved the matching of the interaction rhythm, reduced user attention loss, realized closed-loop evaluation and linkage control driven by stage nodes, and enhanced the effect of immersive interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_6
    Figure QLYQS_6
Patent Text Reader

Abstract

The invention relates to a man-machine interaction method and system, in particular to a virtual human interaction control method and system in an immersion interaction space, and the method comprises the steps: constructing the immersion space, collecting posture, gazing, voice rhythm and other multi-mode signals, and outputting scores and confidence coefficients; according to weight fusion calculation, a stage comprehensive score is obtained, so that the fusion weight of the next stage is adaptively adjusted along with the user state, a control parameter set of the next interaction stage is obtained through calculation, and the control parameter set is used for controlling the environment linkage parameters of the immersion space and the expression parameters of the virtual human. The system comprises an immersion space subsystem, a multi-modal detection subsystem, a stage scoring and weight self-adaption subsystem and an interaction and environment regulation subsystem. According to the method, stage node driven closed-loop evaluation and linkage control are achieved, the interaction rhythm matching performance is high, and user attention is not prone to losing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a human-computer interaction method and system, and more particularly to a virtual human interaction control method and system in an immersive interactive space. Background Technology

[0002] Existing virtual human interaction systems rely heavily on single modalities or subjective questionnaires to determine changes in user engagement and focus. Furthermore, the interaction and environmental parameters are mostly statically configured, lacking closed-loop evaluation and linkage control driven by stage nodes. This can easily lead to problems such as mismatched interaction rhythms and easy loss of user attention.

[0003] In view of this, the above issues were studied in depth, which led to this case. Summary of the Invention

[0004] The purpose of this invention is to provide a virtual human interaction control method and system in an immersive interactive space with high interaction rhythm matching and low user attention loss.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for controlling the interaction of a virtual human in an immersive interactive space includes the following steps performed sequentially: S1, Construct an immersive space, and arrange a display carrier for displaying virtual humans within the immersive space; S2, during the interaction between the user and the virtual human, the user's posture state, gaze state and speech prosody state are collected and synchronously quantized to obtain posture information, gaze information and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence. Within the same sliding window, the gaze information is used as input to calculate the corresponding gaze score and gaze confidence. Within the same sliding window, the speech prosody information is used as input to calculate the corresponding speech score and speech confidence. S3, at the end of each stage of interaction, the posture score, the gaze score and the speech score are obtained and the stage comprehensive score is calculated by weighted fusion. The weakness strength is calculated according to the posture score, gaze score and speech score respectively, the weakness mode is determined, and the weakness mode is updated with weights according to each weakness strength and the confidence level corresponding to the weakness mode. S4. Based on the stage comprehensive score, the posture score and its corresponding weight, the gaze score and its corresponding weight, and the voice score and its corresponding weight, calculate the control parameter set for the next interaction stage, and use the control parameter set to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

[0006] As an improvement of the present invention, in step S1, an edge transition band is generated on the outer edge of the display screen of the display carrier. The width of the edge transition band is 2%-8% of the short side length of the display carrier. The display content of the edge transition band is formed by fusing the wall texture sampling map at the corresponding position of the immersive space with the virtual background edge map on the display screen. Simultaneously, the indoor illuminance and color temperature of the immersive space are acquired by an ambient light sensor and output to the display carrier for color mapping; the brightness, color temperature and / or angle of the physical lights in the immersive space are adjusted by a lighting controller to make the virtual scene lighting consistent with the real lighting.

[0007] As an improvement of the present invention, in step S1, a real space coordinate system is established based on the real space of the immersive space, a screen coordinate system is established based on the display carrier, and a virtual scene camera coordinate system is established based on the virtual camera in the immersive space. The real space coordinate system, the screen coordinate system, and the virtual scene camera coordinate system are mapped to each other. The wall boundaries, ground planes, and reference points of the immersive space are obtained using low-visibility markers, depth scanning, and / or structured light. The pose and projection parameters of the virtual camera are calculated so that the perspective lines of the background inside the screen of the display carrier are geometrically continuous with the real decoration at the screen edge.

[0008] As an improvement of the present invention, a virtual interactive panel is arranged in the immersive space and is fixed relative to the virtual human.

[0009] As an improvement of the present invention, in step S3, before calculating the stage comprehensive score, it is first determined whether each confidence level is insufficient. When the confidence level of a certain modality is insufficient, the score corresponding to that confidence level does not participate in or is weakened in the fusion calculation and weight update.

[0010] As an improvement of the present invention, in step S2, the posture state is divided into three types: open posture, closed posture, and neutral posture. When acquiring the posture state, human body detection and tracking are performed on the video frame, key points of the upper body are extracted, the probability of arm crossing, upper limb abduction angle, torso forward tilt angle, and hand position are calculated, and the probability of open posture is output. Neutral posture probability and closed attitude probability Within the time window T, the average probability of each pose is taken, according to the formula... Calculate the pose score and obtain the pose confidence based on the keypoint missing rate and / or occlusion rate. : , in Keypoint missing rate, This refers to the occlusion rate.

[0011] As an improvement of the present invention, in step S2, the area that the user can gaze at during interaction is divided into multiple gaze regions. When collecting gaze status, the face is detected and key points are located, the head pose is estimated and gaze vector is obtained by combining eye region features. The time proportion of the gaze vector falling into different gaze regions is determined. The gaze score is calculated within the same sliding window, and the gaze confidence is calculated based on the eye region visibility and head deflection.

[0012] As an improvement of the present invention, in step S2, the speech prosody information is segmented by VAD, and the speech proportion, mean and variance of fundamental frequency F0, mean and variance of energy RMS, speech rate, number of pauses and average pause duration, and jitter features are extracted. Prosody stability is obtained by analyzing the mean and variance of fundamental frequency F0, the mean and variance of energy RMS, and speech rate. Pause quality is obtained by analyzing the number of pauses, average pause duration, and jitter features. The speech proportion, the prosody stability, and the pause quality are normalized and weighted and fused to obtain a speech score. The speech confidence is calculated based on the signal-to-noise ratio, echo, and / or saturation.

[0013] A virtual human interaction control system for an immersive interactive space includes: The immersive space subsystem includes a real space and a display carrier arranged within the real space for displaying virtual humans; The multimodal detection subsystem is used to collect the user's posture state, gaze state, and speech prosody state during the interaction between the user and the virtual human, and to simultaneously quantize them to obtain posture information, gaze information, and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence. Within the same sliding window, the gaze information is used as input to calculate the corresponding gaze score and gaze confidence. Within the same sliding window, the speech prosody information is used as input to calculate the corresponding speech score and speech confidence. The stage scoring and weight adaptation subsystem is used to obtain the posture score, gaze score, and speech score at the end of each stage of interaction, and calculate the stage comprehensive score by fusing them according to weights. Based on the posture score, gaze score, and speech score, the bottleneck mode is determined, and the bottleneck strength is calculated. , and according to Update the weights, where i is the modality index. The modal score of the i-th mode at the end of the current phase. The threshold for determining the weakest link For the update increment of the fusion weights of the i-th modality, For learning rate or update coefficient, To prevent small positive constants with a denominator of 0, Let i be the confidence level of the i-th mode, and then... Upper and lower bound constraints are applied and normalization is performed so that the fusion weights in the next stage can be adaptively adjusted according to the user's state. This refers to the weight of the i-th mode, while simultaneously misleading the weight update by controlling the noise mode with the confidence level corresponding to the weakest mode; and The interaction and environment adjustment subsystem is used to calculate the control parameter set for the next interaction stage based on the stage comprehensive score, the posture score and corresponding weight, the gaze score and corresponding weight, and the voice score and corresponding weight, and to use the control parameter set to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

[0014] As an improvement of the present invention, the immersive space subsystem further includes a virtual interaction panel that is fixed relative to the virtual human.

[0015] By adopting the above technical solution, the present invention has the following beneficial effects: 1. This invention divides the interaction process into multiple task stages. During each stage of the interaction, multimodal signals such as posture, gaze, and speech prosody are collected, and scores and confidence levels are output. Based on these scores, weights are updated and control parameter sets for the next interaction stage are generated. It has closed-loop evaluation and linkage control driven by stage nodes, with high matching of interaction rhythm and less user attention loss.

[0016] 2. The present invention uses a virtual interactive panel to present example videos, key points or task prompts, and can provide a stable reference for gaze area recognition. Detailed Implementation

[0017] The present invention will be further described below with reference to specific embodiments.

[0018] This embodiment provides a virtual human interaction control method in an immersive interactive space, wherein the virtual human can be set according to the actual use scenario. In this embodiment, the CCBT virtual human (i.e., a virtual human used for computerized cognitive behavioral therapy) is used as an example for illustration.

[0019] The virtual human interaction control method provided in this embodiment uses the immersion consistency calibration result as the constraint on the virtual scene camera and display output, and uses the multimodal score triggered by the stage node as the only update time of the interaction control parameter set. It also requires that the update of the interaction parameter set and the environmental linkage update be completed within the same control cycle, so that the immersion consistency chain, the multimodal perception chain and the stage closed-loop control chain are strongly coupled in terms of data and timing. Specifically, it includes the following steps performed in sequence: S1, construct an immersive space, and arrange display carriers for displaying virtual humans in the immersive space. Of course, the immersive space also includes spatial sensors, calibration calculation units, lighting controllers and rendering output units.

[0020] The display carrier is one or more of the following: large screen, splicing screen, and projection screen. The display plane of the display carrier is installed flush with the wall of the immersive space, with a flushing error of ≤(1-5)mm. A dark seam light-shielding strip is set around the display plane. The dark seam light-shielding strip is preferably made of black low-reflection material (reflectivity ρ≤0.05), with a width of 5-20mm and a thickness of 1-5mm. A dark seam is formed between the dark seam light-shielding strip and the wall to absorb edge light leakage and block the high brightness of the splicing seam, reducing the visibility of the boundary.

[0021] Preferably, to reduce the user's perception of the display boundary and avoid the virtual display content being disconnected from the real environment, this embodiment also generates an edge transition band around the outer edge of the display screen on the display carrier. The width of the edge transition band is 2%-8% of the short side length of the display carrier. The display content of the edge transition band is formed by fusing the wall texture sampling map of the corresponding position of the immersive space and the display carrier with the virtual background edge map on the display screen. The specific fusion method is as follows: S1.11 Sampling: The texture of the wall and its decoration is captured by a camera, and then processed through white balance and... Correction yields wall texture maps .

[0022] S1.12 Extension: Extract edge maps from the pixel areas of the virtual background at the edge of the display carrier's screen. .

[0023] S1.13 Blending: For pixel p in the edge transition zone, output color: ,in, Represents edge mapping The pixel value at pixel position p, Represents wall texture mapping The pixel value at pixel position p To reduce the noise level from 1 to 0 along the normal direction of the edge transition zone (linear, cosine, or Gaussian attenuation can be used), low-amplitude noise jitter and edge spectrum matching are added to suppress visible "band marks" at the fusion boundary. Of course, if the fusion effect is good, edge spectrum matching can be omitted, i.e., it is optional.

[0024] To ensure consistency in spatial style between the real space and the virtual display space, in this embodiment, the hard furnishing materials, soft furnishing style, and color system of the real space are kept consistent with the digital scene of the virtual human background (i.e. the scene displayed through the display carrier) (e.g., wall texture, floor material, furniture style and placement), so that the virtual background visually extends into the real space.

[0025] To further avoid the disconnect between the virtual display content and the real environment, this embodiment also performs space-image alignment calibration. A real-space coordinate system is established based on the real space of the immersive space, a screen coordinate system is established based on the display carrier, and a virtual scene camera coordinate system is established based on the virtual camera in the immersive space. The real-space coordinate system, screen coordinate system, and virtual scene camera coordinate system are mapped to each other. Low-visibility markers, depth scanning, and / or structured light are used to acquire the wall boundaries, ground plane, and reference points of the immersive space (specific reference points can be selected according to actual needs). The pose and projection parameters of the virtual camera are calculated to ensure that the perspective lines of the background within the screen of the display carrier are geometrically continuous with the real-world decoration at the screen edges. Specifically, this includes the following steps: S1.21 Acquisition: Place low-visibility markers around the screen or on the wall and acquire N frames of RGB / depth data.

[0026] S1.22 Feature Extraction: Extract the marked corner points and decorative line features from the RGB data, fit the ground plane and wall plane from the depth data, and extract the intersection line and boundary point set.

[0027] S1.23 Solving for initial values: Obtain the initial values ​​of the camera's extrinsic parameters using PnP or planar constraints. , The camera's intrinsic parameters can be retrieved from the factory calibration or estimated online. .

[0028] S1.24 Error Minimization: Perform nonlinear optimization (such as LM) to minimize the following objective function: ; The first item is reprojection error, the second is decoration straight-line consistency error, and the third is wall / floor plane constraint error. "" indicates summing over all corresponding point pairs used for calibration; the set of corresponding point pairs can be represented as , Let i be the coordinates of the i-th three-dimensional point (or three-dimensional feature point). To and The corresponding two-dimensional pixel observation point; For the camera intrinsic parameter matrix, Let the rotation matrix be the camera's extrinsic parameters. The translation vector of the camera's extrinsic parameters; This is a projection function used to project three-dimensional points onto the pixel plane after transformation by extrinsic parameters and mapping by intrinsic parameters to obtain pixel coordinates. Represents the Euclidean norm; This is the linear consistency error term for decoration. This refers to the wall / floor planar constraint error term. and These are the error term weighting coefficients, used to balance the contribution of each error term to the overall objective function.

[0029] Furthermore, the abbreviated form of the projection function Equivalent to the unfolded form, preferably satisfying: .

[0030] S1.25 Acceptance Threshold: If any of the following thresholds are met, the application passes: (1) The average reprojection error is ≤ (1-3)px; (2) The geometric misalignment between real and virtual objects at the edge of the screen is ≤ (2-8) mm (based on the statistical data of the boundary-aligned measurement point set).

[0031] If the conditions are not met, return to S1.21 to increase the collection or adjust the marker distribution.

[0032] S1.26 Solidification: Write the mapping relationship between K, R, t and each coordinate system into the configuration, and verify it at a low frequency (such as once every time the machine is turned on or once a day) during operation.

[0033] To further ensure consistency in spatial style between the real space and the virtual display space, in this embodiment, an ambient light sensor acquires the indoor illuminance and color temperature of the immersive space and outputs it to the display carrier for color mapping; and a lighting controller adjusts the brightness, color temperature, and / or angle of the physical lights in the immersive space to ensure that the lighting in the virtual scene matches the lighting in the real world. Specifically, this includes: (1) Illuminance Lux and color temperature CCT are obtained through an ambient light sensor; (2) The rendering end of the display carrier performs color mapping: white balance and / or color temperature mapping is performed on the virtual background to make the screen output CCT' close to the ambient CCT; (3) The lighting controller of the immersive space outputs the physical lighting parameters (brightness, color temperature and / or angle) and specifies the linkage timing constraints: when the virtual scene lighting parameters or background main color change exceeds the threshold (e.g. ΔCCT or Δbrightness exceeds the set ratio), the physical lighting adjustment is completed within the same control cycle (e.g. ≤200ms), otherwise it enters the degradation strategy (e.g. only adjust the screen mapping and freeze the lights).

[0034] In addition, this embodiment also arranges a virtual interactive panel (displayed on the screen of the display carrier) that is relatively fixed to the virtual human in the immersive space, which is used to synchronously play real-person consultation example videos, key knowledge points or task prompts. Since the position of the virtual interactive panel is fixed relative to the virtual human, eye-tracking AOI can be used to stably identify "looking at the virtual human or looking at the small screen".

[0035] S2 divides the entire interaction process between the user and the virtual human into multiple stages. The specific stage division method can be determined according to actual needs. In this embodiment, the entire interaction process is divided into multiple stages such as assessment and alliance building, problem identification, cognitive reconstruction, and behavioral experimentation according to the CCBT process. The virtual human is used as the consultant and the main interaction subject. Consultation explanation, demonstration, questioning, feedback and review are organized according to the CCBT process nodes, so that the psychological intervention task can be implemented in the immersive space with an executable interaction process. In addition to the virtual consultant, professional real-person demonstration consultation / psychological education videos or task prompts are played synchronously using the virtual interactive panel. A dual information focus structure of "virtual consultant - auxiliary small screen" is formed in the same immersive space to support the cognitive load distribution and attention guidance of different task stages.

[0036] In each stage of the user-virtual human interaction process, the user's posture, gaze, and speech prosody are collected and synchronously quantified to obtain posture, gaze, and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence; the gaze information is used as input to calculate the corresponding gaze score and gaze confidence; and the speech prosody information is used as input to calculate the corresponding speech score and speech confidence. In other words, within each interaction task stage, the user's posture, gaze, and speech prosody are monitored. The state and prosodic state are quantified synchronously, and two types of results are output: modal scores and modal confidence scores. The modal scores are posture scores, gaze scores, and speech scores, all normalized to the same dimension of 0–100 for cross-modal fusion and stage evaluation. The modal confidence scores are posture confidence, gaze confidence, and speech confidence, normalized to 0–1 for a "gating" mechanism: when a modal confidence score is insufficient, that modal score is not involved in subsequent fusion and weight updates, to suppress noisy modalities from misleading system decisions. Weakening participation in fusion calculation means multiplying the modal fusion weight by its confidence score. Then normalize and participate in the summation; and when season The specific judgment method is as follows: the confidence level of the i-th mode is... With the corresponding preset confidence threshold When comparing, < If the confidence level of the i-th mode is deemed insufficient, the score corresponding to that confidence level will not participate in the fusion calculation and weight update, i.e., let... Or weaken participation, that is to say ,in, It refers to the update amount of the weight of the i-th mode (the increment of the weight in this window / this stage). This refers to the effective weights used for fusion calculation after weakening. This refers to the weight of the i-th mode. This refers to the confidence level (0–1) of the i-th mode, representing the reliability of that mode in the current window / stage. It should be noted that the above score and confidence level must be calculated within the same sliding time window and read at the stage node trigger time for subsequent control decisions.

[0037] The posture score and posture confidence are obtained as follows: Posture states are divided into three types: open posture, closed posture, and neutral posture. Open postures include upper limb abduction / body extension, etc.; closed postures include crossed arms / curled up / occluded, etc.; neutral postures are in between. When collecting the posture states, human detection and tracking are performed on video frames, key points of the upper body (such as head, shoulder, elbow, wrist, hip, etc.) are extracted, the probability of arm crossing, upper limb abduction angle, trunk forward tilt angle, and hand position are calculated, and the probability of open posture is output. Neutral posture probability and closed attitude probability Within the time window T, the average probability of each pose is taken, according to the formula... Calculate the pose score and obtain the pose confidence based on the keypoint missing rate and / or occlusion rate. : , in Keypoint missing rate, This refers to the occlusion rate.

[0038] It should be noted that, since the probabilities of the three types of poses satisfy... ,therefore Although it is not explicitly stated in the above formula, its increase will lead to and The proportion of [something] decreases, thus lowering the posture score, and [something else]. It can also be used for attitude category determination and policy triggering, for example when When the average value within the time window T exceeds the preset threshold, it is determined that the user is in a closed posture and corresponding interactive guidance is triggered or used as a basis for determining the weakness.

[0039] The specific way in which the modality score does not participate or is weakened in fusion and weight update is as follows: when the attitude confidence level Below the gate threshold If the key point missing rate or occlusion rate is too high, causing the pose estimation to be unreliable, then the effective weights of the pose modality in the fusion will not participate in the fusion calculation of the stage comprehensive score; if the pose estimation can still provide limited reference, then the effective weights of the pose modality in the fusion will be weakened according to the confidence level, and then normalized together with the weights of other modalities before participating in the fusion calculation.

[0040] The gaze score and gaze confidence are obtained by dividing the area that the user can gaze at during interaction into multiple gaze areas. In this embodiment, the gaze area AOI is defined as follows: AOI-Avatar is the virtual human area (face / upper body), AOI-Panel is the virtual interactive panel area (in-world panel), and AOI-Other is other areas (environment / ground / non-task area, etc.). At the same time, a "target gaze AOI" is defined for each task stage (for example, the gaze target is Panel in the stage of playing the real person example; the gaze target is Avatar in the dialogue question stage) to provide criteria for subsequent reminders and zooming. The target gaze AOI is determined by the interaction task stage, such as "content presentation stage → Panel" and "dialogue stage → Avatar".

[0041] When collecting gaze data, the face is detected and key points are located. The head pose is estimated and combined with eye region features to obtain the gaze vector. The proportion of time the gaze vector falls into different gaze regions is determined, and within the same sliding window, the formula is applied. The fixation score is calculated, where The cumulative duration of gaze falling within the AOI within the time window. The total duration of the time window. The gaze score can be used to characterize the degree of user attention matching to the current interaction node.

[0042] Fixation confidence is calculated based on eye visibility and head deflection. In this embodiment, gaze confidence is determined by at least the following factors: (1) Visibility and occlusion rate of the eye area (glare from glasses, bangs occlusion, etc.); (2) Head posture angle and degree of facial lateralization (the gaze estimation is unreliable when the head is turned too much). (3) Face detection and eye tracking stability (continuous frame loss / drift); (4) Image quality factors (blur, underexposure).

[0043] When the gaze confidence is lower than the gating threshold: the trigger threshold for deviation determination can be appropriately relaxed, or only "head orientation" can be used as a weak substitute, and at this time the gaze score is not used for weight adaptive update, or its influence is significantly reduced, that is, its participation in fusion calculation and weight update is weakened.

[0044] When the user's gaze is detected to deviate from the target area of ​​the current stage, the virtual human can issue a voice reminder and perform adaptive emphasis on the virtual interactive panel and / or the virtual human's related interface (such as interface magnification, salience enhancement, prompt animation or sound effects, etc.) to achieve closed-loop attention regulation of "detection-judgment-intervention-recovery".

[0045] The specific way in which the modality score does not participate or is weakened in fusion and weight update is as follows: when the fixation confidence... Below the gate threshold If the eye region is not visible / severely occluded, the head posture is excessively turned, face detection or eye tracking fails continuously, or the image quality is poor, resulting in unreliable gaze estimation, then the effective weight of the gaze modality in the fusion will not participate in the fusion calculation of the stage comprehensive score; if the gaze estimation can still provide limited reference, then the effective weight of the gaze modality in the fusion will be weakened according to the confidence level, and normalized together with the weights of other modalities before participating in the fusion calculation.

[0046] The speech score and speech confidence are obtained as follows: VAD segmentation is performed on the speech prosody information; speech proportion, mean and variance of fundamental frequency (F0), mean and variance of energy (RMS), speech rate, number of pauses and average pause duration, and jitter features are extracted. Prosodic stability is obtained based on the mean and variance of fundamental frequency (F0), the mean and variance of energy (RMS), and speech rate; pause quality is obtained based on the number of pauses, average pause duration, and jitter features. The speech proportion, prosodic stability, and pause quality are normalized and weighted to obtain the speech score. Speech confidence is calculated based on signal-to-noise ratio, echo, and / or saturation. .

[0047] In other words, the prosodic score can be composed of at least the following dimensions: (1) Voice proportion Within the sliding window, the percentage of the total duration for the audio segment; (2) Prosodic stability: reflects the degree of fluctuation of prosodic elements such as fundamental frequency, energy, and speech rate within a short window; the more drastic the fluctuation, the worse the stability. (3) Quality of pause This reflects whether the distribution of pause durations is within a reasonable range; excessively long or fragmented pauses usually reduce quality.

[0048] The specific formula for calculating the speech prosody score used in this embodiment is as follows: .

[0049] in, , , These represent the weighting coefficients corresponding to the proportion of speech, prosodic stability, and pause quality, respectively. ; This represents the normalization function, used to map the corresponding feature to the interval [0,1]. It represents the proportion of speech segments within the sliding window to the total duration; stability represents prosodic stability, used to characterize the degree of fluctuation of prosodic elements such as fundamental frequency F0, energy RMS, and speech rate within a fixed time window; The pause quality is used to characterize the reasonableness of the pause distribution reflected by the number of pauses, average pause duration, and / or jitter characteristics.

[0050] To make speech score calculation reproducible, a function is defined to normalize the feature x to [0,1] as follows: .

[0051] in: This represents the result after normalizing the feature x, with a value range of [0,1]; x represents the speech feature to be normalized (e.g., the value of speechratio, stability, or pausequality); L represents the low threshold of the feature; H represents the high threshold of the feature, and... ; Represents the truncation function: when Take 0, when Take y, when Set to 1. The proportion of valid frames for speech features within the sliding window is lower than the threshold. At the same time, reduce voice confidence. Furthermore, gating or weakening of speech modalities are performed during stage fusion and weight updates.

[0052] In this embodiment, prosodic stability can be achieved as follows: within a fixed time window, the short-term fluctuations of indicators such as fundamental frequency (F0), energy (RMS / loudness), and speech rate (word / syllable rate) are statistically analyzed; if the fluctuations are too large, it indicates incoherent or tense expression; if the fluctuations are small, it indicates more stable expression. To accommodate different individuals, the user's baseline fluctuation range can be estimated during the "initial calibration period" or "the first few stages," and stability can be evaluated subsequently based on relative deviations.

[0053] The pause quality in this embodiment The quality is measured by the proportion of pauses falling within the target range. A reasonable pause range can be set (e.g., 0.2–1.5 seconds, configurable and adjustable in stages). If most pauses fall within the reasonable range, the quality is high. If the proportion of long pauses (e.g., >2–3 seconds) increases or frequent extremely short pause fragments occur, the quality is low. This metric can be linked to the percentage of voice input to avoid being misled by a single metric when the user is silent.

[0054] In this embodiment, each speech feature is truncated and segmented to a scale of 0–100. A configurable low and high threshold is defined for each feature; features below the low threshold are recorded as low scores, and features above the high threshold are recorded as high scores, with a linear transition in the intermediate range. Outliers (pops, saturation, distortion) are either removed or assigned a low confidence level. This achieves normalization and scale uniformity.

[0055] The specific way in which the modality score does not participate or is weakened in fusion and weight update is as follows: when the speech confidence score Below the gate threshold If the proportion of valid speech frames within the sliding window is lower than the threshold, or if speech modalities are continuously missing for more than a preset duration (e.g., 1–3 seconds), resulting in unreliable continuous speech feature extraction, then the effective weights of the speech modalities in the fusion process will not participate in the fusion calculation of the stage comprehensive score. If the speech features can still provide limited reference, then the effective weights of the speech modalities in the fusion process will be weakened according to the confidence level and normalized together with the weights of other modalities before participating in the fusion calculation.

[0056] S3, at the end of each interaction phase, the posture score, gaze score, and speech score are obtained and weighted to calculate the phase composite score. The calculation formula is as follows: , Among them, among them, Indicates attitude score The corresponding weighting coefficients, Indicates fixation score The corresponding weighting coefficients, Indicates speech score The corresponding weight coefficients, and the initial weights .

[0057] Before calculating the overall score for each stage, it is necessary to determine whether the confidence levels are insufficient. If the confidence level of a certain modality is insufficient, the score corresponding to that confidence level will not participate in or will be weakened in the fusion calculation and weight update. The specific determination method is as follows: the confidence level of the i-th modality is... With the corresponding preset confidence threshold When comparing, < The confidence level of the i-th mode is deemed insufficient.

[0058] To refine the assessment and target the reinforcement of the weakest modality, the weakness strength is calculated based on posture score, gaze score, and speech score respectively. The weakest modality is identified, which refers to the modality with the least ideal interaction state among the three modalities (gestural, gaze, and speech) at the end of the current interaction phase (or within the statistical period at the end of the phase). This modality indicates the most significant interaction deficiency or attentional deviation by the user, serving as the basis for adjusting the interaction strategy and adaptive fusion weights in the next phase. The aforementioned least ideal modality refers to the modality that meets the following criteria at the end of the current interaction phase: In the modal analysis, the strength of the short slab The largest mode, of which, The modality score of the i-th modality at the end of the current stage is preferably in the range of [0, 100] (a higher score indicates a more ideal interaction state of the modality). The threshold for determining the weakest link is used to characterize the minimum score requirement for "achieving an acceptable level of interaction", and its value range is preferably [0, 100]; i is the modality index, which includes at least the posture modality, gaze modality and voice prosody modality. To ensure that the larger value is obtained during the calculation. To avoid a negative weak point in the strength. If there is no condition that satisfies this requirement... If the modes are such that the mode with the lowest score is identified as the weakest mode, then the mode with the lowest score will be identified as the weakest mode.

[0059] After identifying the weakest mode, to make the next stage's overall score more sensitive to the weakest dimension, this embodiment also updates the weights of the weakest modes based on the strength of each weakest mode and the confidence level corresponding to that mode. Specifically, the update increment of the fusion weights for the i-th mode is defined as... And calculate according to the following formula: ,in: This represents the update increment for the fusion weights of the i-th modality; The learning rate or update coefficient is used to control the magnitude of weight updates. Its value can be preset by the system or determined through experimental calibration; The strength of the short plate in the i-th mode; This is the sum of the modal weakness intensities within this stage (e.g., summation of three modes), used to convert the severity of a single modal weakness into a relative proportion; To prevent small positive constants with a denominator of 0, When all modes have no weak links (i.e.) This is used to avoid calculation errors caused by a denominator of 0 and to maintain numerical stability; The confidence level of the i-th mode is used to suppress the misleading influence of noisy modes on weight updates. The preferred value range is [0,1]. The higher the confidence level, the more reliable the score of the mode.

[0060] Then update the weights. ,in Let be the fusion weight of the i-th modality in the stage comprehensive score fusion calculation, used to characterize the proportion of the contribution of this modality's score to the stage comprehensive score. By increasing the fusion weight of the weakest modality, the stage comprehensive score becomes more sensitive to the weakest modality, thereby driving the next stage interaction strategy to prioritize intervention and guidance targeting the weakest dimension. Subsequently, ... Set upper and lower bound constraints ,in and These are the lower and upper limits of the weights, used to prevent a mode from failing during fusion due to an excessively small weight, or the system from becoming overly sensitive to noise and becoming unstable due to an excessively large weight. , These parameters can be pre-set as system configurable parameters by the deployment scenario, or determined through experimental calibration using samples collected in the target scenario, ensuring that the stability of the phase-based comprehensive score and the accuracy of strategy triggering meet preset indicators. After updating the weights, normalization is performed, and upper and lower bound constraints are re-executed after normalization. If necessary, the above constraint and normalization steps are repeated to ensure the weights meet the requirements. and This allows the fusion weights in the next stage to adaptively adjust according to the user's state, while simultaneously suppressing the misleading influence of noisy modes on weight updates through the confidence levels corresponding to each mode. When the confidence level of the i-th mode... Below the confidence threshold When this occurs, the mode is determined to be a low-confidence mode, and gating suppression is applied: the weight update increment of this mode is set. This avoids the scoring distortion caused by noise from misleading the identification of weaknesses and the updating of weights.

[0061] In addition, it is possible to obtain a comprehensive score for each stage of the entire process. The total score for the entire CCBT experience is obtained by weighted summation (with stage weights set according to stage importance). And record the average score and trend of each dimension throughout the process.

[0062] S4 calculates the control parameter set for the next interaction stage based on the stage comprehensive score, posture score and corresponding weight, gaze score and corresponding weight, and voice score and corresponding weight. The control parameter set includes: question openness parameter, question complexity and decomposition granularity parameter, interaction rhythm parameter (speech rate / pause / confirmation frequency), virtual human expression intensity parameter (facial expression / nodding / gesture / gaze duration), and optional environmental linkage parameters (light brightness / color temperature, background dynamic range, etc.). Then, the control parameter set is used to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

[0063] In one embodiment, let the stage comprehensive score be... Set threshold (For example ).when At that time, the intensity of the push will be reduced (problem complexity will be reduced by one level, and the interaction pace will be reduced by 10% to 25%) and the frequency of confirmation / prompts will be increased (by 10% to 30%). Maintain or make minor adjustments; when When the openness and intensity of interaction are increased (openness increased by 1 level, interaction pace increased by 5% to 15%), the frequency of prompts is decreased (reduced by 5% to 20%). When the weakest modality is gaze / voice / gesture, priority is given to enhancing prompts and guidance, reducing complexity and increasing confirmation, and reducing the amplitude of movements and enhancing encouraging expressions, respectively.

[0064] Phase Overall Score The scores are used to determine the overall intensity of the push. If the score is low, the intensity of the push is reduced and the confirmation is increased. If the score is high, the openness is increased and the push is accelerated appropriately. The scores of each dimension are used to determine the direction of reinforcement. For example, if the gaze score is low, visual guidance is enhanced and screen noise is reduced. If the voice prosody score is low, the pace is slowed down and the response burden is reduced. If the posture score is low, gentler guidance is used and direct access to sensitive points is reduced, so as to achieve separate control of intensity and direction.

[0065] Because the control module adjusts the lighting and / or color mapping simultaneously when adjusting the virtual human's performance or rhythm, ensuring that the physical space and virtual background remain consistent, it can prevent attention shifts caused by sudden environmental changes, thereby enhancing immersion and focus.

[0066] This embodiment also provides a virtual human interaction control system in an immersive interactive space for implementing the above method, including an immersive space subsystem, a multimodal detection subsystem, a stage scoring and weight adaptation subsystem, and an interaction and environment adjustment subsystem.

[0067] The immersive space subsystem includes a real space, a display carrier arranged in the real space to display virtual humans, a virtual interaction panel fixed relative to the virtual humans, a display carrier, a space sensor, a calibration calculation unit, a lighting controller, and a rendering output unit.

[0068] The multimodal detection subsystem is used to collect the user's posture state, gaze state, and speech prosody state during the interaction between the user and the virtual human, and to simultaneously quantize them to obtain posture information, gaze information, and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence. Within the same sliding window, the gaze information is used as input to calculate the corresponding gaze score and gaze confidence. Within the same sliding window, the speech prosody information is used as input to calculate the corresponding speech score and speech confidence.

[0069] The stage scoring and weighted adaptive subsystem is used to obtain posture score, gaze score, and speech score at the end of each stage of interaction, and to fuse them according to weights to calculate a comprehensive stage score. Based on the posture score, gaze score, and speech score, the bottleneck mode is determined, and the bottleneck strength is calculated. , and according to Update the weights, and then... Upper and lower bound constraints are imposed and normalization is performed to enable the fusion weights in the next stage to adaptively adjust according to the user's state. At the same time, the weight update is misled by the noise mode with the confidence level corresponding to the short board mode.

[0070] The interaction and environment adjustment subsystem is used to calculate the control parameter set for the next interaction stage based on the stage comprehensive score, posture score and corresponding weight, gaze score and corresponding weight, and voice score and corresponding weight, and to use the control parameter set to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

[0071] The specific operation mode of the virtual human interaction control system provided in this embodiment is the same as that provided in this embodiment, and will not be repeated here.

[0072] The method and system provided in this embodiment are a scheme for constructing spatial consistency, quantifying multimodal states, and controlling closed-loop stages of virtual human interaction processes in an immersive interactive space. It forms a quantifiable and verifiable chain of immersive consistency through "spatial-screen alignment calibration + borderless display edge generation + lighting synchronization / color mapping." During the interaction, multimodal signals such as posture, gaze, and speech prosody are collected, and scores and confidence levels are output. At the end of each interaction task stage, evaluation and control decisions are triggered, outputting the parameterized control set for the next stage, which is executed synchronously with the environmental consistency maintenance module. This allows for a comprehensive judgment of the user's state (such as willingness to communicate / attention / participation) and stage matching degree, and dynamically adjusts the virtual human's question openness, feedback intensity, pace, and task difficulty accordingly, forming adaptive interactive control.

[0073] The present invention has been described in detail above with reference to specific embodiments. However, the implementation of the present invention is not limited to the above-described embodiments. Those skilled in the art can make various modifications to the present invention based on the prior art, and these modifications all fall within the protection scope of the present invention.

Claims

1. A method for controlling virtual human interaction in an immersive interactive space, characterized in that, The following steps are performed sequentially: S1, Construct an immersive space, and arrange a display carrier for displaying virtual humans within the immersive space; S2, during the interaction between the user and the virtual human, the user's posture state, gaze state and speech prosody state are collected and synchronously quantized to obtain posture information, gaze information and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence. Within the same sliding window, the gaze information is used as input to calculate the corresponding gaze score and gaze confidence. Within the same sliding window, the speech prosody information is used as input to calculate the corresponding speech score and speech confidence. S3, at the end of each stage of interaction, the posture score, the gaze score and the speech score are obtained and the stage comprehensive score is calculated by weighted fusion. The weakness strength is calculated according to the posture score, gaze score and speech score respectively, the weakness mode is determined, and the weakness mode is updated with weights according to each weakness strength and the confidence level corresponding to the weakness mode. S4. Based on the stage comprehensive score, the posture score and corresponding weight, the gaze score and corresponding weight, and the voice score and corresponding weight, calculate the control parameter set for the next interaction stage, and use the control parameter set to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

2. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S1, an edge transition band is generated around the outer edge of the display screen on the display carrier. The width of the edge transition band is 2%-8% of the short side length of the display carrier. The display content of the edge transition band is formed by merging the wall texture sampling map at the corresponding position of the immersive space with the virtual background edge map on the display screen. Simultaneously, the indoor illuminance and color temperature of the immersive space are acquired by an ambient light sensor and output to the display carrier for color mapping; the brightness, color temperature and / or angle of the physical lights in the immersive space are adjusted by a lighting controller to make the lighting of the virtual scene consistent with the lighting of the real world.

3. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S1, a real space coordinate system is established based on the real space of the immersive space, a screen coordinate system is established based on the display carrier, and a virtual scene camera coordinate system is established based on the virtual camera in the immersive space. The real space coordinate system, the screen coordinate system, and the virtual scene camera coordinate system are mapped to each other. The wall boundaries, ground plane, and reference points of the immersive space are obtained using low-visibility markers, depth scanning, and / or structured light. The pose and projection parameters of the virtual camera are calculated so that the perspective lines of the background inside the screen of the display carrier are geometrically continuous with the real decoration at the screen edge.

4. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, A virtual interactive panel, relatively fixed to the virtual human, is arranged within the immersive space.

5. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S3, before calculating the stage comprehensive score, it is first determined whether each confidence level is insufficient. When the confidence level of a certain modality is insufficient, the score corresponding to that confidence level will not participate in or will be weakened in the fusion calculation and weight update.

6. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S2, the posture state is divided into three types: open posture, closed posture, and neutral posture. When acquiring the posture state, human detection and tracking are performed on the video frame, key points of the upper body are extracted, the probability of arm crossing, upper limb abduction angle, torso forward tilt angle, and hand position are calculated, and the probability of open posture is output. Neutral posture probability and closed attitude probability Within the time window T, the average probability of each pose is taken, according to the formula... Calculate the pose score and obtain the pose confidence based on the keypoint missing rate and / or occlusion rate. : , in Keypoint missing rate, This refers to the occlusion rate.

7. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S2, the area that the user can gaze at during interaction is divided into multiple gaze regions. When collecting gaze status, the face is detected and key points are located, the head pose is estimated and gaze vector is obtained by combining eye region features. The time proportion of the gaze vector falling into different gaze regions is determined. The gaze score is calculated within the same sliding window, and the gaze confidence is calculated based on eye region visibility and head deflection.

8. The virtual human interaction control method in an immersive interactive space as described in claim 1, characterized in that, In step S2, the speech prosody information is segmented using VAD, and the speech proportion, mean and variance of fundamental frequency F0, mean and variance of energy RMS, speech rate, number of pauses and average pause duration, and jitter features are extracted. Prosody stability is obtained based on the mean and variance of fundamental frequency F0, the mean and variance of energy RMS, and speech rate. Pause quality is obtained based on the number of pauses, average pause duration, and jitter features. The speech proportion, prosody stability, and pause quality are normalized and weighted and fused to obtain a speech score. The speech confidence is calculated based on the signal-to-noise ratio, echo, and / or saturation.

9. A virtual human interaction control system for an immersive interactive space, characterized in that, include: The immersive space subsystem includes a real space and a display carrier arranged within the real space for displaying virtual humans; The multimodal detection subsystem is used to collect the user's posture state, gaze state, and speech prosody state during the interaction between the user and the virtual human, and to simultaneously quantize them to obtain posture information, gaze information, and speech prosody information. Within the same sliding window, the posture information is used as input to calculate the corresponding posture score and posture confidence. Within the same sliding window, the gaze information is used as input to calculate the corresponding gaze score and gaze confidence. Within the same sliding window, the speech prosody information is used as input to calculate the corresponding speech score and speech confidence. The stage scoring and weight adaptation subsystem is used to obtain the posture score, gaze score, and speech score at the end of each stage of interaction, and calculate the stage comprehensive score by fusing them according to weights. Based on the posture score, gaze score, and speech score, the bottleneck mode is determined, and the bottleneck strength is calculated. , and according to Update the weights, where i is the modality index. The modal score of the i-th mode at the end of the current phase. The threshold for determining the weakest link For the update increment of the fusion weights of the i-th modality, For learning rate or update coefficient, To prevent small positive constants with a denominator of 0, Let i be the confidence level of the i-th mode, and then... Upper and lower bound constraints are applied and normalization is performed so that the fusion weights in the next stage can be adaptively adjusted according to the user's state. This refers to the weight of the i-th mode, while simultaneously misleading the weight update by controlling the noise mode with the confidence level corresponding to the weakest mode; and The interaction and environment adjustment subsystem is used to calculate the control parameter set for the next interaction stage based on the stage comprehensive score, the posture score and corresponding weight, the gaze score and corresponding weight, and the voice score and corresponding weight, and to use the control parameter set to control the environmental linkage parameters of the immersive space and the expression parameters of the virtual human.

10. The virtual human interaction control system in an immersive interactive space as described in claim 9, characterized in that, The immersive space subsystem also includes a virtual interaction panel that is relatively fixed to the virtual human.

Citation Information

Patent Citations

  • Student participation degree analysis method in talent social practice teaching based on VR

    CN118674168A

  • Virtual human real-time generation method and system based on expression control embedding space

    CN121437697A

  • Method and system for automatically adapting teaching atmosphere in immersive teaching environment

    CN121455324A

  • Digital human interaction action generation method and device, electronic equipment and storage medium

    CN121524918A

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20180204111A1