VR rendering virtual examiner multi-mode reaction generation system for interview speech

Through multi-modal input analysis and virtual examiner behavior generation system, the problem of insufficient response authenticity in the virtual interview system is solved, and the high coordination and precise pressure adjustment of expressions, voice, and body movements are achieved, which improves the authenticity and effect of interview training.

CN120339478APending Publication Date: 2025-07-18SHANGHAI YANXI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508664.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing virtual interview speech system, virtual examiners have insufficient authenticity in response, poor multimodal coordination, and inaccurate pressure adjustment, resulting in unreal user experience.

Method used

The multi-modal input analysis module is used to analyze the candidate's status in real time, and combine the virtual examiner's behavior generation module and core algorithm module, including the emotional expression collaborative generation network, real-time interruption decision-making system and pressure adjustment mechanism to achieve high coordination and dynamic pressure adjustment of expressions, voice, and body movements.

Benefits of technology

It improves the realism and multimodal coordination of virtual examiners, provides a natural and smooth interruption mechanism, realizes 5-level precise pressure control, and improves the immersion and effect of interview training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339478A_ABST
    Figure CN120339478A_ABST
Patent Text Reader

Abstract

The invention discloses a VR rendering virtual examiner multi-modal reaction generation system for interview speech, and relates to the technical field of artificial intelligence and man-machine interaction, and the system comprises a multi-modal input analysis module which is used for analyzing examinee state data in real time; the virtual examiner behavior generation module is used for cooperatively generating expressions, voices and limb actions; and the core algorithm module comprises an emotion expression collaborative generation network, a real-time interruption decision system and a pressure regulation mechanism. According to the VR rendering virtual examiner multi-modal reaction generation system for interview speech, high collaboration of expression-voice-limb movement of a virtual examiner is realized, the sense of reality is improved, the voice interaction performance is improved, a natural and smooth interruption mechanism is provided, the pressure simulation effect is improved, and the experience of a user is improved. The pressure gradient control is matched with the multi-sensory collaborative pressure adjusting device to realize five-level accurate pressure control, so that the interview training effect is improved, the immersion is improved, and the real interview speech performance of the user is favorably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and human-computer interaction, and specifically provides a VR rendering virtual examiner multimodal reaction generation system for interview speeches. Background Technique

[0002] In the current process of preparing for interview speech training, the traditional offline speech simulation training method still dominates. This method usually involves organizing a group of real audiences, conducting face-to-face speech practice, and relying on manually provided feedback. However, this method has some obvious deficiencies, including high costs, time limitations, inability to repeat specific speech scenarios, and difficulty in simulating the real feelings under high-pressure environments. At the same time, speech systems based on VR virtual reality technology have gradually emerged. They use advanced virtual reality devices to construct various speech scenarios, such as simulated meeting rooms or auditoriums, bringing a visually immersive experience to users. Nevertheless, these systems still lack dynamic interaction capabilities and are currently unable to implement intelligent questioning by the audience or provide real-time stress feedback.

[0003] The existing virtual interview speech systems face a series of technical challenges: First of all, the reactions of virtual examiners often appear single and cannot exhibit multimodal interaction methods like in real interviews, including the coordinated use of expressions, voices, and body languages; Secondly, the interruption mechanism is usually handled too rigidly and cannot naturally simulate the intervention behavior of the interviewer in real interviews; Furthermore, the system lacks dynamic adaptability in stress regulation and cannot be adjusted in real time according to the performance of the examinee; Finally, the out-of-sync problem of multimodal output leads to an insufficiently real user experience and affects the overall effect of the system.

[0004] In view of this, in response to the above problems, in-depth research has been carried out, and thus this case has emerged. Summary of the Invention

[0005] The purpose of the present invention is to provide a VR rendering virtual examiner multimodal reaction generation system for interview speeches to solve the technical problems of insufficient authenticity of virtual examiner reactions, poor multimodal coordination, and inaccurate stress regulation in the existing virtual interview systems mentioned in the above background technique.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A VR rendering virtual examiner multimodal reaction generation system for interview speeches, including: A multimodal input analysis module for real-time analysis of examinee status data; A virtual examiner behavior generation module for collaborative generation of expressions, voices, and body movements; The core algorithm module includes an emotion expression collaborative generation network, a real-time interruption decision-making system, and a stress regulation mechanism.

[0007] Preferably, the multi-modal input parsing module real-time detects and analyzes the examinee's state including stress level, answer quality, and limb unnaturalness index, while high-frequency processes data on facial expression changes, speech intonation, and eye movement trajectories, and the low-latency interruption mechanism processes abnormal events.

[0008] Preferably, the reaction decision center of the virtual examiner behavior generation module includes a digital human multi-dimensional synthesis module, a scenario dynamic regulation module, and an interruption / questioning strategy generation module; The digital human multi-dimensional synthesis module performs multi-dimensional synthesis on the digital human's facial expression, speech, and body; The scenario dynamic regulation module controls the scene lighting, environmental acoustic effects, and perspective transformation, and the interruption / questioning strategy generation module performs real-time intervention based on the uncertainty threshold.

[0009] Preferably, the emotion expression collaborative generation network realizes synchronous mapping of facial expression parameters, speech prosody parameters, and body movement parameters for cross-modal consistency control, and generates an expression-speech-gesture three-dimensional mapping system through cross-modal emotion expression.

[0010] Preferably, the facial expression parameters are realized by the facial expression animation engine based on the 52-point facial muscle drive of ARKit and the SMPL-X skeletal animation, combined with micro-expression fusion technology and facial expression dynamics control.

[0011] Preferably, the speech prosody parameters are realized by high-emotion-color TTS combined with an emotion prosody library to achieve real-time interruption ability and speech rate dynamic control.

[0012] Preferably, the body movement parameters are realized by a hierarchical skeletal control system based on a full-body movement library, including the generation of fine hand movements and gaze behaviors.

[0013] Preferably, the real-time interruption decision-making system calculates the interruption timing based on the increasing conditional probability P(interrupt|context)=P_base+α·Q_drop+β·T_off+γ·S_stress.

[0014] Preferably, the stress regulation mechanism includes a five-level stress model and a closed-loop control based on physiological feedback, and the stress regulation mechanism realizes collaborative dynamic adjustment in vision, hearing, and space based on multi-sensory collaboration technology.

[0015] Preferably, the generation method of the reaction generation system includes the steps of: A. Receiving and parsing the examinee's multi-modal input data; B. Generating collaborative feedback based on the emotional state and intensity; C. Perform real-time interruption according to the situation analysis result; D. Dynamically adjust the interview pressure level.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The VR rendering virtual examiner multi-modal reaction generation system for interview speech: Achieve a high degree of coordination of the virtual examiner's expressions - voices - body movements, enhancing the sense of reality, including a facial expression driving device of the expression animation system based on ARKit 52 Blendshape, realizing precise control of 52-point micro-expressions, and the authenticity of emotional expression is increased by 65%; Improve the voice interaction performance, the maximum voice - lip synchronization error is < 3 frames, provide a natural and smooth interruption mechanism, the real-time interruption mechanism cooperates with a multi-threshold dynamic intervention decision-making system, and the interruption response time is < 100 ms; Improve the stress simulation effect, simulate a real interview stress scenario, the stress gradient control cooperates with a multi-sensory collaborative stress regulation device to achieve 5-level precise stress control, and the simulation authenticity reaches 85%; Improve the interview training effect, improve the immersion, the similarity between the simulated training heart rate change and the real interview is 82%, improve the training effectiveness, and help to improve the user's real interview speech performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the overall process of the generation system of the present invention; Figure 2 It is a schematic diagram of the overall system architecture of the present invention; Figure 3 It is a schematic diagram of the multi-modal collaborative generation process of the present invention; Figure 4 It is a schematic diagram of the interruption decision logic of the present invention; Figure 5 It is a schematic diagram of the stress gradient control of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figures 1-5 , the present invention provides a technical solution: A VR rendering virtual examiner multi-modal reaction generation system for interview speech, including: ① A multi-modal input parsing module, used to analyze the candidate's state data in real time; The multi-modal input parsing module for improving the user's real interview speech performance detects and analyzes the candidate's state in real time, including stress level, answer quality, and body unnaturalness index. At the same time, it processes facial expression changes, speech intonation, and eye movement trajectory data at high frequency, and the low-latency interruption mechanism processes abnormal events.

[0020] Real-time detection and analysis: Evaluation engine data reception: Stress level 1-5, answer quality 0-100, body unnaturalness index; High-frequency input processing: Facial expression changes 60Hz, speech intonation 16kHz, eye movement trajectory 60Hz; Low-latency interruption mechanism: Abnormal event priority sorter, giving priority to processing stuttering / off-topic behaviors.

[0021] ② The virtual examiner behavior generation module is used to collaboratively generate facial expressions, voices, and body movements; The reaction decision center of the virtual examiner behavior generation module for improving the user's real interview speech performance includes a digital human multi-dimensional synthesis module, a scenario dynamic regulation module, and an interruption / questioning strategy generation module; The digital human multi-dimensional synthesis module for improving the user's real interview speech performance performs multi-dimensional synthesis on the digital human's facial expressions, voices, and bodies. The scenario dynamic regulation module controls scene lighting, environmental acoustic effects, and perspective transformation, and the interruption / questioning strategy generation module makes real-time intervention based on the uncertainty threshold.

[0022] Generation decision center: Digital human multi-dimensional synthesis: Facial expressions, voices, bodies, latency < 16.7ms; Scenario dynamic regulation: Scene lighting, environmental acoustic effects, perspective transformation; Interruption / questioning strategy generation: Real-time intervention mechanism based on the uncertainty threshold.

[0023] ③ The core algorithm module includes an emotion expression collaborative generation network, a real-time interruption decision system, and a stress regulation mechanism.

[0024] The emotion expression collaborative generation network for improving the user's real interview speech performance realizes the synchronous mapping of facial expression parameters, speech prosody parameters, and body movement parameters for cross-modal consistency control, through a cross-modal emotion expression collaborative generation facial expression-voice-gesture three-dimensional mapping system.

[0025] Facial expression animation engine: Based on ARKit 52-point facial muscle drive + SMPL-X skeleton animation; Micro-expression fusion technology: Supports 7 basic emotions + 15 compound emotions; Facial Dynamics Control: Precise control (±8ms) of the time curve of eyebrow raising / mouth corner upward turning.

[0026] Speech Synthesis and Processing: High-emotion-color TTS: Adaptive Flow-TTS model; Real-time Interruption Ability: Intervention delay < 50ms at any point in a sentence; Emotional Prosody Library: 12 special intonations for examiners such as questioning / doubting / dissatisfaction / urging, etc. Dynamic Speech Rate Control: Automatically match according to the examinee's speech rate (in the range of 0.8 - 1.2 times the speed).

[0027] Body Movement Generation: Hierarchical Skeleton Control System: Full-body Movement Library: 10 basic postures such as leaning forward / backward / crossing arms, etc. Fine Hand Movements: High-frequency detailed movements such as pen tip tapping / file flipping, etc. Gaze Behavior Generation: A line-of-sight interaction feedback system based on eye tracking.

[0028] Cross-modal Consistency Control: def generate_coherent_response(emotional_state, intensity): # Map emotional intensity to the facial muscle parameter space facial_params = emotion_to_blendshapes(emotional_state, intensity) # Synchronously generate voice prosody parameters voice_params = emotion_to_prosody(emotional_state, intensity) # Collaboratively generate body movements gesture_params = emotion_to_posture(emotional_state, intensity) # Cross-modal synchronization control return SynchronizedOutput(facial_params, voice_params, gesture_params) The real-time interruption decision-making system for improving users' real interview speech performance calculates the interruption timing based on the incremental conditional probability P(interrupt|context) = P_base + α·Q_drop + β·T_off + γ·S_stress.

[0029] P_base: Basic interruption probability (0.05 - 0.15) Q_drop: Degree of decline in answer quality T_off: Topic deviation degree S_stress: Current stress level Interruption execution control: Pre-load 3 - 5 interruption voice / action combinations; Collision avoidance algorithm: Ensure that the interruption point is a natural speech pause > 150ms; Priority control: Emergency correction > Questioning and follow-up > General feedback.

[0030] The stress adjustment mechanism for improving users' real interview speech performance includes a five-level stress model and closed-loop control based on physiological feedback, and the stress adjustment mechanism realizes collaborative dynamic adjustment in vision, hearing, and space based on multi-sensory collaboration technology.

[0031] Five-level stress model: L1 (Warm-up): Positive and encouraging feedback, line-of-sight contact rate 60%; L2 (Basic): Neutral expression, routine questions; L3 (Medium pressure): Doubtful expression + crossed arms, fixed line of sight; L4 (High pressure): Multiple examiners lean forward synchronously + frequently record behaviors; L5 (Extreme): Sudden environmental interference + examiners whispering to each other + clock speeding up.

[0032] Progressive stress adaptation: Closed-loop control based on physiological feedback; Automatic fallback mechanism: Detect that the extreme stress lasts > 45 seconds and automatically degrade.

[0033] Multi-sensory collaboration technology: Vision: Dynamically adjust the color temperature from 2700K to 6500K; Hearing: The environmental acoustic engine includes: background conversation / chair movement / clock countdown; Space: Dynamically adjust the distance of the examiner, virtual distance sense 0.5 - 3.5m.

[0034] In summary, the generation method of the VR-rendered virtual examiner multi-modal reaction generation system for interview speeches includes the steps: A. Receive and parse the multi-modal input data of the examinee; B. Generate collaborative feedback based on the emotional state and intensity; C. Perform real-time interruption according to the situation analysis result; D. Dynamically adjust the interview stress level.

[0035] The following are some examples of specific application scenarios: Scenario 1: Facial-voice collaborative feedback when the examinee digresses When it is detected that the examinee's answer deviates from the topic and the similarity is <0.6, the system generates collaborative feedback: Facial expression: The eyebrows are raised by 0.7 amplitude, the eyes are slightly squinted by 0.4 degree, and the transition time is 350 ms; Voice: "Can you specifically explain the relationship between this and the requirements of the question?" The examinee's intonation curve is [0, +2, +3, +1, 0, -2]; Body movement: The sequence of leaning forward + supporting the chin with the hand, and the action emphasis point is 1.2 seconds.

[0036] Specific generation process: { "facial": { "brow_raise": 0.7, / / Amplitude of eyebrow raising "eye_squint": 0.4, / / Degree of slight eye squinting "transition_time": 350 / / Expression transition time (ms) }, "speech": { "text": "Can you specifically explain the relationship between this and the requirements of the question?", "pitch_curve": [0, +2, +3, +1, 0, -2], / / Intonation curve "speed_factor": 0.9 / / The speech speed is slowed down by 10% }, "gesture": { "sequence": ["lean_forward", "hand_chin"], / / Leaning forward + supporting the chin with the hand "emphasis_point": 1.2 / / Action emphasis time point (seconds) } } Example 2: Multi-examiner stress scenario When the stress level reaches level 4 and lasts for >20 seconds: Examiner: 40%+ reduction in eye contact + frequent note-taking behavior; Deputy Examiner A: Leaning forward + whispering to each other; Deputy Examiner B: Obvious watch-checking action + soft coughing; Environment: 6000K cold light + 8dB increase in background noise; Interface: Timer color turns red + beating speed increases by 15%.

[0037] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art. Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A VR rendering virtual examiner multimodal response generation system for interview speeches, characterized in that, Including: A multi-modal input parsing module for real-time analysis of examinee status data; A virtual examiner behavior generation module for collaborative generation of expressions, voices, and body movements; A core algorithm module, including an emotion expression collaborative generation network, a real-time interruption decision-making system, and a stress regulation mechanism.

2. The VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 1, characterized in that: The multi-modal input parsing module real-time detects and analyzes the examinee status, including stress level, answer quality, and body unnaturalness index. Meanwhile, it high-frequency processes data such as expression changes, speech intonation, and eye movement trajectories, and a low-latency interruption mechanism processes abnormal events.

3. The VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 1, wherein: The reaction decision-making center of the virtual examiner behavior generation module includes a digital human multi-dimensional synthesis module, a scenario dynamic regulation module, and an interruption / questioning strategy generation module; The digital human multi-dimensional synthesis module performs multi-dimensional synthesis on the expressions, voices, and bodies of digital humans; The scenario dynamic regulation module controls scene lighting, environmental acoustic effects, and perspective transformation. The interruption / questioning strategy generation module performs real-time intervention based on an uncertainty threshold.

4. A VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 1, characterized in that: The emotion expression collaborative generation network realizes synchronous mapping of expression parameters, speech prosody parameters, and body movement parameters for cross-modal consistency control, and generates an expression-voice-gesture three-dimensional mapping system through cross-modal emotion expression collaboration.

5. The VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 4, characterized in that: The expression parameters are realized through an expression animation engine based on 52-point facial muscle drive of ARKit and SMPL-X skeletal animation, combined with micro-expression fusion technology and expression dynamics control.

6. The VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 4, wherein: The speech prosody parameters are realized through high-emotion-color TTS combined with an emotion prosody library to achieve real-time interruption ability and dynamic control of speech rate.

7. A VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 4, characterized in that: The body movement parameters are realized through a hierarchical bone control system based on a full-body movement library, including generation of fine hand movements and fixation behaviors.

8. The VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 1, characterized in that: The real-time interruption decision-making system calculates the interruption timing based on the increasing conditional probability P(interrupt|context)=P_base+α·Q_drop+β·T_off+γ·S_stress.

9. The multi-modal response generation system for VR-rendered virtual examiners in an interview speech according to claim 1, characterized in that: The stress regulation mechanism includes a five-level stress model and closed-loop control based on physiological feedback, and the stress regulation mechanism realizes collaborative dynamic adjustment in vision, hearing, and space based on multi-sensory collaboration technology.

10. A VR rendering virtual examiner multimodal response generation system for interview speeches according to claim 1, characterized in that: The generation method of the reaction generation system includes the steps of: A. Receiving and parsing examinee multi-modal input data; B. Generating collaborative feedback based on the emotional state and intensity; C. Performing real-time interruption according to the scenario analysis result; D. Dynamically adjusting the interview stress level.