Real-time pair training system and method based on multi-modal large model

By leveraging the synergistic effect of a multi-source sensor array and a multimodal alignment engine, combined with a deep learning model, the problem of insufficient multimodal signal capture in existing technologies is solved, enabling accurate, real-time, and adaptive evaluation of behavioral states and providing efficient personalized feedback for training.

CN121256265APending Publication Date: 2026-01-02HENAN JINMINGYUAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511446845.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies struggle to fully capture multimodal signals in behavioral state-based training that simulates real-world interactive scenarios. The evaluation results are rigid and lack adaptability, failing to provide accurate and personalized real-time feedback.

Method used

A multi-source sensor array is used to capture visual, auditory, and motion modal data. A multimodal alignment engine is used to achieve spatiotemporal alignment of the data. A comprehensive behavior evaluation index is generated using a 3D convolutional network, spatiotemporal joint convolution, and a Gated Recurrent Unit model. The evaluation criteria are optimized by combining dynamic threshold learning.

Benefits of technology

It achieves accurate synchronous acquisition and fusion of multimodal data, provides more comprehensive, real-time and adaptive behavioral state assessment, improves the accuracy and adaptability of assessment, and supports personalized real-time feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256265A_ABST
    Figure CN121256265A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time pair practice training system and method based on a multi-modal large model, and relates to the technical field of large model pair practice. A multi-source sensor array comprising a visual capturing module, an auditory processing module and an action sensing module is deployed, and space-time synchronization of sensor data is realized by using a multi-modal alignment engine; generating a visual space-time vector reflecting a facial dynamic state and a limb time sequence rule; generating a fundamental frequency of the voice signal and an auditory feature vector of formant dynamic change; and generating a pressure sensitive feature vector fusing the hand motion trail and the contact force data. And carrying out weighted fusion on the feature vectors, calculating a comprehensive behavior evaluation index, comparing the comprehensive behavior evaluation index with a preset positive threshold value and a preset negative threshold value, and judging a user behavior state. And calculating a prediction error according to the actual state feedback of the user, and adjusting a positive threshold and a negative threshold in real time by using a dynamic threshold learning updating mechanism to realize adaptive optimization of an evaluation standard.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large model training, in particular to a real-time training system and method based on a multi-modal large model. BACKGROUND

[0002] In the current large model training field, especially in behavior state training that needs to simulate real interaction scenarios, it is crucial to accurately and timely assess the behavior state of the trainee and provide feedback. Existing technologies have significant limitations.

[0003] Traditional evaluation methods often rely on single modal data, making it difficult to fully capture the complex multi-modal signals naturally present in human interaction, such as facial micro-expressions, body movements, speech acoustic features, semantic content, and hand operation force, resulting in one-sided and distorted behavior state judgments.

[0004] The evaluation process usually lacks effective real-time fusion mechanisms, making it difficult to achieve accurate spatio-temporal alignment of perception data from different sources and different time characteristics, affecting the synchronization and accuracy of analysis. In addition, evaluation standards are often statically preset and cannot adapt to individual differences of different users, changes in task scenarios, or behavior pattern evolution caused by trainee skill improvement, resulting in rigid and poorly adaptive evaluation results that cannot provide precise and personalized real-time feedback for training. Therefore, a technical solution is needed that can overcome these deficiencies and achieve more comprehensive, real-time, and adaptive evaluation of the behavior state of users during training. SUMMARY

[0005] The purpose of the present application is to provide a real-time training system and method based on a multi-modal large model to solve the problems existing in the prior art.

[0006] To achieve the above purpose, the present application provides the following technical solution: a real-time training method based on a multi-modal large model, comprising the following steps: Step S100: Deploy a multi-source sensor array to capture visual, auditory, and motion modal data. The visual capture sensor module captures facial AU motion vectors and body motion vectors; the auditory processing sensor module extracts F0 and formant features; the motion perception sensor module records gesture three-dimensional trajectories and contact force feedback, and a multi-modal alignment engine unifies the timestamps of each sensor; Step S100 includes the following steps: Step S101: Deploy the multi-source sensor array to capture visual, auditory, and motion data in real time during the interaction process; The visual capture sensor module uses a high frame rate stereo camera group to track key visual information in real time. The visual capture sensor module includes a facial motion sensor and a limb motion vector sensor. The facial motion sensor captures the activation intensity and changes of AU to reflect subtle facial expression states, forming AU motion vectors. The AU is the smallest unit of facial muscle movement that constitutes an expression. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and forms limb motion vectors. The auditory processing sensor module achieves precise acquisition and separation of sound signals through a microphone array, separating the target speech signal from the environmental background noise to obtain a pure speech signal; acoustic features are extracted from the pure speech signal to obtain F0 and formant, where F0 is the lowest frequency component generated when the vocal cords vibrate, and the formant is the concentrated energy frequency band generated by resonance in the speech signal; The motion sensing sensor module integrates an inertial measurement module as a motion quantification tool. Through the built-in accelerometer and gyroscope, it records the three-dimensional trajectory of the gesture movement and the contact force feedback in real time. The three-dimensional trajectory reflects the changes in spatial position in real time, and the contact force feedback reflects the pressure intensity during the gesture operation. Step S102: The multimodal alignment engine eliminates the spatiotemporal deviation between different modal data through time synchronization. The multimodal alignment engine uses a precise time protocol to uniformly calibrate the acquisition timestamps of each sensor, controlling the time error of different modal data within the range, so that the different modal data are strictly aligned in the time dimension, thereby improving the accuracy of multimodal data fusion and the reliability of analysis results. Step S200: A 3D convolutional network processes AU motion vectors to capture transient features of micro-expressions; spatiotemporal joint convolution processes limb motion vectors to maintain the continuity of actions and outputs visual spatiotemporal vectors; a temporal convolutional network extracts dynamic features of F0 and formants, encodes and parses semantics through Phoneme-Level Transformer, and outputs auditory feature vectors; a pressure-sensitive feature vector is generated through a GatedRecurrent Unit model. Step S200 includes the following steps: Step S201: Capture the transient characteristics of micro-expressions by performing temporal convolution on the AU motion vector using a 3D convolutional network; preserve the continuity of movement by performing spatiotemporal joint convolution on the limb motion vector, and jointly output the k-th visual spatiotemporal vector A. k The visual spatiotemporal vector is a feature matrix vector that includes the dynamics of facial micro-expressions and the temporal patterns of body movements; Step S202: Extract the dynamic changes of F0 and formants using a temporal convolutional network, and then use Phoneme-Level Transformer encoding to parse the speech content semantics and output the k-th auditory feature vector B. k The auditory feature vector is a comprehensive feature that integrates F0, formants, and semantic content of speech. Step S203: The trajectory features output by the inertial measurement module and the contact force feedback data are fused using the GatedRecurrent Unit model to generate the k-th pressure-sensitive feature vector C. k The pressure-sensitive feature vector makes the action feature include the motion trajectory and associate it with force information. It is a comprehensive feature that includes the long-term trajectory pattern of the gesture and the dynamics of the force.

[0007] Step S300: The visual spatiotemporal vector, auditory feature vector and pressure-sensitive feature vector are weighted and fused. The fusion process is achieved by weighted summation, and a comprehensive behavioral assessment index for behavioral state assessment is output. The comprehensive behavioral evaluation index mentioned in step S300 is: Step S300: Weight and fuse the multimodal feature vectors, and calculate the comprehensive behavioral evaluation index using the weighted summation formula. The multimodal feature vectors include visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors; ; In the formula, This represents the comprehensive behavioral evaluation index for the k-th group. The weights of the k-th visual spatiotemporal vectors; The weights of the k-th group of auditory feature vectors; The weights for the k-th pressure-sensitive feature vectors are set by professionals.

[0008] Step S400: Compare the comprehensive behavior evaluation index with the preset positive and negative thresholds to determine the user's status. Based on the actual status reported by the user, map it to a binary label, calculate the status prediction error, and dynamically adjust the positive and negative thresholds through the threshold learning update formula. Step S400 includes the following steps: Step S401: Combine the comprehensive behavioral assessment indicators with the positive threshold. With negative threshold The comparison determines whether the overall behavior is in a positive or negative state: ; In the formula, when the comprehensive behavior evaluation index is in the fuzzy decision zone, the state judgment is neutral. The setting of the fuzzy decision zone increases the tolerance for uncertain states and avoids arbitrary judgments. Step S402: Map user feedback to binary labels y, calculate the current state prediction error e, and learn and adjust the positive and negative thresholds in real time through threshold learning update formula based on the actual positive and negative states of user feedback. The binary label y: ; The state prediction error e: ; The threshold learning and update formula is as follows: ; In the formula, This represents the positive threshold for the (t+1)th update. This represents the negative threshold for the (t+1)th update. This represents the positive threshold for the t-th update. This represents the negative threshold after t updates. It is the historical weight decay factor. , It is the error compensation coefficient, N H N L e is the current number of positive and negative state samples. t This represents the state prediction error after the t-th update, allowing the threshold to adaptively approximate the user's true state and avoid oscillations.

[0009] A real-time training system based on a multimodal large model, comprising a multimodal data acquisition and synchronization module, a multimodal feature extraction and fusion module, a comprehensive behavior evaluation index calculation module, and a behavior decision-making and learning optimization module; The multimodal data acquisition and synchronization module captures the original visual, auditory, and motion signals during user interaction through three types of sensor arrays, and achieves spatiotemporal consistency of cross-modal data by relying on the multimodal alignment engine. The multimodal feature extraction and fusion module transforms the data obtained by the multimodal data acquisition and synchronization module into a high-order fusion feature vector, captures instantaneous changes in facial expressions and movements to generate a visual spatiotemporal vector, captures the frequency and frequency band of sound to generate an auditory feature vector, and captures the associated trajectory and force of gestures to generate a pressure-sensitive feature vector. The comprehensive behavior evaluation index calculation module receives visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors generated by the multimodal feature extraction and fusion module. It assigns feature weights to these three types of feature vectors, performs standardization on each feature vector to eliminate dimensional differences, and then multiplies each vector by its corresponding weight coefficient. Finally, it adds the weighted results to generate a real-time comprehensive behavior evaluation index. The behavior decision-making and learning optimization module compares the comprehensive behavior evaluation index with preset positive and negative thresholds to determine whether the user's overall behavior state is positive, negative, or neutral. At the same time, it calculates the state prediction error based on the user's actual feedback state and dynamically adjusts the positive and negative thresholds through threshold learning updates.

[0010] The multimodal data acquisition and synchronization module includes a visual capture unit, an auditory processing unit, a motion perception unit, and a multimodal alignment unit. The visual capture unit uses a high frame rate stereo camera group to simultaneously track facial micro-expressions and limb movements through a facial motion sensor and a limb motion vector sensor. The facial motion sensor generates motion vectors by capturing the activation intensity and changes of facial muscle movements. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and constitutes limb motion vectors. The auditory processing unit is equipped with a microphone array to separate the target human voice from the environmental background noise in the sound field and extract the core acoustic features of the pure speech, including the fundamental frequency components that reflect the pitch and the energy concentration frequency band that reflects the resonance characteristics of the speech. The motion sensing unit integrates a high-sensitivity inertial sensor, and combines an accelerometer and a gyroscope to synchronously record the motion trajectory of the gesture in three-dimensional space and the real-time pressure feedback during the operation process, forming motion quantification data that combines spatial path and force changes. The multimodal alignment unit performs global timestamp calibration on all sensor acquisition nodes, strictly aligning the time of visual images, sound signals and motion trajectories to construct unified multimodal data.

[0011] The multimodal feature extraction and fusion module includes a visual feature generation unit, an auditory feature generation unit, and a motion feature generation unit; The visual feature generation unit uses a spatiotemporal convolutional network to fuse the instantaneous facial expressions and the coherence of body movements output by the visual capture unit into a unified visual spatiotemporal vector. The auditory feature generation unit extracts the dynamic changes of the output data of the auditory processing unit through a temporal convolutional network, and then uses Phoneme-Level Transformer encoding to parse the semantic content of the speech content and output auditory feature vectors. The motion feature generation unit fuses the trajectory features output by the motion perception unit with the contact force feedback data through a Gated Recurrent Unit model to generate a pressure-sensitive feature vector.

[0012] The decision-making and learning optimization module includes a state decision-making unit, a feedback learning unit, and a threshold learning adaptation unit. The state decision unit compares the comprehensive behavior evaluation index with a preset threshold. If the index is higher than the positive threshold, the user is determined to be in a positive state. If the index is lower than the negative threshold, the user is determined to be in a negative state. If the index is between the two thresholds, it is marked as a fuzzy decision zone. When the comprehensive behavior evaluation index is in the fuzzy decision zone, it is determined to be in a neutral state. The feedback learning unit establishes a correlation channel between the user's actual state and the system's judgment. When the user reports their true state, the feedback learning unit converts it into a binary signal: a positive state is marked as positive feedback, and a negative state is marked as negative feedback. The state prediction error is calculated by comparing the judgment results of the state decision unit. The threshold learning and adaptation unit drives threshold updates based on the state prediction error. The threshold learning and adaptation unit continuously collects and updates historical threshold data: the positive threshold converges to the average level of positive user behavior samples; the negative threshold synchronously approaches the mean of negative behavior samples, and the historical threshold weights prevent overcorrection in a single feedback.

[0013] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention achieves precise synchronous acquisition and spatiotemporal alignment of visual, auditory, and motion data through the synergistic effect of a multi-source sensor array and a multimodal alignment engine, thereby improving the accuracy and reliability of multimodal data fusion.

[0014] 2. This invention extracts facial micro-expressions, body movements, speech acoustics, and hand trajectory force information into feature vectors with clear behavioral semantics, overcoming the limitations of single-modal analysis, and enabling a more comprehensive and in-depth portrayal of user behavior.

[0015] 3. This invention employs a weighted fusion strategy to generate comprehensive behavioral evaluation indicators, quantitatively reflecting the overall user status and providing a unified and objective evaluation basis for real-time decision-making. Dynamic threshold learning, driven by user feedback, adaptively adjusts positive and negative thresholds, enabling the evaluation criteria to continuously approximate the true distribution of user behavior, smoothly adapting to individual differences and behavioral pattern evolution, thus solving the evaluation rigidity problem caused by static thresholds. The fuzzy decision zone enhances the inclusiveness of uncertain states, and combined with a closed-loop feedback optimization process, it improves the accuracy of behavioral status judgment while ensuring real-time response, providing more precise and efficient intelligent evaluation support for training and practice. Attached Figure Description

[0016] Figure 1 This is a schematic diagram illustrating the steps of the real-time training method based on a multimodal large model according to the present invention. Figure 2 This is a schematic diagram of the structure of the real-time training system based on a multimodal large model according to the present invention; Figure 3 This is a schematic diagram of the multimodal data acquisition and synchronization module structure of the real-time training system based on a multimodal large model according to the present invention. Figure 4 This is a schematic diagram of the multimodal feature extraction and fusion module structure of the real-time training system based on a multimodal large model according to the present invention; Figure 5 This is a schematic diagram of the comprehensive behavioral evaluation index calculation module of the real-time training system based on a multimodal large model according to the present invention. Figure 6 This is a schematic diagram of the behavioral decision-making and learning optimization module structure of the real-time peer training system based on a multimodal large model according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example: Figures 1-6 As shown, this invention provides a real-time training system and method based on a multimodal large model, which sets up a user to train against the system, and the system monitors the user's emotional state in real time.

[0019] Step S100: Deploy a multi-source sensor array to capture visual, auditory, and motion modal data. The visual capture sensor module captures facial AU motion vectors and limb motion vectors; the auditory processing sensor module extracts F0 and formant features; the motion perception sensor module records the three-dimensional trajectory of gestures and contact force feedback; and the multimodal alignment engine unifies the timestamps of each sensor. Step S100 includes the following steps: Step S101: Deploy the multi-source sensor array to capture data from different modalities of vision, hearing, and motion during the interaction process in real time; The visual capture sensor module uses a high frame rate stereo camera group to track key visual information in real time. The visual capture sensor module includes a facial motion sensor and a limb motion vector sensor. The facial motion sensor captures the activation intensity and changes of AU to reflect subtle facial expression states, forming AU motion vectors. The AU is the smallest unit of facial muscle movement that constitutes an expression. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and forms limb motion vectors. The auditory processing sensor module achieves precise acquisition and separation of sound signals through a microphone array, separating the target speech signal from the environmental background noise to obtain a pure speech signal; acoustic features are extracted from the pure speech signal to obtain F0 and formant, where F0 is the lowest frequency component generated when the vocal cords vibrate, and the formant is the concentrated energy frequency band generated by resonance in the speech signal; The motion sensing sensor module integrates an inertial measurement module as a motion quantification tool. Through the built-in accelerometer and gyroscope, it records the three-dimensional trajectory of the gesture movement and the contact force feedback in real time. The three-dimensional trajectory reflects the changes in spatial position in real time, and the contact force feedback reflects the pressure intensity during the gesture operation. Step S102: The multimodal alignment engine eliminates the spatiotemporal deviation between different modal data through time synchronization. The multimodal alignment engine uses a precise time protocol to uniformly calibrate the acquisition timestamps of each sensor, controlling the time error of different modal data within the range, so that the different modal data are strictly aligned in the time dimension, thereby improving the accuracy of multimodal data fusion and the reliability of analysis results. Step S200: By capturing AU motion vectors and combining them with limb motion vectors, output a visual spatiotemporal vector that integrates facial dynamics and limb temporal patterns; by dynamically changing F0 and formants, output an auditory feature vector that integrates acoustics and semantics; fuse hand movement trajectory with contact force data to generate a pressure-sensitive feature vector that includes gesture trajectory and force dynamics. Step S200 includes the following steps: Step S201: Capture the transient characteristics of micro-expressions by performing temporal convolution on the AU motion vector using a 3D convolutional network; preserve the continuity of movement by performing spatiotemporal joint convolution on the limb motion vector, and jointly output the k-th visual spatiotemporal vector A. k The visual spatiotemporal vector is a feature matrix vector that includes the dynamics of facial micro-expressions and the temporal patterns of body movements; Step S202: Extract the dynamic changes of F0 and formants using a temporal convolutional network, and then use Phoneme-Level Transformer encoding to parse the speech content semantics and output the k-th auditory feature vector B. kThe auditory feature vector is a comprehensive feature that integrates F0, formants, and semantic content of speech. Step S203: The trajectory features output by the inertial measurement module and the contact force feedback data are fused using the GatedRecurrent Unit model to generate the k-th pressure-sensitive feature vector C. k The pressure-sensitive feature vector makes the action feature include the motion trajectory and associate it with force information. It is a comprehensive feature that includes the long-term trajectory pattern of the gesture and the dynamics of the force.

[0020] Step S300: The visual spatiotemporal vector, auditory feature vector and pressure-sensitive feature vector are weighted and fused. The fusion process is achieved by weighted summation, and a comprehensive behavioral assessment index for behavioral state assessment is output. The comprehensive behavioral evaluation index mentioned in step S300 is: Step S300: Weight and fuse the multimodal feature vectors, and calculate the comprehensive behavioral evaluation index using the weighted summation formula. The multimodal feature vectors include visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors; ; In the formula, This represents the comprehensive behavioral evaluation index for the k-th group. The weights of the k-th visual spatiotemporal vectors; The weights of the k-th group of auditory feature vectors; The weights for the k-th pressure-sensitive feature vectors are set by professionals.

[0021] Example 1: Setting weights in the 5th group of practice The values ​​are 0.4, 0.3, and 0.3 respectively. The L2 norms of the fifth group of visual spatiotemporal vectors A5, B5, and C5 are set to... , , The comprehensive behavioral assessment index was calculated. =0.61.

[0022] Step S400: Compare the comprehensive behavior evaluation index with the preset positive and negative thresholds to determine the user's status. Based on the actual status reported by the user, map it to a binary label, calculate the status prediction error, and dynamically adjust the positive and negative thresholds through the threshold learning update formula. Step S400 includes the following steps: Step S401: Combine the comprehensive behavioral assessment indicators with the positive threshold. With negative threshold The comparison determines whether the overall behavior is in a positive or negative state: ; In the formula, when the comprehensive behavior evaluation index is in the fuzzy decision zone, the state judgment is neutral. The setting of the fuzzy decision zone increases the tolerance for uncertain states and avoids arbitrary judgments. Step S402: Map user feedback to binary labels y, calculate the current state prediction error e, and learn and adjust the positive and negative thresholds in real time through threshold learning update formula based on the actual positive and negative states of user feedback. The binary label y: ; The state prediction error e: ; The threshold learning and update formula is as follows: ; In the formula, This represents the positive threshold for the (t+1)th update. This represents the negative threshold for the (t+1)th update. This represents the positive threshold for the t-th update. This represents the negative threshold after t updates. It is the historical weight decay factor. , It is the error compensation coefficient, N H N L e is the current number of positive and negative state samples. t This represents the state prediction error after the t-th update, allowing the threshold to adaptively approximate the user's true state and avoid oscillations.

[0023] Example 2: Setting the positive threshold to 0.65 and the negative threshold to 0.35, the following results are obtained. The state is determined to be neutral, and the user feedback is set to a negative state, i.e., y=-1, e=0.26. The updated negative threshold is obtained. =0.41.

[0024] A real-time training system based on a multimodal large model, comprising a multimodal data acquisition and synchronization module, a multimodal feature extraction and fusion module, a comprehensive behavior evaluation index calculation module, and a behavior decision-making and learning optimization module; The multimodal data acquisition and synchronization module captures the original visual, auditory, and motion signals during user interaction through three types of sensor arrays, and achieves spatiotemporal consistency of cross-modal data by relying on the multimodal alignment engine. The multimodal feature extraction and fusion module transforms the data obtained by the multimodal data acquisition and synchronization module into a high-order fusion feature vector, captures instantaneous changes in facial expressions and movements to generate a visual spatiotemporal vector, captures the frequency and frequency band of sound to generate an auditory feature vector, and captures the associated trajectory and force of gestures to generate a pressure-sensitive feature vector. The comprehensive behavior evaluation index calculation module receives visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors generated by the multimodal feature extraction and fusion module. It assigns feature weights to these three types of feature vectors, performs standardization on each feature vector to eliminate dimensional differences, and then multiplies each vector by its corresponding weight coefficient. Finally, it adds the weighted results to generate a real-time comprehensive behavior evaluation index. The behavior decision-making and learning optimization module compares the comprehensive behavior evaluation index with preset positive and negative thresholds to determine whether the user's overall behavior state is positive, negative, or neutral. At the same time, it calculates the state prediction error based on the user's actual feedback state and dynamically adjusts the positive and negative thresholds through threshold learning updates.

[0025] The multimodal data acquisition and synchronization module includes a visual capture unit, an auditory processing unit, a motion perception unit, and a multimodal alignment unit. The visual capture unit uses a high frame rate stereo camera group to simultaneously track facial micro-expressions and limb movements through a facial motion sensor and a limb motion vector sensor. The facial motion sensor generates motion vectors by capturing the activation intensity and changes of facial muscle movements. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and constitutes limb motion vectors. The auditory processing unit is equipped with a microphone array to separate the target human voice from the environmental background noise in the sound field and extract the core acoustic features of the pure speech, including the fundamental frequency components that reflect the pitch and the energy concentration frequency band that reflects the resonance characteristics of the speech. The motion sensing unit integrates a high-sensitivity inertial sensor, and combines an accelerometer and a gyroscope to synchronously record the motion trajectory of the gesture in three-dimensional space and the real-time pressure feedback during the operation process, forming motion quantification data that combines spatial path and force changes. The multimodal alignment unit performs global timestamp calibration on all sensor acquisition nodes, strictly aligning the time of visual images, sound signals and motion trajectories to construct unified multimodal data.

[0026] The multimodal feature extraction and fusion module includes a visual feature generation unit, an auditory feature generation unit, and a motion feature generation unit; The visual feature generation unit uses a spatiotemporal convolutional network to fuse the instantaneous facial expressions and the coherence of body movements output by the visual capture unit into a unified visual spatiotemporal vector. The auditory feature generation unit extracts the dynamic changes of the output data of the auditory processing unit through a temporal convolutional network, and then uses Phoneme-Level Transformer encoding to parse the semantic content of the speech content and output auditory feature vectors. The motion feature generation unit fuses the trajectory features output by the motion perception unit with the contact force feedback data through a Gated Recurrent Unit model to generate a pressure-sensitive feature vector.

[0027] The decision-making and learning optimization module includes a state decision-making unit, a feedback learning unit, and a threshold learning adaptation unit. The state decision unit compares the comprehensive behavior evaluation index with a preset threshold. If the index is higher than the positive threshold, the user is determined to be in a positive state. If the index is lower than the negative threshold, the user is determined to be in a negative state. If the index is between the two thresholds, it is marked as a fuzzy decision zone. When the comprehensive behavior evaluation index is in the fuzzy decision zone, it is determined to be in a neutral state. The feedback learning unit establishes a correlation channel between the user's actual state and the system's judgment. When the user reports their true state, the feedback learning unit converts it into a binary signal: a positive state is marked as positive feedback, and a negative state is marked as negative feedback. The state prediction error is calculated by comparing the judgment results of the state decision unit. The threshold learning and adaptation unit drives threshold updates based on the state prediction error. The threshold learning and adaptation unit continuously collects and updates historical threshold data: the positive threshold converges to the average level of positive user behavior samples; the negative threshold synchronously approaches the mean of negative behavior samples, and the historical threshold weights prevent overcorrection in a single feedback.

[0028] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A real-time training method based on a multimodal large model, characterized by: The real-time training method includes the following steps: Step S100: Deploy a multi-source sensor array to capture visual, auditory, and motion modal data. The visual capture sensor module captures facial AU motion vectors and limb motion vectors; the auditory processing sensor module extracts F0 and formant features, where F0 is the lowest frequency component generated when the vocal cords vibrate; the motion perception sensor module records the three-dimensional trajectory of the gesture and the contact force feedback; and the multi-modal alignment engine unifies the timestamps of each sensor. Step S200: A 3D convolutional network processes AU motion vectors to capture transient features of micro-expressions; spatiotemporal joint convolution processes limb motion vectors to maintain the continuity of actions and outputs visual spatiotemporal vectors; a temporal convolutional network extracts dynamic features of F0 and formants, encodes and parses semantics through Phoneme-Level Transformer, and outputs auditory feature vectors; a pressure-sensitive feature vector is generated through a Gated RecurrentUnit model. Step S300: The visual spatiotemporal vector, auditory feature vector and pressure-sensitive feature vector are weighted and fused. The fusion process is achieved by weighted summation, and a comprehensive behavioral assessment index for behavioral state assessment is output. Step S400: Compare the comprehensive behavioral evaluation index with the preset positive and negative thresholds to determine the user's status. Based on the actual status reported by the user, map it to a binary label, calculate the status prediction error, and dynamically adjust the positive and negative thresholds through the threshold learning update formula.

2. The real-time training method based on a multimodal large model according to claim 1, characterized in that: Step S100 includes the following steps: Step S101: Deploy the multi-source sensor array to capture data from different modalities of vision, hearing, and motion during the interaction process in real time; The visual capture sensor module uses a high frame rate stereo camera group to track key visual information in real time. The visual capture sensor module includes a facial motion sensor and a limb motion vector sensor. The facial motion sensor captures the activation intensity and changes of AU to reflect subtle facial expression states, forming AU motion vectors. The AU is the smallest unit of facial muscle movement that constitutes an expression. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and forms limb motion vectors. The auditory processing sensor module achieves precise acquisition and separation of sound signals through a microphone array, separating the target speech signal from the environmental background noise to obtain a pure speech signal; acoustic features are extracted from the pure speech signal to obtain F0 and formant, where the formant is the concentrated energy frequency band generated by resonance in the speech signal; The motion sensing sensor module integrates an inertial measurement module as a motion quantification tool. Through the built-in accelerometer and gyroscope, it records the three-dimensional trajectory of the gesture movement and the contact force feedback in real time. The three-dimensional trajectory reflects the changes in spatial position in real time, and the contact force feedback reflects the pressure intensity during the gesture operation. Step S102: The multimodal alignment engine eliminates the spatiotemporal deviation between different modal data through time synchronization. The multimodal alignment engine uses a precise time protocol to uniformly calibrate the acquisition timestamps of each sensor.

3. The real-time training method based on a multimodal large model according to claim 2, characterized in that: Step S200 includes the following steps: Step S201: Capture the transient characteristics of micro-expressions by performing temporal convolution on the AU motion vector using a 3D convolutional network; preserve the continuity of movement by performing spatiotemporal joint convolution on the limb motion vector, and jointly output the k-th visual spatiotemporal vector A. k The visual spatiotemporal vector is a feature matrix vector that includes the dynamics of facial micro-expressions and the temporal patterns of body movements; Step S202: Extract the dynamic changes of F0 and formants using a temporal convolutional network, and then use Phoneme-Level Transformer encoding to parse the speech content semantics and output the k-th auditory feature vector B. k The auditory feature vector is a comprehensive feature that integrates F0, formants, and semantic content of speech. Step S203: The trajectory features output by the inertial measurement module and the contact force feedback data are fused using the GatedRecurrent Unit model to generate the k-th pressure-sensitive feature vector C. k The pressure-sensitive feature vector makes the action feature include the motion trajectory and associate it with force information. It is a comprehensive feature that includes the long-term trajectory pattern of the gesture and the dynamics of the force.

4. The real-time training method based on a multimodal large model according to claim 3, characterized in that: The comprehensive behavioral evaluation index in step S300 is: Step S300: Weight and fuse the multimodal feature vectors, and calculate the comprehensive behavioral evaluation index using the weighted summation formula. The multimodal feature vectors include visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors; ; In the formula, This represents the comprehensive behavioral evaluation index for the k-th group. The weights of the k-th visual spatiotemporal vectors; The weights of the k-th group of auditory feature vectors; The weights for the k-th pressure-sensitive feature vectors are set by professionals.

5. The real-time training method based on a multimodal large model according to claim 4, characterized in that: Step S400 includes the following steps: Step S401: Combine the comprehensive behavioral assessment indicators with the positive threshold. With negative threshold The comparison determines whether the overall behavior is in a positive or negative state: ; In the formula, when the comprehensive behavior evaluation index is in the fuzzy decision zone, the state is judged as neutral. Step S402: Map user feedback to binary labels y, calculate the current state prediction error e, and learn and adjust the positive and negative thresholds in real time through threshold learning update formula based on the actual positive and negative states of user feedback. The binary label y: ; The state prediction error e: ; The threshold learning and update formula is as follows: ; In the formula, This represents the positive threshold for the (t+1)th update. This represents the negative threshold for the (t+1)th update. This represents the positive threshold for the t-th update. This represents the negative threshold after t updates. It is the historical weight decay factor. , It is the error compensation coefficient, N H N L e is the current number of positive and negative state samples. t This represents the state prediction error after the t-th update.

6. A real-time peer-to-peer training system based on a multimodal large model, characterized in that: The real-time training system includes a multimodal data acquisition and synchronization module, a multimodal feature extraction and fusion module, a comprehensive behavior evaluation index calculation module, and a behavior decision-making and learning optimization module. The multimodal data acquisition and synchronization module captures the original visual, auditory, and motion signals during user interaction through three types of sensor arrays, and achieves spatiotemporal consistency of cross-modal data by relying on the multimodal alignment engine. The multimodal feature extraction and fusion module transforms the data obtained by the multimodal data acquisition and synchronization module into a high-order fusion feature vector, captures instantaneous changes in facial expressions and movements to generate a visual spatiotemporal vector, captures the frequency and frequency band of sound to generate an auditory feature vector, and captures the associated trajectory and force of gestures to generate a pressure-sensitive feature vector. The comprehensive behavior evaluation index calculation module receives visual spatiotemporal vectors, auditory feature vectors, and pressure-sensitive feature vectors generated by the multimodal feature extraction and fusion module. It assigns feature weights to these three types of feature vectors, performs standardization on each feature vector to eliminate dimensional differences, and then multiplies each vector by its corresponding weight coefficient. Finally, it adds the weighted results to generate a real-time comprehensive behavior evaluation index. The behavior decision-making and learning optimization module compares the comprehensive behavior evaluation index with preset positive and negative thresholds to determine whether the user's overall behavior state is positive, negative, or neutral. At the same time, it calculates the state prediction error based on the user's actual feedback state and dynamically adjusts the positive and negative thresholds through threshold learning updates.

7. The real-time training system based on a multimodal large model according to claim 6, characterized in that: The multimodal data acquisition and synchronization module includes a visual capture unit, an auditory processing unit, a motion perception unit, and a multimodal alignment unit. The visual capture unit uses a high frame rate stereo camera group to simultaneously track facial micro-expressions and limb movements through a facial motion sensor and a limb motion vector sensor. The facial motion sensor generates motion vectors by capturing the activation intensity and changes of facial muscle movements. The limb motion vector sensor calculates the spatial displacement, movement speed and direction of each joint of the limb through stereo vision, quantifies the dynamic characteristics of limb movements, and constitutes limb motion vectors. The auditory processing unit is equipped with a microphone array to separate the target human voice from the environmental background noise in the sound field and extract the core acoustic features of the pure speech, including the fundamental frequency components that reflect the pitch and the energy concentration frequency band that reflects the resonance characteristics of the speech. The motion sensing unit integrates a high-sensitivity inertial sensor, and combines an accelerometer and a gyroscope to synchronously record the motion trajectory of the gesture in three-dimensional space and the real-time pressure feedback during the operation process, forming motion quantification data that combines spatial path and force changes. The multimodal alignment unit performs global timestamp calibration on all sensor acquisition nodes, strictly aligning the time of visual images, sound signals and motion trajectories to construct unified multimodal data.

8. The real-time training system based on a multimodal large model according to claim 6, characterized in that: The multimodal feature extraction and fusion module includes a visual feature generation unit, an auditory feature generation unit, and a motion feature generation unit; The visual feature generation unit uses a spatiotemporal convolutional network to fuse the instantaneous facial expressions and the coherence of body movements output by the visual capture unit into a unified visual spatiotemporal vector. The auditory feature generation unit extracts the dynamic changes of the output data of the auditory processing unit through a temporal convolutional network, and then uses Phoneme-Level Transformer encoding to parse the semantic content of the speech content and output auditory feature vectors. The motion feature generation unit fuses the trajectory features output by the motion perception unit with the contact force feedback data through a Gated Recurrent Unit model to generate a pressure-sensitive feature vector.

9. The real-time training system based on a multimodal large model according to claim 6, characterized in that: The decision-making and learning optimization module includes a state decision-making unit, a feedback learning unit, and a threshold learning adaptation unit. The state decision unit compares the comprehensive behavior evaluation index with a preset threshold. If the index is higher than the positive threshold, the user is determined to be in a positive state. If the index is lower than the negative threshold, the user is determined to be in a negative state. If the index is between the two thresholds, it is marked as a fuzzy decision zone. When the comprehensive behavior evaluation index is in the fuzzy decision zone, it is determined to be in a neutral state. The feedback learning unit establishes a correlation channel between the user's actual state and the system's judgment. When the user reports the true state, the feedback learning unit converts it into a binary signal: a positive state is marked as positive feedback, and a negative state is marked as negative feedback. The state prediction error is calculated by comparing the judgment result of the state decision unit. The threshold learning and adaptation unit drives threshold updates based on the state prediction error. The threshold learning and adaptation unit continuously collects and updates historical threshold data: the positive threshold converges to the average level of positive user behavior samples; the negative threshold synchronously approaches the mean of negative behavior samples, and the historical threshold weights prevent overcorrection in a single feedback.