Analysis method and device based on multi-modal information, equipment and medium
By integrating multimodal information processing of visual, sound and physiological signals, generating standard multimodal information and performing feature extraction and comparison, it solves the subjectivity and lag problems of traditional training evaluation methods, and realizes real-time and objective evaluation of training effects. It is suitable for the fields of financial technology and medical health.
Patent Information
- Application Number
- CN202510826681.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing training effectiveness evaluation methods rely on manual judgment, are highly subjective, and lack real-time, quantitative evaluation methods. They are unable to meet the needs of high-demand industries for refined evaluation of training quality, especially in the fields of financial technology and healthcare. Traditional methods are unable to capture the physiological state and emotional changes of trainees, resulting in delayed feedback and affecting training efficiency and adaptability.
By obtaining the visual information, sound information and physiological indicator information of the target object, preprocessing and feature extraction of multimodal information are performed to generate standard multimodal information. Analysis is performed based on multi-dimensional features to generate state representation, which is then compared with the preset reference benchmark to generate analysis results.
It achieves an objective and comprehensive assessment of the status of trainees during the training process, provides quantifiable, comparable and explainable analysis results, improves the real-time and scientific nature of training feedback, and is suitable for high-demand industries such as financial technology and medical health.
Smart Images

Figure CN120746783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an analysis method, apparatus, device, and storage medium based on multimodal information. Background Art
[0002] Against the backdrop of accelerating digital transformation, businesses and institutions are increasingly focused on training quality. As a result, how to scientifically, comprehensively, and accurately evaluate training effectiveness has become a widespread concern. Traditional training effectiveness evaluation methods rely primarily on self-assessment by trainees, instructor ratings, or manager observation. These methods rely heavily on human judgment and are inherently subjective, making it difficult to guarantee objectivity and consistency.
[0003] In the fintech sector, financial institutions often lack real-time, quantitative assessment methods for employee risk awareness training, compliance education, and customer service enhancement. This is especially true in high-frequency, high-intensity business scenarios, where trainees' stress levels, attention spans, and emotional states often directly impact training effectiveness. However, these factors are difficult to capture or quantify using traditional assessment methods, resulting in delayed feedback on training quality and an inability to adjust training strategies in a timely manner.
[0004] In the healthcare sector, the physiological state and behavioral performance of medical staff during courses like professional skills training, psychological stress intervention, or patient communication skills development significantly influence their absorption. However, traditional methods often overlook trainees' facial expressions, behavioral details, and voice changes during training, making it impossible to systematically and objectively assess their psychological state and cognitive load. Furthermore, the healthcare industry places higher demands on the accuracy and scientific nature of assessment results, and existing methods that rely on questionnaires or teacher-graded assessments lack accuracy and data depth.
[0005] Furthermore, existing training evaluation methods mostly occur after training has concluded. This delayed feedback mechanism prevents timely identification and correction of problems, hindering training efficiency and adaptability. The data collected during the evaluation process lacks multi-dimensional insights, typically focusing only on superficial indicators such as academic performance, failing to gain insight into key factors such as deeper cognitive states, emotional changes, and attention spans. Furthermore, traditional methods generally lack automation and intelligent support, making the evaluation process cumbersome and inefficient, making them unsuitable for large-scale, real-time training scenarios.
[0006] To sum up, existing technologies in training effect evaluation have problems such as strong subjectivity, untimely feedback, single evaluation dimension, lack of quantitative basis and cumbersome process. It is difficult to meet the needs of high-demand industries for refined evaluation of training quality. There is an urgent need for a more objective, comprehensive and real-time intelligent evaluation method. Summary of the Invention
[0007] The main purpose of the present invention is to provide an analysis method, device, equipment and storage medium based on multimodal information, aiming to solve the technical problem that the existing technology is unable to achieve multimodal fusion based on visual information, sound information and physiological indicator information, and comprehensively extract multi-dimensional features to generate a representation of the target object state, thereby resulting in a lack of objectivity and completeness in the evaluation results.
[0008] To achieve the above objectives, the present invention provides an analysis method based on multimodal information, comprising:
[0009] Acquire multimodal perceptual information of the target object, including visual information, sound information, and physiological indicator information;
[0010] Preprocessing the multimodal perception information to generate standard multimodal information;
[0011] Extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information;
[0012] Analyze the multi-dimensional features to generate a state representation of the target object;
[0013] Comparing the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0014] An analysis result of the target object is generated according to the status difference result.
[0015] Furthermore, to achieve the above-mentioned object, the present invention provides an analysis device based on multimodal information, comprising:
[0016] A multimodal acquisition module is used to obtain multimodal perception information of the target object, including visual information, sound information, and physiological indicator information;
[0017] A multimodal preprocessing module, configured to preprocess the multimodal perception information to generate standard multimodal information;
[0018] A multi-dimensional feature extraction module, configured to extract multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information;
[0019] A state modeling module, configured to analyze the multi-dimensional features and generate a state representation of the target object;
[0020] A state comparison module is used to compare the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0021] The analysis generation module is used to generate an analysis result of the target object according to the state difference result.
[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an analysis program based on multimodal information stored in the memory and executable on the processor. When the analysis program based on multimodal information is executed by the processor, the steps of the analysis method based on multimodal information as described above are implemented.
[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an analysis program based on multimodal information is stored. When the analysis program based on multimodal information is executed by a processor, the steps of the analysis method based on multimodal information as described above are implemented.
[0024] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses an analysis method, device, equipment and medium based on multimodal information, including: obtaining visual information, sound information and physiological indicator information of the target object; preprocessing the multimodal perception information to generate standard multimodal information; extracting facial representation features, behavioral representation features, voice representation features and physiological representation features from the standard multimodal information; generating a state representation of the target object based on multi-dimensional features; comparing the state representation with a preset reference benchmark to generate a state difference result; and generating an analysis result of the target object based on the state difference result. The present invention constructs standard multimodal information by fusing visual, sound and physiological signals, extracts multi-dimensional features and generates a state representation, and then compares it with a reference benchmark to obtain a difference result, so that the analysis result has the characteristics of being quantifiable, comparable and interpretable, thereby achieving an objective evaluation of the state of the target object, overcoming the problems of the existing technology in which the evaluation method is highly subjective and covers a single dimension. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0026] Figure 1 Schematic diagram of an application environment of an analysis method based on multimodal information in one embodiment of the present invention;
[0027] Figure 2 1 is a flow chart of an embodiment of an analysis method based on multimodal information according to the present invention;
[0028] Figure 3 Schematic diagram of functional modules of a preferred embodiment of an analysis device based on multimodal information of the present invention;
[0029] Figure 4A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0030] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0032] The analysis method based on multimodal information provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain the visual information, sound information and physiological indicator information of the target object through the user terminal; pre-process the multimodal perception information to generate standard multimodal information; extract facial representation features, behavioral representation features, voice representation features and physiological representation features from the standard multimodal information; generate the state representation of the target object based on the multi-dimensional features; compare the state representation with a preset reference benchmark to generate a state difference result; and generate an analysis result of the target object based on the state difference result. The present invention constructs standard multimodal information by fusing visual, sound and physiological signals, extracts multi-dimensional features and generates a state representation, and then compares it with the reference benchmark to obtain a difference result, so that the analysis result has the characteristics of quantification, comparability and interpretability, thereby realizing an objective evaluation of the state of the target object and overcoming the problems of strong subjectivity and single coverage dimension of the evaluation method in the prior art. The user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented by an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0033] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the multimodal information analysis method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0034] like Figure 2 As shown, the analysis method based on multimodal information proposed in the present invention includes the following steps:
[0035] S10, acquiring multimodal perception information of the target object including visual information, sound information, and physiological indicator information;
[0036] In this embodiment, acquiring multimodal perceptual information of the target object is fundamental to achieving state analysis and assessment. Visual information is captured by an image acquisition module connected to the processing system. The image acquisition module typically includes a high-definition camera or a depth camera to record facial and body posture images. Facial images capture detailed features such as facial expression changes, eye movements, and mouth movements, while limb images record dynamic information of key areas such as the head, shoulders, and arms. The image acquisition process must ensure the stability of the frame rate and resolution. Generally, a resolution of 1080P or higher and a capture frequency of no less than 30 frames per second are used to ensure the accuracy of subsequent analysis.
[0037] The acquisition of sound information relies on the synergy between a microphone array and an audio acquisition device. The microphone array provides multi-channel sound source localization and separation, ensuring the extraction of pure speech signals from the target subject in multi-person environments. The acquisition process requires simultaneous signal processing, including echo suppression and background noise reduction. The target subject's speech information, including acoustic characteristics, intonation, speech rate, and prosody, is typically acquired at a sampling rate of 16kHz or higher to ensure high-quality speech data for subsequent acoustic modeling and analysis.
[0038] Physiological indicator information is collected through an integrated sensor module, which can be composed of a wearable device, a handheld terminal, or an embedded monitoring unit. Heart rate data is typically acquired by a photoplethysmography (PPG) sensor, which is highly real-time and non-invasive. Skin conductivity data is obtained through electrode sensing and reflects the level of sympathetic nerve activity. To improve data collection accuracy, the data sampling frequency of different indicators should be set. For example, the sampling frequency of heart rate data can be set to 64Hz to 128Hz, and the sampling frequency of skin conductivity data can be between 10Hz and 32Hz.
[0039] To ensure comparability and synchronization of information from different sources, visual, audio, and physiological data are time-stamped upon collection. These timestamps are generated based on the system clock, with millisecond accuracy, and are used for subsequent alignment analysis. The system implements a time synchronization control strategy for all data input channels, ensuring a consistent time base for data from different modalities, thus avoiding errors caused by timing drift during feature extraction.
[0040] In specific implementations, the image acquisition module can be deployed in a fixed location for remote monitoring or integrated into mobile devices to flexibly capture the target object's movements. To improve acquisition coverage, the visual acquisition module can be configured with multi-angle cameras for image stitching and angle compensation. Microphone arrays can be deployed in conjunction with sound source localization algorithms to enhance speech recognition in noisy environments. Audio data synchronization can be precisely controlled based on audio frame header encoding to ensure alignment with video data frames.
[0041] Physiological data collection can be performed using wearable wristbands based on Bluetooth communication or remote health monitoring terminals connected via Wi-Fi. The edge computing module performs preliminary filtering and anomaly removal on the collected data. During the multimodal information aggregation phase, all data is first synchronized by a unified data time alignment module before being merged into a structured multimodal perception data structure for downstream processing modules.
[0042] In terms of deployment strategy, visual and audio information collection components can be configured with permission levels and data encryption mechanisms to prevent raw data leakage. In user-oriented scenarios, the device's current collection status can also be displayed through the user interface, improving transparency and operational controllability.
[0043] Example: In the healthcare field, in clinical rehabilitation scenarios, cameras and microphones can be used to capture the patient's facial expressions, vocalization status, and body movement trajectories during training. Combined with the heart rate and skin conductivity collected by wearable devices, the patient's emotional changes and neural activation level can be identified, providing a basis for doctors to adjust rehabilitation plans.
[0044] In the financial technology business field, in remote training or interview scenarios, the user's facial state, voice expression and physiological signals are collected through terminal devices to judge the user's emotional fluctuations, attention changes and tension level throughout the task process, which is used to assist in interview scoring or training feedback mechanism, and improve the scientific nature of decision-making and the objectivity of evaluation.
[0045] This embodiment integrates visual information, sound information, and physiological indicator information and uniformly adds high-precision timestamps to them, thereby establishing a collaborative analysis basis for multimodal data, improving the temporal consistency and expression integrity of state recognition, and providing stable and comprehensive data support for subsequent feature extraction and state assessment processes.
[0046] S20, preprocessing the multimodal perception information to generate standard multimodal information;
[0047] In this embodiment, multimodal perception information often has problems such as inconsistent data quality, time axis misalignment, and non-uniform format during the original acquisition stage due to different signal sources, equipment differences, and environmental interference. Therefore, it needs to be preprocessed to ensure the feasibility and accuracy of subsequent feature extraction and analysis. The preprocessing operation first includes frame selection, image enhancement, and standardized size reconstruction of visual information. Frame selection uses the sampling rate control module to screen key frames with obvious movements and expression changes in continuous video frames, and removes blurred frames and duplicate frames. Image enhancement includes brightness normalization, contrast adaptive adjustment, and noise filtering. Common algorithms include histogram equalization, CLAHE (contrast limited adaptive histogram equalization), and Gaussian filters. To ensure a unified processing flow, the image data needs to be converted to a standard size format, for example, uniformly scaled to 224×224 pixels, and converted to a fixed channel order such as RGB three channels.
[0048] Sound information preprocessing primarily involves denoising, speech framing, and time-frequency feature conversion. The denoising process uses algorithms such as spectral subtraction and Wiener filtering to remove background noise while preserving the effective speech components. Speech signals are typically framed using a 25-ms window length and a 10-ms frame shift. Mel-frequency cepstral coefficients (MFCCs) or mel-frequency spectrograms are then extracted using a Mel filter bank to form a unified spectral representation. To enhance the accuracy of time-frequency analysis, the short-time Fourier transform (STFT) can be used as a basic transformation tool.
[0049] Physiological indicator information preprocessing includes outlier removal, filtering and smoothing, and time resampling. Heart rate signals are significantly affected by motion artifacts, so a combination of median filtering and low-pass filtering is used to smooth the signal. Extreme values are removed within a reasonable range, such as 40 to 180 bpm. Skin conductivity data requires baseline calibration, and its fluctuation range is calculated using sliding window statistics to form a standardized curve. The sampling frequency of each physiological signal is unified through linear interpolation or resampling to achieve cross-modal time alignment.
[0050] After completing preprocessing of their own dimensions, all modal data must enter the fusion and standardization phase. This phase includes time synchronization, data alignment, format conversion, and numerical normalization. Time synchronization maps all modal data to a unified time axis based on a timestamp mechanism; data alignment pairs sequences according to the nearest neighbor time matching principle; format conversion transforms data into a unified vector or tensor expression format based on the data structure requirements; and normalization uses Z-score standardization or Min-Max scaling to unify the value ranges of different modalities, ultimately outputting standard multimodal information with a clear structure, consistent timing, and standardized numerical values.
[0051] In one implementation, the image processing module uses image libraries such as OpenCV for frame preprocessing and image enhancement, with enhancement parameters automatically adjusted based on the acquisition environment. If the ambient lighting changes dramatically, a light sensing module can be integrated to dynamically adjust image brightness and contrast. The speech processing component integrates a real-time speech preprocessing model deployed on the end-side or edge device and utilizes a GPU to accelerate STFT and MFCC conversion operations. To address background noise, a multi-channel adaptive filter can be introduced to improve the speech signal-to-noise ratio.
[0052] Physiological signal acquisition devices can connect to the physiological monitoring platform to obtain complete timestamps and sampling parameter settings. After data upload, it is first stored in a local cache, pre-processed by the edge computing module, and then sent to the analysis center. During the aggregation phase, the central control scheduling unit manages the data of different modalities, controls the delay tolerance of data frames, and uses the dynamic time warping (DTW) algorithm to improve the stability of cross-modal alignment. The normalization process can update normalization parameters in real time based on historical sample statistics to adapt to individual differences.
[0053] Example description: In a medical and health training scenario, when medical staff participate in surgical operation standard training, the system collects visual information (such as eye movement trajectory and facial expressions), sound information (such as speech speed and tone), and physiological indicator information (such as heart rate changes and skin electrical response) during the operation. When executing this processing step, the collected surgical operation video is stabilized and the facial area is extracted. At the same time, the breathing noise and environmental noise in the recording are filtered out, and the heart rate data is synchronized and sorted according to the time window, so that all kinds of information are consistent in time dimension and spatial structure, forming multimodal sample data in a standard format. This processing provides input for the subsequent precise extraction of operation tension, concentration level and other states, and then evaluates the trainees' stress response and mastery of operation standards.
[0054] In a fintech training scenario, a relationship manager participates in a remote risk identification training course. The system collects real-time eye contact, voice intonation, and physiological signals from a smart bracelet during simulated conversations. During this processing step, visual data undergoes background removal to preserve facial and upper body images. The voice channel enhances language expression features through segmentation and echo suppression. Physiological data undergoes heart rate filtering and outlier processing. These multi-source data, after being aligned at a unified frame rate and standardizedly encoded, form standard multimodal information with consistent structure. This provides a high-quality data foundation for subsequent analysis of the manager's judgment and reaction speed, voice tone, and psychological stability during training conversations, thereby enabling quantification of training feedback and optimization of evaluation paths.
[0055] This embodiment, through unified preprocessing of visual, acoustic, and physiological signals, effectively eliminates interference factors, synchronizes time dimensions, and unifies data structures, thereby improving the accuracy, stability, and processing efficiency of subsequent multimodal feature extraction and state assessment. Standardized multimodal information, serving as structured input to the analysis system, establishes a reliable bridge between raw sensory data and the intelligent assessment module, significantly enhancing the overall system's data processing reliability and practicality.
[0056] S30, extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information;
[0057] In this embodiment, the process of extracting multidimensional features from standard multimodal information relies on structured modeling and feature fusion mechanisms tailored to different modal content. Facial representation features are a set of quantifiable parameters extracted from facial region images in visual information. These typically include the coordinates of facial key points, activation patterns of facial muscle groups, changes in eye movement trajectories, gaze direction heatmaps, and more. They are derived from facial action unit (APU) recognition models or expression classification networks based on convolutional neural networks. Behavioral representation features are a set of features formed by identifying behavioral attributes such as body movement patterns, posture stability, and repetitive movement frequency. They are often modeled and acquired using skeleton tracking algorithms, temporal posture encoders, or optical flow analysis algorithms, with a particular focus on upper body dynamics and hand interaction gestures. Speech representation features are audio structural parameters constructed based on speech signals, including Mel-Frequency Cepstral Coefficients (MFCCs), pitch contours, speech rate and duration, loudness energy, and pause structure. These features support the judgment of speech rhythm, tension, and speech clarity. Physiological characterization features come from the dynamic feature extraction process of standardized heart rate data, skin conductance values, respiratory rhythm and other signals. They are mainly expressed through signal micro-change period analysis, short-time energy statistics and peak frequency calculation to reflect the response level of psychological and physical states.
[0058] Multidimensional feature extraction uses standard multimodal information as input. It is crucial to ensure that the different modal data are aligned in time and have a unified spatial coordinate framework. Therefore, the various features are fused and encoded using time windows, normalized scale transformations, and simultaneous data masking. This entire process relies on a multi-channel parallel feature extraction framework or a joint embedding network architecture to ensure that the low-dimensional features of facial, behavioral, speech, and physiological data can be expressed in a unified tensor structure and passed to subsequent analysis stages.
[0059] By deploying a local data processing platform with integrated multimodal input interfaces and combining it with a real-time synchronization mechanism, standard multimodal information can be modularly distributed, with each modal data being fed into its corresponding feature extraction sub-model. Facial images are first captured and normalized using a face detection and alignment model. They are then fed into a trained expression recognition model to obtain expression classification probabilities and keypoint response intensities, forming facial representation features. Behavioral data uses a deep gesture recognition network to identify hand movement sequences and encode them into two-dimensional temporal coordinate features. A sliding window is then used to calculate movement amplitude and regularity to form behavioral representation features. Speech data, after noise suppression and endpoint detection, is fed into an acoustic feature encoder to extract acoustic descriptors of short speech frames. A speech rate estimation algorithm is also used to analyze intonation stability. Physiological data is extracted from the standardized continuous signal sequence using a sliding window, including heart rate variability, skin conductance trend curve features, and respiratory rhythm frequency index. The sampling rate and encoding dimensions are then unified. All extracted features are then spliced and normalized in time series to form a multidimensional feature set for analysis.
[0060] It is also possible to build a data processing architecture based on edge devices, pre-install lightweight feature extraction modules on the acquisition terminals of cameras, microphones, and physiological acquisition devices, and adapt to the embedded environment through model pruning and quantization technology, so that feature extraction can be achieved on-side computing without relying on cloud computing resources.
[0061] Example: In a healthcare training scenario, new nurses undergo standardized nursing process assessment training. The system collects standard multimodal information from their operations and extracts multidimensional features. Facial features reflect their concentration and emotional stability, behavioral features record the continuity of their gestures and body posture during ward rounds, voice features indicate the consistency of their speaking speed and tone when communicating with patients, and physiological features are used to identify whether their stress response is excessive during different nursing links. The extracted features are analyzed by the state recognition module and then fed back to the training system, helping managers adjust the content and methods of training in real time.
[0062] In a FinTech training scenario, financial advisors are trained in remote client communication simulations. The system processes standard multimodal information from these simulated interactions, extracting facial features to monitor emotional control during semantic expression, behavioral features to identify the appropriateness of body posture and gestures during communication, voice features to quantify changes in speech rate, pauses, and tone intensity, and physiological features to determine whether responses experience stress fluctuations. These features provide objective data support for the training platform, enabling quantitative assessment of the maturity and stability of communication strategies.
[0063] This embodiment utilizes a multidimensional feature extraction mechanism to construct high-fidelity representation vectors representing the target subject's facial state, behavioral dynamics, voice changes, and physiological responses from a unified multimodal format. This allows for signal decoupling and structural enhancement between different modalities, enabling subsequent state recognition models to possess greater expressiveness and analytical accuracy. The consistent processing of multimodal features ensures the analysis system operates stably in dynamic scenarios, improving the efficiency and accuracy of training behavior recognition and feedback generation.
[0064] S40, analyzing based on the multi-dimensional features to generate a state representation of the target object;
[0065] In this embodiment, analysis based on multi-dimensional features means taking structured data obtained from facial representation features, behavioral representation features, voice representation features, and physiological representation features as input, establishing a cross-modal fusion mechanism and a temporal modeling process, and performing overall expression modeling of the target object's cognitive, emotional, behavioral, and physiological response states in the current scenario, thereby generating a comprehensive output result that can reflect its current physical and mental state.
[0066] A state representation is a combination of multidimensional variables, typically expressed as a data vector containing continuous values or discrete labels, and includes aggregated indicators reflecting factors such as the target subject's emotional stability, attention concentration, interactive participation, and physiological stress level. The generation process of state representations typically relies on neural network models, time series analysis models, or expert rule systems to fuse, calculate, and attribute input features. Fusion models can include attention mechanisms, Transformer structures, or graph neural networks to enhance information relevance between different modalities. State attribution mechanisms can be based on label-supervised classifiers or use unsupervised clustering for category discovery, outputting representational state vectors or category labels.
[0067] This analysis process includes not only the statistical analysis and fusion of static features but also the dynamic identification and structural modeling of behavioral trends, emotional fluctuations, changes in speech rhythm, and physiological fluctuations. For example, multi-channel time series models, such as LSTM or TCN networks, can be constructed to learn the evolutionary patterns between features in the temporal dimension. Alternatively, multi-task learning mechanisms can be employed to output multiple state dimensions in parallel. The resulting state representation should exhibit cross-modal consistency, temporal stability, and task interpretability.
[0068] Based on a unified state analysis module, facial, behavioral, speech, and physiological features can be input and, after feature concatenation, fed into a multimodal fusion model. For example, in a multi-branch convolutional neural network, each modal feature is first extracted through independent convolution paths, then cross-fused in a shared layer. Nonlinear mapping is performed through a multi-layer perceptron to output a state representation vector. This vector can be divided into several dimensions, such as emotional fluctuation score, attention response level, behavioral consistency index, and physiological stress intensity, to comprehensively characterize the current comprehensive state of the target object.
[0069] Expert knowledge graphs can also be incorporated into the analysis process to match these features with established training target state indicators, generating state labels through conditional reasoning. For example, if the attention indicator falls below a set threshold and mood swings are severe, a combined state label of "attention loss + mood swings" can be added to the state representation for subsequent feedback decision modules to call.
[0070] To accommodate a variety of training scenarios, the output format of state representations can be configured as a configurable structure, including numerical scores, binary state labels, trend curves, or multi-label classification vectors. This module supports customizing state definition criteria based on the type of training project, thereby improving the versatility and adaptability of the analysis process.
[0071] Example: In healthcare training scenarios, this system can be used to assess changes in residents' states during simulated first aid training. The system extracts multi-dimensional features in real time and generates a state representation based on an analysis model. If facial features indicate tension, behavioral features identify slow movements, speech rhythm is disordered, and physiological features show a significant increase in heart rate, the state representation will output "high stress state" and "operational hysteresis" labels. This label is then used to trigger calming instructions or adjust training difficulty, enabling contextual intelligent intervention.
[0072] In FinTech training scenarios, this can be used to assess relationship managers' performance in simulated financial advisory sessions. The state representation process analyzes facial emotional stability, behavioral continuity, vocal clarity, and physiological stress responses during interactions. If the state representation indicates high attentional focus but weak vocal emotional output, the feedback system can flag this as "lack of expressive power," allowing for subsequent specialized training designed to enhance customer trust. This mechanism effectively assists managers in implementing targeted improvement plans, enhancing the scientific nature and individual variability of training optimization.
[0073] By inputting multi-dimensional features into a unified analysis model and generating a structured state representation, this embodiment achieves precise modeling of the target subject's current comprehensive state, providing clearer and more interpretable representations of individual behaviors and reactions. This process opens up cross-links between multimodal information, breaking through the limitations of traditional single-data source evaluation. This enables the training system to rapidly respond based on structured state information, providing more personalized and real-time feedback, significantly improving the scientific nature of training outcomes and the efficiency of adjustments.
[0074] S50, comparing the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0075] In this embodiment, the state comparison process relies on mapping the currently generated state representation into a standardized state assessment space and performing vector-level or structural difference calculations with one or more pre-built reference benchmark sets in this space. The reference benchmark is usually constructed based on historical efficient training samples, expert experience modeling, or goal achievement state abstraction, and is represented as a structured template containing multiple state dimensions and their weight distributions. For example, the reference benchmark can be represented as a multi-dimensional state vector, where each dimension corresponds to attention level, emotional stability, physiological activation, and behavioral consistency.
[0076] The comparison process first requires establishing a metric function for the state vector to quantify the degree of difference between the current state and the reference benchmark. This metric can be implemented based on Euclidean distance, cosine similarity, or Mahalanobis distance. Euclidean distance is suitable for quantifying differences in scalar state variables, cosine similarity is suitable for expressing trend consistency, and Mahalanobis distance is suitable for modeling scenarios where the state dimensions have a covariance structure. For categorical state labels, an evaluation mechanism in the form of a confusion matrix or cross-entropy loss is used to quantify the error between the current state prediction result and the reference state label.
[0077] To support fine-grained comparison, the state representation structure can be further decomposed into different levels, such as basic physiological state, complex emotional characteristics, and comprehensive cognitive indicators. A weighted structure is introduced during the comparison, assigning different importance factors to different dimensions. This approach allows the system to emphasize attention and cognitive state in training courses and emotional stability and behavioral control in emotion regulation courses.
[0078] After completing the difference measurement, the comparison module generates a state difference result, which includes at least one difference score metric and its corresponding state dimension annotation. The output can be a difference matrix, vector distance value, state offset label, or deviation level label (e.g., mild, moderate, severe deviation). The state difference result not only measures the degree of deviation between current performance and the desired state but also serves as a key input signal for adjusting intervention strategies or triggering personalized feedback mechanisms.
[0079] Furthermore, to ensure dynamic adaptability of the comparison, the comparison processing module should support dynamic configuration and adjustment of reference benchmarks. Different reference benchmark groups can be dynamically loaded based on the system's task type, user characteristics, or course stage. Furthermore, the reference state distribution can be adaptively updated using user data over the long term, enabling continuous optimization and transfer and generalization capabilities.
[0080] In actual deployment, a state comparison module can be built based on a deep feature matching framework. For example, a multilayer perceptron can be used to score the match between the state representation and the reference state vector, outputting a difference rating. Alternatively, a Siamese network architecture can be used to embed the state representation and the reference state, and then calculate the distance in the embedding space as a difference metric.
[0081] When addressing different course objectives, course-specific benchmarks can be introduced. For example, for a coping training course, the key benchmark might be set as a state of high attention and low anxiety; whereas for an expressive course, the benchmark might be set as a state of moderate emotional activation and natural behavioral tension. Benchmarks can be pre-configured in the model or automatically generated by the system based on clustering of high-performing learners.
[0082] In addition, in multi-person comparative training or group courses, a group baseline distribution can be constructed, and the variational inference mechanism can be used to model the degree of deviation between the current state and the group distribution boundary, and to identify individuals with significant deviations to indicate the need for personalized intervention.
[0083] Example: In psychological and emotional management training in healthcare, the system can establish a reference baseline based primarily on emotional stability and breathing rhythm. During meditation training, the system continuously generates a representation of the trainee's current state and compares it with this baseline. If the discrepancy indicates increased emotional fluctuations or uneven breathing, the system can prompt the instructor to strengthen emotional comfort guidance or adjust the training pace.
[0084] In high-pressure interview simulation training for the financial sector, a benchmark is established, including high voice clarity, stable facial expressions, and moderate physiological activation. If, during a student's simulated answering, the system detects that physiological stress indicators representing their current state are significantly higher than the benchmark, and their voice and intonation are disorganized, this discrepancy triggers a feedback mechanism that provides relaxation guidance and records the problematic scenario for subsequent review and optimization.
[0085] This embodiment introduces a multi-dimensional state representation and a comparison mechanism with a structured reference benchmark, enabling objective analysis of differences in the states of participants during training. Compared to evaluation methods that rely solely on single-point behaviors or results, this comparison mechanism can quantify the degree of differences in attention, emotions, physiological states, and behavioral characteristics in real time. This supports high-frequency state assessment and dynamic strategy adjustment, improving training adaptability and effectiveness stability.
[0086] S60: Generate an analysis result of the target object according to the state difference result.
[0087] In this embodiment, generating the analysis results of the target object refers to further extracting its multi-dimensional deviation indicators based on the state difference results formed by the comparison in the previous stage, and generating a targeted comprehensive description or optimization suggestion set through decision logic or mapping structure. The state difference results are usually structured outputs, which may include numerical difference indicators (such as attention deviation values, physiological state deviation values, etc.), classification labels (such as behavioral disorder levels) or level identifications (such as mild, moderate, and severe abnormalities). The purpose of further analysis based on such results is to convert the deviation information into understandable, executable, and traceable feedback output.
[0088] The analysis results can be generated by a set of predefined policy rules or a learning-based generative model. Predefined policy rules refer to the state-to-recommendation mapping logic set within the system. For example, if attention deviation exceeds a certain threshold and mood swings are frequent, feedback such as "Recommend a five-minute break and refocus" may be output. Alternatively, if physiological indicators deviate from the reference range for a long period of time and cognitive efficiency decreases, a warning message indicating "Risk of cognitive overload" may be generated.
[0089] Another approach is to use a multi-output model for joint reasoning. This involves building a neural network structure that takes in multiple state-differentiating features and outputs results across multiple analysis dimensions. This model can simultaneously predict feedback framework selection, intervention recommendations, and optimization priorities, achieving a closed-loop approach to the entire process.
[0090] Analysis results typically include multiple aspects of information, including problem attribution, performance assessment, risk level, and recommended actions. For example, in a cognitive course, the analysis might indicate "significant fluctuations in attention, suspected to be due to external interference"; in emotion regulation training, the analysis might state "emotions are becoming increasingly aggressive, requiring a moderate reduction in stimulating interactive content."
[0091] Furthermore, to enhance the contextual consistency of analysis results, the system should support combining historical state difference results with the long-term characteristic trajectory of the target object to enable trend deviation analysis and phased target adjustment. The system should also support loading different analysis rule sets or model structures for different training tasks, roles, and phases, thereby enabling flexible deployment of module-level analysis engines.
[0092] In practice, a rule system based on a tree-like decision path can be used, using each differential feature in the state difference results as an input node for a judgment condition, triggering the corresponding condition's result expression downward. For example, if a deviation in voice features and an increase in facial tension are detected, an emotional intervention suggestion path will be triggered, and the corresponding analysis result text will be output.
[0093] Alternatively, a knowledge graph structure can be used to represent analysis logic, where the state difference dimension is used as a node and connected to nodes such as behavioral recommendations and risk warnings through relationship edges, thus forming a scalable semantic reasoning structure. Dynamically bind the current object node in the graph and quickly generate a set of analysis results through graph path retrieval.
[0094] The analysis results can be output as structured documents, summary prompts, or parameterized control instructions. The output format is configured based on subsequent feedback or visual display requirements. For example, in an AI-assisted teaching system, the analysis results can be converted into directly readable teaching control parameters, such as "enter the buffer phase" or "reduce the difficulty of the current task," thereby enabling automated content adjustment.
[0095] In multi-role training scenarios, analysis results of different granularities can also be generated based on status difference results. For example, targeted feedback prompts can be generated for trainees, while visual trend reports and individual performance warning information can be generated for training instructors, thus achieving hierarchical feedback and role differentiation support.
[0096] Example: In an emotion regulation training course in the healthcare field, the system detected that the trainee's physiological indicators remained high and his expression was tense during a long task. After comparing the status to form a difference result, the analysis module further combined the trainee's past performance to generate an analysis result: "The subject is in a chronic high-activation state. It is recommended to arrange a relaxation guidance module and reduce the task load." This result can be directly called by the system to switch course content.
[0097] In high-pressure expression training courses in the financial field, if the system detects that the target subject has disordered speech rhythm and incoherent behavior in the simulated defense, and the state difference results significantly deviate from the expression fluency benchmark, the analysis module can output: "The current defense behavior has multiple interruption characteristics. It is recommended to insert an expression rhythm training module and prompt to reduce multi-threaded expression structure" to guide personalized expression ability optimization training.
[0098] This embodiment further analyzes the status difference results and converts them into specific structured results or recommended information, enabling continuous assessment and intelligent feedback on the target subject's training status. This processing approach enhances the practicality of status perception results, accurately attributing deviations from target performance based on multi-dimensional features and generating interpretable and actionable operational recommendations or risk warnings, effectively supporting the achievement of intelligent teaching, personalized coaching, and adaptive training goals.
[0099] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses an analysis method, device, equipment and medium based on multimodal information, including: obtaining visual information, sound information and physiological indicator information of a target object; preprocessing the multimodal perception information to generate standard multimodal information; extracting facial representation features, behavioral representation features, voice representation features and physiological representation features from the standard multimodal information; generating a state representation of the target object based on multidimensional features; comparing the state representation with a preset reference benchmark to generate a state difference result; and generating an analysis result of the target object based on the state difference result. The present invention constructs standard multimodal information by fusing visual, sound and physiological signals, extracts multidimensional features and generates a state representation, and then compares it with a reference benchmark to obtain a difference result, so that the analysis result has the characteristics of being quantifiable, comparable and interpretable, thereby achieving an objective evaluation of the state of the target object, overcoming the problems of the existing technology in which the evaluation method is highly subjective and covers a single dimension.
[0100] In one embodiment, the above step S10 includes:
[0101] S101, starting a high-definition image sensor connected to a processing module to capture a facial image sequence and a limb movement image sequence of the target object as the visual information;
[0102] S102, starting a microphone array connected to the processing module to collect the target object's vocal audio stream as the sound information;
[0103] S103, monitoring and acquiring heart rate time series data and skin conductivity change data of the target subject as the physiological indicator information through an integrated physiological sensor module worn by or in contact with the target subject;
[0104] S104: adding a unified timestamp to the visual information, sound information, and physiological indicator information to generate the multimodal perception information.
[0105] In this embodiment, a multi-channel acquisition method is used to synchronously acquire different data sources, constructing a temporally consistent multimodal perceptual information set for subsequent high-dimensional behavioral state modeling. Visual information is acquired via a high-definition image sensor, which, coupled with a processing module, provides real-time data stream integration, enabling the capture of facial image sequences and upper or full-body dynamic sequences of the target subject under uninterrupted sampling. Facial images are primarily used for subsequent emotion recognition, gaze direction determination, and muscle tension extraction, while body movement images assist in determining postural stability, concentration level, and nonverbal behavioral response characteristics.
[0106] Sound information is captured via a microphone array. The array's structural design supports multi-directional sound localization and noise rejection. The captured audio stream contains not only the speech content but also key elements of speech expression, such as intonation, speaking rate, rhythm, and pauses. This information is then used to construct speech feature vectors for metrics such as emotional tone extraction and clarity of expression. The microphone array supports highly sensitive input channels and maintains high-fidelity time resolution through system clock synchronization.
[0107] Physiological indicators are collected through wearable or contact-based physiological sensor modules, typically finger-clip or wrist-worn. Heart rate time series data can be used to monitor states such as psychological tension and task stress responses. Skin conductivity changes reflect sympathetic nerve activity and are widely used to monitor emotional arousal and stress responses. This type of data acquisition equipment requires a high sampling rate and robustness against motion interference to accommodate the diverse postures and behaviors encountered in dynamic training environments.
[0108] All collected information must be timestamped to form a time-consistent data structure. This unified timestamp not only serves as the foundation for inter-modal alignment but also determines the accuracy of subsequent analytical capabilities, such as synchronous feature extraction, behavioral peak comparison, and dynamic response modeling. Timestamps are typically derived from a unified system clock signal within the processing module and injected into the raw data stream during data packaging, ultimately forming multimodal perception information with modal synergy, temporal alignment, and the potential for feature fusion.
[0109] Through the above-mentioned synchronous acquisition and unified structural processing, it is possible to ensure that different types of perception information have a fusion basis in the spatial and temporal dimensions, providing a stable and high-quality data source input for subsequent behavior recognition and state analysis.
[0110] This embodiment collects visual information, audio information, and physiological indicator information, and unifies timestamps to form multimodal perception information. This allows for comprehensive perception of the target subject during the training process, covering multiple dimensions such as cognitive behavior, emotional expression, and physiological response. Facial images and body movements provide the basis for observing overt behavior, voice signals record expression patterns and language fluency, and physiological data reveals internal stress and emotional changes. The unified time alignment of the three types of data enables a high degree of synchronization between subsequent feature extraction and state modeling, significantly improving the accuracy of subsequent analysis and the effectiveness of intervention recommendations.
[0111] In one embodiment, the above step S20 includes:
[0112] S201, performing noise filtering operations and invalid data segment removal operations on the visual information, sound information, and physiological indicator information in the multimodal perception information, respectively, to generate preliminary pure visual information, preliminary pure sound information, and preliminary pure physiological indicator information;
[0113] S202, performing voice activity detection on the preliminary clean sound information, and extracting valid voice information segments from the preliminary clean sound information;
[0114] S203, performing data format unification conversion processing and data value range normalization adjustment processing on the preliminary clean visual information, the valid voice information segment, and the preliminary clean physiological indicator information, respectively, to generate visual information, valid voice information segment, and physiological indicator information with a unified format and normalized range;
[0115] S204 , based on a unified timestamp, synchronously aligning the visual information, valid voice information segments, and physiological indicator information with a unified format and normalized range to generate standard multimodal information.
[0116] In this embodiment, multimodal sensory information is used as input, and data cleaning, structural regularization, and alignment and synchronization operations are performed on the visual channel, speech channel, and physiological indicator channel, respectively, to obtain high-quality analysis data with a unified structure. First, the original image sequence, speech waveform, and physiological time series data are used as input. Two types of data cleaning operations are performed on each channel: noise filtering and invalid segment removal. The noise filtering operation constructs filtering models based on channel characteristics: for visual information, median filtering or convolution edge-preserving filtering methods are used to remove image blur and interfering background; for audio information, high-frequency noise or current noise is removed based on frequency domain analysis; and for physiological indicator information, transient pulse interference is removed through low-pass filtering. Invalid data segments are identified by setting thresholds based on indicators such as duration, intensity, or rate of change. For example, black frames, still frame sequences, silent periods, or heart rate flat bands are all considered invalid and removed.
[0117] After initial purification, voice activity detection is performed on the sound information. This process uses energy envelope analysis, speech spectral entropy, or endpoint detection algorithms to extract valid speech segments, removing non-semantic segments such as breathing, silence, and coughing. These extracted valid speech information segments form the analysis input basis for the language channel, improving the accuracy of subsequent speech feature extraction.
[0118] Next, the data structures of the three channels are formatted and range-normalized. Format unification includes standardization of image encoding formats, alignment of audio sampling rates and encoding formats, and matrixing of physiological data storage structures to ensure consistent data input interfaces. Range normalization compresses image pixel values, audio amplitudes, and physiological data amplitudes into a unified range through linear or Z-score normalization, preventing intermodal feature magnitude differences from affecting feature fusion weights.
[0119] After format and amplitude standardization, the data from each channel is synchronized and aligned based on a unified timestamp. This timestamp, derived from the system clock information uniformly injected during the original acquisition phase, is used to achieve cross-modal binding of image frames, speech frames, and physiological samples on the timeline. Specifically, interpolation or segmentation operations are used to ensure that the three data channels provide input of the same length within the same time window, forming a one-to-one corresponding fusion unit. The resulting data structure is standard multimodal information, which exhibits input stability, semantic integrity, and temporal synchronization.
[0120] This embodiment optimizes the input basis of multimodal analysis from three dimensions: data quality, structural consistency, and modal synergy, by cleaning, standardizing, and aligning the raw perceptual information. Noise filtering and invalid segment removal increase the effective information ratio of the data, voice activity detection improves the semantic density of the language channel, format unification and range normalization solve the problem of data source heterogeneity, and unified timestamp alignment ensures the temporal correlation between features. This processing chain enhances the fusion capability between modalities and the stability of data-driven analysis, providing input data support with reasonable structure, strong comparability, and consistent timeliness for subsequent state modeling and analysis.
[0121] In one embodiment, the above step S30 includes:
[0122] S301, inputting visual information in the standard multimodal information into a facial feature extraction module;
[0123] S302, identifying and quantifying facial key point position information and facial muscle activity information from the visual information by the facial feature extraction module, and combining the facial key point position information and the facial muscle activity information to generate the facial representation features;
[0124] S303, inputting the visual information in the standard multimodal information into a behavior feature extraction module;
[0125] S304, identifying and quantifying body joint motion information and body posture information from the visual information through the behavior feature extraction module, and combining the body joint motion information and the body posture information to generate the behavior representation feature;
[0126] S305, inputting the valid voice information segment in the standard multimodal information into a voice feature extraction module;
[0127] S306, analyzing and extracting acoustic parameter information and prosodic pattern information from the valid speech information segment by the speech feature extraction module, and combining the acoustic parameter information with the prosodic pattern information to generate the speech representation feature;
[0128] S307, inputting the physiological indicator information in the standard multimodal information into a physiological feature extraction module;
[0129] S308, extracting statistical features and time series features including heart rate variability parameters, skin conductivity mean and peak values from the physiological index information through the physiological feature extraction module, and combining the statistical features with the time series features to generate physiological representation features;
[0130] S309 , integrating the facial representation features, the behavioral representation features, the voice representation features, and the physiological representation features to generate the multi-dimensional features.
[0131] In this embodiment, a set of structured features with analytical value is extracted from standard multimodal information to characterize the state expression of the target object in visual, language and physiological dimensions. First, the visual information in the standard multimodal information is input into the facial feature extraction module. The module can identify the facial area and locate the key point position information based on the pre-trained convolutional neural network model. The key point position information includes but is not limited to the two-dimensional or three-dimensional coordinate data of facial landmark geometric nodes such as eyebrows, corners of eyes, corners of mouth, and nose tip. At the same time, facial muscle activity information is extracted. Common methods include muscle activation decoding based on action unit (AU) recognition or epidermal displacement intensity quantification based on optical flow analysis. The two types of information are jointly generated into facial representation features through vector splicing or structural fusion to reflect expression changes and emotional tendencies.
[0132] The same visual information is input into the behavioral feature extraction module, and the processing flow of the behavioral channel focuses on dynamic recognition and structural modeling at the limb level. First, the positions of the main joints of the body, including shoulders, elbows, knees, ankles, etc., are located, and the spatial motion trajectory and rate information of the joints are quantified through posture estimation models (such as OpenPose, HRNet, etc.). Secondly, the overall posture state is inferred, such as posture classification results such as sitting, standing, and leaning forward, as well as the corresponding stability parameters or posture change trends. The above-mentioned body joint motion information and body posture information are combined to generate behavioral representation features, which are suitable for identifying non-verbal behavior indicators such as attention state and degree of behavioral engagement.
[0133] Valid speech segments from the standard multimodal information are fed into the speech feature extraction module. Speech processing focuses on extracting phonetic features related to emotional state, expression intensity, and communication rhythm. Acoustic parameter information includes but is not limited to Mel-Frequency Cepstral Coefficients (MFCCs), spectrogram features, fundamental frequency variations, and energy envelopes. Prosodic pattern information can be characterized through frame-level speech rate variations, syllable duration, and pause distribution. The fusion of these two types of information constitutes speech representation features, which can further reflect psychological and behavioral attributes such as intonation, tension, and initiative.
[0134] Physiological indicators from the standard multimodal information are fed into the physiological feature extraction module. The raw heart rate and skin conductance data are windowed and segmented to extract statistical and temporal features. Heart rate variability parameters include SDNN, RMSSD, and PNN50, which represent the fluctuation of sympathetic and parasympathetic nerve activity. Skin conductance features include mean level and high-response peak frequency, reflecting the subject's physiological activation level per unit time. The output of this channel is a physiological feature, characterized by both long-term trend stability and instantaneous state volatility.
[0135] Finally, the four types of feature vectors are integrated in the time dimension and sample dimension to generate a multi-dimensional feature set with a unified structure to support subsequent state modeling operations.
[0136] This embodiment extracts explanatory and discriminative representational features from different modal information, establishes a feature set for four major channels: facial, behavioral, voice, and physiological, and can achieve a three-dimensional portrayal of the state of the target object. This feature extraction method has three advantages: first, a fusion processing strategy is adopted in different channels to jointly express static structure and dynamic changes, thereby improving the integrity of the features; second, the feature composition has highly quantifiable properties, which facilitates subsequent algorithm analysis and model training; third, the multimodal feature structure provides a semantically clear and structurally consistent foundation for subsequent feature fusion and state modeling, enabling the system to make accurate inferences about complex behavioral states, thereby improving the real-time and accuracy of multi-dimensional evaluation.
[0137] In one embodiment, the above step S40 includes:
[0138] S401, performing emotional state analysis based on the facial representation features in the multi-dimensional features to generate emotional state information;
[0139] S402, performing attention level analysis based on the behavior representation feature in the multi-dimensional feature to generate attention level information;
[0140] S403, performing sentiment analysis based on the speech representation features in the multi-dimensional features to generate sentiment information;
[0141] S404, performing physiological activation analysis based on the physiological representation features in the multi-dimensional features to generate physiological activation information;
[0142] S405, integrating the emotional state information, attention level information, emotional tendency information, and physiological activation information to generate a comprehensive state representation;
[0143] S406: Optimize the conflict indicators in the comprehensive state representation to generate the state representation of the target object.
[0144] In this embodiment, multi-dimensional features are used as input to construct a structured reasoning path for the psychological, physiological, and behavioral states of the target object, and aggregation and consistency optimization are performed between multi-channel state dimensions to ultimately form a unified state representation result. First, facial representation features are fed into the emotional state analysis module as input. The extracted facial key point trajectories and muscle activity maps are mapped and matched with the emotional facial action patterns in the pre-trained emotion recognition model to generate emotional state information containing emotion categories (such as joy, tension, confusion, anger) and confidence scores. During the recognition process, a temporal smoothing mechanism can be introduced to improve the stability of the recognition results of expression changes and suppress short-term interference.
[0145] Behavioral characteristics are fed into the attention level analysis module. By analyzing behavioral signals such as body joint stability, posture deviation range, and frequency of gaze changes, combined with behavioral focus feature templates established during training, the module infers the target subject's level of attention during the current period. This output information is typically a hierarchical classification (e.g., highly focused, moderately distracted, completely disengaged) or a continuous indicator (e.g., a concentration score) to assist in determining learning engagement.
[0146] The speech representation features are fed into the sentiment analysis module. Using the extracted acoustic parameters and prosodic structure, the module compares these features with the sample's positive, negative, and neutral sentiment label feature vectors in the emotional speech modeling pipeline to extract the attitude bias in the semantic expression. This module can identify semantic features such as urgency, encouragement, and resistance in the tone of voice, providing the system with a judgment of communication motivation based on the expression.
[0147] Physiological characteristics are then fed into the Physiological Activation Analysis module, which focuses on analyzing sympathetic nerve activation and emotional arousal. By identifying indicators such as heart rate fluctuations and changes in galvanic skin response frequency, the module calculates overall activation levels and generates physiological activation information. This information serves as a key basis for determining physiological states such as psychological stress, tension, and alertness.
[0148] After completing the generation of the state information of each channel, the system aggregates the emotional state information, attention level information, emotional tendency information and physiological activation information to construct a comprehensive state representation with a consistent structure. Each state dimension in this representation structure constitutes a vector field, and the overall structure can be represented by multi-dimensional vector splicing, graph structure expression or hierarchical nested tensor representation. On this basis, in response to possible cross-channel state conflicts (such as high behavioral concentration but indifferent voice performance), a conflict indicator optimization mechanism is introduced to perform weighted reconciliation or confidence difference adjustment on the state conflict areas, and finally output a semantically clear and structurally complete state representation of the target object.
[0149] This embodiment achieves multi-dimensional distributed state modeling of emotional state, attention level, emotional tendency and physiological activation by guiding the representation features from multiple channels of vision, speech and physiology to independent state analysis pathways. The fusion processing not only improves the integrity of state recognition, but also supports consistency judgment and deviation correction between multiple channels. The state conflict optimization mechanism further enhances the system's ability to parse semantic contradictions under complex state expressions, so that the final generated state representation can take into account the coordination of individual subjective performance and objective response. Through this mechanism, the actual state of the target object during the training process can be captured more accurately, supporting the subsequent intelligent feedback and strategy adjustment modules, and improving the accuracy and real-time performance of the overall system's understanding of the trainee's state.
[0150] In one embodiment, the above step S50 includes:
[0151] S501, according to the current analysis task type, select and load preset reference benchmarks corresponding to the analysis task type from a preset benchmark library, wherein the preset reference benchmarks include an emotional state benchmark value, an attention level benchmark value, an emotional tendency benchmark value, and a physiological activation degree benchmark value;
[0152] S502, decomposing the state representation into an emotional state component, an attention level component, an emotional tendency component, and a physiological activation component;
[0153] S503, determining a first difference value between the emotional state component and the emotional state reference value;
[0154] S504, determining a second difference between the attention level component and the attention level reference value;
[0155] S505, determining a third difference value between the emotion tendency component and the emotion tendency reference value;
[0156] S506, determining a fourth difference between the physiological activation component and the physiological activation reference value;
[0157] S507 , weightedly combining the first difference value, the second difference value, the third difference value, and the fourth difference value to generate the state difference result.
[0158] In this embodiment, after completing the comprehensive modeling of the target object's state, it is necessary to perform a structured comparison with the predefined performance benchmark to identify the deviation between its current state and the expected training performance. First, based on the type of the current analysis task, the system selects a reference sample set that matches the task attributes from the pre-built benchmark library. The analysis task type can be "concentration training", "emotional regulation ability training", "language expression ability training" or "psychological stability assessment", etc. The system calls the corresponding reference benchmark template through the task context or task identifier. Each reference benchmark structure contains four types of state benchmark values, namely emotional state benchmark value, attention level benchmark value, emotional tendency benchmark value and physiological activation benchmark value. These values are derived from historical excellent sample statistics, behavioral psychology standard model or training goal setting, and have a clear semantic range and indicator distribution.
[0159] Subsequently, the state representation vector structure is deconstructed at the component level to map out the four state dimension components of the current sample. The components on each dimension maintain the same data format and standardized units as the baseline value, ensuring semantic consistency and computational legitimacy of the difference value calculation. In the emotional state dimension, the system compares the relative deviation between the emotion category confidence distribution of the current sample and the reference value to generate a first difference value. The difference may include emotion category deviation, intensity deviation, emotional instability, etc. In the attention dimension, the distance relationship between the current attention score or level and the target reference value is compared to generate a second difference value. In the emotional tendency dimension, the deviation between the positivity or engagement in the speech expression and the standard expression level is compared to form a third difference value. In the physiological activation dimension, the degree of deviation between activation parameters such as heart rate variability and galvanic skin response and the target value is compared to generate a fourth difference value.
[0160] The four difference values are weighted and combined to produce a unified state difference result. This weighting mechanism supports the configuration of task sensitivity parameters. For example, in expression training tasks, the weight of emotional tendency differences can be increased, and in psychological adjustment training, the influence of physiological activation deviations can be enhanced. The comprehensive operation can adopt linear weighting, nonlinear combination, or normalized fusion strategies. The output state difference result can be used for subsequent feedback strategy selection, risk assessment, or personalized training path recommendation.
[0161] This embodiment differentiates the current state representation with a standard benchmark structure set for a specific analysis task, so that the state evaluation has a clear reference object and avoids the inconsistency problem caused by relying on human subjective experience. The deviation value of each state indicator is obtained based on the quantitative method of structural alignment, ensuring the reproducibility and traceability of the evaluation. The weighted fusion mechanism further improves the system's responsiveness to task objectives and the evaluation accuracy, enabling the evaluation process to dynamically adjust the focus according to the task emphasis, effectively improving the pertinence and effectiveness of training feedback.
[0162] In one embodiment, the above step S60 includes:
[0163] S601, comparing the state difference result with a preset state difference threshold to determine the state performance level;
[0164] S602, selecting a corresponding feedback framework according to the status performance level;
[0165] S603, generating an emotional state optimization solution based on the first difference value in the state difference result;
[0166] S604, generating an attention level optimization solution based on a second difference value in the state difference result;
[0167] S605, generating a sentiment tendency optimization solution based on the third difference value in the state difference result;
[0168] S606, generating a physiological state optimization plan based on a fourth difference value in the state difference result;
[0169] S607: Integrate the emotional state optimization scheme, the attention level optimization scheme, the affective tendency optimization scheme, and the physiological state optimization scheme into the feedback framework to generate the analysis result.
[0170] In this embodiment, based on the state difference results generated in the previous stage, it is necessary to construct a personalized state optimization analysis path to support subsequent intervention or guidance. In terms of operation, the system first evaluates the difference level of the current state difference results by quantitatively comparing the comprehensive difference score with the state difference threshold predefined by the system. The state difference threshold can be set to multiple level intervals, such as "excellent", "good", "to be improved", and "significant deviation", and each level corresponds to a different feedback strategy strength and content structure. The system outputs the corresponding state performance level based on the comparison results, thereby providing a judgment basis for the next step of feedback mechanism selection.
[0171] Based on the performance level of the status, the system calls the corresponding feedback framework from the policy library. This feedback framework includes pre-set interaction structures, suggestion semantic templates, display methods, and personalized recommendation interfaces, providing flexible adaptability. High-performance feedback frameworks focus on positive motivation and reinforcement strategies, while low-performance feedback frameworks include intervention suggestions, guidance content, and behavior correction task lists.
[0172] The system generates optimization plans for each of the four dimensional difference values in the state difference results. In the process of generating the emotional state optimization plan, based on the specific deviation direction and magnitude of the first difference value, a list of emotion regulation techniques is matched, such as deep breathing exercises, cognitive reconstruction guidance, or emotion recognition training modules. In terms of attention level optimization, the second difference value is used to guide the selection of focus training tasks, meditation guidance tasks, or audio-visual concentration challenges. For the third difference value, an emotional tendency optimization strategy is generated, such as voice adjustment training, expression structure drills, or interactive language evaluation mechanisms. For the fourth difference value, combined with the directionality of excessive or low physiological activation, moderate exercise suggestions, relaxation training guidance, or physiological stability monitoring plans are generated.
[0173] After generating the four optimization solutions, the system structured them into the selected feedback framework, generating analysis results tailored to the current target audience. These results include not only structured status evaluation data but also actionable training paths and feedback, supporting automatic training platform push, coach intervention, or phased training optimization.
[0174] During processing, to ensure data integrity and privacy, the system encrypts and protects key data related to the target object. First, after collecting visual information, audio information, and physiological indicators, the processing module performs type-based encryption and encapsulation on the raw data of these three types of information. Because visual information contains identity features such as facial images and body posture, it is encrypted using a symmetric encryption algorithm such as AES (Advanced Encryption Standard). Keys are dynamically generated by the system through the key management module and are periodically rotated and updated to reduce the risk of key leakage.
[0175] A multi-layered encryption strategy is employed for voice information, particularly the valid voice segments extracted after voice activity detection. TLS (Transport Layer Security) is used to establish an end-to-end encrypted channel during transmission. Voice segments are encrypted using a one-time key during local caching, with configured key lifetimes and access permissions. Physiological indicators, including heart rate time series data and skin conductivity change data, are highly sensitive and employ lightweight asymmetric encryption using the ECC (Elliptic Curve Cryptography) algorithm. These data are then bound to a unique user identity through a user authentication token.
[0176] Before all encryption operations, the system uniquely identifies and encodes the data, ensuring that even after encryption, it can still be associated with the original timestamp. This identification structure includes the data type tag, acquisition time, source module number, and encryption round information, supporting cross-node heterogeneous module identification and invocation. To ensure link security during the processing process, all inter-module data invocations pass through a unified data access gateway, which determines invocation permissions using a two-factor authentication mechanism based on an access control list (ACL) and a role permission table.
[0177] During multimodal feature extraction, state analysis, and result feedback, encrypted data is locally decrypted and subjected to least-privilege computation. A decryption token is temporarily generated only within the current task execution window and immediately destroyed upon task completion. All intermediate computation results are also desensitized, with no original traceable information retained, preventing unauthorized identification or data reconstruction during subsequent use.
[0178] Furthermore, the entire encryption policy execution process is logged and periodically verified by the system security audit module. This includes encryption algorithm call records, key lifecycle management logs, abnormal access interception reports, and data call path summaries. By embedding data encryption mechanisms throughout the entire data lifecycle, a closed-loop data security structure is established.
[0179] By embedding encryption mechanisms at each stage of multimodal data collection, storage, transmission, and processing, the risks of illegal access, theft, or leakage that may occur during the data processing chain can be effectively blocked. The use of a multi-algorithm combination strategy improves the encryption adaptability of different types of data, avoiding insufficient encryption strength or system performance bottlenecks. Combined with dynamic key management and a least-privilege computing strategy, it ensures that sensitive data is minimized and decrypted only in necessary scenarios, controlling potential leakage paths at the source. This is particularly suitable for training data analysis processes that contain individual facial, physiological, and voice features. While improving the system's intelligent processing capabilities, it also achieves strict protection of personal privacy data, enhancing the system's compliance and credibility in sensitive application areas such as medical training and human-computer interaction assessment.
[0180] To illustrate this, a bank's new employee credit assessment training program deployed a variety of information collection devices to provide personalized training adjustments and status feedback. These included a high-definition camera array mounted at the front of the classroom, a surrounding array of far-field microphone modules, and wristband-mounted physiological monitoring devices worn by employees. These devices collected visual, auditory, and physiological signals from participants during training interactions. The system first used image sensors to capture facial image sequences and body movement sequences of trainees, recording their dynamic visual behavior as they watched videos, participated in group tasks, or presented questions. Simultaneously, the microphone array captured speech stream data during speaking, including audio content of responses to questions, presentations, and discussions. Furthermore, the wearable device captured real-time heart rate variability and skin conductivity fluctuations to reflect physiological responses to tension, concentration, or anxiety. The system then accurately timestamped the data from each channel, generating structured multimodal perception information encompassing visual, auditory, and physiological information, and establishing a time-consistent data input channel.
[0181] The system then performs multi-level preprocessing on this multimodal perception information. For image data, blurry frames are removed and camera noise is filtered to improve the accuracy of facial and posture analysis. For voice data, background noise and invalid segments are filtered out through speech enhancement technology. For heart rate and skin conductance data, sudden changes, outliers, and non-contact periods are removed through a smoothing algorithm. The processed visual, audio, and physiological data are converted into a unified data format and numerically normalized based on a fixed sampling window to eliminate distribution differences in data from different modalities. The system then performs time synchronization alignment based on the unified timestamps of each channel to complete the construction of standardized multimodal data and form a standardized information input set for subsequent feature extraction.
[0182] Based on the standard data, the system then proceeds to the multi-dimensional feature extraction phase. First, the facial recognition module extracts the locations of facial key points and muscle change trajectories from the image sequences, constructing a vector that quantifies the trainee's facial changes and reflects their facial expressions during different training tasks. Next, the body posture recognition model analyzes movement amplitude, sitting stability, and hand movement frequency, extracting the locations of body joints and posture change trajectories to form behavioral feature representations. Simultaneously, the speech data is fed into the acoustic feature extractor and prosodic analysis network to extract speech rhythm and emotional parameters such as speaking rate, pitch, and pause intervals, further constructing a speech representation feature vector. Finally, the system extracts statistical and temporal metrics such as average heart rate, heart rate variability, and skin conductance mean and peak frequency from the physiological data to generate a quantifiable physiological state description vector. All four feature types are encoded and integrated into a unified multi-dimensional feature set, which is used to characterize the individual's state expression across multiple dimensions during training.
[0183] The system performs state inference based on the extracted multi-dimensional features. Facial feature analysis is used to determine whether the student is currently in an emotional state such as pleasure, tension, or confusion. Behavioral features are used to identify the student's concentration level, such as whether the student frequently turns the body or whether the student's attention is focused on the lecturer. Voice features are used to identify the emotional tendencies in language expression, such as the degree of positivity in the tone and willingness to communicate. Physiological features are used to determine the student's level of psychological activation, including anxiety, stress, or calmness. The above-mentioned multiple dimensional states are integrated into a comprehensive state representation, and conflicting or redundant information (such as a facial smile but abnormal skin conductance) is optimized within the model to output a unified psychological and physiological state mapping result.
[0184] The system loads the state benchmark that matches the current task from the reference database based on the task type configured on the training management platform (such as answering knowledge points and scenario analysis discussions). The benchmark includes the emotional stability, attention threshold, language and emotional activity, and physiological activation interval that should be achieved under normal conditions. The state representation is broken down into four categories of indicators, which are compared with the reference benchmarks to calculate the degree of difference. For example, if the current heart rate deviates from the standard by more than 20%, a physiological difference offset record is generated; if the emotional score is lower than the threshold of 75 points under normal conditions, it is marked as emotionally substandard. The system weights and summarizes the four types of differences to generate a unified state difference result, identifying the current employee's comprehensive deviation degree and key risk dimensions.
[0185] After generating the status difference results, the system starts the analysis and feedback process. First, based on the difference value and the set threshold, it determines whether the employee's status meets the standard and is divided into excellent, average or attention levels; then, the corresponding feedback template is retrieved according to the different levels. The system automatically generates dimensional optimization suggestions based on each type of difference result. For example, if there is attention deviation, it is recommended to assign short video learning tasks instead of document-style reading; if there is emotional deviation, it is recommended to use gamification interaction modules to relieve tension; if there is negative voice performance, it is recommended to join a group discussion to enhance expression motivation; if there is a deviation in physiological indicators, it is recommended to arrange a short break or flexibly adjust the rhythm. All dimensional optimization suggestions are encapsulated in the feedback framework, returned to the training system as analysis results, and pushed to the instructor terminal for tracking and intervention arrangements.
[0186] Throughout the training process, the system ensures information security through end-to-end data encryption. Raw data is encrypted locally on the collection device, and the transmission channel utilizes a TLS encryption layer. Multimodal data is encrypted using a symmetric encryption algorithm before storage. Training and inference tasks run in controlled execution containers and are protected from external access. System logs are stored in an encrypted format to prevent unauthorized access or secondary use of employee status, behavior, language, and physiological data involved in banking training, meeting the financial industry's requirements for data sensitivity and security compliance.
[0187] In the healthcare sector, during pre-job training for newly hired nurses at a Class A tertiary hospital, a multimodal perception component was deployed to monitor the nurses' psychological and physiological states in real time during training, in order to comprehensively assess their response capabilities in emergency tasks such as cardiopulmonary resuscitation, venipuncture, and critically ill patient identification. Specifically, the system activated a high-definition image acquisition module located above the simulated ward to record the trainees' facial expressions and body movements in real time as they performed simulated tasks. Simultaneously, a positioning microphone array was used to capture voice and audio data from the nurses' interactions with virtual patients or rehearsals of procedures. Furthermore, physiological sensors worn by the trainees on their wrists or earlobes captured continuous heart rate fluctuations and skin conductance fluctuations during training. These sensory data, derived from visual, auditory, and physiological channels, were uniformly timestamped and encapsulated into structured multimodal perception information, which served as input for the intelligent state analysis process.
[0188] The system first preprocesses the multimodal data. Visual channel data undergoes blur detection, redundant frame removal, and illumination normalization to output a clear and stable image sequence. Voice data uses the VAD (Voice Activity Detection) model to remove silent segments and background noise. Physiological data uses a filter to smooth abnormal transitions and remove invalid intervals caused by shedding or loosening. After various data types are converted to a standard format and normalized to a unit, they are synchronized in a unified time dimension using a timestamp alignment mechanism to generate a standardized multimodal input sequence, providing a standardized data foundation for subsequent state modeling.
[0189] The system then performs multidimensional feature extraction on each type of standardized data. First, the facial feature extraction module analyzes the trajectory of key facial points and muscle activity in the image sequence, such as eyebrow lift, eye tension, and changes in the corners of the mouth, to generate facial features that reflect emotional state. The behavioral feature extraction module extracts parameters for gesture and movement stability based on limb joint positions, movement speed, and body center of gravity shift, which are used to evaluate behavioral performance. The speech feature extraction module identifies acoustic parameters such as speech rate, intonation, rhythm, and pauses during nurse responses to determine language emotion and cognitive stability. The physiological feature extraction module extracts HRV (heart rate variability) parameters, skin conductance response peak frequency, and temporal stability indicators from heart rate and skin conductance sequences, forming a set of physiological features that reflect psychological workload and physiological activation state. After dimensional unification and embedding, these four features form a multidimensional feature vector that comprehensively depicts the state of the trained nurses.
[0190] The system performs state determination analysis based on the extracted multi-dimensional features. Emotional state is inferred from facial features, such as facial tension indicating anxiety; attention level is reflected by behavioral characteristics, such as prolonged gaze deviation and relaxed posture, which indicate distraction; emotional tendency is identified through voice pitch stability and rhythmic rhythm, such as rapid speech indicating excessive stress; and physiological activation is determined by decreased HRV and increased skin conductance, indicating tension or excessive stress. The system integrates the results of these four types of analysis to establish a comprehensive state representation. A built-in conflict indicator optimization module eliminates interference from conflicting indicators (e.g., behavioral relaxation but physiological tension), outputting a more consistent state representation.
[0191] To assess whether the nurse's current state meets the training objectives, the system matches the corresponding benchmark parameter group in the reference benchmark library from the task type (such as high-voltage emergency handling) and extracts the standard emotion value, attention threshold, emotional activity standard, and physiological stability target range. The system deconstructs the current state representation into four components and compares them with the standard value item by item to calculate the deviation difference. For example, if the heart rate variability is less than 20% of the benchmark lower limit, a "physiological activation difference" is generated; if the speech prosody performance differs significantly from the benchmark, an "emotional tendency difference" is generated. All difference values are weighted and fused through weight setting to generate a state difference result that characterizes the degree and direction of deviation of the current state.
[0192] Based on the above status difference results, the system starts a personalized analysis and feedback mechanism. According to the difference value classification standard, it is determined that the nurse belongs to the normal response, mild tension or high-risk concern level, and the adaptive feedback template is selected from the feedback library. A detailed optimization plan is generated for the source of deviation of each type of indicator. For example, if the heart rate is abnormal, a breathing relaxation training course is recommended; if the emotion is depressed, a video guidance module with positive emotion induction is recommended; if the attention is not focused, content structured exercises with interactive rhythm control are pushed; if the voice expression is low, targeted communication drill simulation tasks are pushed. The above optimization suggestions are assembled and encapsulated in the feedback framework by module, and the analysis results for the nurse himself and the training instructor are generated and pushed synchronously for the formulation of personalized supplementary training paths or task adjustments.
[0193] To ensure the compliance and privacy of the physiological and behavioral data collected by nurses during training, the system incorporates multiple data encryption mechanisms starting at the perception layer. Visual and audio information is locally encrypted by edge devices at the acquisition stage and then transmitted through TLS before being transmitted to the central processing system. All physiological data is stored in structured fields using symmetric key encryption. The training model and state analysis module run in a sandbox execution container to prevent model theft or sensitive data leakage. Data processing logs utilize field desensitization strategies, and a record-keeping mechanism for access control is in place.
[0194] This embodiment maps the state difference values to specific feedback levels and intervention strategies, so that the analysis results can not only stay at the evaluation level, but also extend to action-oriented optimization solutions, thus realizing a closed-loop process from evaluation to improvement. The independent generation and fusion processing of the four-dimensional optimization solutions ensure the targeted and detailed feedback, while maintaining structural consistency and coordination of intervention rhythm. The integrated feedback framework is deployable and highly interpretable, which is conducive to the rapid completion of dynamic adjustment and phased optimization of individual status in training scenarios, significantly improving the response speed and intervention quality of personalized training.
[0195] In one embodiment, a multimodal information-based analysis device is provided, which corresponds one-to-one to the multimodal information-based analysis method in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multimodal information analysis device of the present invention. It includes a multimodal acquisition module 10, a multimodal preprocessing module 20, a multidimensional feature extraction module 30, a state modeling module 40, a state comparison module 50, and an analysis generation module 60. Each functional module is described in detail below:
[0196] The multimodal acquisition module 10 is used to obtain multimodal perception information of the target object, including visual information, sound information, and physiological indicator information;
[0197] A multimodal preprocessing module 20, configured to preprocess the multimodal perception information to generate standard multimodal information;
[0198] A multi-dimensional feature extraction module 30 is used to extract multi-dimensional features including facial representation features, behavioral representation features, voice representation features and physiological representation features from the standard multimodal information;
[0199] A state modeling module 40 is configured to generate a state representation of the target object by analyzing the multi-dimensional features;
[0200] A state comparison module 50 is used to compare the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0201] The analysis generation module 60 is configured to generate an analysis result of the target object according to the state difference result.
[0202] In one embodiment, the multimodal acquisition module 10 is specifically configured to:
[0203] Starting a high-definition image sensor connected to the processing module to capture a facial image sequence and a limb movement image sequence of the target object as the visual information;
[0204] activating a microphone array connected to the processing module to collect an audio stream of a sound emitted by the target object as the sound information;
[0205] monitoring and acquiring the target subject's heart rate time series data and skin conductivity change data as the physiological indicator information through an integrated physiological sensor module worn or in contact with the target subject;
[0206] A unified timestamp is added to the visual information, sound information, and physiological indicator information to generate the multimodal perception information.
[0207] In one embodiment, the multimodal preprocessing module 20 is specifically configured to:
[0208] Performing noise filtering and invalid data segment removal operations on the visual information, sound information, and physiological indicator information in the multimodal perception information, respectively, to generate preliminary pure visual information, preliminary pure sound information, and preliminary pure physiological indicator information;
[0209] performing voice activity detection on the preliminary clean sound information, and extracting valid voice information segments from the preliminary clean sound information;
[0210] Performing data format unification conversion processing and data value range normalization adjustment processing on the preliminary pure visual information, the effective voice information segment, and the preliminary pure physiological indicator information, respectively, to generate visual information, effective voice information segment, and physiological indicator information with unified format and normalized range;
[0211] Based on a unified timestamp, visual information, valid voice information segments, and physiological indicator information with unified format and normalized range are synchronously aligned to generate standard multimodal information.
[0212] In one embodiment, the multi-dimensional feature extraction module 30 is specifically configured to:
[0213] Inputting visual information in the standard multimodal information into a facial feature extraction module;
[0214] Identifying and quantifying facial key point position information and facial muscle activity information from the visual information by the facial feature extraction module, and combining the facial key point position information and the facial muscle activity information to generate the facial representation features;
[0215] Inputting the visual information in the standard multimodal information into a behavior feature extraction module;
[0216] Identifying and quantifying body joint motion information and body posture information from the visual information through the behavior feature extraction module, and combining the body joint motion information and the body posture information to generate the behavior representation feature;
[0217] Inputting the valid speech information segments in the standard multimodal information into a speech feature extraction module;
[0218] Analyzing and extracting acoustic parameter information and prosodic pattern information from the valid speech information segment by the speech feature extraction module, and combining the acoustic parameter information with the prosodic pattern information to generate the speech representation feature;
[0219] Inputting the physiological indicator information in the standard multimodal information into a physiological feature extraction module;
[0220] Extracting statistical features and time series features including heart rate variability parameters, skin conductivity mean and peak values from the physiological index information through the physiological feature extraction module, and combining the statistical features with the time series features to generate physiological representation features;
[0221] The facial representation features, the behavioral representation features, the voice representation features, and the physiological representation features are integrated to generate the multi-dimensional features.
[0222] In one embodiment, the state modeling module 40 is specifically configured to:
[0223] Performing emotional state analysis based on facial representation features in the multi-dimensional features to generate emotional state information;
[0224] Performing attention level analysis based on the behavioral representation features in the multi-dimensional features to generate attention level information;
[0225] Performing sentiment analysis based on the speech representation features in the multi-dimensional features to generate sentiment information;
[0226] Performing physiological activation analysis based on the physiological representation features in the multi-dimensional features to generate physiological activation information;
[0227] fusing the emotional state information, attention level information, affective tendency information, and physiological activation information to generate a comprehensive state representation;
[0228] Optimizing the conflict indicators in the comprehensive state representation to generate the state representation of the target object.
[0229] In one embodiment, the state comparison module 50 is specifically configured to:
[0230] According to the current analysis task type, select and load preset reference benchmarks corresponding to the analysis task type from a preset benchmark library, wherein the preset reference benchmarks include an emotional state benchmark value, an attention level benchmark value, an emotional tendency benchmark value, and a physiological activation benchmark value;
[0231] Decomposing the state representation into an emotional state component, an attention level component, an affective tendency component, and a physiological activation component;
[0232] determining a first difference value between the emotional state component and the emotional state reference value;
[0233] determining a second difference between the attention level component and the attention level reference value;
[0234] determining a third difference value between the sentiment tendency component and the sentiment tendency reference value;
[0235] determining a fourth difference between the physiological activation component and the physiological activation baseline value;
[0236] The first difference value, the second difference value, the third difference value and the fourth difference value are weighted and integrated to generate the state difference result.
[0237] In one embodiment, the analysis and generation module 60 is specifically configured to:
[0238] Comparing the status difference result with a preset status difference threshold to determine the status performance level;
[0239] Selecting a corresponding feedback framework according to the state performance level;
[0240] generating an emotional state optimization solution based on a first difference value in the state difference result;
[0241] generating an attention level optimization scheme based on a second difference value in the state difference result;
[0242] generating a sentiment tendency optimization scheme based on a third difference value in the state difference result;
[0243] generating a physiological state optimization plan based on a fourth difference value in the state difference result;
[0244] The emotional state optimization scheme, the attention level optimization scheme, the affective tendency optimization scheme and the physiological state optimization scheme are integrated into the feedback framework to generate the analysis results.
[0245] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an analysis method based on multimodal information.
[0246] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. Among them, the processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a method for analyzing multimodal information.
[0247] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0248] Acquire multimodal perceptual information of the target object, including visual information, sound information, and physiological indicator information;
[0249] Preprocessing the multimodal perception information to generate standard multimodal information;
[0250] Extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information;
[0251] Analyze the multi-dimensional features to generate a state representation of the target object;
[0252] Comparing the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0253] An analysis result of the target object is generated according to the status difference result.
[0254] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0255] Acquire multimodal perceptual information of the target object, including visual information, sound information, and physiological indicator information;
[0256] Preprocessing the multimodal perception information to generate standard multimodal information;
[0257] Extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information;
[0258] Analyze the multi-dimensional features to generate a state representation of the target object;
[0259] Comparing the state representation of the target object with a preset reference benchmark to generate a state difference result;
[0260] An analysis result of the target object is generated according to the status difference result.
[0261] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0262] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0263] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0264] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An analysis method based on multimodal information, characterized in that: The following steps are involved: Acquire multimodal perceptual information of the target object, including visual information, sound information, and physiological indicator information; Preprocessing the multimodal perception information to generate standard multimodal information; Extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information; Analyze the multi-dimensional features to generate a state representation of the target object; Comparing the state representation of the target object with a preset reference benchmark to generate a state difference result; An analysis result of the target object is generated according to the status difference result.
2. The analysis method based on multimodal information according to claim 1, characterized in that: Acquire multimodal perceptual information of the target object, including visual information, sound information, and physiological indicator information, including: Starting a high-definition image sensor connected to the processing module to capture a facial image sequence and a limb movement image sequence of the target object as the visual information; activating a microphone array connected to the processing module to collect an audio stream of a sound emitted by the target object as the sound information; monitoring and acquiring the target subject's heart rate time series data and skin conductivity change data as the physiological indicator information through an integrated physiological sensor module worn or in contact with the target subject; A unified timestamp is added to the visual information, sound information, and physiological indicator information to generate the multimodal perception information.
3. The analysis method based on multimodal information according to claim 1, characterized in that: Preprocessing the multimodal perception information to generate standard multimodal information includes: Performing noise filtering and invalid data segment removal operations on the visual information, sound information, and physiological indicator information in the multimodal perception information, respectively, to generate preliminary pure visual information, preliminary pure sound information, and preliminary pure physiological indicator information; performing voice activity detection on the preliminary clean sound information, and extracting valid voice information segments from the preliminary clean sound information; Performing data format unification conversion processing and data value range normalization adjustment processing on the preliminary pure visual information, the effective voice information segment, and the preliminary pure physiological indicator information, respectively, to generate visual information, effective voice information segment, and physiological indicator information with unified format and normalized range; Based on a unified timestamp, visual information, valid voice information segments, and physiological indicator information with unified format and normalized range are synchronously aligned to generate standard multimodal information.
4. The analysis method based on multimodal information according to claim 1, characterized in that: Extracting multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information includes: Inputting visual information in the standard multimodal information into a facial feature extraction module; Identifying and quantifying facial key point position information and facial muscle activity information from the visual information by the facial feature extraction module, and combining the facial key point position information and the facial muscle activity information to generate the facial representation features; Inputting the visual information in the standard multimodal information into a behavior feature extraction module; Identifying and quantifying body joint motion information and body posture information from the visual information through the behavior feature extraction module, and combining the body joint motion information and the body posture information to generate the behavior representation feature; Inputting the valid speech information segments in the standard multimodal information into a speech feature extraction module; Analyzing and extracting acoustic parameter information and prosodic pattern information from the valid speech information segment by the speech feature extraction module, and combining the acoustic parameter information with the prosodic pattern information to generate the speech representation feature; Inputting the physiological indicator information in the standard multimodal information into a physiological feature extraction module; Extracting statistical features and time series features including heart rate variability parameters, skin conductivity mean and peak values from the physiological index information through the physiological feature extraction module, and combining the statistical features with the time series features to generate physiological representation features; The facial representation features, the behavioral representation features, the voice representation features, and the physiological representation features are integrated to generate the multi-dimensional features.
5. The analysis method based on multimodal information according to claim 1, characterized in that: Analyzing the multi-dimensional features to generate a state representation of the target object includes: Performing emotional state analysis based on facial representation features in the multi-dimensional features to generate emotional state information; Performing attention level analysis based on the behavioral representation features in the multi-dimensional features to generate attention level information; Performing sentiment analysis based on the speech representation features in the multi-dimensional features to generate sentiment information; Performing physiological activation analysis based on the physiological representation features in the multi-dimensional features to generate physiological activation information; fusing the emotional state information, attention level information, affective tendency information, and physiological activation information to generate a comprehensive state representation; Optimizing the conflict indicators in the comprehensive state representation to generate the state representation of the target object.
6. The analysis method based on multimodal information according to claim 1, characterized in that: Comparing the state representation of the target object with a preset reference benchmark to generate a state difference result, including: According to the current analysis task type, select and load preset reference benchmarks corresponding to the analysis task type from a preset benchmark library, wherein the preset reference benchmarks include an emotional state benchmark value, an attention level benchmark value, an emotional tendency benchmark value, and a physiological activation benchmark value; Decomposing the state representation into an emotional state component, an attention level component, an affective tendency component, and a physiological activation component; determining a first difference value between the emotional state component and the emotional state reference value; determining a second difference between the attention level component and the attention level reference value; determining a third difference value between the sentiment tendency component and the sentiment tendency reference value; determining a fourth difference between the physiological activation component and the physiological activation baseline value; The first difference value, the second difference value, the third difference value and the fourth difference value are weighted and integrated to generate the state difference result.
7. The analysis method based on multimodal information according to claim 1, characterized in that: Generate analysis results of the target object based on the status difference results, including: Comparing the status difference result with a preset status difference threshold to determine the status performance level; Selecting a corresponding feedback framework according to the state performance level; generating an emotional state optimization solution based on a first difference value in the state difference result; generating an attention level optimization scheme based on a second difference value in the state difference result; generating a sentiment tendency optimization scheme based on a third difference value in the state difference result; generating a physiological state optimization plan based on a fourth difference value in the state difference result; The emotional state optimization scheme, the attention level optimization scheme, the affective tendency optimization scheme and the physiological state optimization scheme are integrated into the feedback framework to generate the analysis results.
8. An analysis device based on multimodal information, characterized in that: The multimodal information-based analysis device includes: A multimodal acquisition module is used to obtain multimodal perception information of the target object, including visual information, sound information, and physiological indicator information; A multimodal preprocessing module, configured to preprocess the multimodal perception information to generate standard multimodal information; A multi-dimensional feature extraction module, configured to extract multi-dimensional features including facial representation features, behavioral representation features, voice representation features, and physiological representation features from the standard multimodal information; A state modeling module, configured to analyze the multi-dimensional features and generate a state representation of the target object; A state comparison module is used to compare the state representation of the target object with a preset reference benchmark to generate a state difference result; The analysis generation module is used to generate an analysis result of the target object according to the state difference result.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an analysis program based on multimodal information stored in the memory and capable of running on the processor. When the analysis program based on multimodal information is executed by the processor, the steps of the analysis method based on multimodal information as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores an analysis program based on multimodal information, and when the analysis program based on multimodal information is executed by the processor, the steps of the analysis method based on multimodal information as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Information processing method and device for psychological crisis screening
CN120998534A
Information processing methods and devices for psychological crisis screening
CN120998534B
Student psychological state monitoring method and system based on comprehensive data
CN121030784A
Multi-mode body feeling information processing method and system for physiological state evaluation
CN121400782A
A multi-modal somatosensory information processing method and system for physiological state assessment
CN121400782B