A non-contact mental capability assessment method
By combining multispectral cameras and millimeter-wave radar with sound sensors to acquire multimodal data, and then performing data fusion and machine learning, the problems of light source interference and insufficient heart rate signal quality in existing technologies have been solved, enabling accurate assessment of psychological abilities.
Patent Information
- Application Number
- CN202510346551.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Existing non-contact multimodal signal acquisition and evaluation technologies cannot be widely applied in real-world scenarios, mainly because visual sensors are greatly affected by light sources, facial expression recognition accuracy is low, and the quality of millimeter-wave radar heart rate signals cannot meet analysis requirements.
Behavioral cues are acquired using a multispectral camera, non-behavioral cues are acquired using millimeter-wave radar and a sound sensor, and fused data is generated through data fusion. This data is then combined with a machine learning model to assess psychological abilities.
It enables accurate and stable assessment of psychological abilities in real-world scenarios, solves the applicability and accuracy problems of traditional methods, and possesses the practicality of non-contact multimodal assessment technology.
Smart Images

Figure CN119970040B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of psychological ability monitoring technology, and in particular to a non-contact psychological ability assessment method. Background Technology
[0002] In the current field of psychological or emotional monitoring and assessment, evaluation methods that use specific data and multimodal data fusion as data sources have various technical directions. In terms of acquiring multimodal signals, contact-based, wearable, and non-contact methods have emerged, and their application scope is expanding, showing significant value in areas such as public health, health education, human resources, consumer healthcare, and even military training. Consequently, higher standards are being demanded of the objectivity, accuracy, and practicality of psychological capacity monitoring and assessment technologies.
[0003] Currently, non-contact multimodal signal sensors mainly include visual sensors (cameras or webcams), millimeter-wave radar, and recording devices. Firstly, image signal acquisition and analysis techniques based on visual sensors primarily include ippG technology (extracting pulse signal features from time-series images) and facial expression recognition technology (extracting facial expression features through facial recognition). However, ippG technology is significantly limited by the influence of light sources in the application environment and noise from head displacement, hindering its widespread application. Facial expression recognition technology, on the other hand, exhibits varying accuracy probabilities for different expressions, with even greater errors in recognizing expressions during conversation, thus limiting its widespread applicability. Secondly, millimeter-wave radar-based physiological signal acquisition and analysis techniques are limited by the entanglement of heart rate and respiratory fluctuations. While heart rate can be extracted theoretically and in ideal environments, the quality of the extracted heart rate signal fails to meet the requirements for analyzing heart rate variability. Both of these non-contact multimodal signal acquisition and evaluation techniques face bottlenecks preventing their widespread application in real-world scenarios. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a non-contact method for assessing psychological abilities.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A non-contact psychological ability assessment method includes:
[0007] Behavioral cue information is obtained using multispectral cameras and RGB cameras;
[0008] Utilize millimeter-wave radar and sound sensors to acquire non-behavioral cues;
[0009] The behavioral clues and non-behavioral clues are fused together to obtain fused data.
[0010] The psychological ability assessment indicators were determined based on the fused data.
[0011] Preferably, the step of acquiring behavioral cue information using a multispectral camera and an RGB camera includes:
[0012] The multispectral camera is used to acquire facial images and extract key points of the eyes and mouth, wherein the facial images are two near-infrared images;
[0013] The RBG camera is used to acquire panoramic images of the human body and extract the activity features of the torso, hands, and legs;
[0014] The pixel brightness difference is obtained by comparing the differences between the two near-infrared images;
[0015] The sweating characteristics are determined based on the pixel brightness differences;
[0016] The behavioral cue information is determined based on the activity characteristics of the eyes, mouth, torso, hands, and legs, as well as the sweating characteristics.
[0017] Preferably, the calculation expression for the sweating characteristic is:
[0018]
[0019] Where ΔB represents the sweating characteristic, B NIR1 B represents the brightness value of the NIR1 spectrum. NIR2 The value represents the brightness of the NIR2 spectrum, k1 and k2 are the first and second fitting parameters, respectively, α and β are the first and second exponential parameters, respectively, and γ is the third exponential parameter.
[0020] Preferably, the method of acquiring non-behavioral cues using millimeter-wave radar and a sound pickup sensor includes:
[0021] Using millimeter-wave radar to acquire physiological signal data;
[0022] A time-series curve is obtained based on the physiological signal data;
[0023] Based on the aforementioned time-series curve, the fusion signal curve of heart rate and respiration;
[0024] The fused signal curve is subjected to Fourier transform to obtain a neural dynamic spectrum.
[0025] Calculate the frequency band data of the neural dynamic spectrogram to obtain the first multi-band data;
[0026] The first spectral variation characteristics are determined based on the multi-band data;
[0027] The sound rhythm spectrum is obtained by acquiring audio signals using a sound pickup sensor and extracting features.
[0028] The power spectral density map of the sound rhythm spectrum is obtained by performing a Fourier transform on the sound rhythm spectrum.
[0029] The second multi-band data is calculated based on the power spectral density map.
[0030] The first multi-band data and the second multi-band data are fused to obtain fused features;
[0031] The non-behavioral cues are determined based on the fusion features.
[0032] Preferably, the calculation expression for the neural dynamic spectrogram is:
[0033]
[0034] Preferably, the multi-band data includes: first band data, second band data, and third band data, wherein the expressions for the first band data, the second band data, and the third band data are respectively:
[0035]
[0036] Among them, P VLF P LF and P HF These are data from the first frequency band, the second frequency band, and the third frequency band, respectively.
[0037] Preferably, the calculation expression for the second multi-band data is:
[0038]
[0039] Among them, PSD sound (f) is the power spectral density plot. This is the second multi-band data.
[0040] Preferably, the calculation expression for the fusion feature is:
[0041]
[0042] Where F represents the fusion feature, and w1 and w2 are the first and second weight coefficients, respectively. For fourth frequency band data, This is data from the fifth frequency band.
[0043] The present invention discloses the following technical effects:
[0044] This invention provides a non-contact psychological ability assessment method, comprising: acquiring behavioral cue information using a multispectral camera and an RGB camera; acquiring non-behavioral cues using millimeter-wave radar and a sound sensor; fusing the behavioral cue information and the non-behavioral cues to obtain fused data; and determining psychological ability assessment indicators based on the fused data. This invention addresses the issue of broad applicability by using multispectral imaging technology, which is unaffected by changes in scene light sources; (2) employing a heart rate and respiration fusion spectral analysis method, which solves the problem of loss of variation characteristics caused by separating heart rate signals in millimeter-wave radar signal analysis, making it usable in practical work; (3) using more accurate and stable behavioral data, such as blinking, eye area, limb movements, and sweating trends, as the data source for the behavioral cue model, and fusing the heart rate & respiratory rate fusion spectrum (HRV-R) with the sound rhythm spectrum as the data source for the non-behavioral cue model for multimodal fusion and analysis, solving the problem that traditional facial expressions and physiological signals cannot provide in-depth, multi-dimensional evaluation of psychological abilities and phenomena; thus realizing the true practicality and effectiveness of non-contact multimodal assessment technology. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 A flowchart of a non-contact psychological ability assessment method provided in an embodiment of the present invention;
[0047] Figure 2 This is a schematic diagram of a non-contact psychological ability assessment method strategy provided by an embodiment of the present invention;
[0048] Figure 3 The image diagrams of five spectra provided for embodiments of the present invention are shown. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] like Figure 1-2 As shown, the present invention provides a non-contact psychological ability assessment method, comprising:
[0052] Step 100: Use multispectral cameras and RGB cameras to acquire behavioral cue information;
[0053] Step 200: Use millimeter-wave radar and sound sensors to obtain non-behavioral cues;
[0054] Step 300: The behavioral clue information and the non-behavioral clues are fused to obtain fused data;
[0055] Specifically, behavioral cues (such as facial expressions and body language) of interviewees are acquired using multispectral and RGB cameras, while non-behavioral cues (such as heart rate, respiratory rate, and voice features) are collected using millimeter-wave radar and audio sensors. First, the multi-source data is synchronized and preprocessed in time, extracting visual, physiological, and voice features separately. Then, early fusion, late fusion, or hybrid fusion strategies are employed to fuse behavioral and non-behavioral cues at multiple levels, generating high-dimensional feature vectors. Next, machine learning or deep learning models are used to analyze the fused data to infer the interviewees' emotional state, level of anxiety, and other psychological characteristics. Finally, the model output provides support for interview decision-making.
[0056] Step 400: Determine the psychological ability assessment indicators based on the fused data.
[0057] Specifically, by using multispectral cameras and RBG camera components in non-contact devices to obtain features of the eyes, mouth, torso, hands, legs, and sweating, a model containing multiple conversational behavior parameters, namely a behavioral cue model, is constructed.
[0058] The psychological ability assessment indicators include: personality adaptability, emotional stability, willpower, memory ability, language comprehension, personality bias, and stress level.
[0059] Physiological signals and conversation audio signals are obtained through millimeter-wave radar and sound sensor components. A spectrum analysis model, i.e., non-behavioral cues, is constructed by using a joint multi-task learning method, which includes the distribution of conversation neural activity and conversation stress.
[0060] Multimodal data fusion is performed using the two data models of behavioral cues and non-behavioral cues obtained above. Deep learning models, namely Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM), are used to extract and learn the behavioral cues and non-behavioral cues features related to psychological abilities and obtain a psychological ability assessment model.
[0061] By using simulation technology, the target subject in a real conversation is modeled and simulated to enable real-time monitoring, prediction, and evaluation of the target subject's psychological abilities during the conversation.
[0062] Furthermore, the acquisition of behavioral cue information using multispectral cameras and RGB cameras includes:
[0063] The multispectral camera is used to acquire facial images and extract key points of the eyes and mouth, wherein the facial images are two near-infrared images;
[0064] The RBG camera is used to acquire panoramic images of the human body and extract the activity features of the torso, hands, and legs;
[0065] The pixel brightness difference is obtained by comparing the differences between the two near-infrared images;
[0066] The sweating characteristics are determined based on the pixel brightness differences;
[0067] The behavioral cue information is determined based on the activity characteristics of the eyes, mouth, torso, hands, and legs, as well as the sweating characteristics.
[0068] Specifically, using a near-focus and far-focus synchronous image acquisition structure and device, human images are acquired simultaneously, at the same frame rate, with the same pixel count, and in the same space. From this, the extracted features and trends of eye, mouth, torso, hand, and leg movements maintain a sequence of images completely consistent with the dynamic behavior of humans in a real conversation scene; specifically including:
[0069] By equipping a multispectral camera with a close-focus lens, clear facial images can be obtained, and key points and variation features of the eyes and mouth can be extracted; by equipping an RGB camera with a telephoto lens, panoramic images of the human body can be obtained, thereby extracting the activity features and trends of the torso, hands, and legs.
[0070] By connecting a multispectral camera and an RBG camera through a pulse generator, a sequence of images with the same sequence, frame rate, and pixel count from both cameras can be obtained simultaneously.
[0071] Through structured design, the multispectral camera and the RBG camera are positioned and combined to obtain a sequence of images of the two cameras in the same space.
[0072] A 3D digital human is created using 3D modeling technology, resulting in real-time behavioral depictions, abnormal behavior indicators, and behavioral change trend data models that are consistent with the real individual behavior in real conversation scenarios.
[0073] More specifically, by acquiring two identical near-infrared facial images of the same size and pixel count using a multispectral camera, and comparing the differences between the two near-infrared images, the differences in pixel brightness caused by minute amounts of sweat in the human body can be obtained, thus revealing the characteristics and trends of human sweating, including:
[0074] like Figure 3 As shown, five spectral images—visible light images R, G, and B, and two near-infrared (NIR) images (spectral ranges 700-800 and 800-900)—are obtained using a 5-channel multispectral camera and a 3MOS dedicated lens. Based on the characteristics of the dual-band infrared images and the brightness changes caused by human sweating, a model containing multiple parameters is constructed. This model combines Planck's radiation law, the basic principles of infrared radiation, and image processing techniques to evaluate the characteristics and trends of human sweating during conversation.
[0075] Specifically, the calculation expression for the sweating characteristic is as follows:
[0076]
[0077] Where ΔB represents the sweating characteristic, B NIR1 B represents the brightness value of the NIR1 spectrum (700-800nm). NIR2 The value represents the brightness of the NIR2 spectrum (800-900nm). k1 and k2 are the first and second fitting parameters, respectively, used to adjust the weight of the brightness difference between the two bands. α and β are the first and second exponential parameters, respectively, used to describe the nonlinear relationship of brightness values in different bands. γ is the third exponential parameter, used to describe the nonlinear relationship of the brightness difference ratio.
[0078] Furthermore, the acquisition of non-behavioral cues using millimeter-wave radar and audio sensors includes:
[0079] Using millimeter-wave radar to acquire physiological signal data;
[0080] A time-series curve is obtained based on the physiological signal data;
[0081] Specifically, firstly, the millimeter-wave radar emits electromagnetic waves and receives reflected signals to collect minute motion data of the target object; secondly, the original signal is denoised, and Fourier transform or wavelet transform is used to extract frequency band components related to physiological signals, and respiratory signals and heartbeat signals are separated from them; then, the extracted signals are sampled in time sequence to generate discrete data points, and are transformed into continuous time-series curves through smoothing and time-frequency analysis.
[0082] Based on the aforementioned time-series curve, the fusion signal curve of heart rate and respiration;
[0083] The fused signal curve is subjected to Fourier transform to obtain a neural dynamic spectrum.
[0084] Calculate the frequency band data of the neural dynamic spectrogram to obtain the first multi-band data;
[0085] The first spectral variation characteristics are determined based on the multi-band data;
[0086] Specifically, the physiological signals from the millimeter-wave radar spectrum, namely the fused signal curves of heart rate and respiration, are subjected to Fourier transform to obtain a neural dynamic spectrum, which serves as a physiological state characteristic, including:
[0087] By preprocessing and extracting features from the raw physiological signals obtained from millimeter-wave radar, and considering that filtering and separation of heart rate and respiratory signals reduce the variation characteristics of the body's natural rhythms, we plotted the raw signals containing both heart rate and respiration into a time-series curve to obtain a fused signal curve of heart rate and respiration. By performing Fourier transform on the obtained fused signal curve of heart rate and respiration, we obtained the HRV-R spectrum and calculated the data for each frequency band. Thus, we can observe the spectral variation characteristics caused by changes in neural activity and heart rate-respiration rate resonance.
[0088] The sound rhythm spectrum is obtained by acquiring audio signals using a sound pickup sensor and extracting features.
[0089] The power spectral density map of the sound rhythm spectrum is obtained by performing a Fourier transform on the sound rhythm spectrum.
[0090] The second multi-band data is calculated based on the power spectral density map.
[0091] The first multi-band data and the second multi-band data are fused to obtain fused features;
[0092] The non-behavioral cues are determined based on the fusion features.
[0093] Specifically, multimodal spectral fusion involves fusing the heart rate and respiratory rate fused spectrum (HRV-R) with the sound rhythm spectrum and using this as a non-behavioral cue feature. This feature is then input into a machine learning model for analyzing psychological abilities based on physiological states, i.e., non-behavioral cues. This includes:
[0094] Plot the power spectral density of the sound rhythm spectrum using audio signals.
[0095] By combining spectral data from two different sources—heart rate and respiratory rate fused spectra (HRV-R)—a fused feature is obtained. This fused feature F is then input into a machine learning model to predict and analyze changes in psychological capabilities caused by physiological states.
[0096] Specifically, the calculation expression for the neural dynamic spectrogram is as follows:
[0097]
[0098] Where PSD(f) represents the power spectral density at frequency f.
[0099] This represents the normalization factor, used to normalize the power spectral density. N is the length of the signal. ||RR(f)||^ 2 Let RR(f) represent the power spectrum of the fused signal at frequency f. RR(f) is the frequency domain representation obtained through Fourier transform. ||RR(f)|| represents its amplitude, which is squared to obtain the power spectrum.
[0100] Here, RR(f) represents the frequency domain representation of the signal that incorporates heart rate (HR) and respiratory rate (RR).
[0101] Specifically, the multi-band data includes: first band data, second band data, and third band data, wherein the expressions for the first band data, second band data, and third band data are as follows:
[0102]
[0103] Among them, P VLF P LF and P HF These are data from the first frequency band, the second frequency band, and the third frequency band, respectively.
[0104] Specifically,
[0105] Among them, PSD sound (f) is the power spectral density plot. This is the second multi-band data.
[0106] Specifically, the calculation expression for the fused feature is as follows:
[0107]
[0108] Where F represents the fusion feature, and w1 and w2 are the first and second weight coefficients, respectively. For fourth frequency band data, This is data from the fifth frequency band.
[0109] Furthermore, the process for determining the indicators for assessing psychological abilities is as follows:
[0110] By leveraging simulation platforms such as Unity3D, a virtual interactive system architecture was built to simulate behavioral and non-behavioral cues in a virtual space. Specifically, this includes constructing a digital twin model, replicating the target object's behavioral trajectory and non-behavioral cue features during a conversation within the software system, and performing real-time simulation. This enables the digital twin application—an interactive display model—of the entire conversation process, encompassing both behavioral and non-behavioral elements. A digital twin virtual simulation interactive system capable of reflecting individual behavior and physiological states in real time was constructed, achieving seamless integration and interaction between the physical and digital worlds.
[0111] The real-time assessment and calculation module uses digital tags to represent psychological abilities: personality adaptability, emotional stability, willpower, memory, and language comprehension. It generates shortcut keys, which, when pressed during a conversation, provide analytical indicators for each psychological ability. Specifically, this includes:
[0112] Through model construction and validation, assessment indicators for psychological abilities were obtained, including personality adaptability, emotional stability, willpower, memory ability, language comprehension, personality bias, and stress level.
[0113] By establishing a one-to-one correspondence between shortcut commands and psychological ability assessment indicators, the system can immediately calculate the corresponding assessment indicators and display them on the digital twin's visualization interface to obtain psychological ability assessment and interpretation. The corresponding commands correspond to the assessment indicators, and each corresponding command is implemented through different shortcut keys.
[0114] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0115] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A non-contact mental capability evaluation method, characterized by, The method comprises the following steps: acquiring behavior clue information by using a multi-spectrum camera and an RGB camera; acquiring non-behavior clues by using a millimeter wave radar and a sound pickup sensor; performing data fusion on the behavior clue information and the non-behavior clues to obtain fusion data; determining a psychological ability evaluation index according to the fusion data; the step of acquiring behavior clue information by using a multi-spectrum camera and an RGB camera comprises the following steps: acquiring a face image by using the multi-spectrum camera and extracting eye and mouth key points, wherein the face image is two near-infrared images; acquiring a human body panoramic image by using the RGB camera and extracting activity features of a torso, hands and legs; performing difference comparison according to the two near-infrared images to obtain pixel brightness differences; determining sweat features according to the pixel brightness differences; determining the behavior clue information according to the eye and mouth key points, the activity features of the torso, hands and legs and the sweat features; the step of acquiring non-behavior clues by using a millimeter wave radar and a sound pickup sensor comprises the following steps: acquiring physiological signal data by using the millimeter wave radar; obtaining a time series curve according to the physiological signal data; obtaining a heart rate and respiration fusion signal curve according to the time series curve; performing Fourier transformation on the fusion signal curve to obtain a neural dynamic frequency spectrum graph; calculating each frequency band data of the neural dynamic frequency spectrum graph to obtain first multi-frequency band data; determining a first frequency spectrum change feature according to the multi-frequency band data; acquiring an audio signal by using the sound pickup sensor and performing feature extraction to obtain a sound rhythm frequency spectrum; performing Fourier transformation on the sound rhythm frequency spectrum to obtain a power spectral density graph of the sound rhythm frequency spectrum; calculating second multi-frequency band data according to the power spectral density graph; performing fusion on the first multi-frequency band data and the second multi-frequency band data to obtain fusion features; determining the non-behavior clues according to the fusion features.
2. The non-contact mental capability evaluation method according to claim 1, wherein the calculation expression of the sweat features is: ; wherein B is a sweat characteristic, NIR1 B represents a brightness value of the NIR1 spectrum, NIR2 B represents a brightness value of the NIR2 spectrum, k1 and k2 are a first and a second fitting parameter, respectively, and a and b are a first and a second exponential parameter, respectively, and g is a third exponential parameter.
3. The non-contact mental capability assessment method of claim 1, wherein, the calculation expression of the neural dynamic frequency spectrum graph is: ; wherein, is the power spectral density at frequency N is the signal length, is the representation of the signal fused with heart rate (HR) and respiration rate (RR) in the frequency domain.
4. The non-contact mental capability assessment method of claim 1, wherein, the first multi-frequency band data comprises first frequency band data, second frequency band data and third frequency band data, wherein the expressions of the first frequency band data, the second frequency band data and the third frequency band data are respectively: ; wherein, , and are first, second and third frequency band data, respectively.
5. The non-contact mental capability assessment method of claim 4, wherein, the calculation expression of the second multi-frequency band data is: ; ; wherein, is a power spectral density plot, and are both second multi-band data.
6. The non-contact mental capability assessment method of claim 5, wherein, the calculation expression of the fusion features is: ; Wherein, F is a fusion feature, w1 and w2 are respectively a first weight coefficient and a second weight coefficient, is fourth frequency band data, is fifth frequency band data.
Citation Information
Patent Citations
Health assessment method for calculating individual psychological abnormality based on emotion data
CN115312195A
Psychological state sensing method and system and readable storage medium
CN116077062A