A non-contact multi-modal data acquisition and emotion monitoring method, device, system, equipment, medium and product

CN121196548BActive Publication Date: 2026-07-24程乃俊
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
程乃俊
Filing Date
2025-10-16
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing psychological assessment and emotion recognition solutions suffer from problems such as inconvenient device wearing, psychological interference, difficulty in recognizing multimodal data, limited generalization feature extraction capabilities, and low accuracy in intelligent classification and recognition.

Method used

The system employs millimeter-wave bio-radar modules and audio/video acquisition modules to collect human vital signs and audio/video signals non-contactly. It then uses a combination of empirical mode decomposition and multi-signal classification algorithms for signal separation and feature extraction, and combines respiratory rate, heart rate, audio features, and video features for emotional state analysis.

Benefits of technology

It achieves real, fast, and effective non-contact multimodal data collection, provides high-quality data support, establishes high-precision and high-reliability emotional state analysis for artificial intelligence emotion early warning and assessment models, and improves the naturalness of human-computer interaction and the effectiveness of mental health assessment tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121196548B_ABST
    Figure CN121196548B_ABST
Patent Text Reader

Abstract

The application discloses a kind of non-contact multimodal data acquisition and emotion monitoring method, device, system, equipment, medium and product, it is related to emotional recognition technical field.The method is first to start man-machine interaction program and target personnel are interviewed, and receive by millimeter wave biological radar module and audio and video acquisition module in the process of interview the human vital sign signal and audio and video signal collected by target personnel, then on the one hand to human vital sign signal sequentially pre-processing, signal separation processing and frequency estimation processing, obtain respiratory rate and heartbeat frequency, on the other hand to audio and video signal carries out feature extraction processing, obtains audio feature and video feature, finally according to respiratory rate, heartbeat frequency, audio feature and video feature, comprehensive analysis obtains the emotional state of target personnel, thereby can provide high-quality data support for subsequent artificial intelligence emotion early warning and evaluation model establishment, realize higher accuracy and higher reliability of emotional state analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of emotion recognition technology, specifically relating to a non-contact multimodal data acquisition and emotion monitoring method, device, system, equipment, medium, and product. Background Technology

[0002] Military missions are often characterized by sudden events, perilous tasks, harsh environments, and complex nature. Soldiers involved in these missions are prone to psychological problems, which can affect their operational capabilities and the achievement of mission objectives. If these problems are not addressed and managed, they can develop into mental illnesses, impacting the physical and mental health of soldiers and jeopardizing the safety and stability of the military. Furthermore, given the varying degrees of stress experienced by people living in high-altitude areas, it is necessary to conduct early screening and objective assessments, develop effective prevention strategies, and implement and monitor these strategies. Therefore, conducting psychological assessments for groups such as soldiers and people living in high-altitude areas is crucial for developing targeted intervention programs to safeguard their health, work performance, and mission completion capabilities.

[0003] Traditional psychological assessment methods primarily rely on questionnaires. Currently, with technological advancements, emotion recognition methods based on physiological parameters and behavioral patterns have been gradually introduced, and significant progress has been made in objectifying assessment criteria. However, the collection of physiological and behavioral parameters typically requires the wearer to wear additional data collection devices (such as electrodes and related equipment for collecting ECG and EEG signals), which not only presents inconvenience but also, to some extent, can cause additional psychological interference to the test subjects. Furthermore, existing psychological assessment and emotion recognition methods still suffer from limitations such as difficulty in recognizing and covering multimodal data, limited generalization feature extraction capabilities, and low accuracy in intelligent classification and recognition.

[0004] Therefore, how to conduct non-contact multimodal data collection on the population in a real, fast and effective manner and complete emotion recognition and monitoring, so as to provide high-quality data support for the establishment of subsequent artificial intelligence emotion early warning and assessment models and achieve higher accuracy and higher reliability of emotion state analysis, has become an urgent research topic for those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a non-contact multimodal data acquisition and emotion monitoring method, device, system, computer equipment, computer-readable storage medium, and computer program product to solve the problems of existing psychological assessment and emotion recognition schemes, such as inconvenient wearing of acquisition devices, additional psychological interference caused by wearing them, difficulty in recognizing and covering multimodal data, limited generalization feature extraction capabilities, and / or low accuracy of intelligent classification and recognition.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, a non-contact multimodal data acquisition and emotion monitoring method is provided, executed by a computer device that is communicatively connected to a millimeter-wave bio-radar module and an audio / video acquisition module, including: The human-computer interaction program is initiated to conduct an interview with the target personnel, and during the interview, the program receives human vital signs signals collected by the millimeter-wave bio-radar module and audio-visual signals collected by the audio-visual acquisition module. The human vital signs signal is preprocessed to obtain the phase change information of the mixed signal; Based on the phase change information, the mixed signal is separated using the ensemble empirical mode decomposition algorithm to obtain the respiratory signal and the heartbeat signal. The respiratory signal is processed by a multiple signal classification algorithm to estimate the frequency, and the heartbeat signal is also processed by the same algorithm to estimate the frequency, thus obtaining the heartbeat frequency. The audio and video signals are subjected to feature extraction processing to obtain audio features and video features; The emotional state of the target person is obtained through comprehensive analysis based on the breathing rate, heart rate, audio characteristics, and video characteristics.

[0007] Based on the above-mentioned invention, a novel solution is provided that enables real, rapid, and effective non-contact multimodal data collection from a population and the completion of emotion recognition and monitoring. This involves first initiating a human-computer interaction program to interview the target individual. During the interview, the system receives vital signs signals collected by a millimeter-wave bio-radar module and audio / video signals collected by an audio / video acquisition module. Then, the vital signs signals are preprocessed, separated, and frequency estimated sequentially to obtain respiratory and heart rates. Simultaneously, the audio / video signals undergo feature extraction to obtain audio and video features. Finally, based on the respiratory rate, heart rate, audio features, and video features, the target individual's emotional state is comprehensively analyzed. This non-contact acquisition of multiple physiological parameters, including respiration, heart rate, micro-expressions, and voice, provides high-quality data support for the subsequent establishment of AI-based emotion warning and assessment models, achieving higher accuracy and reliability in emotion state analysis, facilitating practical application and promotion.

[0008] In one possible design, the human vital signs signal is preprocessed to obtain phase change information of the mixed signal, including: The human vital signs signal is subjected to DC component removal and distance dimension FFT transformation in sequence to obtain the FFT transformation result; Based on the FFT transform results, the phase is solved using the arctangent method to obtain the phase solution; Based on the phase solution results, the signal is enhanced by coherent accumulation to obtain the signal enhancement result; The signal enhancement result is subjected to phase unwinding and phase difference processing to obtain the phase change information of the mixed signal.

[0009] In one possible design, after obtaining the respiratory and heartbeat signals and before performing frequency estimation processing on the heartbeat signal using the multi-signal classification algorithm, the method further includes: Based on the respiratory signal, the least mean square algorithm is used to adaptively cancel the respiratory harmonics in the heartbeat signal to obtain a new heartbeat signal.

[0010] In one possible design, feature extraction processing is performed on the audio and video signals to obtain video features, including: The audio and video signals are processed by video frame extraction to obtain a first static image sequence containing multiple video frames; Based on the inter-frame pixel differences, redundant frames are removed from the first static image sequence to obtain the second static image sequence. The second static image sequence is subjected to data denoising and / or data enhancement to obtain a third static image sequence. The data denoising includes median filtering, Gaussian filtering and / or time smoothing filtering. The data enhancement includes random rotation of video frames, random flipping of video frames, random adjustment of video frame brightness, random adjustment of video frame contrast, random scaling of video frames and / or random translation of video frames. Face detection and alignment are performed on each video frame in the third static image sequence to obtain the fourth static image sequence. High-level visual features of the fourth static image sequence are extracted using a convolutional neural network, and the extraction results are standardized and / or PCA dimensionality reduction processed to obtain video features.

[0011] Secondly, a non-contact multimodal data acquisition and emotion monitoring device is provided, which is arranged in a computer device that is respectively connected to a millimeter-wave bio-radar module and an audio and video acquisition module. It includes a multi-source signal collection module, a signal preprocessing module, a signal separation processing module, a frequency estimation processing module, a feature extraction processing module, and an emotion state analysis module. The multi-source signal collection module is used to initiate a human-computer interaction program to conduct an interview with the target personnel, and during the interview, it receives human vital signs signals collected by the millimeter-wave bio-radar module towards the target personnel and audio and video signals collected by the audio and video acquisition module towards the target personnel. The signal preprocessing module is communicatively connected to the multi-source signal collection module and is used to preprocess the human vital signs signal to obtain the phase change information of the mixed signal. The signal separation and processing module is communicatively connected to the signal preprocessing module, and is used to perform signal separation processing on the mixed signal based on the phase change information using a set empirical mode decomposition algorithm to obtain respiratory signal and heartbeat signal; The frequency estimation processing module is communicatively connected to the signal separation processing module. It is used to perform frequency estimation processing on the respiratory signal using a multiple signal classification algorithm to obtain the respiratory frequency, and also to perform frequency estimation processing on the heartbeat signal using the multiple signal classification algorithm to obtain the heartbeat frequency. The feature extraction and processing module is communicatively connected to the multi-source signal collection module and is used to perform feature extraction processing on the audio and video signals to obtain audio features and video features; The emotion state analysis module is communicatively connected to the frequency estimation processing module and the feature extraction processing module, respectively, and is used to comprehensively analyze the emotion state of the target person based on the breathing frequency, the heart rate, the audio features and the video features.

[0012] Thirdly, the present invention provides a non-contact multimodal data acquisition and emotion monitoring system, including a millimeter-wave bio-radar module, an audio and video acquisition module, and a computer device; The millimeter-wave bio-radar module is communicatively connected to the computer device and is used to collect human vital signs signals of the target personnel in a non-contact manner, and transmit the collection results to the computer device. The audio and video acquisition module is communicatively connected to the computer device and is used to non-contactly acquire the audio and video signals of the target person and transmit the acquisition results to the computer device. The computer device is used to perform the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect or any possible design in the first aspect.

[0013] In one possible design, the millimeter-wave bio-radar module includes an antenna unit, a radio frequency transceiver unit, an intermediate frequency amplification and filtering unit, a digital-to-analog converter unit, a data processing unit, a waveform modulation unit, and a crystal oscillator unit. The antenna unit, the radio frequency transceiver unit, the intermediate frequency amplification and filtering unit, the digital-to-analog converter unit, and the data processing unit are electrically connected in sequence. The clock signal input terminal of the digital-to-analog converter unit and the reference signal input terminal of the waveform modulation unit are respectively electrically connected to the signal output terminal of the crystal oscillator unit. The waveform modulation unit employs a frequency synthesizer with a variable division ratio to provide an analog baseband signal to the radio frequency transceiver unit and a sampling trigger signal to the digital-to-analog converter unit. The period of the sampling trigger signal is an integer multiple of the period of the analog baseband signal and is consistent with the chirp period of the millimeter-wave bio-radar module.

[0014] Fourthly, the present invention provides a computer device comprising a storage module, a processing module, and a transceiver module connected in sequence for communication, wherein the storage module is used to store a computer program, the transceiver module is used to send and receive messages, and the processing module is used to read the computer program and execute the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect or any possible design in the first aspect.

[0015] Fifthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect or any possible design within the first aspect.

[0016] In a sixth aspect, the present invention provides a computer program product, including a computer program or instructions, wherein the computer program or instructions, when executed by a computer, implement the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect or any possible design in the first aspect.

[0017] The beneficial effects of the above scheme are: (1) This invention creatively provides a new solution for non-contact multimodal data collection of people and completion of emotion recognition and monitoring in a real, fast and effective manner. That is, firstly, a human-computer interaction program is started to conduct an interview with the target person. During the interview, the human vital signs signals collected by the millimeter-wave bio-radar module and the audio and video signals collected by the audio and video acquisition module are received. Then, on the one hand, the human vital signs signals are preprocessed, signal separation and frequency estimation are performed sequentially to obtain the respiratory rate and heart rate. On the other hand, the audio and video signals are feature extracted to obtain audio features and video features. Finally, based on the respiratory rate, heart rate, audio features and video features, the emotional state of the target person is comprehensively analyzed. Thus, by integrating non-contact collection of multiple physiological parameters such as breathing and heart rate, micro-expression and voice, high-quality data support can be provided for the establishment of subsequent artificial intelligence emotion warning and assessment models, and higher accuracy and higher reliability of emotion state analysis can be achieved. (2) Since the multimodal data collection method is non-contact, it can improve the naturalness and intimacy of human-computer interaction, and is of great significance for optimizing mental health assessment tools and enhancing public safety monitoring systems, which is convenient for practical application and promotion. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the non-contact multimodal data acquisition and emotion monitoring method provided in the embodiments of this application.

[0020] Figure 2 This is a schematic diagram of the system architecture of the millimeter-wave bio-radar module provided in the embodiments of this application.

[0021] Figure 3 This is an example diagram of the hardware circuit of a millimeter-wave bio-radar module provided in an embodiment of this application.

[0022] Figure 4 Example diagram of the finished PCB board of the millimeter-wave bio-radar module provided in this application embodiment, wherein, Figure 4 Image (a) shows a front view example of the finished PCB board. Figure 4 (b) shows an example of the back side of the finished PCB board.

[0023] Figure 5This is a schematic diagram of the structure of the non-contact multimodal data acquisition and emotion monitoring device provided in the embodiments of this application.

[0024] Figure 6 This is a schematic diagram of the structure of the non-contact multimodal data acquisition and emotion monitoring system provided in the embodiments of this application.

[0025] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0027] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.

[0028] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0029] Example like Figure 1As shown, the non-contact multimodal data acquisition and emotion monitoring method provided in the first aspect of this embodiment can be executed, but is not limited to, by a computer device with certain computing resources and communicatively connected to a millimeter-wave bio-radar module and an audio-visual acquisition module. The millimeter-wave bio-radar module is used to non-contactly acquire the human vital signs signals of the target person and transmit the acquisition results to the computer device; the audio-visual acquisition module is used to non-contactly acquire the audio-visual signals of the target person and transmit the acquisition results to the computer device. The target person is the individual subject to psychological assessment and emotion recognition, specifically, but not limited to, military personnel or residents of high-altitude areas. Considering that heart rate variability (HRV) and facial expressions are the physiological basis for revealing emotional experiences (e.g., the heart and physiological feedback such as heart rate changes affect an individual's perception and evaluation of emotional stimuli), and also considering that the temporal profile of emotional experiences can be observed through changes in HRV and facial expressions, heart rate and facial expressions play an important role in psychological assessment. They not only help identify and assess an individual's emotional state and regulation ability, but also provide important information about the physiological basis of emotional experiences. Therefore, it is necessary to collect the aforementioned vital signs signals containing heart rate information and the aforementioned audio-visual signals containing facial expressions.

[0030] The non-contact acquisition of human vital signs signals from a target individual refers to the ability to detect various physiological signals from a living organism without any electrodes or sensors contacting the organism, and at a considerable distance. This not only provides a non-invasive, convenient, and widely applicable method for physiological signal detection, but also makes it possible to detect physiological signals in special situations, such as monitoring vital signs of critically ill patients with extensive burns or trauma, infectious disease patients, and mental patients. Compared with existing infrared, ultrasound, and laser technologies, the greatest advantage of non-contact physiological signal detection using the millimeter-wave bio-radar module lies in its strong penetrating power. It can penetrate non-metallic obstacles such as clothing or bedding and enter human tissue through the air-skin interface, thereby monitoring the displacement and velocity information of various tissues caused by heart and lung pulsation, achieving the goal of simultaneously detecting information such as respiration, heartbeat, blood flow, and intestinal peristalsis. Meanwhile, the millimeter-wave bio-radar module uses microwave circuits to implement the sensor circuitry, requiring no special materials or processes. This allows for miniaturization and high integration through CMOS (Complementary Metal Oxide Semiconductor) integrated design. Furthermore, it requires low radiated power to detect human micro-vibration signals and produces no electromagnetic pollution. These advantages make millimeter-wave bio-radar technology potentially applicable in medical diagnosis, health monitoring, and disaster relief. Specifically, the millimeter-wave bio-radar module can be implemented using, but is not limited to, a continuous-wave bio-radar with a zero-IF structure. The front-end RF circuit of this continuous-wave bio-radar typically includes transceiver antennas, matching circuits, oscillators, mixers, RF amplifiers, and analog filters. The working principle of the RF front-end of the continuous-wave bio-radar is as follows: when the human chest cavity is used as the radar target, the displacement of the chest wall proportionally modulates the radar carrier phase. Ideally, through phase demodulation, time-varying phase information proportional to the time-varying displacement of the chest wall can be obtained, allowing the detection of information related to respiration and heartbeat.

[0031] The millimeter-wave bio-radar module is used to acquire physiological data such as respiratory rate and heart rate by collecting human vital signs signals. Through signal amplification, filtering, and level shifting, it ensures that the acquired signals meet the requirements of ADC (Analogue-to-Digital Conversion). Simultaneously, it automatically adjusts the transmitted signal strength and acquisition interval to adapt to different environmental needs, enabling the subsequent effective separation and reconstruction of respiratory and heartbeat signals using the Ensemble Empirical Mode Decomposition (EEMD) algorithm. Preferably, as follows... Figure 2As shown, the millimeter-wave bio-radar module includes, but is not limited to, an antenna unit, a radio frequency (RF) transceiver unit, an intermediate frequency (IF) amplification and filtering unit, a digital-to-analog converter (DAC) unit, a data processing unit, a waveform modulation unit, and a crystal oscillator unit. The antenna unit, RF transceiver unit, IF amplification and filtering unit, DAC unit, and data processing unit are electrically connected sequentially. The waveform modulation unit provides an analog baseband signal to the RF transceiver unit. Due to the combined limitations of area and size within a compact space in radar module design, and because angle measurement capability is not required in this embodiment, the antenna unit can be implemented using a single-transmitter, single-receiver antenna. The specific functions of the RF transceiver unit, IF amplification and filtering unit, DAC unit, and data processing unit can be derived conventionally from existing continuous-wave bio-radar technology and will not be elaborated here. Furthermore, the IF amplification and filtering unit, DAC unit, and data processing unit can be centrally implemented using a single general-purpose microcontroller to maximize the functional potential of the integrated components in the module. In pulse radar theory, coherent accumulation of multiple received pulse signals is a crucial method for improving the signal-to-noise ratio (SNR) of the received signal. This method is also applicable to the millimeter-wave bio-radar module. Therefore, to achieve coherent reception and accumulation of the acquired signal by the millimeter-wave bio-radar module, the precise synchronization of the waveform modulation unit and the digital-to-analog converter unit is of utmost importance. This requirement can be specifically defined as follows: under the entire sawtooth wave chirp cycle of a frame of modulation waveform (the chirp cycle refers to the period in which the signal frequency changes linearly with time in a linear frequency modulated continuous wave radar, usually composed of idle time, start time, transmission, and sampling processes), it should be ensured that the intermediate frequency signals acquired in each sawtooth wave chirp cycle are completely identical in phase relative relationship, and the start and end times of sampling of the intermediate frequency signal by the digital-to-analog converter unit in each sawtooth wave chirp cycle should be completely consistent. Therefore, to meet the aforementioned requirements, the millimeter-wave bio-radar module needs to be specifically designed in the following two aspects: First, the clock signals of the waveform modulation unit and the digital-to-analog converter unit should be from the same source; second, the sampling trigger of the digital-to-analog converter unit should be directly driven by the chirp generation indication signal of the waveform frequency modulation unit. That is, the clock signal input terminal of the digital-to-analog converter unit and the reference signal input terminal of the waveform modulation unit are respectively electrically connected to the signal output terminal of the crystal oscillator unit. The waveform modulation unit is also used to provide a sampling trigger signal for the digital-to-analog converter unit. The period of the sampling trigger signal is an integer multiple of the period of the analog baseband signal and is consistent with the chirp period of the millimeter-wave bio-radar module.

[0032] In such Figure 2In the system architecture shown, the waveform modulation unit specifically employs a frequency synthesizer with a variable division ratio to provide the analog baseband signal and the sampling trigger signal, respectively. The input reference signal of the waveform modulation unit is provided by the crystal oscillator unit. When each chirp occurrence indication signal is used as the sampling trigger signal for the digital-to-analog converter (DAC), since the sampling action of the DAC occurs at the changing edge of the sampling clock, the time difference between this clock changing edge and the sampling trigger edge must be considered. This results in the sampling start time of the DAC in each chirp cycle still not being completely determined (it also depends on the phase relationship between the sampling clock used by the DAC and the reference signal of the waveform modulation unit). When the sampling clock of the digital-to-analog converter (DAC) and the reference clock of the waveform modulation unit are not from the same source, their phase relationship is uncertain because they are independent of each other. The time difference between the clock transition edge and the sampling trigger edge will therefore change cyclically and drift randomly over time. This will cause the sampling start time of the DAC to be inconsistent in each chirp cycle, destroying the phase coherence of the intermediate frequency signal. However, when the sampling clock of the DAC and the reference clock of the waveform modulation unit are from the same source, the time difference between the clock transition edge and the sampling trigger edge will be completely fixed. It must be an integer multiple of the period of the reference source signal. Thus, the sampling time of the DAC in each chirp cycle is at a completely consistent relative time, and the phase coherence of the acquired intermediate frequency signal will be fully guaranteed.

[0033] The millimeter-wave bio-radar module can be specifically adopted as follows: Figure 3 The hardware circuit shown, in detail, includes the following: the RF transceiver unit can use a 24GHz radar sensor chip, model BGT24LTR11 (this chip integrates a single-transmit, single-receive RF channel and a voltage-controlled oscillator, with approximately 6dBm transmit power and approximately 10dB single-sideband noise); the waveform modulation unit can use a frequency synthesizer chip, model ADF4158 (this chip operates on the principle of a phase-locked loop circuit with a programmable division ratio, integrating different forms of frequency modulation waveform generation functions such as sawtooth waves, triangular waves, and frequency shift keying); the general-purpose microcontroller can be a general-purpose embedded microcontroller, model STM32F303 (this chip integrates a digital-to-analog converter, operational amplifier, and fast digital signal processing algorithm core); and the power supply section can use a linear regulator chip, such as the LP5907, a low-noise linear regulator chip with high PSRR (Power Supply Rejection Ratio). Figure 3As shown, the key design features of the millimeter-wave bio-radar module are highlighted with bolded lines: The output signal of the voltage-controlled oscillator inside the BGT24LTR11 chip is connected to the RF input terminal RFin of the phase-locked loop circuit inside the ADF4158 chip after being divided by a 16:1 frequency divider. The charge pump output terminal CP of the ADF4158 chip is connected to the tuning voltage terminal V_TUNE of the voltage-controlled oscillator via a loop filter, forming a complete feedback control to drive the RF transceiver chip BGT24LTR11 to generate a precise linear frequency-modulated continuous wave signal. The hardware structure of the millimeter-wave bio-radar module features a three-stage cascaded multi-negative feedback active high-pass filter as the intermediate frequency amplification and filtering unit (its main function is to amplify the amplitude of the intermediate frequency signal to meet the sampling requirements of the subsequent analog-to-digital conversion unit, while filtering out low-frequency interference introduced by sawtooth wave frequency modulation). This active high-pass filter is implemented based on the integrated operational amplifier built into the STM32F303 microcontroller and external resistor-capacitor components. The elimination of a separate operational amplifier chip effectively reduces hardware area and cost. To meet the phase coherent acquisition requirements, connection W1 connects the Chirp generation indicator signal of the ADF4158 chip to the external interrupt trigger terminal of the STM32F303 chip. The system clock of the STM32F303 chip and the reference frequency input terminal REFin of the ADF4158 chip are from the same source as the X1 crystal oscillator (i.e., the crystal oscillator unit) through connection W2. The sampling start and end of the analog-to-digital converter unit is controlled by the microcontroller. Connection W1 enables the microcontroller to accurately start the sampling of the analog-to-digital converter unit when each frequency-modulated Chirp cycle occurs. The sampling clock of the analog-to-digital converter unit is obtained by dividing the system clock of the microcontroller. Connection W2 ensures that the sampling clock is in phase with the reference clock generated by the frequency-modulated Chirp cycle. Therefore, a preparation process exists between the microcontroller's startup trigger and the actual start of sampling by the analog-to-digital conversion unit. This preparation process introduces an "on-sampling" delay (the duration of which is determined by the product of the total number of program instructions and the instruction cycle). Furthermore, the microcontroller instruction cycle depends on the system clock cycle; the "on-sampling" delay is only completely consistent under different frequency chirp cycles when the system clock and the waveform modulation unit's reference clock are from the same source. In addition, the millimeter-wave bio-radar module can be specifically implemented using PCB (Printed Circuit Board) level integration. All hardware circuitry, including the antenna, is integrated onto a 3-cm diameter circular Rogers 4350B PCB board, resulting in a highly compact structural form factor. Figure 4 As shown; and for ease of testing, a corresponding shell structure can also be made for the millimeter-wave bio-radar module.

[0034] The audio and video acquisition module can specifically utilize a Logitech C1000e camera equipped with high resolution and HDR (High Dynamic Range Imaging) technology to achieve non-contact acquisition of audio and video signals from the target person. Specifically, on one hand, the camera features a large Sony CMOS sensor combined with advanced digital overlay HDR technology. These features ensure that even in low-light conditions, the camera can still capture high-definition facial images. Furthermore, the camera supports multiple field-of-view angles and high-definition digital zoom, along with autofocus and automatic light correction functions, thus flexibly adapting to changes in lighting and environment, providing ideal conditions for accurately capturing facial expressions. On the other hand, the camera has a built-in high-quality stereo microphone, enabling clear and natural acquisition of the user's voice information, ensuring the multi-dimensional integrity of physiological information acquisition, and thus guaranteeing high-fidelity acquisition of natural speech. The aforementioned camera needs to be securely mounted on a fixed bracket to ensure the stability and consistency of the image frame during data acquisition. This consistency not only effectively reduces image shifts caused by changes in camera position but also provides a stable visual foundation for accurate facial feature analysis. A stable compositional environment allows subsequent steps to focus on analyzing facial micro-expressions and other relevant features without being affected by changes in camera angle or position, thereby improving the accuracy and reliability of data analysis. Furthermore, the audio / video acquisition module can preferably be used in conjunction with a lifting rod structure to support flexible height adjustment, adapting to different heights and usage scenarios, thus improving operational convenience and data acquisition accuracy.

[0035] In addition, the acquisition software used in conjunction with the millimeter-wave bio-radar module and the audio-visual acquisition module can be designed to support functions such as information input, module settings (which enable users to easily configure and manage signal acquisition and monitor the physiological data acquisition process in real time), pause and playback during data acquisition (which enable users to adjust or review data during acquisition to ensure data accuracy and integrity), and tag recording (which helps users mark important moments for subsequent data analysis and retrospection). This will greatly improve the efficiency and management capabilities of data acquisition, thereby achieving precise control of the acquisition process and efficient storage of acquired data, providing strong technical support for the management and research of multiple physiological information.

[0036] The computer equipment mentioned may specifically include, but is not limited to, servers, personal computers (PCs, which are multi-purpose computers of a size, price, and performance suitable for personal use; desktops, laptops, mini-laptops, tablets, and ultrabooks are all considered personal computers), smartphones, personal digital assistants (PDAs), or wearable devices. Figure 1 As shown, the non-contact multimodal data acquisition and emotion monitoring method includes, but is not limited to, the following steps S1 to S6.

[0037] S1. Initiate the human-computer interaction program to conduct an interview with the target personnel, and during the interview, receive human vital signs signals collected by the millimeter-wave bio-radar module towards the target personnel and audio and video signals collected by the audio and video acquisition module towards the target personnel.

[0038] In step S1, the interview process can be, but is not limited to, using a question-and-answer format based on a questionnaire or pre-set questions. Therefore, the human-computer interaction program can be conventionally written by combining question-and-answer methods with virtual digital human technology. Additionally, before the interview, it is necessary to ensure that the positional relationship between the data acquisition device and the target person meets the acquisition requirements. For example, the millimeter-wave bio-radar module and the audio / video acquisition module should be facing the target person and approximately 0.5 meters away, and the target person's head and face should be accurately centered in the video frame.

[0039] S2. The human vital signs signal is preprocessed to obtain the phase change information of the mixed signal.

[0040] In step S2, since the human vital signs signal is a mid-frequency mixed signal containing signals such as respiration, heartbeat, blood flow, intestinal peristalsis, and other signals, and obtained by ADC sampling, it is necessary to preprocess the human vital signs signal to obtain the phase change information of the mixed signal in order to achieve effective separation and reconstruction of the respiration and heartbeat signals. Specifically, preprocessing the human vital signs signal to obtain the phase change information of the mixed signal includes, but is not limited to, the following steps S21 to S24.

[0041] S21. The human vital signs signal is sequentially processed by DC component removal and distance dimension FFT transformation to obtain the FFT transformation result.

[0042] In step S21, the purpose of the DC component removal process is to eliminate the interference of DC offset phenomenon on the signal caused by hardware heating and other reasons. Specifically, it can be achieved, but is not limited to, through conventional digital filtering methods. The range dimension FFT (Fast Fourier Transform) transformation process is a key step in radar signal processing for extracting target range information. Its core is to determine the distance between the target and the radar by analyzing the spectral characteristics of each pulse signal. Therefore, it can also be conventionally implemented based on existing technologies.

[0043] S22. Based on the FFT transform result, the phase is solved by arctangent method to obtain the phase solution result.

[0044] S23. Based on the phase solution results, the signal is enhanced by coherent accumulation to obtain the signal enhancement result.

[0045] In step S23, the coherent accumulation method is a key technology in radar signal processing, which refers to the use of the phase information of the signal to superimpose multiple echo signals in phase, thereby achieving a significant improvement in the signal-to-noise ratio.

[0046] S24. Perform phase unwinding and phase difference processing on the signal enhancement result to obtain the phase change information of the mixed signal.

[0047] In step S24, the specific process of phase unwinding and phase difference processing is as follows: first, phase unwinding is performed, and then a difference operation is performed on the new signal to obtain the phase difference signal. Because the phase being solved has a phase entanglement problem, phase unwinding and phase difference steps are required in phase extraction to obtain accurate phase change information.

[0048] S3. Based on the phase change information, the mixed signal is separated using the ensemble empirical mode decomposition algorithm to obtain the respiratory signal and the heartbeat signal.

[0049] In step S3, the Ensemble Empirical Mode Decomposition (EEMD) algorithm is a method for decomposing a signal into multiple intrinsic mode functions (each intrinsic mode function represents an intrinsic oscillation mode of the signal and satisfies two conditions: extreme points and zero-crossing points alternate in number or differ by at most one; at any point, the mean of the envelope defined by the local maximum and local minimum must be zero). Since the respiratory signal and the heartbeat signal have different frequency ranges (i.e., the respiratory signal's frequency range is 0.1–0.6 Hz, while the heartbeat signal's frequency range is 0.8–2.0 Hz), the EEMD algorithm can be used to separate these two signals.

[0050] S4. The respiratory signal is processed by a multi-signal classification algorithm to estimate the frequency and obtain the respiratory rate. The heartbeat signal is also processed by the multi-signal classification algorithm to estimate the frequency and obtain the heartbeat rate.

[0051] In step S4, the Multiple Signal Classification (MUSIC) algorithm is a type of spatial spectrum estimation algorithm. Its idea is to use the covariance matrix of the received data for eigenvalue decomposition to separate the signal subspace and noise subspace. It then uses the orthogonality between the signal direction vector and the noise subspace to construct a spatial scanning spectrum, performing a global search for spectral peaks to achieve signal frequency estimation. Therefore, the respiratory frequency and the heartbeat frequency can be obtained separately. Furthermore, considering that respiratory harmonics can interfere with the heartbeat signal, to eliminate this interference, preferably, after obtaining the respiratory signal and heartbeat signal and before using the MUSIC algorithm to perform frequency estimation on the heartbeat signal, the method also includes, but is not limited to: using the least mean square algorithm to adaptively cancel the respiratory harmonics in the heartbeat signal based on the respiratory signal, obtaining a new heartbeat signal.

[0052] S5. Perform feature extraction processing on the audio and video signals to obtain audio features and video features.

[0053] In step S5, the audio features are used to reflect the voice characteristics of the target person during the interview and can be extracted conventionally. The video features are used to reflect the facial micro-expression characteristics of the target person during the interview; in order to preserve key facial information, preferably, the audio and video signals are subjected to feature extraction processing to obtain video features, including but not limited to the following steps S51 to S55.

[0054] S51. Perform video frame extraction processing on the audio and video signals to obtain a first static image sequence containing multiple video frames.

[0055] In step S51, the video frame extraction process aims to convert dynamic video into a sequence of static images for subsequent image analysis and processing. Specifically, the OpenCV video processing library can be used, and depending on the specific task requirements, a frame rate of 3 frames per second is selected for frame extraction. (This frame rate setting balances detail capture and data volume control: on the one hand, a frame rate of 3 frames per second can effectively capture important details and dynamic changes in the video, ensuring data integrity; on the other hand, appropriately reducing the frame rate can avoid generating too much redundant data, reduce storage space occupation, and improve the efficiency of subsequent data processing.) This frame extraction method provides an efficient and accurate foundation for further image analysis and algorithm applications.

[0056] S52. Based on the inter-frame pixel differences, the first static image sequence is subjected to redundant frame removal processing to obtain the second static image sequence.

[0057] In step S52, it is considered that some frames in the first static image sequence may have redundancy, such as almost no change between adjacent frames. This not only wastes storage space but also increases the computational burden. In order to effectively reduce redundancy and reduce computational and storage requirements, this step adopts a method based on inter-frame differences to optimize the data processing flow. Specifically, the pixel difference between adjacent frames is calculated, and a preset threshold is used to determine whether the adjacent subsequent frame is redundant. If the pixel difference between adjacent frames is less than this threshold, the subsequent frame is considered a redundant frame and can be safely excluded. In addition, considering that the resolution of the captured video is generally 1920×1080 (i.e., 2,073,600 pixels), the initial threshold can be set to a small percentage of the total number of pixels, namely 2% (this threshold setting ensures that sufficient video details are retained while effectively reducing redundant frames, thereby optimizing the efficiency of data processing and storage requirements).

[0058] S53. Perform data denoising and / or data enhancement processing on the second static image sequence to obtain a third static image sequence. The data denoising processing includes, but is not limited to, median filtering, Gaussian filtering and / or time smoothing filtering. The data enhancement processing includes, but is not limited to, random rotation of video frames, random flipping of video frames, random adjustment of video frame brightness, random adjustment of video frame contrast, random scaling of video frames and / or random translation of video frames.

[0059] In step S53, the data denoising process and the data augmentation process are key steps in improving the generalization ability of the emotion recognition model. Considering that noise problems are often encountered during video data processing, such as background noise and lighting changes, these noises may affect the accuracy of subsequent analysis; therefore, effective denoising is necessary. Specifically, methods such as median filtering and Gaussian filtering can be used to remove spatial noise. Median filtering is used to effectively remove isolated noise points, while Gaussian filtering is used to smooth the image and reduce the impact of continuous noise. Simultaneously, to address transient noise, temporal smoothing filtering can also be applied to further reduce instantaneous noise in the video frame sequence. In addition, regarding data augmentation, to improve the robustness and generalization ability of the model, a series of augmentation operations can be performed on the video frames: the random rotation and random flipping of the video frames can help the model adapt to different viewpoints and orientations, thereby improving its adaptability to image transformations; the random brightness and random contrast adjustment of the video frames are used to simulate different shooting conditions, enhancing the model's performance under various lighting conditions; the random scaling and random translation of the video frames are used to further improve the model's adaptability to changes in scale and position.

[0060] S54. Perform face detection and alignment processing on each video frame in the third static image sequence to obtain the fourth static image sequence.

[0061] In step S54, the face detection and alignment process specifically includes the following steps: First, using mature computer vision tools such as OpenCV and Dlib, the positions of faces in the image are quickly identified; then, by identifying and locating key points of the face (such as eyes, nose, and corners of the mouth), geometric transformations are performed (the goal of which is to correct for differences in face rotation, scaling, and position, so that all detected faces are aligned at the same angle and position), in order to effectively reduce feature differences caused by head movement or posture changes, thereby improving the model's accuracy in analyzing facial expressions and micro-expressions; finally, the faces are standardized so that key facial regions (such as eyes and mouth) can be presented more consistently, laying a solid foundation for subsequent feature extraction and sentiment analysis. Therefore, based on the aforementioned face detection and alignment process, not only can data consistency be improved, but the robustness of the analysis can also be enhanced, especially when dealing with diverse expressions and subtle changes in expression, ensuring that the model can accurately capture effective emotional signals.

[0062] S55. Use a convolutional neural network to extract high-level visual features from the fourth static image sequence, and perform standardization and / or PCA dimensionality reduction on the extraction results to obtain video features.

[0063] In step S55, the convolutional neural network can, but is not limited to, employ the VGG16 network architecture, which is capable of image classification and recognition tasks. (Its architecture follows a principle of simplicity, using a fixed 3×3 small convolutional kernel in each convolutional layer, and using max pooling layers between all convolutional layers to reduce the size of the feature map, thereby gradually extracting high-level features of the image.) Its 13 convolutional layers can extract features from low to high levels, such as edges, textures, shapes, and semantic information, enabling effective representation learning for complex image data. Therefore, it is suitable for step S55. Specifically, the VGG16 network architecture can be pre-trained on the ImageNet image set, and then the pre-trained weights can be used for transfer learning to apply it to various image classification or feature extraction tasks. Furthermore, the normalization process further ensures the efficiency of the emotion recognition model when training and recognizing using feature data, while the PCA dimensionality reduction process reduces the complexity of the feature space while retaining key information; both can be conventionally implemented using existing techniques.

[0064] S6. Based on the breathing rate, heart rate, audio features, and video features, the emotional state of the target person is obtained through comprehensive analysis.

[0065] In step S6, the specific comprehensive analysis method may include, but is not limited to, importing the breathing frequency, heart rate, audio features, and video features into an emotion recognition model pre-trained based on a machine learning algorithm, so as to output the identified emotional state of the target person. The machine learning algorithm may include, but is not limited to, random forest algorithms, etc., and the emotion recognition model may be conventionally trained based on a certain amount of sample data (its model inputs are breathing frequency, heart rate, audio features, and video features, and its model outputs are these features and pre-labeled emotion state labels).

[0066] Therefore, based on the non-contact multimodal data acquisition and emotion monitoring method described in steps S1 to S6 above, a new scheme is provided that can realistically, quickly, and effectively collect non-contact multimodal data from a population and complete emotion recognition and monitoring. This involves first initiating a human-computer interaction program to interview the target individual. During the interview, the system receives vital signs signals collected by a millimeter-wave bio-radar module and audio / video signals collected by an audio / video acquisition module. Then, the vital signs signals are preprocessed, separated, and frequency estimated sequentially to obtain respiratory and heart rates. Simultaneously, the audio / video signals undergo feature extraction to obtain audio and video features. Finally, based on the respiratory rate, heart rate, audio features, and video features, the target individual's emotional state is comprehensively analyzed. This non-contact acquisition of multiple physiological parameters, including respiration, heart rate, micro-expressions, and voice, provides high-quality data support for the subsequent establishment of AI-based emotion warning and assessment models, achieving higher accuracy and reliability in emotion state analysis, facilitating practical application and promotion.

[0067] like Figure 5 As shown, the second aspect of this embodiment provides a virtual device for implementing the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect. The device is arranged in a computer device that is communicatively connected to a millimeter-wave bio-radar module and an audio-visual acquisition module. It includes a multi-source signal collection module, a signal preprocessing module, a signal separation processing module, a frequency estimation processing module, a feature extraction processing module, and an emotion state analysis module. The multi-source signal collection module is used to initiate a human-computer interaction program to conduct an interview with the target personnel, and during the interview, it receives human vital signs signals collected by the millimeter-wave bio-radar module towards the target personnel and audio and video signals collected by the audio and video acquisition module towards the target personnel. The signal preprocessing module is communicatively connected to the multi-source signal collection module and is used to preprocess the human vital signs signal to obtain the phase change information of the mixed signal. The signal separation and processing module is communicatively connected to the signal preprocessing module, and is used to perform signal separation processing on the mixed signal based on the phase change information using a set empirical mode decomposition algorithm to obtain respiratory signal and heartbeat signal; The frequency estimation processing module is communicatively connected to the signal separation processing module. It is used to perform frequency estimation processing on the respiratory signal using a multiple signal classification algorithm to obtain the respiratory frequency, and also to perform frequency estimation processing on the heartbeat signal using the multiple signal classification algorithm to obtain the heartbeat frequency. The feature extraction and processing module is communicatively connected to the multi-source signal collection module and is used to perform feature extraction processing on the audio and video signals to obtain audio features and video features; The emotion state analysis module is communicatively connected to the frequency estimation processing module and the feature extraction processing module, respectively, and is used to comprehensively analyze the emotion state of the target person based on the breathing frequency, the heart rate, the audio features and the video features.

[0068] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect, and will not be repeated here.

[0069] like Figure 6 As shown, the third aspect of this embodiment provides a physical system for implementing the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect, including a millimeter-wave bio-radar module, an audio and video acquisition module, and a computer device; The millimeter-wave bio-radar module is communicatively connected to the computer device and is used to collect human vital signs signals of the target personnel in a non-contact manner, and transmit the collection results to the computer device. The audio and video acquisition module is communicatively connected to the computer device and is used to non-contactly acquire the audio and video signals of the target person and transmit the acquisition results to the computer device. The computer device is used to perform the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect.

[0070] In one possible design, the millimeter-wave bio-radar module includes an antenna unit, a radio frequency transceiver unit, an intermediate frequency amplification and filtering unit, a digital-to-analog converter unit, a data processing unit, a waveform modulation unit, and a crystal oscillator unit. The antenna unit, the radio frequency transceiver unit, the intermediate frequency amplification and filtering unit, the digital-to-analog converter unit, and the data processing unit are electrically connected in sequence. The clock signal input terminal of the digital-to-analog converter unit and the reference signal input terminal of the waveform modulation unit are respectively electrically connected to the signal output terminal of the crystal oscillator unit. The waveform modulation unit employs a frequency synthesizer with a variable division ratio to provide an analog baseband signal to the radio frequency transceiver unit and a sampling trigger signal to the digital-to-analog converter unit. The period of the sampling trigger signal is an integer multiple of the period of the analog baseband signal and is consistent with the chirp period of the millimeter-wave bio-radar module.

[0071] The working process, working details and technical effects of the aforementioned system provided in the third aspect of this embodiment can be found in the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect, and will not be repeated here.

[0072] like Figure 7 As shown, the fourth aspect of this embodiment provides a computer device for executing the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect. The device includes a storage module, a processing module, and a transceiver module connected in sequence. The storage module stores a computer program, the transceiver module sends and receives messages, and the processing module reads the computer program and executes the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect. Specifically, the storage module may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the processing module may, but is not limited to, use a microprocessor of the STM32F105 series. Furthermore, the computer device may also include, but is not limited to, a power supply module, a display screen, and other necessary components.

[0073] The working process, working details and technical effects of the aforementioned computer device provided in the fourth aspect of this embodiment can be found in the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect, and will not be repeated here.

[0074] This fifth aspect of the embodiment provides a computer-readable storage medium storing instructions comprising the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0075] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fifth aspect of this embodiment can be found in the non-contact multimodal data acquisition and emotion monitoring method described in the first aspect, and will not be repeated here.

[0076] The sixth aspect of this embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implement the non-contact multimodal data acquisition and emotion monitoring method as described in the first aspect. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0077] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A non-contact multimodal data acquisition and emotion monitoring method, characterized in that, This is performed by a computer device that is communicatively connected to both the millimeter-wave bio-radar module and the audio / video acquisition module, including: The human-computer interaction program is initiated to conduct an interview with the target personnel, and during the interview, the program receives human vital signs signals collected by the millimeter-wave bio-radar module and audio-visual signals collected by the audio-visual acquisition module. The human vital signs signal is preprocessed to obtain the phase change information of the mixed signal; Based on the phase change information, the mixed signal is separated using the ensemble empirical mode decomposition algorithm to obtain the respiratory signal and the heartbeat signal. Based on the respiratory signal, the respiratory harmonics in the heartbeat signal are adaptively canceled using the least mean square algorithm to obtain a new heartbeat signal. A multiple signal classification algorithm is used to perform frequency estimation processing on the respiratory signal to obtain the respiratory rate, and to perform frequency estimation processing on the new heartbeat signal to obtain the heartbeat rate. The audio and video signals are subjected to feature extraction processing to obtain audio features and video features. Specifically, this includes: performing video frame extraction processing on the audio and video signals to obtain a first static image sequence containing multiple video frames; performing redundant frame removal processing on the first static image sequence based on inter-frame pixel differences to obtain a second static image sequence; performing data denoising processing and data augmentation processing on the second static image sequence to obtain a third static image sequence, wherein the data denoising processing includes median filtering, Gaussian filtering, or time smoothing filtering, and the data augmentation processing includes random rotation of video frames, random flipping of video frames, random adjustment of video frame brightness, random adjustment of video frame contrast, random scaling of video frames, or random translation of video frames; performing face detection and alignment processing on each video frame in the third static image sequence to obtain a fourth static image sequence; using a convolutional neural network to extract high-level visual features from the fourth static image sequence, and performing standardization processing or PCA dimensionality reduction processing on the extraction results to obtain video features. The emotional state of the target person is obtained through comprehensive analysis based on the breathing rate, heart rate, audio characteristics, and video characteristics.

2. The non-contact multimodal data acquisition and emotion monitoring method according to claim 1, characterized in that, The human vital signs signals are preprocessed to obtain phase change information of the mixed signal, including: The human vital signs signal is subjected to DC component removal and distance dimension FFT transformation in sequence to obtain the FFT transformation result; Based on the FFT transform results, the phase is solved using the arctangent method to obtain the phase solution; Based on the phase solution results, the signal is enhanced by coherent accumulation to obtain the signal enhancement result; The signal enhancement result is subjected to phase unwinding and phase difference processing to obtain the phase change information of the mixed signal.

3. A non-contact multimodal data acquisition and emotion monitoring device, characterized in that, The computer device, which is connected to the millimeter-wave bio-radar module and the audio and video acquisition module respectively, includes a multi-source signal collection module, a signal preprocessing module, a signal separation processing module, a frequency estimation processing module, a feature extraction processing module, and an emotion state analysis module. The multi-source signal collection module is used to initiate a human-computer interaction program to conduct an interview with the target personnel, and during the interview, it receives human vital signs signals collected by the millimeter-wave bio-radar module towards the target personnel and audio and video signals collected by the audio and video acquisition module towards the target personnel. The signal preprocessing module is communicatively connected to the multi-source signal collection module and is used to preprocess the human vital signs signal to obtain the phase change information of the mixed signal. The signal separation and processing module is communicatively connected to the signal preprocessing module. It is used to perform signal separation processing on the mixed signal according to the phase change information using the ensemble empirical mode decomposition algorithm to obtain a respiratory signal and a heartbeat signal. Based on the respiratory signal, it uses the least mean square algorithm to perform adaptive noise cancellation processing on the respiratory harmonics in the heartbeat signal to obtain a new heartbeat signal. The frequency estimation processing module is communicatively connected to the signal separation processing module, and is used to perform frequency estimation processing on the respiratory signal to obtain the respiratory frequency, and to perform frequency estimation processing on the new heartbeat signal to obtain the heartbeat frequency, using a multiple signal classification algorithm. The feature extraction and processing module, communicatively connected to the multi-source signal collection module, is used to perform feature extraction processing on the audio and video signals to obtain audio features and video features. Specifically, it includes: performing video frame extraction processing on the audio and video signals to obtain a first static image sequence containing multiple video frames; performing redundant frame removal processing on the first static image sequence based on inter-frame pixel differences to obtain a second static image sequence; performing data denoising and data enhancement processing on the second static image sequence to obtain a third static image sequence, wherein the data denoising processing includes median filtering, Gaussian filtering, or time smoothing filtering; and the data enhancement processing includes random rotation of video frames, random flipping of video frames, random adjustment of video frame brightness, random adjustment of video frame contrast, random scaling of video frames, or random translation of video frames; performing face detection and alignment processing on each video frame in the third static image sequence to obtain a fourth static image sequence; and using a convolutional neural network to extract high-level visual features from the fourth static image sequence, and performing standardization or PCA dimensionality reduction processing on the extraction results to obtain video features. The emotion state analysis module is communicatively connected to the frequency estimation processing module and the feature extraction processing module, respectively, and is used to comprehensively analyze the emotion state of the target person based on the breathing frequency, the heart rate, the audio features and the video features.

4. A non-contact multimodal data acquisition and emotion monitoring system, characterized in that, It includes millimeter-wave bio-radar modules, audio and video acquisition modules, and computer equipment; The millimeter-wave bio-radar module is communicatively connected to the computer device and is used to collect human vital signs signals of the target personnel in a non-contact manner, and transmit the collection results to the computer device. The audio and video acquisition module is communicatively connected to the computer device and is used to non-contactly acquire the audio and video signals of the target person and transmit the acquisition results to the computer device. The computer device is used to execute the non-contact multimodal data acquisition and emotion monitoring method as described in any one of claims 1 to 2.

5. The non-contact multimodal data acquisition and emotion monitoring system according to claim 4, characterized in that, The millimeter-wave bio-radar module includes an antenna unit, a radio frequency transceiver unit, an intermediate frequency amplification and filtering unit, a digital-to-analog converter unit, a data processing unit, a waveform modulation unit, and a crystal oscillator unit. The antenna unit, the radio frequency transceiver unit, the intermediate frequency amplification and filtering unit, the digital-to-analog converter unit, and the data processing unit are electrically connected in sequence. The clock signal input terminal of the digital-to-analog converter unit and the reference signal input terminal of the waveform modulation unit are respectively electrically connected to the signal output terminal of the crystal oscillator unit. The waveform modulation unit employs a frequency synthesizer with a variable division ratio to provide an analog baseband signal to the radio frequency transceiver unit and a sampling trigger signal to the digital-to-analog converter unit. The period of the sampling trigger signal is an integer multiple of the period of the analog baseband signal and is consistent with the chirp period of the millimeter-wave bio-radar module.

6. A computer device, characterized in that, The device includes a storage module, a processing module, and a transceiver module that are sequentially connected in communication. The storage module is used to store computer programs, the transceiver module is used to send and receive messages, and the processing module is used to read the computer programs and execute the non-contact multimodal data acquisition and emotion monitoring method as described in any one of claims 1 to 2.

7. A computer-readable storage medium, characterized in that... The computer-readable storage medium stores instructions that, when executed on a computer, perform the non-contact multimodal data acquisition and emotion monitoring method as described in any one of claims 1 to 2.

8. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instructions are executed by the computer, they implement the non-contact multimodal data acquisition and emotion monitoring method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • CN119318486A