A subject data acquisition system based on multi-modal interaction

By using a multi-mode interactive system consisting of a flexible piezoelectric thin film array, a millimeter-wave radar module, and a directional microphone array, the problems of spatiotemporal alignment and noise suppression in physiological data acquisition in traditional devices are solved, achieving high signal-to-noise ratio physiological feature capture and low power consumption monitoring.

CN122420330APending Publication Date: 2026-07-17BEIJING YAOHAI NINGKANG PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YAOHAI NINGKANG PHARMACEUTICAL TECHNOLOGY CO LTD
Filing Date
2026-04-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional single-modal contact wearable devices have poor continuity in physiological data acquisition, while non-contact multi-sensor solutions suffer from inaccurate spatiotemporal alignment and poor noise resistance in complex environments, resulting in a high false alarm rate.

Method used

By employing a flexible piezoelectric thin film array, a millimeter-wave radar module, and a directional microphone array, a multi-modal interaction system is constructed through an edge computing gateway to achieve spatiotemporal alignment between sensors and suppression of environmental noise. Power consumption is reduced by utilizing confidence assessment of multimodal data and adaptive acquisition logic.

Benefits of technology

It achieves high signal-to-noise ratio acquisition and cross-validation of weak physiological signals from multiple sources, reduces system power consumption and computing power consumption, and improves the accuracy and reliability of data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122420330A_ABST
    Figure CN122420330A_ABST
Patent Text Reader

Abstract

This invention discloses a subject data acquisition system based on multi-mode interaction. The system includes an edge computing gateway, and a flexible piezoelectric thin film array, a millimeter-wave radar module, and a directional microphone array, all connected to the edge computing gateway. The edge computing gateway is configured to control the millimeter-wave radar module to operate at a low duty cycle and extract the compression profile in basic scanning mode. When the piezoelectric signal envelope variance exceeds the limit, it triggers entry into a directional diagnostic mode, calculating the geometric centroid coordinates and extracting the main peak of ventricular ejection. Based on these geometric centroid coordinates, the beamforming weight vector of the directional microphone array is updated to extract the audio signal sequence. The cross-correlation function between the main peak and the Doppler phase signal trough of the chest wall displacement is calculated to obtain the time delay parameter. After sliding window compensation alignment and feature vector extraction, the results are input into the model to output the state confidence assessment result. This application can achieve high signal-to-noise ratio and low false alarm rate for accurate spatiotemporal alignment of multi-source heterogeneous physiological signals and imperceptible monitoring of abnormal states.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things (IoT) data acquisition, and specifically refers to a subject data acquisition system based on multimodal interaction. Background Technology

[0002] In health monitoring scenarios, traditional single-modal contact wearable devices suffer from poor continuity in physiological data acquisition due to the susceptibility of optocoupler failure and low subject compliance. Existing non-contact multi-sensor solutions often employ a logic of independent data acquisition by each module and post-processing on the server side. However, due to the lack of underlying dynamic hierarchical mechanisms and spatial anchoring constraints, it is difficult to achieve precise spatiotemporal alignment between different modalities. Processing high-frequency, multi-dimensional sampling data generates massive computational loads, and heterogeneous data exhibits a high false alarm rate under complex environmental noise interference. It is also difficult to effectively integrate the physical correlations between multiple sensing dimensions, resulting in spatiotemporal misalignment and poor noise resistance. Summary of the Invention

[0003] This application provides a subject data acquisition system based on multi-modal interaction, which solves the problems of inaccurate spatiotemporal alignment of multi-source weak physiological signals and difficulty in suppressing environmental noise in non-contact or weak-contact environments.

[0004] This application provides a subject data acquisition system based on multi-modal interaction, including an edge computing gateway, and a flexible piezoelectric thin film array, a millimeter-wave radar module, and a directional microphone array, all connected to the edge computing gateway. The edge computing gateway is configured to execute control logic, which includes:

[0005] In the basic scanning mode, the millimeter-wave radar module is controlled to operate at a low duty cycle, and the two-dimensional pressure profile and turning frequency are extracted through the flexible piezoelectric thin film array.

[0006] If the variance of the piezoelectric signal envelope in adjacent periods exceeds the preset environmental noise threshold, it triggers the entry into the directional diagnostic mode, activating the full sampling state of the millimeter-wave radar module and the directional microphone array.

[0007] The geometric centroid coordinates are calculated based on the four adjacent nodes with the largest output voltage amplitude in the flexible piezoelectric thin film array. The piezoelectric signal sequence in the region corresponding to the geometric centroid coordinates is then subjected to bandpass digital filtering to extract the main peak of ventricular ejection.

[0008] The beamforming weight vector of the directional microphone array is calculated and updated based on the geometric centroid coordinates, and the three-dimensional spatial orientation of the pickup main lobe is locked to the facial projection area determined based on the geometric centroid coordinates in order to extract the audio signal sequence.

[0009] The cross-correlation function between the extracted ventricular ejection action peak and the trough in the chest wall displacement Doppler phase signal sequence extracted by the millimeter-wave radar module is calculated to obtain the time delay parameter.

[0010] The audio signal sequence and the chest wall displacement Doppler phase signal sequence are aligned using the time delay parameter through sliding window compensation. Feature vectors are extracted and input into the multimodal cross-attention fusion model to output the state confidence evaluation result.

[0011] This application embodiment constructs a physical feedback loop of a flexible piezoelectric thin film array, millimeter-wave radar, and microphone array at the sensing front end, dynamically adjusting the spatial beam and sampling parameters of the high-precision sensor using macroscopic position information from a low-frequency sensor. Based on the confidence index extracted from multimodal data, adaptive degradation or upgrade acquisition logic is executed, reducing system power consumption and computing power usage while achieving high signal-to-noise ratio capture and cross-validation of abnormal physiological characteristics. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the hardware structure of the subject data acquisition system based on multi-modal interaction provided in an embodiment of the present invention.

[0013] Figure 2 This is a flowchart illustrating the basic scanning and mode switching logic performed by the edge computing gateway provided in this embodiment of the invention.

[0014] Figure 3 This is a flowchart illustrating the execution of signal alignment and feature processing logic in the directional diagnostic mode provided in this embodiment of the invention.

[0015] [Explanation of Labels in the Attached Image]

[0016] In the diagram: 101 - Edge computing gateway, 102 - Flexible piezoelectric thin film array, 103 - Millimeter wave radar module, 104 - Directional microphone array. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions of this application, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0018] Figure 1 This is a schematic diagram of the hardware topology of a subject data acquisition system based on multi-modal interaction provided in an embodiment of the present invention. Figure 1 As shown, the subject data acquisition system provided in this application includes a multi-layered sensing and computing topology network in its underlying hardware architecture. The system mainly includes a core control device, namely an edge computing gateway 101, as well as a flexible piezoelectric thin film array 102, a millimeter-wave radar module 103, and a directional microphone array 104, which are respectively signal-connected to the edge computing gateway 101.

[0019] Specifically, the flexible piezoelectric film array 102 is a polyvinylidene fluoride flexible piezoelectric film array laid inside the subject's mattress or under the sheet. In this embodiment, the flexible piezoelectric film array 102 consists of 64 piezoelectric sensing nodes arranged in a matrix, with the center-to-center distance between adjacent nodes set to 5 cm. The dense matrix arrangement can fully cover the subject's main trunk range of motion. Through the piezoelectric effect of this film, local body movement changes under passive pressure and mechanical vibrations of the body surface caused by ventricular ejection can be converted into voltage signals.

[0020] The millimeter-wave radar module 103 is installed in the ceiling area directly above the bed and is configured as a frequency-modulated continuous wave millimeter-wave radar module. The module operates in the frequency band between 60 GHz and 64 GHz, and the antenna configuration uses a three-transmit, four-receive array. The millimeter-wave radar module 103 radiates frequency-modulated continuous waves downwards and receives echoes reflected from the human chest wall and limbs. It utilizes the echo time difference and Doppler frequency shift mechanism to obtain the absolute distance to the human target and the minute cardiopulmonary displacement characteristics caused by respiration.

[0021] A directional microphone array 104 is positioned near the headboard and configured as a quaternary microelectromechanical system (MEMS) directional microphone array. This array is responsible for picking up ambient sound waves in real time, as well as the faint frictional sound waves generated by the subject's breathing airflow through narrowed airways.

[0022] In terms of hardware communication interface and main control unit design, the edge computing gateway 101 integrates a field-programmable gate array (FPGA) and a high-performance microprocessor chip. The edge computing gateway 101 is connected to the built-in analog-to-digital converter (ADC) of the flexible piezoelectric thin film array 102 via a serial peripheral interface bus. Simultaneously, the edge computing gateway 101 possesses high-speed network switching capabilities, receiving digitized high-frequency data streams output from the millimeter-wave radar module 103 and the directional microphone array 104 via an Ethernet interface.

[0023] To ensure the stable operation of the entire distributed sensing system during long-endurance monitoring, this embodiment also includes a power management module. The power management module integrates a low-ripple voltage regulator circuit and an automatic switching scheme between AC / battery dual-mode power supply, providing a clean power supply with a high signal-to-noise ratio for highly sensitive sensors and preventing abrupt changes in the acquisition baseline due to AC power fluctuations. In addition, the system is equipped with a communication transceiver module to support data synchronization. This module includes a low-power communication module and a wireless LAN module, ensuring highly reliable uploading of heterogeneous data.

[0024] It should be noted that the specific number of flexible thin-film nodes and the sensor spacing in the above embodiments are merely preferred examples, and the scale of the underlying array in this application is not limited thereto. Those skilled in the art can adopt alternative solutions such as increasing the number of nodes or adjusting the array topology, depending on the specific size of the bed or the body type of the applicable population, to achieve the micro-vibration acquisition purpose of this application.

[0025] Figure 2 This is a flowchart illustrating the basic scanning and mode switching logic performed by the edge computing gateway according to an embodiment of the present invention. Combined with... Figure 1 and Figure 2 As shown, the microprocessor chip inside the edge computing gateway 101 runs fixed control logic and adaptively switches to different environmental conditions and event stages.

[0026] In the basic scan mode control S201, the system operates in basic scan mode by default. The edge computing gateway 101 issues control commands to control the millimeter-wave radar module 103 to operate in a low duty cycle state. When operating at a low duty cycle, the radar transmitter transmits pulses with longer intervals, outputting only low-resolution macroscopic range envelope information for bed status detection. At the same time, the edge computing gateway 101 shuts down the power supply rail of the directional microphone array 104 to reduce background power consumption. In basic scan mode, the system only retains the flexible piezoelectric thin film array 102 to continuously operate at a low sampling rate of 10 Hz. The edge computing gateway 101 uses this 10 Hz serial peripheral interface data stream to analyze the two-dimensional spatial energy distribution matrix of the piezoelectric signal in 64 nodes in real time, thereby extracting the two-dimensional pressure profile and macroscopic turning frequency of the subject on the bed surface.

[0027] In noise-triggered monitoring and judgment S202, the edge computing gateway 101 runs a noise-triggered monitoring thread to determine whether deep detection is required. The edge computing gateway 101 continuously calculates and extracts the instantaneous voltage signal sequence of the flexible piezoelectric thin film array 102 within the current time window. The instantaneous voltage signal sequence is analyzed and reconstructed using Hilbert transform to obtain the envelope of the low-frequency physiological signal.

[0028] The system specifically performs the following operations: It performs a Fast Fourier Transform (FFT) on the instantaneous voltage signal sequence in the time domain, transforming it into the frequency domain; it constructs a single-sideband spectrum filter in the frequency domain, forcing the amplitude of the negative frequency components in this frequency domain to zero, and multiplies the amplitude of the positive frequency components by a weighting constant k to compensate for the energy loss caused by truncation; it performs an inverse Fast Fourier Transform on the filtered single-sideband spectrum to obtain the reconstructed analytic signal sequence. By calculating the complex modulus of this reconstructed analytic signal sequence, it generates an envelope signal amplitude array. In a preferred embodiment, the above reconstruction and envelope calculation logic is implemented through the following formula:

[0029]

[0030] Where Z(n) represents the reconstructed analytical signal sequence obtained through the Hilbert transform mechanism, and X(n) represents the time-domain instantaneous voltage signal sequence extracted from the flexible piezoelectric film. H\{·\} represents the standard Hilbert transform operator operation, and j represents the imaginary unit of the complex plane. E(n) represents the amplitude of the envelope signal representing the slowly varying profile of the signal. The weighting constant k is preferably 2 in this embodiment.

[0031] Further, after obtaining the envelope signal amplitude array, the edge computing gateway 101 calculates the variance value of the aforementioned signal envelope within multiple consecutive sliding time windows. The system divides the envelope signal amplitude array into M mutually overlapping sub-intervals of equal length, where M is a positive integer. For each sub-interval, the system accumulates all envelope signal amplitude values ​​within the sub-interval and divides them by the number of sampling points N in that sub-interval to obtain the sub-interval mean. Then, the difference between each envelope signal amplitude value and the corresponding sub-interval mean is squared and summed, then divided by N-1 to obtain the sub-interval variance distribution matrix representing the intensity of local fluctuations. In a preferred embodiment, the above variance matrix calculation logic is implemented using the following formula:

[0032] Where μm represents the arithmetic mean parameter of the m-th envelope sub-interval, and N represents the total number of sampling points within the preset sub-interval. E m,i Let σm represent the specific envelope signal amplitude value at index i in the m-th subinterval. σm2 represents the estimated variance of the signal envelope in this subinterval. The variance values ​​of all subintervals are combined to form the variance distribution matrix.

[0033] Edge computing gateway 101 compares the absolute value of the variance difference between adjacent sliding time windows with a preset environmental noise threshold in real time. Since normal breathing or low-noise environments often exhibit stationary random processes, the difference in their envelope variance fluctuates within a small range. If a subject experiences paroxysmal violent movements, convulsions, or severe panting, the variance in a local time period will increase. If edge computing gateway 101 determines that the absolute value of the variance difference between adjacent periods exceeds the preset environmental noise threshold, it generates an interrupt wake-up command and sends it to the internal core controller, triggering the system to enter directional diagnostic mode. In this mode, the system sends a control signal to the Ethernet bus, activating the millimeter-wave radar module 103 and the directional microphone array 104 to enter a high-frequency full-sampling state. By employing the aforementioned conditional triggering mechanism based on the envelope variance of low-frequency piezoelectric signals, macroscopic turning and abrupt changes in body kinetic energy levels can be used as signal sentinels, avoiding redundant power consumption caused by the uninterrupted full-scale operation of high-dimensional heterogeneous sensors, and extending the continuous standby and working life of the edge monitoring system.

[0034] Once the system is activated and enters the directional diagnostic mode, multidimensional physiological parameters flood into the edge computing gateway 101. Due to the different physical locations of the sensors and the impedance differences of the signal propagation medium, spatiotemporal coupling and alignment of the concurrent data is the foundation for achieving subsequent fusion and identification. Figure 3 This is a flowchart illustrating the execution of signal alignment and feature processing logic in the directional diagnostic mode provided in this embodiment of the invention.

[0035] Combination Figure 3 As shown, in the centroid localization and cardiac wave extraction S301, in order to establish a unified spatial reference point for multi-source data, the edge computing gateway 101 locates the four adjacent nodes with the largest output voltage amplitude based on the matrix information returned by the flexible piezoelectric thin film array 102, and calculates the geometric centroid coordinates based on the physical array coordinates of these four nodes. This geometric centroid represents the dynamic equilibrium center of the subject's current trunk weight distribution. The edge computing gateway 101 extracts the piezoelectric signal sequence of the physical region corresponding to the geometric centroid coordinates and inputs the sequence into a pre-configured bandpass digital filter from 0.25 Hz to 2.5 Hz. This bandpass frequency band can filter out high-frequency noise in the environment and low-frequency bed static pressure interference, thereby filtering out a pure cardiac impact map signal and locking in the main peak waveform features representing the ventricular ejection action.

[0036] In the dynamic sound source suppression and pickup S302, the edge computing gateway 101 executes a dynamic sound source suppression and pickup process guided by the aforementioned physical centroid coordinates to address spatial noise interference. The system establishes an absolute coordinate system mapping matrix between the two-dimensional pressure plane where the flexible piezoelectric film array 102 is located and the three-dimensional space where the directional microphone array 104 is positioned at the head of the bed. The calculated dynamic geometric centroid coordinates are substituted into the system's built-in human body prior scale model to estimate the three-dimensional spatial coordinates of the sound source emanating from the subject's face. Based on these three-dimensional spatial coordinates, the time delay difference between the arrival and reception of the received sound wave signal from each independent microphone element in the directional microphone array 104 is calculated. The system generates a desired spatial guidance vector for the sound source target based on this arrival and reception time difference. The target weight vector is calculated in real-time using a minimum variance distortion-free response algorithm, combined with the environmental noise covariance matrix extracted by the directional microphone array 104 in the current environment.

[0037] Specifically, at the execution level of the minimum variance distortionless response algorithm, the system constructs a noise basis. The edge computing gateway 101 collects background acoustic data sequences from each microphone array element in a silent state, constructing an environmental noise sample correlation matrix. To prevent singularity divergence during matrix inversion, the environmental noise sample correlation matrix undergoes eigenvalue decomposition, and a diagonal loading factor is added to its main diagonal to obtain the inverse of the regularized covariance matrix. This inverse of the regularized covariance matrix is ​​then multiplied by the desired spatial steering vector to generate the numerator of the beamforming weights. The conjugate transpose of the spatial steering vector, the inverse of the regularized covariance matrix, and the spatial steering vector itself are then multiplied sequentially to generate a constant scalar denominator. The numerator is divided by the denominator to obtain the normalized optimal complex weight vector. In a preferred embodiment, the above beamforming weight logic is implemented using the following matrix formula:

[0038]

[0039] Among them, W opt R represents the optimal complex weight vector obtained through algorithmic optimization. n -1 Let represent the inverse of the regularized covariance matrix with a diagonal loading factor added. Let A represent the spatial steering vector of the sound-emitting target calculated based on the arrival delay. H This represents the complex conjugate transpose matrix corresponding to the spatial steering vector.

[0040] The optimal complex weight vector W is obtained through calculation. opt Subsequently, the edge computing gateway 101 sends its phase adjustment value and amplitude attenuation coefficient to the data buffer link of each microphone element of the directional microphone array 104. This configuration operation generates a digital filter lens, locking the three-dimensional spatial pointing of the system's main lobe to the facial projection area determined based on the geometric centroid coordinates. Simultaneously, deep valley nulls are formed in other directions, outputting an acoustic reception spectrum with spatial response gain configuration information, accurately extracting the audio signal sequence representing the breathing event. By employing this microphone beamforming method guided by the physical spatial coordinates of a flexible thin film, and utilizing the absolute orientation prior knowledge provided by the piezoelectric array, the acoustic device can achieve clear targeted acquisition of faint breathing sounds without being affected by non-stationary sound wave interference.

[0041] In the cross-correlation delay parameter calculation S303, the time phase misalignment problem of multi-source sensors is resolved. The edge computing gateway 101 calculates the cross-correlation function between the extracted piezoelectric ventricular ejection action peak and the trough in the chest wall displacement Doppler phase signal sequence extracted by the millimeter-wave radar module 103.

[0042] The specific matching processing logic is as follows: Extract the piezoelectric signal waveform segment corresponding to the main peak of cardiac impact, and perform normalization processing to obtain a dimensionless first standard sequence. Extract trough segments from the Doppler phase signal sequence of chest wall displacement representing respiratory fluctuations calculated from millimeter-wave radar, and perform zero-mean processing to eliminate baseline offset artifacts caused by phase-locked loop drift within the radar, obtaining a second standard sequence. In the time domain, use a sliding window algorithm to perform successive sliding inner product operations on the first and second standard sequences to obtain a cross-correlation coefficient sequence reflecting the change in their similarity with time offset. Search for the offset point corresponding to the maximum absolute value in the cross-correlation coefficient sequence, and multiply this offset point number by the system's uniformly preset reference sampling period to obtain the time delay parameter between the two heterogeneous signals. In a preferred embodiment, the above sliding matching calculation logic is implemented through the following formula:

[0043]

[0044]

[0045] Where Rxy(τ) represents the calculated cross-correlation coefficient distribution sequence. norm (n) represents the normalized first standard sequence of heartbeats as the time anchor. ynorm(n-τ) represents the zero-mean radar displacement second standard sequence after a shift of τ steps. τ represents the relative offset point variable during the sliding window traversal process, and K represents the number of discrete data points covered by a single cross-correlation operation. τmax represents the number of time offset points that make the cross-correlation coefficient reach its absolute peak, and T represents the global reference sampling period. Δt is the time delay parameter for eliminating the difference in medium propagation delay. By using the low-frequency and periodic mechanical heartbeat peak as the time anchor to calibrate the radar phase, and utilizing the periodic coupling relationship of heterogeneous physical quantities, the Doppler frequency jump is eliminated, providing a stable time-domain reference system.

[0046] In the sliding window alignment and feature extraction S304, the time delay parameter obtained by cross-correlation is used to perform sliding window compensation alignment on different data streams. The edge computing gateway 101 backslides or forwards the start timestamps of the audio signal sequence and the chest wall displacement Doppler phase signal sequence in the memory cache according to the time delay parameter, ensuring that the physical events represented by the heterogeneous data at the same time section are synchronized. After alignment, the system performs multi-dimensional feature vector extraction.

[0047] The system feeds the aligned audio signal sequence into the front-end acoustic processor. By performing pre-emphasis, framing, Hamming windowing, and discrete cosine transform operations, it extracts the spectral envelope feature matrix representing the acoustic dimension of the vocal organs using the standard Mel-frequency cepstral coefficient algorithm. For the aligned chest wall displacement Doppler phase signal sequence, the system uses a continuous wavelet transform algorithm to expand it and extract the time-frequency energy distribution matrix reflecting the respiratory frequency fluctuation dimension. The system also extracts the relative energy spectral density parameter from the radar signal and the main peak fluctuation variation coefficient of the cardiac impact map extracted by the flexible film. After extracting the multi-source features, the system concatenates the spectral envelope feature matrix, the time-frequency energy distribution matrix, and the correlation coefficient side-by-side along the same time frame dimension to construct a multimodal fusion feature tensor, which serves as the feature vector basis for network inference.

[0048] In the cross-modal fusion and evaluation output S305, the edge computing gateway 101 inputs the constructed multimodal fusion feature tensor to the in-house multimodal cross-attention fusion model to perform health feature analysis and output the state confidence evaluation result for the current subject.

[0049] In the forward inference phase of this model, the input multimodal fusion feature tensor is fed into a parallel adaptive linear mapping layer, where it is dimensionality-reduced into a query matrix, a key matrix containing feature relationships, and a value matrix preserving the original information. In the cross-attention mechanism, the query matrix representing the first modality is multiplied by the transpose of the key matrix representing the second modality. The implicit correlation distribution between the two different physical phenomena is calculated, generating an attention distribution weight matrix. This weight matrix is ​​then used to perform a weighted summation operation on the value matrix corresponding to the second modality, completing the intelligent update of the modal representation features.

[0050] The updated features are fed into a fully connected layer for dimensionality compression. A normalized exponential function is then used to perform regression calculations on the deep modal representation features, outputting a probability value between 0 and 1 to represent an abnormal respiratory state, which serves as the state confidence assessment result. In a preferred embodiment, the above multimodal feature weighted fusion and probability output calculation logic is implemented using the following formula:

[0051]

[0052] Where Attention(Q, K, V) represents the hybrid modality representation feature matrix obtained after weighted update via cross-attention mechanism. Q represents the query matrix vector generated by the first modality feature flow mapping. K represents the key matrix vector extracted by the second modality feature mapping. T Let V be its conjugate transpose. V represents the modulated value matrix vector. d kThis represents the intrinsic channel dimension coefficient of the mapped feature space. Softmax represents the exponential normalization function, which converts the pre-matrix parameters into a standard probability weight distribution. By employing a multimodal cross-attention fusion mechanism, and utilizing the complementary constraint characteristics of heterogeneous physical parameters in representing the same physiological event, false alarm noise from a single sensor is dynamically eliminated, outputting a high-confidence health status confidence level in complex environments.

[0053] In alarm interaction and data upload S306, the edge computing gateway 101 performs interactive operations based on the generated confidence assessment results. After outputting the state confidence assessment results with probability labels, the edge computing gateway 101 sends the results along with compressed radar and audio key slices to the wireless LAN module via the internal system bus. The wireless LAN module reassembles and encapsulates the message according to the time-sequenced format and uploads the information packet to the cloud node. Upon receiving the message, the cloud node triggers the monitoring terminal of the corresponding nursing staff to issue an alarm notification.

[0054] As an explanation of the AI-assisted decision-making process in this embodiment, to improve the generalization ability of the multimodal cross-attention fusion model, the system model is trained based on a pre-collected offline clinical dataset. During training, the cross-entropy loss function is used to calculate the error between the network output probability and the gold standard, and the attention matrix parameters and fully connected layer weights are iteratively optimized by combining backpropagation and adaptive gradient descent algorithms.

[0055] The data acquisition system provided in this application breaks the traditional multi-sensor island effect by establishing a tight physical feedback loop and spatiotemporal coupling mechanism among the heterogeneous sensors at the underlying level. The physical center of gravity anchor point generated by the flexible piezoelectric array provides precise spatial constraint direction for the beamforming of the directional microphone; cross-modal time-domain cross-correlation is performed using the low-frequency piezoelectric main peak as the time anchor point, eliminating the influence of continuous wave radar baseline skipping; and a hardware dynamic hierarchical wake-up mechanism based on energy level variance not only overcomes the data loss problem caused by low wear compliance but also meets the system's low power consumption requirements. This achieves high signal-to-noise ratio and low false alarm rate capture of weak physiological characteristic signals in multimodal non-contact monitoring environments, demonstrating good application potential.

Claims

1. A subject data acquisition system based on multi-modal interaction, comprising an edge computing gateway, and a flexible piezoelectric thin film array, a millimeter-wave radar module, and a directional microphone array respectively connected to the edge computing gateway; characterized in that, The edge computing gateway is configured to execute control logic, which includes: In the basic scanning mode, the millimeter-wave radar module is controlled to operate at a low duty cycle, and the two-dimensional pressure profile and turning frequency are extracted through the flexible piezoelectric thin film array. If the variance of the piezoelectric signal envelope in adjacent periods exceeds the preset environmental noise threshold, it triggers the entry into the directional diagnostic mode, activating the full sampling state of the millimeter-wave radar module and the directional microphone array. The geometric centroid coordinates are calculated based on the four adjacent nodes with the largest output voltage amplitude in the flexible piezoelectric thin film array. The piezoelectric signal sequence in the region corresponding to the geometric centroid coordinates is then subjected to bandpass digital filtering to extract the main peak of ventricular ejection. The beamforming weight vector of the directional microphone array is calculated and updated based on the geometric centroid coordinates, and the three-dimensional spatial orientation of the pickup main lobe is locked to the facial projection area determined based on the geometric centroid coordinates in order to extract the audio signal sequence. The cross-correlation function between the extracted ventricular ejection action peak and the trough in the chest wall displacement Doppler phase signal sequence extracted by the millimeter-wave radar module is calculated to obtain the time delay parameter. The audio signal sequence and the chest wall displacement Doppler phase signal sequence are aligned using the time delay parameter through sliding window compensation. Feature vectors are extracted and input into the multimodal cross-attention fusion model to output the state confidence evaluation result.

2. The subject data acquisition system based on multi-modal interaction as described in claim 1, characterized in that, If the variance of the piezoelectric signal envelope in adjacent periods exceeds a preset environmental noise threshold, triggering the entry into the directional diagnostic mode includes: Extract the instantaneous voltage signal sequence of the flexible piezoelectric thin film array within the current time window; The instantaneous voltage signal sequence is analytically reconstructed using the Hilbert transform to obtain the signal envelope; The variance of the signal envelope is calculated over multiple consecutive sliding time windows. Compare the absolute value of the difference in variance between adjacent sliding time windows with the environmental noise threshold; If the absolute value of the variance difference is greater than the environmental noise threshold, a wake-up command is generated and sent to the edge computing gateway, and the system is switched to the targeted diagnostic mode.

3. The subject data acquisition system based on multi-modal interaction as described in claim 2, characterized in that, The step of using Hilbert transform to perform analytical signal reconstruction of the instantaneous voltage signal sequence to obtain the signal envelope includes: Perform a Fast Fourier Transform on the instantaneous voltage signal sequence to enter the frequency domain; A single-sideband spectrum filter is constructed to set the amplitude of the negative frequency component in the frequency domain to zero and multiply the amplitude of the positive frequency component by a weighting constant k; Perform an inverse fast Fourier transform to obtain the reconstructed analytic signal sequence; Extract the complex modulus of the reconstructed analytical signal sequence to generate an envelope signal amplitude array; The step of calculating the variance of the signal envelope within multiple consecutive sliding time windows includes dividing the envelope signal amplitude array into M overlapping sub-intervals of equal length, where M is a positive integer; For each sub-interval, the sum of the amplitude values ​​of all envelope signals within the sub-interval and the result divided by the number of sampling points N are used to obtain the mean of the sub-interval. Divide the sum of squared differences between the amplitude value of each envelope signal and the mean of the sub-interval by N-1 to obtain the variance distribution matrix of the sub-interval.

4. The subject data acquisition system based on multi-modal interaction as described in claim 1, characterized in that, The calculation and updating of the beamforming weight vector of the directional microphone array based on the geometric centroid coordinates includes: Establish a coordinate system mapping matrix between the two-dimensional plane containing the flexible piezoelectric thin film array and the three-dimensional space containing the directional microphone array; Substitute the geometric centroid coordinates into the prior human body scale model to estimate the three-dimensional spatial coordinates of the facial sound source. The arrival delay time difference of each microphone array element is calculated using the three-dimensional spatial position coordinates; Based on the arrival and reception delay time difference, a spatial steering vector is generated. Combined with the environmental noise covariance matrix of the directional microphone array, the target weight vector is calculated using the minimum variance distortionless response algorithm.

5. The subject data acquisition system based on multi-modal interaction as described in claim 4, characterized in that, The environmental noise covariance matrix of the combined directional microphone array is used to calculate the target weight vector using a minimum variance distortionless response algorithm, including: Collect background acoustic data sequences of each microphone array element under silent conditions, and construct an environmental noise sample correlation matrix; The inverse of the regularized covariance matrix is ​​obtained by performing eigenvalue decomposition on the correlation matrix of the environmental noise samples and adding a diagonal loading factor. The inverse of the regularized covariance matrix is ​​multiplied by the spatial steering vector to generate the numerator term; The denominator term is generated by sequentially multiplying the conjugate transpose of the spatial steering vector, the inverse of the regularized covariance matrix, and the spatial steering vector. The optimal complex weight vector is obtained by dividing the numerator by the denominator. The phase and amplitude of the optimal complex weight vector are respectively configured to each microphone element of the directional microphone array, and an acoustic reception spectrum with spatial response gain configuration information is output.

6. The subject data acquisition system based on multi-modal interaction as described in claim 1, characterized in that, The cross-correlation function between the main peak of ventricular ejection action extracted by the calculation and the trough in the Doppler phase signal sequence of chest wall displacement extracted by the millimeter-wave radar module includes: The piezoelectric signal waveform corresponding to the main peak of the ventricular ejection action is normalized to obtain the first standard sequence. The radar displacement waveform corresponding to the trough in the chest wall displacement Doppler phase signal sequence is processed with zero mean to obtain the second standard sequence. In the time domain, a sliding inner product operation is performed on the first standard sequence and the second standard sequence to obtain a cross-correlation coefficient sequence; The time delay parameter Δt is obtained by multiplying the number of offset points τ at the point of maximum absolute value in the cross-correlation coefficient sequence by the preset sampling period T.

7. The subject data acquisition system based on multimodal interaction as described in claim 1, characterized in that, The extraction of feature vectors includes: The acoustic dimension spectral envelope feature matrix is ​​extracted from the aligned audio signal sequence using the Mel frequency cepstral coefficient algorithm; The time-frequency energy distribution matrix of the respiratory dimension was extracted from the aligned chest wall displacement Doppler phase signal sequence through continuous wavelet transform. The spectral envelope feature matrix and the time-frequency energy distribution matrix are concatenated along the time dimension to construct a multimodal fusion feature tensor, which serves as the feature vector.

8. The subject data acquisition system based on multimodal interaction as described in claim 7, characterized in that, The input to the multimodal cross-attention fusion model outputs the state confidence evaluation result, including: The multimodal fusion feature tensor is mapped to a query matrix, a key matrix, and a value matrix, respectively. The attention distribution weight matrix is ​​generated by performing a dot product operation between the query matrix corresponding to the first modality and the transpose of the key matrix corresponding to the second modality. The modal representation features are updated by weighting and summing the value matrix corresponding to the second modality using the attention distribution weight matrix. The modal characterization features are calculated using a fully connected layer and a normalized exponential function, and the output probability value representing the abnormal respiratory state is used as the state confidence assessment result.

9. The subject data acquisition system based on multi-modal interaction as described in claim 1, characterized in that, The flexible piezoelectric thin film array includes multiple polyvinylidene fluoride piezoelectric sensor nodes arranged in a matrix. The millimeter-wave radar module includes a radio frequency processing unit configured with a frequency-modulated continuous wave transmitting antenna and a receiving antenna. The edge computing gateway integrates a field-programmable gate array and a microprocessor chip.

10. The subject data acquisition system based on multimodal interaction as described in claim 1, characterized in that, The edge computing gateway is equipped with a low-power communication module and a wireless local area network module. The edge computing gateway receives the serial peripheral interface data stream output by the flexible piezoelectric thin film array through the low-power communication module; After generating the state confidence assessment result, the edge computing gateway encapsulates the state confidence assessment result and uploads it to the cloud node through the wireless LAN module in a time-slice format.