Radar-based health monitoring method, device, computer equipment and storage medium

CN122004824BActive Publication Date: 2026-06-12CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGCHUN UNIV OF SCI & TECH
Filing Date
2026-04-16
Publication Date
2026-06-12

Smart Images

  • Figure CN122004824B_ABST
    Figure CN122004824B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of radar signal data processing, and discloses a health monitoring method and device based on a radar, computer equipment and a storage medium. The method obtains phase time sequence data through pre-processing of a chest cavity radar echo signal; obtains initial time-frequency features of a current time step through an initial time-frequency analysis model; further determines a joint state vector, a continuous action vector and a time-frequency quality score of the current time step, and determines dense internal rewards of the current time step according to the joint state vector, the continuous action vector and the time-frequency quality score; iteratively updates the initial time-frequency analysis model by taking the dense internal rewards as reward signals, and obtains a target time-frequency analysis model based on reinforcement learning; and determines a health monitoring result according to target time-frequency features obtained through the target time-frequency analysis model. The application realizes self-adaptive optimization of a hyperwavelet parameter in a time-frequency analysis model, improves the robustness of time-frequency representation, and guarantees the effect of health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar signal data processing technology, and in particular to a radar-based health monitoring method, device, computer equipment, and storage medium. Background Technology

[0002] In the field of smart healthcare and non-contact physiological monitoring, frequency modulated continuous wave (FMCW) radar can be used for health monitoring of heart rate, heart rate intervals, and respiratory rate.

[0003] In existing technologies, classical time-frequency analysis methods (such as short-time Fourier transform) inherently trade off between time and frequency resolution, failing to simultaneously guarantee the accuracy of both. While wavelet-based time-frequency analysis methods can improve the concentration of time-frequency energy to some extent, they rely on empirically selecting and fixing key parameters (such as wavelet cycle number / order related parameters, frequency sampling density, and scale control parameters). When the amplitude-frequency structure of the phase micro-motion signal changes due to individual differences, body motion interference, changes in monitoring distance, or heart rate fluctuations, fixing these parameters can easily lead to time-frequency energy diffusion, sidelobe enhancement, or ridge breakage. This reduces the generalization and robustness of radar signal analysis results in complex application scenarios (such as post-exercise heart rate fluctuations and long-distance monitoring), thereby affecting the reliability of health status indicators such as heart rate and heart rate intervals. Summary of the Invention

[0004] Therefore, it is necessary to provide a radar-based health monitoring method, device, computer equipment, and storage medium to address the aforementioned technical problems and solve the problem of poor robustness of radar signal analysis in radar-based health monitoring.

[0005] A radar-based health monitoring method includes:

[0006] Acquire the thoracic radar echo signal, and preprocess the thoracic radar echo signal to obtain phase time series data;

[0007] The phase time series data is processed by wavelet transform using an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations that update the initial time-frequency analysis model.

[0008] The joint state vector of the current time step is determined based on the phase time series data and the initial time-frequency characteristics; the continuous action vector of the current time step is determined based on the joint state vector; and the time-frequency quality score of the current time step is determined based on the initial time-frequency characteristics.

[0009] The dense internal reward for the current time step is determined based on the joint state vector, the continuous action vector, and the time-frequency quality score.

[0010] The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and after the iterative update is completed, the target time-frequency analysis model based on reinforcement learning is obtained.

[0011] The phase time series data are processed by the target time-frequency analysis model to obtain the target time-frequency characteristics, and the health monitoring results are determined based on the target time-frequency characteristics.

[0012] A radar-based health monitoring device includes:

[0013] The preprocessing module is used to acquire the thoracic radar echo signal, preprocess the thoracic radar echo signal, and obtain phase time series data.

[0014] The initial model analysis module is used to perform wavelet transform processing on the phase time series data through an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations that update the initial time-frequency analysis model.

[0015] The feature analysis module is used to determine the joint state vector of the current time step based on the phase time series data and the initial time-frequency features, determine the continuous action vector of the current time step based on the joint state vector, and determine the time-frequency quality score of the current time step based on the initial time-frequency features.

[0016] The reward analysis module is used to determine the dense internal reward of the current time step based on the joint state vector, the continuous action vector, and the time-frequency quality score.

[0017] The model update module is used to iteratively update the initial time-frequency analysis model using the dense internal reward as a reward signal, and obtain the target time-frequency analysis model based on reinforcement learning after the iterative update is completed.

[0018] The monitoring result determination module is used to perform ultrawavelet transform processing on the phase time series data through the target time-frequency analysis model to obtain the target time-frequency characteristics, and determine the health monitoring results based on the target time-frequency characteristics.

[0019] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the radar-based health monitoring method described above when executing the computer-readable instructions.

[0020] A computer-readable storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the radar-based health monitoring method described above.

[0021] In the aforementioned radar-based health monitoring method, device, computer equipment, and storage medium, the method acquires chest radar echo signals, preprocesses the chest radar echo signals to obtain phase time series data, performs wavelet transform processing on the phase time series data using an initial time-frequency analysis model to obtain initial time-frequency features for the current time step, determines the joint state vector of the current time step based on the phase time series data and initial time-frequency features, determines the continuous action vector of the current time step based on the joint state vector, and determines the time-frequency quality score of the current time step based on the initial time-frequency features, determines the dense internal reward of the current time step based on the joint state vector, continuous action vector, and time-frequency quality score, iteratively updates the initial time-frequency analysis model using the dense internal reward as a reward signal, and obtains a target time-frequency analysis model based on reinforcement learning after the iterative update, performs wavelet transform processing on the phase time series data using the target time-frequency analysis model to obtain target time-frequency features, and determines the health monitoring results based on the target time-frequency features. This invention first obtains initial time-frequency features through an initial time-frequency analysis model. Based on the joint state vector, continuous action vector, and time-frequency quality score, it determines dense internal rewards. Under the constraints of a time-frequency representation quality evaluation mechanism, it learns an internal reward signal that can be progressively fed back, demonstrating the strategy optimization of reinforcement learning and achieving adaptive optimization of key wavelet parameters in the time-frequency analysis model. This invention avoids the problems of insufficient resolution, low energy concentration, and susceptibility of the main ridge line to noise and stray components in the time-frequency representation of heartbeat-related data caused by fixed parameters. It improves the quality and robustness of the time-frequency representation and ensures the effectiveness of health status assessment under complex monitoring conditions. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application environment for a radar-based health monitoring method according to an embodiment of the present invention;

[0024] Figure 2 This is a schematic flowchart of a radar-based health monitoring method according to an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of a model iterative update process for a radar-based health monitoring method according to an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of a reward discovery network structure for a radar-based health monitoring method in one embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram of a strategy optimization network structure for a radar-based health monitoring method in one embodiment of the present invention;

[0028] Figure 6 This is a time-frequency characterization comparison diagram of a radar-based health monitoring method in one embodiment of the present invention;

[0029] Figure 7 This is a schematic diagram of a radar-based health monitoring device according to an embodiment of the present invention;

[0030] Figure 8 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] The radar-based health monitoring method provided in this embodiment can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, various personal computers, laptops, smartphones, tablets, and wearable radar devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0033] In one embodiment, such as Figure 2 As shown, a radar-based health monitoring method is provided, which can be applied to... Figure 1 Taking the server side as an example, the explanation includes the following steps S10-S50.

[0034] S10. Acquire the thoracic radar echo signal, preprocess the thoracic radar echo signal to obtain phase time series data.

[0035] In essence, chest radar echo signals are radar echo signals containing cardiopulmonary activity information collected by illuminating the chest region of the monitored object with an FMCW radar. Chest radar echo signals can be real-time signal data collected by the FMCW radar, historical signal data obtained from a public database, or signal data obtained through interaction with other servers. For example, using a wearable FMCW radar device as a client, the FMCW radar device collects chest radar echo signals in a non-contact manner and sends the collected chest radar echo signals to the server. Phase time series data refers to a continuous phase time series used to characterize minute movements of the chest cavity, obtained after preprocessing the chest radar echo signals.

[0036] In one specific embodiment, during preprocessing, the server first uses coherent accumulation technology to coherently accumulate the thoracic cavity radar echo signal along the slow time dimension to improve the signal-to-noise ratio (SNR) of the target signal. Then, zero-padding and windowing operations are performed on the processed signal, and the range vector corresponding to each receiving antenna channel is obtained through a one-dimensional Fast Fourier Transform (FFT). Next, beamforming processing is performed on the range vectors of multiple receiving antenna channels in the array dimension. By calculating the beam output energy within a preset angle space search range and selecting the angle parameter corresponding to the maximum response, the optimal beam pointing of the target thoracic cavity scattering center is determined, and the target echo signal under this optimal beam is extracted to improve the SNR and anti-interference capability. Finally, the complex phase of the target range unit is extracted, and a phase unwinding algorithm is used to recover the continuous phase time series representing the minute movements of the thoracic cavity, i.e., phase time series data. The phase time series data is represented as... ,in For discrete sampling time.

[0037] S20. Perform ultrawavelet transform processing on the phase time series data through the initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; the time step refers to the number of iterations to update the initial time-frequency analysis model.

[0038] Understandably, the initial time-frequency analysis model is a neural network model that needs to be iteratively updated to convert the phase time series into a super-resolution time-frequency representation. For example... Figure 3As shown, the initial time-frequency analysis model is a pre-constructed adaptive super-resolution time-frequency analysis framework based on deep reinforcement learning, including a time-frequency computation network, a state construction network, a parameter action mapping network, an external sparse evaluation network, a reward discovery network, and a policy optimization network based on policy gradients. Specifically, the time-frequency computation network is used to compute the super-resolution time-frequency representation; the state construction network is used to compute the joint state vector representation; the parameter action mapping network is used to map continuous action vectors to the parameters required for computing the super-resolution time-frequency representation; the external sparse evaluation network is used to compute the quality score of the super-resolution time-frequency representation; the reward discovery network is used to compute dense internal rewards; and the policy optimization network based on policy gradients is used to optimize the policy based on the dense internal rewards and generate optimized continuous action vector representations.

[0039] The server performs super wavelet transform on the phase time series data using the time-frequency calculation network in the initial time-frequency analysis model to obtain the initial time-frequency features for the current time step. The initial time-frequency features for the current time step refer to the super-resolution time-frequency representation corresponding to the phase time series data in the current iteration of updating the initial time-frequency analysis model. One time step corresponds to one initial time-frequency feature.

[0040] In one embodiment, step S20, namely, performing ultrawavelet transform processing on the phase time series data using an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step, includes:

[0041] S201. Obtain the reference period parameter, modulation factor parameter, and order index parameter of the current time step through the initial time-frequency analysis model;

[0042] S202. Based on the preset frequency range parameters, the reference period parameters, the modulation factor parameters, and the order index parameters, the wavelet transform parameters at each center frequency are obtained.

[0043] S203. Perform convolution processing on the ultrawavelet transform parameters and the phase time series data to obtain wavelet convolution results at each center frequency;

[0044] S204. Aggregate the wavelet convolution results of each order at the same center frequency to obtain the initial time-frequency features of the current time step.

[0045] Understandably, the server obtains the reference period parameter (represented as) for the current time step through the initial time-frequency analysis model. ), modulation factor parameter (represented as ) and the order index parameter (represented as When performing a time step, it is necessary to first determine whether the current time step is the first time step. For example... Figure 3 As shown, when the current time step is the first time step, the server extracts the initial values ​​of the reference period parameter, modulation factor parameter, and order index parameter through the parameter action mapping network in the initial time-frequency analysis model, and determines these initial values ​​as the reference period parameter, modulation factor parameter, and order index parameter for the current time step. When the current time step is not the first time step (i.e., other time steps starting from the second time step), it indicates that the policy optimization network based on policy gradient in the initial time-frequency analysis model output a continuous action vector in the previous time step. At this time, the server obtains the reference period parameter, modulation factor parameter, and order index parameter obtained by mapping the continuous action vector of the previous time step through the parameter action mapping network in the initial time-frequency analysis model, and determines these as the reference period parameter, modulation factor parameter, and order index parameter for the current time step.

[0046] In one specific embodiment, the preset frequency range parameter is within the frequency band. Inner interval Constructed center frequency set ,in, The number of center frequencies (dimensionless) is preferred. , , The initial values ​​of the reference period parameter, modulation factor parameter, and order index parameter refer to the reference period parameter, modulation factor parameter, and order index parameter that are preset during the initialization of the initial time-frequency analysis model. The initial value of the reference period parameter is... The initial value of the modulation factor parameter is The initial value of the order index parameter is The server provides each Construct the period number of this order. This yields a periodic family that varies with order. Generate a superwavelet kernel according to frequency order and perform wavelet transform. For arbitrary frequencies... order Define Gaussian scaling parameters Next, a center frequency of [missing information] is constructed. Rank The complex Morlet wavelet kernel function, i.e., the superwavelet transform parameters at each center frequency. ,in, The imaginary unit is used. Then, the wavelet transform parameters and phase time series data are convolved to obtain the wavelet convolution results at each center frequency. ,in, This represents the convolution operation. Finally, for the same... Logarithmic domain aggregation of the multi-order coefficient amplitudes at each point, traversing all center frequencies, yields a super-resolution time-frequency map. ,in, Indicates the number of orders participating in the aggregation (dimensionless). This indicates that the amplitude is obtained by modulo operation on a complex number. This is a numerical stability constant (dimensionless, usually a very small positive number) used to avoid instability in logarithmic calculations. This represents a super-resolution time-frequency plot. Indicates the current time step The initial time-frequency characteristics, i.e. the super-resolution time-frequency representation output by the time-frequency calculation network based on super wavelet transform.

[0047] This embodiment implements ultrawavelet transform based on multi-parameter synthesis, which can more accurately match the characteristics of signals in different frequency bands. This allows ultrawavelet transform to more effectively extract the feature information of signals in each frequency band, improving the resolution and accuracy of time-frequency analysis. At the same time, convolution operations can highlight the feature information of signals within a specific time and frequency range, suppressing interference from noise and other irrelevant components, thereby improving the clarity and reliability of time-frequency features.

[0048] S30. Determine the joint state vector of the current time step based on the phase time series data and the initial time-frequency characteristics, determine the continuous action vector of the current time step based on the joint state vector, and determine the time-frequency quality score of the current time step based on the initial time-frequency characteristics.

[0049] Understandably, such as Figure 3 As shown, after the server obtains the initial time-frequency features through the time-frequency calculation network in the initial time-frequency analysis model, it analyzes the phase time series data and the initial time-frequency features through the state construction network in the initial time-frequency analysis model to obtain the current time step. The joint state vector (represented as) Of the four components of the joint state vector, This represents the power spectrum eigenvector, reflecting the energy distribution characteristics of the signal in the frequency domain; This represents the envelope statistics vector, reflecting the overall amplitude variation level of the signal; Rényi entropy is used to effectively measure the concentration of time-frequency energy distribution; the smaller the entropy value, the more concentrated the energy. (representing the initial time-frequency characteristics), then the joint state vector is analyzed through the policy gradient-based policy optimization network in the initial time-frequency analysis model to obtain the continuous action vector of the current time step (represented as...). ,in, , representing a three-dimensional action vector from a normalized continuous action space. On the other hand, the initial time-frequency features are analyzed through an external sparse evaluation network in the initial time-frequency analysis model to obtain the time-frequency quality score for the current time step (represented as...). ).

[0050] In one embodiment, step S30, namely determining the joint state vector of the current time step based on the phase time series data and the initial time-frequency characteristics, includes:

[0051] S301. Perform Fourier transform and normalization on the phase time series data to obtain the power spectrum feature vector;

[0052] S302. Perform envelope transformation on the phase time series data to obtain an envelope statistical vector;

[0053] S303. Perform entropy analysis on the initial time-frequency characteristics according to a preset order to obtain the Rényi entropy;

[0054] S304. Generate the joint state vector for the current time step based on the initial time-frequency characteristics, the power spectrum feature vector, the envelope statistical vector, and the Rényi entropy.

[0055] Understandably, in one specific embodiment, such as Figure 3 As shown, the state construction network in the initial time-frequency analysis model uses the signal frequency domain power spectrum, envelope statistics, Rényi entropy, and initial time-frequency features to form a joint observation. It jointly models the statistical features of the time and frequency domains with the super-resolution time-frequency structure information, enabling the subsequent policy optimization network to simultaneously perceive the signal's energy distribution, time-frequency clustering, and structural change characteristics in a unified state space. The server analyzes the phase time series data and initial time-frequency features through the state construction network in the initial time-frequency analysis model to determine the joint state vector for the current time step. First, the server analyzes the phase time series data... After performing a discrete Fourier transform to obtain its frequency domain representation, the power spectrum is calculated. Select the front of the preset frequency band Each frequency component is normalized, and the optimal frequency component is selected. The power spectrum eigenvector is obtained. ,in This is the constant for the numerical stability term (dimensionless, usually taken as a very small positive number). Then, the phase time series data are calculated. Instantaneous amplitude envelope And calculate its statistic as the envelope feature. ,in, This represents the number of discrete sampling times. Next, the initial time-frequency features are... Normalization to probability distribution At the preset order (Preferred) Under the condition of ), calculate its Rényi entropy. ,in, A smaller entropy value indicates a more concentrated energy. Finally, the initial time-frequency features, power spectrum feature vector, envelope statistics vector, and Rényi entropy are concatenated in a preset order to construct a joint state vector. .

[0056] This embodiment uses a multi-dimensional splicing and fusion method to jointly model the statistical features of the time domain, frequency domain, and time-frequency domain with the super-resolution time-frequency structure information to obtain a joint state vector, which provides a data foundation for subsequent model updates. This enables the strategy optimization module to simultaneously perceive the energy distribution, time-frequency aggregation, and structural change characteristics of the signal in a unified state space.

[0057] In one embodiment, step S30, namely determining the time-frequency quality score of the current time step based on the initial time-frequency characteristics, includes:

[0058] S305. Calculate time-frequency quality index values ​​based on the initial time-frequency characteristics. The time-frequency quality index values ​​include time-frequency energy concentration, time resolution, frequency resolution, and ridge sharpness.

[0059] S306. The time-frequency energy concentration, time resolution, frequency resolution and ridge sharpness are weighted and calculated according to preset weight parameters to obtain a time-frequency quality score.

[0060] Understandably, time-frequency quality index values ​​are quantitative values ​​used to evaluate the overall performance of initial time-frequency features in a specified dimension (such as time and frequency). Preset weighting parameters refer to parameters that are pre-set to characterize the different levels of importance of time-frequency energy concentration, temporal resolution, frequency resolution, and ridge sharpness in the calculation of time-frequency quality scores.

[0061] In one specific embodiment, such as Figure 3 As shown, the time-frequency quality index values ​​include time-frequency energy concentration (expressed as...). ), temporal resolution (expressed as ), frequency resolution (expressed as ) and ridge line clarity (expressed as The server analyzes the initial time-frequency characteristics through the external sparse evaluation network in the initial time-frequency analysis model to obtain the time-frequency quality score for the current time step. The external sparse evaluation network receives data at each time step sequentially. Corresponding super-resolution time-frequency plot ,in, Indicates a round The number of time steps included. During round execution, the quality score corresponding to the time-frequency graph is progressively cached. A round refers to the stage where the initial time-frequency analysis model is iteratively updated until the stopping condition is reached; a round includes at least one time step. For any time step within a round... Corresponding initial time-frequency features An external sparse evaluation network is used to calculate the corresponding time-frequency quality index. Specifically, the initial time-frequency characteristics are converted into time-frequency energy concentration using a formula. The initial time-frequency features are converted into time-resolution indicators. The initial time-frequency characteristics are converted into frequency resolution indicators. ,in Calculated for standard deviation. Ridge sharpness is calculated at each time index. At that point, select the frequency index with the highest energy. As a frequency ridge index sequence, and with the reciprocal of the standard deviation of the frequency ridge index sequence defining ridge sharpness. The external sparse evaluation network linearly combines the quality index values ​​according to preset weight parameters to calculate the time-frequency quality score at the current time step. The preset weight parameters are: , , , Time-frequency quality score Used to characterize the overall quality of the super-resolution time-frequency map at the current time step, in rounds During execution, the time-frequency quality scores for all time steps are stored sequentially to form rounds. Internal quality scoring sequence Furthermore, in the current round At the end of the round, The time-frequency quality scores of all time steps within the round are aggregated to obtain the round-level quality score. Round-level quality scores are used to reflect the overall performance of the current round's super-resolution time-frequency analysis results. An external sparse evaluation network calculates the round-level quality score for the current round. Best quality rating compared to previous rounds Compare the results and output the round sparse evaluation value at the end of the round. The round-based sparsity evaluation value is represented as follows: ,when Simultaneously update the best quality rating of previous rounds. .

[0062] This embodiment converts the initial time-frequency features into energy concentration, time-frequency resolution, and ridge sharpness, which can comprehensively evaluate the effect of time-frequency characterization from different perspectives. At the same time, the weighted calculation reflects the importance of different perspectives, ensuring the rationality of the time-frequency quality score.

[0063] S40. Determine the dense internal reward for the current time step based on the joint state vector, the continuous action vector, and the time-frequency quality score.

[0064] Understandably, the server analyzes the joint state vector, continuous action vector, and time-frequency quality score through the reward discovery network in the initial time-frequency analysis model to obtain the current time step. Intensive internal rewards (represented as) This embodiment uses a reward discovery network to jointly model state features, action features, and time-frequency representation features, thereby learning dense internal reward signals associated with "state-action pairs." Dense internal rewards refer to the immediate reward signals generated internally by the system at each time step within the reinforcement learning framework, which are directly related to the learning process.

[0065] In one embodiment, the joint state vector includes an initial time-frequency feature component, a power spectrum feature component, an envelope statistical component, and a Rényi entropy component; step S40, namely determining the dense internal reward of the current time step based on the joint state vector, the continuous action vector, and the time-frequency quality score, includes:

[0066] S401. Perform convolution feature extraction on the initial time-frequency feature components to obtain the first state feature vector;

[0067] S402. Perform fully connected feature extraction on the power spectrum feature components, envelope statistics components and Rényi entropy components to obtain the second state feature vector;

[0068] S403. The first state feature vector and the second state feature vector are fused to obtain a state embedding vector;

[0069] S404. Encode the continuous action vector to obtain the action embedding vector;

[0070] S405. Determine the initial dense internal reward based on the state embedding vector and the action embedding vector;

[0071] S406. Determine the state value based on the state embedding vector, determine the round sparsity evaluation value based on the time-frequency quality score, and correct the initial dense internal reward based on the round sparsity evaluation value and the state value to obtain the dense internal reward.

[0072] Understandably, in one specific embodiment, such as Figure 4 As shown, the reward discovery network in the initial time-frequency analysis model includes an action encoding unit, a state encoding unit, a reward regression unit, and a value evaluation unit. The state encoding unit includes a convolutional branch for extracting time-frequency image features and a multilayer perceptron branch for extracting vector observation features. The action encoding unit is the action encoding branch for extracting action features. The reward regression unit is the reward regression network that fuses the above features and outputs a dense internal reward. The value evaluation unit outputs the state value. The first state feature vector refers to the result of feature extraction from the initial time-frequency feature components in the joint state vector. The second state feature vector refers to the result of feature extraction from the power spectrum feature components, envelope statistics components, and Rényi entropy components in the joint state vector.

[0073] First, the server uses the convolutional branch in the state coding unit to... Initial time-frequency characteristic components Convolutional feature extraction is performed to obtain the first state feature vector. On the other hand, the multilayer perceptron branch in the state coding unit is used to... The power spectrum feature components, envelope statistics components, and Rényi entropy components in the first state are extracted using a fully connected feature extraction method to obtain the second state feature vector. The state coding unit then fuses and encodes the first and second state feature vectors to obtain the state embedding vector. .in, Represents a state-coded network. This represents the state embedding vector. Then, the server uses the action encoding unit to process the continuous action vectors. Encoding is performed to obtain action embedding vectors. .in, This represents the action encoding network. Next, the server uses a reward regression unit to fuse the state embedding vector and the action embedding vector, and outputs the initial dense internal reward. .in, This is a non-linear regression mapping used to map fused features to a scalar reward output. . Able to perform different actions in the same state The advantages and disadvantages of each state result in a differentiated output. Finally, the server analyzes the state embedding vector through a value evaluation unit and outputs the state value. .in, This is a value evaluation mapping used to provide a baseline for state value and training constraint signals. The reward discovery network utilizes the sparse evaluation values ​​at the end of each epoch. Construct a training supervision mechanism to enable value evaluation units and reward regression units to update collaboratively, thereby increasing the initial dense internal reward. It can characterize the trend of how actions contribute to the improvement of round-level quality. Specifically, the supervisory objective for constructing the value assessment unit is... And by minimizing value deviation This enables the updating of value assessment units. Among them, This is a time-step representation of the round-based sparse evaluation value, typically set to 0 within the round and to [value missing] at the end of the round. ; This is the discount factor; The terminator is the stop marker; stopgrad indicates the control loss gradient, L value This represents the loss value of the value assessment unit. Simultaneously, an advantage monitoring quantity induced by round-based sparse evaluation values ​​is constructed. Furthermore, the reward regression unit is trained under dominance constraints to output discriminative dense internal rewards, thereby enabling actions with different parameters to produce distinguishable reward responses under the same state. The training objective of the reward regression unit can be expressed as: L reward This represents the loss value of the reward regression unit. Through the training method described above, the reward discovery network can achieve high performance in only a few rounds with minimal evaluation. Under these conditions, a dense internal reward function with fine-grained discriminative ability for evaluating the quality of parameterized actions is learned. After training or during training, the reward regression unit in the reward discovery network modifies the initial dense internal reward at each time step and outputs the dense internal reward. .

[0074] This embodiment fuses the first and second state feature vectors into a state embedding vector, integrating state information from different sources or dimensions, reducing data dimensionality, making subsequent processing more efficient, and reducing computational resource consumption. Encoding converts continuous action vectors into action embedding vectors, enabling the model to better process action information and improve training performance and generalization ability. This embodiment uses a reward discovery network to jointly model state features, action features, and time-frequency representation features, thereby learning dense internal reward signals associated with state-action pairs. Therefore, it can continuously characterize the impact of actions with different parameters on the quality of time-frequency representations within a round, effectively alleviating the credit allocation difficulty caused by relying solely on round-level sparse feedback.

[0075] S50. The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and the target time-frequency analysis model based on reinforcement learning is obtained after the iterative update is completed.

[0076] Understandably, during the iterative update of the initial time-frequency analysis model on the server side, one time step corresponds to one reinforcement learning iteration update. Iteration stops when the number of iterations reaches a preset threshold, completing one round of model iteration update. After the iteration update is complete, the target time-frequency analysis model based on reinforcement learning is obtained. One round includes one or more time steps. The target time-frequency analysis model refers to the neural network model used to convert the phase time series into a super-resolution time-frequency representation after the iteration update. Specifically, during the reinforcement learning iteration update process, it can proceed with only one round of iteration update or, as needed, multiple rounds of iteration update. When multiple rounds of iteration update are required, a threshold for the number of rounds of iteration update can be preset. After completing one round of model iteration update, and when the number of rounds reaches the threshold, the iteration update is confirmed to be complete.

[0077] Specifically, such as Figure 3 As shown, the server uses the dense internal reward of the current time step as a reward signal to iteratively update the initial time-frequency analysis model for the current time step, obtaining the updated time-frequency analysis model for the current time step. It also accumulates the iteration count for the current time step (incrementing by 1) and determines whether the iteration count has reached a preset threshold. If the iteration count reaches the preset threshold, the iterative update is confirmed to be complete, and the target time-frequency analysis model based on reinforcement learning is obtained. If the iteration count has not reached the preset threshold, the server returns the steps of performing ultrawavelet transform processing on the phase time series data to obtain the initial time-frequency features for the next time step, and then performs the iterative update for the next time step. In the next time step, the server obtains the continuous action vector output by the policy optimization network based on policy gradient in the updated time-frequency analysis model at the current time step. Through the parametric action mapping network, it obtains the reference period parameters, modulation factor parameters, and order index parameters for the next time step after mapping this continuous action vector. Based on these parameters, the server performs a wavelet transform on the phase time series data to obtain the initial time-frequency features for the next time step, and further obtains the dense internal reward for the next time step. The dense internal reward of the next time step is used as the reward signal to iteratively update the updated time-frequency analysis model for the current time step, resulting in the updated time-frequency analysis model for the next time step. The number of iterations for the next time step is accumulated (incrementing by 1) until the number of iterations reaches a preset threshold, at which point the iterative update is confirmed to be complete, and the target time-frequency analysis model based on reinforcement learning is obtained.

[0078] In one specific embodiment, such as Figure 5As shown, the server iteratively updates the reward discovery network and the policy optimization network based on policy gradient in the initial time-frequency analysis model. After the iterative update, a target time-frequency analysis model based on reinforcement learning is obtained, and this model is used to perform super-resolution time-frequency analysis on subsequent thoracic radar echo signals. The policy optimization network based on policy gradient adopts an actor-critic structure with a shared feature extraction backbone. The image input is processed by a convolutional network to extract image features, which are concatenated with vector features to form a joint feature representation, and then input to the actor output action and the critic output state value, respectively. The policy optimization network uses dense internal rewards as the reward signal to iteratively update the model parameters. The update method is implemented using probabilistic policy optimization methods, such as Proximal Policy Optimization (PPO). The target time-frequency analysis model is a model updated by Deep Reinforcement Learning (DRL) after updating the initial time-frequency analysis model. It iteratively updates the state value function through dynamic programming to find the optimal policy, demonstrating generalization and robustness in parameter optimization under complex dynamic environments.

[0079] S60. Perform ultrawavelet transform processing on the phase time series data through the target time-frequency analysis model to obtain the target time-frequency characteristics, and determine the health monitoring results based on the target time-frequency characteristics.

[0080] Understandably, the target time-frequency feature refers to the super-resolution time-frequency representation obtained by transforming the phase time series through an iteratively optimized time-frequency analysis model, i.e., an adaptive super-resolution time-frequency representation. The server extracts one or more parameters from the target time-frequency feature, including heart rate, heart interval, respiratory rate, and heart rate variability, as a physiological feature vector. Then, the physiological feature vector is input into a pre-defined health assessment model, which outputs a health level and an abnormal risk score. Finally, based on the health level, abnormal risk score, and warning information, monitoring prompts, warning messages, and intervention suggestions are generated. Health monitoring results are then generated based on these prompts, warning messages, and intervention suggestions and sent to the client of a pre-defined notification recipient.

[0081] In one embodiment, step S60, namely, performing ultrawavelet transform processing on the phase time series data using the target time-frequency analysis model to obtain target time-frequency features, includes:

[0082] S601. Extract the target continuous action vector from the target time-frequency analysis model. The target continuous action vector includes a first action component, a second action component, and a third action component.

[0083] S602. The first action component is mapped according to the preset upper limit value and the preset lower limit value of the cycle to obtain the target reference cycle parameter.

[0084] S603. The second action component is mapped according to the preset modulation upper limit value and the preset modulation lower limit value to obtain the target modulation factor parameter.

[0085] S604. The third action component is mapped according to the preset upper limit value and the preset lower limit value to obtain the target order index parameter;

[0086] S605. Determine the target time-frequency characteristics based on the preset frequency range parameters, the target reference period parameters, the target modulation factor parameters, and the target order index parameters.

[0087] Understandably, the server extracts the target continuous action vector from the target time-frequency analysis model. This target continuous action vector is obtained by the policy optimization network in the target time-frequency analysis model processing the joint state vector from the previous time step. The target continuous action vector includes a first action component (represented as...). ), the second action component (represented as ) and the third action component (represented as There are three dimensions of action vector components.

[0088] In one specific embodiment, the server-side parameter action mapping network is based on a preset period upper limit value. and preset cycle lower limit value For the first action component The target reference period parameters are obtained by performing mapping processing. The mapping method is... Among them, preferred , ; This represents a truncation function used to restrict the output parameters to a valid range. The server-side parameter action mapping network is based on a preset modulation upper limit. and preset modulation lower limit value For the second action component The target modulation factor parameters are obtained by performing mapping processing. The mapping method is... Among them, preferred , The server-side parameter action mapping network is based on a preset upper limit value. and preset order lower limit value For the third action component Perform mapping processing to obtain the order center index. The mapping method is... ,in, This indicates the floor function. Next, index by the order center. Centered on a set of multi-order indices, we construct a set of indices as parameters for the target order index. ,in The preset inter-order offset is typically set to 1. Finally, the server-side time-frequency calculation network obtains the target time-frequency features based on the preset frequency interval parameters, target reference period parameters, target modulation factor parameters, and target order index parameters, through super wavelet transform, wavelet convolution processing, and aggregation processing.

[0089] This embodiment achieves online adaptive adjustment of wavelet parameters by constructing a parameter action mapping mechanism, thereby steadily improving the quality of time-frequency representation. This embodiment dynamically maps the continuous action vector output by the strategy optimization network to key parameters such as the period number benchmark, scale control factor, and order center of the wavelet transform. This allows the wavelet parameters to be updated in real time according to changes in radar phase signal characteristics, avoiding the time-frequency energy diffusion, sidelobe enhancement, and ridge breakage problems caused by fixed parameters in traditional methods. This improves the clustering and recognizability of weak physiological components such as heartbeats in the time-frequency domain.

[0090] In one embodiment, the health monitoring results include heart rate variability monitoring results; step S60, i.e., the above, includes:

[0091] S606. Identify the main ridge of cardiac time frequency based on the target time frequency characteristics, and determine the time interval between adjacent cardiac events based on the main ridge of cardiac time frequency to obtain the cardiac interval sequence;

[0092] S607. Calculate the heart rate variability index value based on the heart rate interval sequence, and determine the heart rate variability monitoring result based on the heart rate variability index value.

[0093] Understandably, heart rate variability (HRV) refers to the minute variations in the time intervals between consecutive heartbeats, reflecting the regulatory function of the cardiac autonomic nervous system. It is an important indicator for assessing cardiovascular health, stress levels, and physical recovery. When the health monitoring result is HRV monitoring, the server identifies the main ridge line of the heartbeat time-frequency pattern based on the target time-frequency characteristics, and determines the time interval between adjacent heartbeat events based on the main ridge line, obtaining a heartbeat interval sequence, and thus determining the HRV monitoring result. Specifically, after locating heartbeat events through the ridge line, the server calculates the time coordinate difference between adjacent events (i.e., the RR interval) for heart rate variability analysis or anomaly detection. For example, the SDNN (standard deviation of all RR intervals) is calculated to assess the overall change in HRV over 24 hours or a specified duration; a larger value indicates a stronger autonomic nervous system regulatory capacity.

[0094] This embodiment acquires chest radar echo signals, preprocesses them to obtain phase time series data, and then performs wavelet transform on the phase time series data using an initial time-frequency analysis model to obtain the initial time-frequency features of the current time step. Based on the phase time series data and the initial time-frequency features, the joint state vector of the current time step is determined, the continuous action vector of the current time step is determined based on the joint state vector, and the time-frequency quality score of the current time step is determined based on the initial time-frequency features. Based on the joint state vector, the continuous action vector, and the time-frequency quality score, the dense internal reward of the current time step is determined. The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and after the iterative update, a target time-frequency analysis model based on reinforcement learning is obtained. The target time-frequency analysis model is then used to perform wavelet transform on the phase time series data to obtain target time-frequency features, and the health monitoring results are determined based on these target time-frequency features. This embodiment first obtains initial time-frequency features through an initial time-frequency analysis model. Based on the joint state vector, continuous action vector, and time-frequency quality score, dense internal rewards are determined. Under the constraints of the time-frequency representation quality evaluation mechanism, a progressively feedbackable internal reward signal is learned, demonstrating the strategy optimization of reinforcement learning and achieving adaptive optimization of key wavelet parameters in the time-frequency analysis model. This embodiment avoids the problems of insufficient resolution, low energy concentration, and susceptibility of the main ridge line to noise and stray components in the time-frequency representation of heartbeat-related data caused by fixed parameters. It improves the quality and robustness of the time-frequency representation and ensures the effectiveness of health status assessment under complex monitoring conditions.

[0095] In one specific embodiment, such as Figure 6 As shown, the top left, bottom left, and top right spectra are time-frequency representations using traditional methods such as short-time Fourier transform, while the bottom right spectra are time-frequency representations using the present invention. It can be seen that as the window length increases from 128 to 512, the STFT exhibits a typical trade-off in time-frequency resolution, with insufficient frequency resolution under short window conditions, making it difficult to form a stable and clear main ridge line for the cardiac frequency energy. The adaptive method of the present invention forms a narrower, more continuous, and more clearly defined high-energy main ridge line in the cardiac frequency band, significantly suppressing background energy rise, sidelobe leakage, and texture pseudostructure. Therefore, the present invention can improve frequency domain focusing ability while maintaining the discernibility of local temporal variations, thereby more accurately characterizing the time-frequency evolution of the cardiac component in millimeter-wave radar phase signals.

[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0097] In one embodiment, a radar-based health monitoring device is provided, which corresponds one-to-one with the radar-based health monitoring method described in the above embodiments. For example... Figure 7 As shown, the radar-based health monitoring device includes a preprocessing module 10, an initial model analysis module 20, a feature analysis module 30, a reward analysis module 40, a model update module 50, and a monitoring result determination module 60. Detailed descriptions of each functional module are as follows:

[0098] Preprocessing module 10 is used to acquire the thoracic radar echo signal, preprocess the thoracic radar echo signal to obtain phase time series data;

[0099] The initial model analysis module 20 is used to perform wavelet transform processing on the phase time series data through the initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; the time step refers to the number of iterations to update the initial time-frequency analysis model.

[0100] The feature analysis module 30 is used to determine the joint state vector of the current time step based on the phase time series data and the initial time-frequency features, determine the continuous action vector of the current time step based on the joint state vector, and determine the time-frequency quality score of the current time step based on the initial time-frequency features.

[0101] Reward analysis module 40 is used to determine the dense internal reward of the current time step based on the joint state vector, the continuous action vector and the time-frequency quality score;

[0102] The model update module 50 is used to iteratively update the initial time-frequency analysis model using the dense internal reward as a reward signal, and obtain the target time-frequency analysis model based on reinforcement learning after the iterative update is completed.

[0103] The monitoring result determination module 60 is used to perform ultrawavelet transform processing on the phase time series data through the target time-frequency analysis model to obtain the target time-frequency characteristics, and determine the health monitoring result based on the target time-frequency characteristics.

[0104] In one embodiment, the initial model analysis module 20 includes:

[0105] The initial parameter acquisition unit is used to acquire the reference period parameter, modulation factor parameter, and order index parameter of the current time step through the initial time-frequency analysis model.

[0106] The wavelet transform parameter determination unit is used to obtain the wavelet transform parameters at each center frequency based on the preset frequency interval parameters, the reference period parameters, the modulation factor parameters, and the order index parameters.

[0107] The convolution processing unit is used to perform convolution processing on the ultrawavelet transform parameters and the phase time series data to obtain wavelet convolution results at each center frequency.

[0108] The aggregation processing unit is used to aggregate the wavelet convolution results of different orders at the same center frequency to obtain the initial time-frequency features of the current time step.

[0109] In one embodiment, the feature analysis module 30 includes:

[0110] The power spectrum feature vector determination unit is used to perform Fourier transform and normalization processing on the phase time series data to obtain the power spectrum feature vector;

[0111] An envelope statistical vector determination unit is used to perform envelope transformation processing on the phase time series data to obtain an envelope statistical vector;

[0112] The entropy analysis unit is used to perform entropy analysis on the initial time-frequency characteristics according to a preset order to obtain the Rényi entropy.

[0113] The state vector generation unit is used to generate a joint state vector for the current time step based on the initial time-frequency features, the power spectrum feature vector, the envelope statistical vector, and the Rényi entropy.

[0114] In one embodiment, the feature analysis module 30 further includes:

[0115] A quality index calculation unit is used to calculate time-frequency quality index values ​​based on the initial time-frequency characteristics. The time-frequency quality index values ​​include time-frequency energy concentration, time resolution, frequency resolution, and ridge sharpness.

[0116] The weighted calculation unit is used to perform weighted calculations on the time-frequency energy concentration, time resolution, frequency resolution and ridge sharpness according to preset weight parameters to obtain a time-frequency quality score.

[0117] In one embodiment, the reward analysis module 40 includes:

[0118] A convolutional feature extraction unit is used to perform convolutional feature extraction on the initial time-frequency feature components to obtain a first state feature vector;

[0119] A fully connected feature extraction unit is used to perform fully connected feature extraction on the power spectrum feature components, envelope statistics components and Rényi entropy components to obtain the second state feature vector.

[0120] A state vector fusion unit is used to fuse the first state feature vector and the second state feature vector to obtain a state embedding vector;

[0121] An action vector processing unit is used to encode the continuous action vectors to obtain action embedding vectors;

[0122] An initial reward determination unit is used to determine an initial dense internal reward based on the state embedding vector and the action embedding vector;

[0123] The reward determination unit is used to determine the state value based on the state embedding vector, determine the round sparse evaluation value based on the time-frequency quality score, and correct the initial dense internal reward based on the round sparse evaluation value and the state value to obtain the dense internal reward.

[0124] In one embodiment, the monitoring result determination module 60 includes:

[0125] The target action vector extraction unit is used to extract the target continuous action vector from the target time-frequency analysis model. The target continuous action vector includes a first action component, a second action component, and a third action component.

[0126] The first mapping processing unit is used to map the first action component according to the preset upper limit value and the preset lower limit value of the period to obtain the target reference period parameter.

[0127] The second mapping processing unit is used to map the second action component according to the preset modulation upper limit value and the preset modulation lower limit value to obtain the target modulation factor parameter;

[0128] The third mapping processing unit is used to map the third action component according to the preset upper limit value and the preset lower limit value of the order to obtain the target order index parameter;

[0129] The target time-frequency characteristic determination unit is used to determine the target time-frequency characteristics based on the preset frequency range parameters, the target reference period parameters, the target modulation factor parameters, and the target order index parameters.

[0130] In one embodiment, the monitoring result determination module 60 further includes:

[0131] The cardiac interval sequence determination unit is used to identify the main ridge of cardiac time frequency based on the target time frequency characteristics, and to determine the time interval between adjacent cardiac events based on the main ridge of cardiac time frequency, thereby obtaining the cardiac interval sequence.

[0132] The heart rate variability monitoring unit is used to calculate the heart rate variability index value based on the heart rate interval sequence, and to determine the heart rate variability monitoring result based on the heart rate variability index value.

[0133] Specific limitations regarding radar-based health monitoring devices can be found in the limitations of radar-based health monitoring methods described above, and will not be repeated here. Each module in the aforementioned radar-based health monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0134] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes a readable storage medium and internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database stores data related to the radar-based health monitoring method. The network interface communicates with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement a radar-based health monitoring method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0135] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor performs the following steps when executing the computer-readable instructions:

[0136] Acquire the thoracic radar echo signal, and preprocess the thoracic radar echo signal to obtain phase time series data;

[0137] The phase time series data is processed by wavelet transform using an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations that update the initial time-frequency analysis model.

[0138] The joint state vector of the current time step is determined based on the phase time series data and the initial time-frequency characteristics; the continuous action vector of the current time step is determined based on the joint state vector; and the time-frequency quality score of the current time step is determined based on the initial time-frequency characteristics.

[0139] The dense internal reward for the current time step is determined based on the joint state vector, the continuous action vector, and the time-frequency quality score.

[0140] The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and after the iterative update is completed, the target time-frequency analysis model based on reinforcement learning is obtained.

[0141] The phase time series data are processed by the target time-frequency analysis model to obtain the target time-frequency characteristics, and the health monitoring results are determined based on the target time-frequency characteristics.

[0142] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which, when executed by one or more processors, perform the following steps:

[0143] Acquire the thoracic radar echo signal, and preprocess the thoracic radar echo signal to obtain phase time series data;

[0144] The phase time series data is processed by wavelet transform using an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations that update the initial time-frequency analysis model.

[0145] The joint state vector of the current time step is determined based on the phase time series data and the initial time-frequency characteristics; the continuous action vector of the current time step is determined based on the joint state vector; and the time-frequency quality score of the current time step is determined based on the initial time-frequency characteristics.

[0146] The dense internal reward for the current time step is determined based on the joint state vector, the continuous action vector, and the time-frequency quality score.

[0147] The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and after the iterative update is completed, the target time-frequency analysis model based on reinforcement learning is obtained.

[0148] The phase time series data are processed by the target time-frequency analysis model to obtain the target time-frequency characteristics, and the health monitoring results are determined based on the target time-frequency characteristics.

[0149] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0151] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A radar-based health monitoring method, characterized in that, include: Acquire the thoracic radar echo signal, and preprocess the thoracic radar echo signal to obtain phase time series data; The phase time series data is processed by super wavelet transform using an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations that update the initial time-frequency analysis model. The joint state vector of the current time step is determined based on the phase time series data and the initial time-frequency characteristics; the continuous action vector of the current time step is determined based on the joint state vector; and the time-frequency quality score of the current time step is determined based on the initial time-frequency characteristics. The dense internal reward for the current time step is determined based on the joint state vector, the continuous action vector, and the time-frequency quality score. The dense internal reward is used as a reward signal to iteratively update the initial time-frequency analysis model, and after the iterative update is completed, the target time-frequency analysis model based on reinforcement learning is obtained. The phase time series data are processed by the target time-frequency analysis model to obtain the target time-frequency characteristics, and the health monitoring results are determined based on the target time-frequency characteristics. The step of performing ultrawavelet transform processing on the phase time series data through an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step includes: The reference period parameters, modulation factor parameters, and order index parameters for the current time step are obtained through the initial time-frequency analysis model. Based on the preset frequency range parameters, the reference period parameters, the modulation factor parameters, and the order index parameters, the wavelet transform parameters at each center frequency are obtained. The wavelet transform parameters and the phase time series data are convolved to obtain the wavelet convolution results at each center frequency; The wavelet convolution results of different orders at the same center frequency are aggregated to obtain the initial time-frequency features of the current time step; The health monitoring results include heart rate variability monitoring results; determining the health monitoring results based on the target time-frequency characteristics includes: The main ridge of cardiac time-frequency is identified based on the target time-frequency characteristics, and the time interval between adjacent cardiac events is determined based on the main ridge of cardiac time-frequency, thus obtaining the cardiac interval sequence. The heart rate variability index value is calculated based on the heart rate interval sequence, and the heart rate variability monitoring result is determined based on the heart rate variability index value.

2. The radar-based health monitoring method as described in claim 1, characterized in that, Determining the joint state vector for the current time step based on the phase time series data and the initial time-frequency characteristics includes: The phase time series data is subjected to Fourier transform and normalization to obtain the power spectrum feature vector; The phase time series data is subjected to envelope transformation to obtain an envelope statistical vector; The initial time-frequency characteristics are analyzed for entropy based on a preset order to obtain the Rényi entropy; Based on the initial time-frequency features, the power spectrum feature vector, the envelope statistics vector, and the Rényi entropy, a joint state vector for the current time step is generated.

3. The radar-based health monitoring method as described in claim 1, characterized in that, The step of determining the time-frequency quality score for the current time step based on the initial time-frequency characteristics includes: The time-frequency quality index value is calculated based on the initial time-frequency characteristics. The time-frequency quality index value includes time-frequency energy concentration, time resolution, frequency resolution, and ridge sharpness. The time-frequency energy concentration, time resolution, frequency resolution, and ridge sharpness are weighted and calculated according to preset weight parameters to obtain a time-frequency quality score.

4. The radar-based health monitoring method as described in claim 1, characterized in that, The joint state vector includes initial time-frequency feature components, power spectrum feature components, envelope statistical components, and Rényi entropy components; The step of determining the dense internal reward for the current time step based on the joint state vector, the continuous action vector, and the time-frequency quality score includes: Convolutional feature extraction is performed on the initial time-frequency feature components to obtain the first state feature vector; Fully connected feature extraction is performed on the power spectrum feature components, envelope statistics components, and Rényi entropy components to obtain the second state feature vector; The first state feature vector and the second state feature vector are fused to obtain a state embedding vector; The continuous action vectors are encoded to obtain action embedding vectors; The initial dense internal reward is determined based on the state embedding vector and the action embedding vector; The state value is determined based on the state embedding vector, the round sparsity evaluation value is determined based on the time-frequency quality score, and the initial dense internal reward is corrected based on the round sparsity evaluation value and the state value to obtain the dense internal reward.

5. The radar-based health monitoring method as described in claim 1, characterized in that, The step of performing ultrawavelet transform processing on the phase time series data through the target time-frequency analysis model to obtain the target time-frequency features includes: The target continuous action vector is extracted from the target time-frequency analysis model. The target continuous action vector includes a first action component, a second action component, and a third action component. The first action component is mapped according to the preset upper limit value and the preset lower limit value of the cycle to obtain the target reference cycle parameter; The second action component is mapped according to the preset modulation upper limit and preset modulation lower limit to obtain the target modulation factor parameter; The third action component is mapped according to the preset upper limit and lower limit of the order to obtain the target order index parameter; The target time-frequency characteristics are determined based on the preset frequency range parameters, the target reference period parameters, the target modulation factor parameters, and the target order index parameters.

6. A radar-based health monitoring device, characterized in that, include: The preprocessing module is used to acquire the thoracic radar echo signal, preprocess the thoracic radar echo signal, and obtain phase time series data. The initial model analysis module is used to perform wavelet transform processing on the phase time series data through an initial time-frequency analysis model to obtain the initial time-frequency characteristics of the current time step; a time step refers to the number of iterations in which the initial time-frequency analysis model is updated. The feature analysis module is used to determine the joint state vector of the current time step based on the phase time series data and the initial time-frequency features, determine the continuous action vector of the current time step based on the joint state vector, and determine the time-frequency quality score of the current time step based on the initial time-frequency features. The reward analysis module is used to determine the dense internal reward of the current time step based on the joint state vector, the continuous action vector, and the time-frequency quality score. The model update module is used to iteratively update the initial time-frequency analysis model using the dense internal reward as a reward signal, and obtain the target time-frequency analysis model based on reinforcement learning after the iterative update is completed. The monitoring result determination module is used to perform ultrawavelet transform processing on the phase time series data through the target time-frequency analysis model to obtain the target time-frequency characteristics, and determine the health monitoring result based on the target time-frequency characteristics; The initial model analysis module includes: The initial parameter acquisition unit is used to acquire the reference period parameter, modulation factor parameter, and order index parameter of the current time step through the initial time-frequency analysis model. The wavelet transform parameter determination unit is used to obtain the wavelet transform parameters at each center frequency based on the preset frequency interval parameters, the reference period parameters, the modulation factor parameters, and the order index parameters. The convolution processing unit is used to perform convolution processing on the ultrawavelet transform parameters and the phase time series data to obtain wavelet convolution results at each center frequency. The aggregation processing unit is used to aggregate the wavelet convolution results of different orders at the same center frequency to obtain the initial time-frequency features of the current time step. The monitoring result determination module includes: The cardiac interval sequence determination unit is used to identify the main ridge of cardiac time frequency based on the target time frequency characteristics, and to determine the time interval between adjacent cardiac events based on the main ridge of cardiac time frequency, thereby obtaining the cardiac interval sequence. The heart rate variability monitoring unit is used to calculate the heart rate variability index value based on the heart rate interval sequence, and to determine the heart rate variability monitoring result based on the heart rate variability index value.

7. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the radar-based health monitoring method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by one or more processors, the one or more processors cause the radar-based health monitoring method as described in any one of claims 1 to 5 to be performed.

Citation Information

Patent Citations

  • Millimeter wave radar breath and heart rate synchronous monitoring method and system

    CN120713487A

  • High-robustness non-contact accurate electrocardiogram monitoring method based on millimeter wave radar

    CN121337367A