A multimodal human physiological abnormality recognition method and device

Through the combination of the YOLOv11 model and the extended long and short-term memory network, the problem of low accuracy and insufficient timeliness of human physiological abnormalities in the prior art is solved, and accurate and timely recognition of multimodal human physiological abnormalities is achieved.

CN120071227BActive Publication Date: 2025-08-08XINXING JIHUA TECHNOLOGY (TIANJIN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510552646.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art lacks a comprehensive analysis of a variety of physiological parameters, postures and behaviors in the recognition of human physiological abnormalities, resulting in low recognition accuracy and difficult timely warning.

Method used

The multimodal human physiological abnormality recognition method is used to detect abnormal postures and physiological image features in the video stream through the YOLOv11 model, and combine the one-dimensional convolution layer and the extended long and short-term memory network modeling timing dependency to generate an alarm signal to identify abnormal states.

Benefits of technology

It realizes accurate and timely recognition of human physiological abnormalities, integrates a variety of physiological image data in the video stream and the timing data of physiological monitors, and improves the accuracy and timeliness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071227B_ABST
    Figure CN120071227B_ABST
Patent Text Reader

Abstract

The present application proposes a multimodal human physiological abnormality recognition method and device. The method comprises: first, obtaining human posture, behavior, physiological image data and time series data of a physiological monitor in a video stream. Then, using the YOLOv11 model, based on human posture, behavior and physiological image data, the target area of abnormal posture, behavior or physiological image features is detected, and the bounding box coordinates and category confidence of these areas are output. Subsequently, the bounding box coordinates and category confidence are convolved and pooled to generate target detection time series feature data. The target detection time series feature data is then combined with the physiological monitoring time series data, key features are extracted through a one-dimensional convolution layer and a pooling layer, and the time series dependency is modeled using an extended long short-term memory network. Finally, based on the key features and time series dependency, a physiological abnormality alarm is triggered or a normal state is displayed to achieve accurate and timely recognition of abnormal physiological states of the human body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of physiological abnormality recognition, and in particular to a multimodal human physiological abnormality recognition method and device. Background Art

[0002] Identifying abnormal physiological states primarily relies on monitoring a single physiological parameter, such as heart rate or blood pressure, or through manual observation by a doctor. However, this approach has limitations, as physiological abnormalities are often accompanied by complex changes in multiple physiological parameters, as well as abnormal posture and behavior. Furthermore, traditional recognition methods lack in-depth analysis and modeling of abnormal physiological states over time, resulting in low recognition accuracy and difficulty in providing timely warnings.

[0003] With the development of computer vision and deep learning technologies, significant progress has been made in image and video-based human behavior analysis. However, applying these technologies to the identification of abnormal physiological states still faces many challenges. For example, how to effectively extract posture, behavior, and physiological image features related to physiological anomalies from video streams, how to perform time series analysis on these features to model their dependencies, and how to accurately trigger alarm signals for abnormal physiological states. Summary of the Invention

[0004] The purpose of this application is to overcome the above-mentioned defects in the prior art and provide a multimodal human physiological abnormality identification method and device.

[0005] This application provides a multimodal human physiological abnormality recognition method, comprising:

[0006] Acquire human posture images, behavior images, physiological image data and time series data of physiological monitors in video streams;

[0007] Based on the human posture image, behavior image and physiological image data, detect the target area of abnormal posture, behavior or physiological image features through the YOLOv11 model, and output the bounding box coordinates and category confidence of the target area;

[0008] Performing convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data;

[0009] For the target detection time series feature data and physiological monitoring time series data, key features are extracted through a one-dimensional convolution layer and a pooling layer, and temporal dependency is modeled through an extended long short-term memory network;

[0010] According to the key features and the timing dependency, an abnormal physiological state alarm signal is triggered or a normal state signal is displayed to realize the recognition of abnormal physiological state of the human body.

[0011] Optionally, triggering a physiological abnormality alarm signal or displaying a normal state signal to realize the recognition of the abnormal physiological state of the human body according to the key features and the time sequence dependency relationship includes:

[0012] When the key features include any two of abnormal heart rate, abnormal respiratory rate, and abnormal blood pressure, and the time dependency lasts for more than 5 seconds, a level 1 alarm signal is triggered;

[0013] When the key features only include abnormal postures or behaviors and the temporal dependency lasts for more than 10 seconds, a secondary warning signal is triggered.

[0014] Optionally, before obtaining the human posture image, behavior image, physiological image data and time series data of the physiological monitor in the video stream, the following steps are further included:

[0015] YUV color space conversion is performed on the original image data in the video stream, contrast-limited adaptive histogram equalization is implemented in the Y channel, bilinear interpolation noise reduction processing is used in the UV channel, and enhanced RGB format image data is output as the human posture image, behavior image and physiological image data.

[0016] Optionally, it also includes:

[0017] Remove convolution kernels with output channel contributions lower than 0.3 in the YOLOv11 model;

[0018] Convert the weight parameters of the extended long short-term memory network from FP32 to INT8 format;

[0019] Use ONNX Runtime tools to encapsulate the pruned and quantized model into a binary inference engine suitable for the ARM architecture.

[0020] Optionally, performing convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data includes:

[0021] Arrange the bounding box coordinates of the target area in timestamp order to generate an original time series sequence, and perform the following processing on the original time series sequence:

[0022] Convert the center coordinates, width, and height of the bounding box to proportional values relative to the image size;

[0023] For sequences shorter than the preset number of frames, zero values are filled to a fixed length, and for sequences longer than the preset number of frames, the end data is truncated;

[0024] The sampling rates of the target detection time series feature data and the physiological monitoring time series data are unified to the same frequency through linear interpolation.

[0025] The present application provides a multimodal human physiological abnormality recognition device, comprising:

[0026] An acquisition module, which acquires human posture images, behavior images, physiological image data and time series data of physiological monitors in the video stream;

[0027] A detection module, based on the human posture image, behavior image and physiological image data, detects target areas with abnormal posture, behavior or physiological image features through the YOLOv11 model, and outputs the bounding box coordinates and category confidence of the target area;

[0028] A feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data;

[0029] An extraction module extracts key features from the target detection time series feature data and the physiological monitoring time series data through a one-dimensional convolution layer and a pooling layer, and models the time series dependency through an extended long short-term memory network;

[0030] The alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency relationship to realize the recognition of abnormal physiological state of the human body.

[0031] Optionally, the alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency to realize the recognition of abnormal physiological state of the human body, including:

[0032] When the key features include any two of abnormal heart rate, abnormal respiratory rate, and abnormal blood pressure, and the time dependency lasts for more than 5 seconds, a level 1 alarm signal is triggered;

[0033] When the key features only include abnormal postures or behaviors and the temporal dependency lasts for more than 10 seconds, a secondary warning signal is triggered.

[0034] Optionally, the acquisition module further includes:

[0035] YUV color space conversion is performed on the original image data in the video stream, contrast-limited adaptive histogram equalization is implemented in the Y channel, bilinear interpolation noise reduction processing is used in the UV channel, and enhanced RGB format image data is output as the human posture image, behavior image and physiological image data.

[0036] Optionally, it also includes: a deployment module that removes convolution kernels with output channel contributions lower than 0.3 in the YOLOv11 model; converts the weight parameters of the extended long short-term memory network from FP32 to INT8 format; and uses the ONNX Runtime tool to encapsulate the pruned and quantized model into a binary inference engine suitable for the ARM architecture.

[0037] Optionally, the feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data, including:

[0038] Arrange the bounding box coordinates of the target area in timestamp order to generate an original time series sequence, and perform the following processing on the original time series sequence:

[0039] Convert the center coordinates, width, and height of the bounding box to proportional values relative to the image size;

[0040] For sequences shorter than the preset number of frames, zero values are filled to a fixed length, and for sequences longer than the preset number of frames, the end data is truncated;

[0041] The sampling rates of the target detection time series feature data and the physiological monitoring time series data are unified to the same frequency through linear interpolation.

[0042] The beneficial effects of this application are:

[0043] The present application provides a multimodal human physiological abnormality recognition method, comprising: obtaining human posture images, behavioral images, physiological image data and time series data of a physiological monitor in a video stream; based on the human posture images, behavioral images and physiological image data, detecting a target area of abnormal posture, behavior or physiological image features through a YOLOv11 model, and outputting the bounding box coordinates and category confidence of the target area; performing convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data; extracting key features from the target detection time series feature data and physiological monitoring time series data through a one-dimensional convolution layer and a pooling layer, and modeling time series dependencies through an extended long short-term memory network; based on the key features and time series dependencies, triggering a physiological abnormality alarm signal or displaying a normal state signal to realize the recognition of human physiological abnormality. The present application realizes accurate and timely recognition of human physiological abnormality by fusing multiple physiological image data and time series data of a physiological monitor in a video stream, detecting abnormal features using a YOLOv11 model, and modeling time series dependencies through an extended long short-term memory network. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a schematic diagram of the multimodal human physiological abnormality recognition process of this application. DETAILED DESCRIPTION

[0045] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that various forms of implementation of the present disclosure are not limited to the embodiments set forth herein. Rather, the embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0046] This application is implemented using computers with high-performance computing capabilities to ensure computing power, and high-performance server clusters can also be used.

[0047] To ensure that this application is applicable to more application scenarios and has better compatibility with the selected YOLOv11, the deep learning framework used is PyTorch1.9, which is compatible with YOLOv11 and not lower than PyTorch1.8.

[0048] The software languages are Python 3.10 and C++, and Python 3.9 and above can meet the requirements. The graphics card is NVIDIA's GEFORCE RTX 40 series GPU with 32GB of memory. The installed library is Ultralytics. The operating system is Ubuntu 20, and Windows or macOS can also be used. The specific system will be determined by the system used by the device where the algorithm is finally deployed.

[0049] like Figure 1 As shown, the present application provides a multimodal human physiological abnormality recognition method, comprising:

[0050] S101, obtaining human posture images, behavior images, physiological image data and time series data of a physiological monitor in a video stream;

[0051] We have collected various image datasets related to human posture, behavior, and physiology, including abnormal heart rate, abnormal breathing rate, abnormal blood pressure, long-term immobility, falls, abnormal facial and walking, and chest abnormalities, and each data is saved in the corresponding category mentioned above.

[0052] S102, based on the human posture image, behavior image and physiological image data, detecting a target area of abnormal posture, behavior or physiological image features through the YOLOv11 model, and outputting the bounding box coordinates and category confidence of the target area;

[0053] The YOLOv11 model training steps are as follows:

[0054] First, we collected image datasets related to human posture, behavior, and physiology, including various abnormal heart rates, abnormal breathing rates, abnormal blood pressure, long-term immobility, falls, facial and walking abnormalities, and chest abnormalities, from public databases on the Internet containing various complex scenes and specialized camera acquisition equipment of units. Each data set is saved in the corresponding category mentioned above.

[0055] Secondly, outliers, missing values, and conversion features are processed to improve the quality of the dataset and make it more suitable for model training, validation, and testing operations.

[0056] To meet the high-quality data requirements of the YOLOv11 model, after completing dataset preparation, we first examined each dataset type to determine whether it contained images showing normal conditions, such as normal heart rate, normal respiratory rate, normal blood pressure, normal facial, gait, and chest, as well as images showing prolonged movement and no falls or falls. We also examined whether these conditions were present, whether there were any poor quality images, any duplication, and whether the images were all in RGB format. We removed any such poor quality or non-RGB color images, converted them to RGB, and enhanced them using a YUV color space histogram equalization algorithm to enhance image clarity, preparing for dataset annotation.

[0057] Again, use image annotation software such as LabelImg to add text labels and rectangular boxes or bounding boxes to the key content in the preprocessed dataset images to improve the recognition accuracy, recognition speed and practicality of the algorithm model.

[0058] Among them, text labels are used for target detection tasks, and rectangular boxes or bounding boxes are used for semantic segmentation, instance segmentation, etc.

[0059] In the labeling of the data set, choose to use the software installed LabelImg to perform corresponding labeling on each type of preprocessed data set and save the final labeling file in XML format.

[0060] According to the txt format annotation file required by the YOLOv11 model, the generated xml format annotation file needs to be converted into a txt format annotation file.

[0061] Divide the dataset into training and test sets, or training, validation, and test sets to further improve the model's generalization capabilities and evaluate and optimize the performance of the algorithm. The training and test sets are used for performance evaluation, while the training, validation, and test sets are used for both evaluation and tuning and selecting the optimal model structure.

[0062] YOLOv11 was evaluated, tuned, and the optimal model structure was selected to reduce model complexity. Therefore, various annotated files, such as heart rate, respiratory rate, blood pressure, prolonged immobility, face and walking, and chest, were divided into training, validation, and test sets at a ratio of 70%, 15%, and 15%, respectively.

[0063] In the YOLOv11 model training phase, first load the pre-trained model, and then use the training script provided by the YOLOv11 model to specify parameters such as the training set, image size, and learning rate. Then, start the cyclic training for each type of training set.

[0064] Finally, to reduce model overfitting, we adjusted the corresponding model parameters by adjusting the learning rate, optimizer, number of training rounds, and other parameters to improve the anomaly detection rate. We used the validation set for each category to determine whether the corresponding YOLOv11 model had converged. This strategy served as a basis for determining whether the performance of each category model met the well-trained standard; ultimately, the corresponding model converged.

[0065] If the model does not converge, it will continue to train within the adjusted number of training rounds until it converges.

[0066] Furthermore, the performance indicators when loading the test set are evaluated, and the accuracy of identifying anomalies is used to evaluate the trained various YOLOv11 models.

[0067] If overfitting occurs at this stage, meaning that anomaly recognition accuracy is poor, data augmentation (increasing the training set) and L2 normalization can be used to improve model accuracy and address overfitting. Overfitting occurs because similar data from the test set is grouped into the validation set.

[0068] To adapt to the subsequent processing of the xLSTM model, when the recognition accuracy of each type of anomaly meets expectations, the category confidence of each data and the center point coordinates, width and height of the bounding box, the relative position of key targets in the image, and other feature information such as the abnormal posture, behavior and physiological image features are obtained.

[0069] S103, performing convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data;

[0070] Convert the feature information of the test set output on the model into a data format suitable for xLSTM processing.

[0071] The feature information of each data in various categories is arranged into a time series in chronological order and then normalized so that the xLSTM network can better process the data.

[0072] At the same time, since the time series length of some data is shorter than that of most other data, the sequence length of these data is padded; otherwise, the sequence length of these data is truncated to ensure that it is consistent with the length of most shorter data.

[0073] S104, extracting key features from the target detection time series feature data and the physiological monitoring time series data through a one-dimensional convolution layer and a pooling layer, and modeling the time series dependency through an extended long short-term memory network;

[0074] Through CNN partial convolution and pooling operations, local features in time series data are extracted and redundant data is removed.

[0075] Through multiple convolutional and pooling layers of the CNN part, multiple reshaping and convolution operations are performed on each time series data and each raw time series data obtained from the physiological monitor. The final output contains a four-dimensional feature map of abnormalities in heart rate, respiratory rate, blood pressure, long-term immobility, facial and walking, chest abnormalities, etc.

[0076] Reshaping means changing the shape of the obtained feature map data through the reshaping operation and converting it into a two-dimensional sequence input suitable for processing as an xLSTM model.

[0077] The four-dimensional feature maps obtained above are reshaped into a three-dimensional sequence consisting of three parts: the number of sequences in the image sequence in the batch (sequence 1), the number of feature maps (sequence 2), and the product of the height, width and number of channels of the feature map (sequence 3).

[0078] S105 : triggering a physiological abnormality alarm signal or displaying a normal state signal according to the key features and the time sequence dependency to realize recognition of the physiological abnormality of the human body.

[0079] xLSTM network modeling:

[0080] The modeling process mainly includes four parts: data set division, training, debugging, and evaluation, namely:

[0081] First, the three-dimensional sequence data were divided into training set, validation set and test set at a ratio of 70%, 15% and 15% respectively;

[0082] By setting parameters such as the number of training rounds and learning rate in the xLSTM model, we began to continuously train each type of data and identify whether there were any abnormal physiological states. Then, by adjusting parameters such as the learning rate, optimizer, and number of training rounds in the model, we adjusted the xLSTM model to improve the recognition rate and used the validation set of each type to determine whether the model had converged, which served as the standard for when to stop training.

[0083] After detecting abnormal physiological states in the training set, the test set needs to be used to evaluate and select the best network architecture.

[0084] Using the test set, test and evaluate the best xLSTM model performance:

[0085] Through evaluation, it was determined that the test set can stably and accurately identify abnormal physiological states of the human body caused by irregular heart rate, irregular breathing rate, abnormal blood pressure, long-term immobility, abnormal facial and walking posture, and chest abnormalities, triggering abnormal physiological state alarm signals or displaying normal state signals to realize the recognition of abnormal physiological states of the human body.

[0086] Integrate the best-performing model algorithm into the actual application environment (such as a server or embedded device) for real-time or batch processing. To meet the low memory and application requirements of the target device, perform the following operations on the deployed model before deployment:

[0087] Pruning, removing unimportant weights or connections in the model to reduce model complexity;

[0088] Quantization, converting the floating-point weights and activation values in the model into low-precision integer form;

[0089] Model conversion: If the target deployment environment is not suitable for the framework used during training, appropriate tools must be used to convert the best-performing model into a format that can be executed in the actual application environment. For the PyTorch framework used in this application, tools such as TorchScript or ONNX can be used to convert the evaluated YOLOv11 and CNN-xSLTM models into a format suitable for execution in the actual application environment.

[0090] Select an embedded device or server and install the cross-compilation toolchain:

[0091] Choosing an embedded device or server mainly means ensuring that it has sufficient computing resources and storage space to support the operation of the model based on the assessed scale;

[0092] Installing a cross-compilation toolchain is mainly to compile the model and deep learning framework into binary files that can be run in the actual application environment. If the target device is a server or a high-performance environment, cross-compilation is not required; instead, you need to install the deep learning framework PyTorch and its corresponding dependencies.

[0093] At this point, you can decide to run the converted model in a corresponding manner based on the specific application scenario and requirements.

[0094] For example: 1) In scenarios requiring real-time or batch data processing, the task of identifying abnormal physiological states can be performed by sequentially loading the YOLOv11 and CNN-xSLTM models and calling the corresponding inference interfaces to execute the inference service. 2) In restricted application environments, the task of identifying abnormal physiological states can be performed by directly running the converted YOLOv11 and CNN-xSLTM models in sequence. The CNN-xSLTM model is a hybrid time series modeling architecture that combines the aforementioned one-dimensional convolutional neural network (CNN) with the extended long short-term memory network (xLSTM).

[0095] The present application provides a multimodal human physiological abnormality recognition device, comprising:

[0096] An acquisition module, which acquires human posture images, behavior images, physiological image data and time series data of physiological monitors in the video stream;

[0097] A detection module, based on the human posture image, behavior image and physiological image data, detects target areas with abnormal posture, behavior or physiological image features through the YOLOv11 model, and outputs the bounding box coordinates and category confidence of the target area;

[0098] A feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data;

[0099] An extraction module extracts key features from the target detection time series feature data and the physiological monitoring time series data through a one-dimensional convolution layer and a pooling layer, and models the time series dependency through an extended long short-term memory network;

[0100] The alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency relationship to realize the recognition of abnormal physiological state of the human body.

[0101] Furthermore, the alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency to realize the recognition of abnormal physiological state of the human body, including:

[0102] When the key features include any two of abnormal heart rate, abnormal respiratory rate, and abnormal blood pressure, and the time dependency lasts for more than 5 seconds, a level 1 alarm signal is triggered;

[0103] When the key features only include abnormal postures or behaviors and the temporal dependency lasts for more than 10 seconds, a secondary warning signal is triggered.

[0104] Furthermore, the acquisition module further includes:

[0105] YUV color space conversion is performed on the original image data in the video stream, contrast-limited adaptive histogram equalization is implemented in the Y channel, bilinear interpolation noise reduction processing is used in the UV channel, and enhanced RGB format image data is output as the human posture image, behavior image and physiological image data.

[0106] Furthermore, it also includes: a deployment module to remove convolution kernels with output channel contributions lower than 0.3 in the YOLOv11 model; converting the weight parameters of the extended long short-term memory network from FP32 to INT8 format; and using the ONNXRuntime tool to encapsulate the pruned and quantized model into a binary inference engine suitable for the ARM architecture.

[0107] Furthermore, the feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data, including:

[0108] Arrange the bounding box coordinates of the target area in timestamp order to generate an original time series sequence, and perform the following processing on the original time series sequence:

[0109] Convert the center coordinates, width, and height of the bounding box to proportional values relative to the image size;

[0110] For sequences shorter than the preset number of frames, zero values are filled to a fixed length, and for sequences longer than the preset number of frames, the end data is truncated;

[0111] The sampling rates of the target detection time series feature data and the physiological monitoring time series data are unified to the same frequency through linear interpolation.

[0112] The above description of the embodiments is intended to facilitate understanding and application of the present invention by those skilled in the art. It will be readily apparent to those skilled in the art that various modifications to the above embodiments can be made, and the general principles described herein can be applied to other embodiments without requiring inventive effort. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art based on the present disclosure are intended to fall within the scope of protection of the present invention.

Claims

1. A multimodal human physiological abnormality recognition method, characterized in that: include: Acquire human posture images, behavior images, physiological image data and time series data of physiological monitors in video streams; Based on the human posture image, behavior image and physiological image data, detect the target area of abnormal posture, behavior or physiological image features through the YOLOv11 model, and output the bounding box coordinates and category confidence of the target area; Performing convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data; For the target detection time series feature data and physiological monitoring time series data, key features are extracted through a one-dimensional convolution layer and a pooling layer, and temporal dependency is modeled through an extended long short-term memory network; According to the key features and the timing dependency, an abnormal physiological state alarm signal is triggered or a normal state signal is displayed to realize the recognition of abnormal physiological state of the human body.

2. A multimodal human physiological abnormality recognition method according to claim 1, characterized in that: According to the key features and the timing dependency, triggering a physiological abnormality alarm signal or displaying a normal state signal to realize the recognition of the abnormal physiological state of the human body includes: When the key features include any two of abnormal heart rate, abnormal respiratory rate, and abnormal blood pressure, and the time dependency lasts for more than 5 seconds, a level 1 alarm signal is triggered; When the key features only include abnormal postures or behaviors and the temporal dependency lasts for more than 10 seconds, a secondary warning signal is triggered.

3. A multimodal human physiological abnormality recognition method according to claim 1, characterized in that: Before obtaining the human posture image, behavior image, physiological image data and time series data of the physiological monitor in the video stream, the following steps are also required: YUV color space conversion is performed on the original image data in the video stream, contrast-limited adaptive histogram equalization is implemented in the Y channel, bilinear interpolation noise reduction processing is used in the UV channel, and enhanced RGB format image data is output as the human posture image, behavior image and physiological image data.

4. A multimodal human physiological abnormality recognition method according to claim 1, characterized in that: Also includes: Remove convolution kernels with output channel contributions lower than 0.3 in the YOLOv11 model; Convert the weight parameters of the extended long short-term memory network from FP32 to INT8 format; Use ONNX Runtime tools to encapsulate the pruned and quantized model into a binary inference engine suitable for the ARM architecture.

5. The multimodal human physiological abnormality recognition method according to claim 1, characterized in that: Perform convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data, including: Arrange the bounding box coordinates of the target area in timestamp order to generate an original time series sequence, and perform the following processing on the original time series sequence: Convert the center coordinates, width, and height of the bounding box to proportional values relative to the image size; For sequences shorter than the preset number of frames, zero values are filled to a fixed length, and for sequences longer than the preset number of frames, the end data is truncated; The sampling rates of the target detection time series feature data and the physiological monitoring time series data are unified to the same frequency through linear interpolation.

6. A multimodal human physiological abnormality recognition device, characterized in that: include: An acquisition module, which acquires human posture images, behavior images, physiological image data and time series data of physiological monitors in the video stream; A detection module, based on the human posture image, behavior image and physiological image data, detects target areas with abnormal posture, behavior or physiological image features through the YOLOv11 model, and outputs the bounding box coordinates and category confidence of the target area; A feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data; An extraction module extracts key features from the target detection time series feature data and the physiological monitoring time series data through a one-dimensional convolution layer and a pooling layer, and models the time series dependency through an extended long short-term memory network; The alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency relationship to realize the recognition of abnormal physiological state of the human body.

7. A multimodal human physiological abnormality recognition device according to claim 6, characterized in that: The alarm module triggers an abnormal physiological state alarm signal or displays a normal state signal according to the key features and the timing dependency to realize the recognition of abnormal physiological state of the human body, including: When the key features include any two of abnormal heart rate, abnormal respiratory rate, and abnormal blood pressure, and the time dependency lasts for more than 5 seconds, a level 1 alarm signal is triggered; When the key features only include abnormal postures or behaviors and the temporal dependency lasts for more than 10 seconds, a secondary warning signal is triggered.

8. The multimodal human physiological abnormality recognition device according to claim 6, characterized in that: The acquisition module also includes: YUV color space conversion is performed on the original image data in the video stream, contrast-limited adaptive histogram equalization is implemented in the Y channel, bilinear interpolation noise reduction processing is used in the UV channel, and enhanced RGB format image data is output as the human posture image, behavior image and physiological image data.

9. The multimodal human physiological abnormality recognition device according to claim 6, characterized in that: Also includes: Deploy the module to remove convolution kernels with output channel contributions lower than 0.3 in the YOLOv11 model; The weight parameters of the extended long short-term memory network are converted from FP32 to INT8 format; the pruned and quantized model is encapsulated into a binary inference engine suitable for the ARM architecture using the ONNX Runtime tool.

10. The multimodal human physiological abnormality recognition device according to claim 6, characterized in that: The feature module performs convolution and pooling processing on the bounding box coordinates and category confidence to generate target detection time series feature data, including: Arrange the bounding box coordinates of the target area in timestamp order to generate an original time series sequence, and perform the following processing on the original time series sequence: Convert the center coordinates, width, and height of the bounding box to proportional values relative to the image size; For sequences shorter than the preset number of frames, zero values are filled to a fixed length, and for sequences longer than the preset number of frames, the end data is truncated; The sampling rates of the target detection time series feature data and the physiological monitoring time series data are unified to the same frequency through linear interpolation.

Citation Information

Patent Citations

  • Cooperative voice gesture generation method and system based on deep learning

    CN119577686A

  • Campus abnormal behavior detection method and device and electronic equipment

    CN119720096A