Method and system for identifying target in emergency care unit based on Internet of Things

Through the YOLOv7 and DeepSORT algorithm combined with Kalman filtering, depth camera and RFID verification target recognition system, the problem of inaccurate target recognition and untimely emergency response in the emergency monitoring room is solved, high-precision target tracking and abnormal detection are achieved, and the efficiency and safety of the emergency monitoring room are improved.

CN120339923AActive Publication Date: 2025-07-18THE FIRST AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE

Patent Information

Application Number
CN202510821849.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The environment of the emergency monitoring room is complex, and traditional monitoring systems are difficult to distinguish between medical staff and patients. Especially when blocking and moving quickly, it is easy to cause the target to be lost or misidentified. Excessive calculation load affects the response speed and accuracy. Difficult to fusion of multimodal data leads to untimely emergency response.

Method used

The YOLOv7 model was used to combine the DeepSORT algorithm and Kalman filtering to track the target through a four-dimensional space-time-semantic synergistic cost matrix, combine the depth camera and RFID verification to build a causal map of medical behavior, and use 3D CNN to analyze patient action sequences and physiological data, and introduce cross-frame attention mechanisms and occlusion density feature fusion strategy.

Benefits of technology

It improves the accuracy of object detection and tracking in complex scenarios, ensures the stability and real-time response capabilities of key targets, realizes intelligent judgment and timely intervention in patients' abnormal behaviors, and improves the efficiency and safety of emergency monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339923A_ABST
    Figure CN120339923A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target identification, in particular to an emergency care unit target identification method and system based on the Internet of Things. The method comprises the following steps: collecting a video stream in an emergency care unit at a fixed frame rate through a camera, and extracting an image frame; performing real-time target detection on each frame in the image frames by adopting a YOLOv7 model, identifying a target patient, associating targets in continuous frames of the video stream through a DeepSORT algorithm, tracking the position of the target patient in combination with Kalman filtering, and constructing a four-dimensional space-time-semantic collaborative cost matrix for optimization in the process of tracking the position of the target patient; and analyzing the action sequence of the target patient in the video stream through the 3D CNN, and judging abnormity in combination with the physiological data of the target patient. According to the design of the invention, a cross-frame attention mechanism and a feature fusion strategy guided by shielding density are introduced, and the detection capability of a moving target in a complex scene is enhanced in a YOLOv7 model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and specifically, to a method and system for target recognition in an emergency intensive care unit based on the Internet of Things. Background Art

[0002] The emergency intensive care unit is an environment with rapid dynamic changes, dense personnel and equipment. In such an environment, traditional monitoring systems are difficult to effectively distinguish medical staff from patients. Especially when there are frequent human body occlusions and rapid movements, it is easy to cause target loss or misidentification; in the emergency scenario, the requirement for the speed of event response is extremely high, and any delay may lead to serious consequences; single video stream analysis may not provide sufficient information to comprehensively understand the patient's status; in a high-density target scenario, in order to ensure the stable operation of the system, the effective utilization of computing resources needs to be considered. Traditional methods may cause excessive computing load due to the simultaneous presence of a large number of targets, affecting the response speed and accuracy of the system. Therefore, a method and system for target recognition in an emergency intensive care unit based on the Internet of Things are provided. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for target recognition in an emergency intensive care unit based on the Internet of Things, so as to solve the problems of inaccurate target recognition, unstable tracking and untimely emergency response caused by the complex environment, high real-time requirement and difficult multi-modal data fusion in the emergency intensive care unit as mentioned in the above background art.

[0004] To achieve the above purpose, the present invention aims to provide a method for target recognition in an emergency intensive care unit based on the Internet of Things, including the following steps: S1. Collect the video stream in the emergency intensive care unit at a fixed frame rate through a camera; S2. Preprocess the collected video stream and extract image frames; S3. Use the YOLOv7 model to perform real-time target detection on each frame in the image frame, identify the target patient, associate the targets in consecutive frames of the video stream through the DeepSORT algorithm, and combine the Kalman filter to track the position of the target patient. During the process of tracking the position of the target patient, a four-dimensional spatio-temporal-semantic collaborative cost matrix is constructed for optimization; S4. Analyze the action sequence of the target patient in the video stream through 3D CNN and judge abnormalities in combination with the physiological data of the target patient.

[0005] As a further improvement of this technical solution, in S3, using the YOLOv7 model to perform real-time target detection on each frame in the image frame and identify the target patient includes the following steps: S3.1. Input the image frame into the YOLOv7 model for forward propagation, and output the position, confidence, and class score of each prediction box for the target patient; S3.2. Use non-maximum suppression to remove overlapping prediction boxes and filter out targets with a confidence higher than the threshold a; S3.3. Among all the prediction boxes retained after screening, further analyze the class information of each prediction box, identify the target patient, and output the final detection result list.

[0006] As a further improvement of this technical solution, in S3.1, inputting the image frame into the YOLOv7 model for forward propagation and outputting the position, confidence, and class score of each prediction box for the target patient includes the following steps: S3.11. In the Backbone of the YOLOv7 model, for each residual block, retain the feature maps corresponding to the past 3 frames during forward propagation, construct a feature cache pool, perform a dot product operation on the feature map of the current frame and the historical frame feature maps to obtain a similarity matrix, and use attention weights to perform weighted summation on the historical features; S3.12. Obtain the depth information in the emergency intensive care unit through a depth camera and calculate the occlusion degree of the local area based on the depth map; S3.13. Use the PANet structure of the YOLOv7 model for multi-scale feature fusion, input the multi-scale fused features into the detection head of the YOLOv7 model, output a prediction tensor containing the bounding box offset, target confidence, and class score, use the anchor box mechanism to decode the offset into the actual box position in the image coordinate space, and calculate the final confidence value in combination with the target confidence and class score.

[0007] As a further improvement of this technical solution, in S3.3, among all the prediction boxes retained after screening, further analyze the class information of each prediction box and identify the target patient, including the following steps: S3.31. Traverse each prediction box and extract the class information; S3.32. Use a predefined class mapping table to map the class ID to a semantic label; S3.33. Filter out the prediction boxes with the class of patient; S3.34. Number the identified target patients and assign a unique ID to each identified patient; S3.35. Output the list of identified target patients as the final detection result list.

[0008] As a further improvement of this technical solution, in S3, associate the targets in consecutive frames of the video stream through the DeepSORT algorithm and track the position of the target patient in combination with Kalman filtering, including the following steps: S3.4. Input the detection box coordinates output by YOLOv7 into the DeepSORT model. For each target detected by the YOLOv7 model in each frame, crop the bounding box area of the target and use the ReID sub-network to extract the appearance ReID embedding feature vector of the target. S3.5. Initialize a Kalman filter for each detected target to predict its state vector in the next frame. S3.6. Using the state estimation obtained by Kalman filter prediction, combined with the cosine distance between the appearance ReID embedding feature vectors, construct a four-dimensional spatio-temporal-semantic collaborative cost matrix. S3.7. Based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-temporal-semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph. S3.8. Output the tracking results with the target patient ID.

[0009] As a further improvement of this technical solution, in S3.6, using the state estimation obtained by Kalman filter prediction, combined with the cosine distance between the ReID embedding vectors, to construct a four-dimensional spatio-temporal-semantic collaborative cost matrix, including the following steps: S3.61. Generate a hybrid ReID feature by combining the texture features of the patient's gown and the device connection status, calculate the cosine similarity between the current frame and the target in the historical trajectory, and set the initial weight to b. S3.62. Calculate the Mahalanobis distance between the targets through the Kalman filter prediction results. S3.63. Check whether the target detected in the current frame carries a specific medical device and compare it with the device connection status in the historical trajectory. If the same device is detected in both the current frame and the historical trajectory, increase the confidence of this matching pair. S3.64. Combining the RFID wristband positioning signal, if the Euclidean distance between the center of the detection box and the RFID coordinates is less than 0.5 meters, increase the confidence of this matching pair. S3.65. When the acceleration of the patient in consecutive frames is greater than the threshold x, reduce the appearance weight and increase the motion weight. S3.66. Synthesize the results of the above steps to construct a four-dimensional cost matrix. The rows in the cost matrix represent each detection box in the current frame, and the columns in the cost matrix represent each tracking trajectory in the previous frame.

[0010] As a further improvement of this technical solution, in S3.7, based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-temporal-semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph, including the following steps: S3.71. Construct a causal graph of medical behaviors in the emergency care unit to real-time analyze the action intentions of medical staff; S3.72. For the target patient carrying first-aid equipment, construct an independent sub-cost matrix; S3.73. Based on the conventional Hungarian algorithm, perform joint verification of the spatial position, device status, and vital signs between a detection box in the current frame and a tracking trajectory in the previous frame; S3.74. Dynamically set the matching threshold according to the regional congestion degree; S3.75. Output the final matching result with priority control.

[0011] As a further improvement of this technical solution, in S4, analyze the action sequence of the target patient in the video stream through 3D CNN and combine the physiological data of the target patient to judge abnormalities, including the following steps: S4.1. According to the tracking ID of the target patient, extract a 3D spatio-temporal cube of a continuous time window from the video stream, and synchronously collect the physiological data of the target patient during the corresponding time period; S4.2. Use the 3D spatio-temporal cube as the input of the 3D CNN model, extract video features in the spatial and temporal dimensions through the convolutional layer of the 3D CNN model, and output the action feature vector of the target patient; S4.3. Use continuous wavelet transform to convert the physiological data of the target patient into a time-frequency diagram, use the time-frequency diagram as the input of the 2D CNN model, extract frequency domain features, and introduce a causal mask; S4.4. Use the cross-attention mechanism to calculate the joint feature vector of the action feature and frequency domain feature of the target patient, and perform anomaly classification through the MLP classifier based on the joint feature vector to trigger the alarm mechanism.

[0012] As a further improvement of this technical solution, in S4.4, use the cross-attention mechanism to calculate the joint feature vector of the action feature and frequency domain feature of the target patient, including the following steps: S4.41. Use cosine similarity to measure the cosine value of the angle between the action feature vector and frequency domain feature vector of the target patient to calculate the similarity score; S4.42. Apply the Softmax function to the similarity score to convert the similarity score into a probability distribution; S4.43. According to the similarity score, perform weighted summation on the action feature vector and frequency domain feature vector to form a new joint feature vector.

[0013] On the other hand, the present invention provides an object recognition system in an emergency care unit based on the Internet of Things, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the above-mentioned object recognition method in the emergency care unit based on the Internet of Things.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. In the object recognition method and system in the emergency care unit based on the Internet of Things, by introducing a cross-frame attention mechanism and a feature fusion strategy guided by occlusion density, the detection ability of moving objects (especially patients) in complex scenarios is enhanced in the YOLOv7 model, effectively alleviating the problems of missed detection and false detection of objects caused by factors such as occlusion by medical staff and changes in light. At the same time, by combining the spatial information obtained by the depth camera and the medical device status verification mechanism, a four-dimensional spatio-temporal-semantic collaborative cost matrix is constructed, and high-precision object tracking is achieved through an improved DeepSORT algorithm. In addition, by using a medical behavior causal graph for matching pre-filtering, the logical reasoning ability and anti-interference ability of the tracking system are further improved, ensuring the ID continuity and tracking stability of key objects (such as critically ill patients) even in dynamic scenarios such as high density or first aid.

[0015] 2. In the object recognition method and system in the emergency care unit based on the Internet of Things, the temporal sequence features of the patient's actions are extracted by 3D CNN, and the time-frequency features generated by continuous wavelet transform of physiological data (heart rate, blood pressure, blood oxygen) are combined. The cross-attention mechanism is used to fuse and analyze multi-modal information, enabling a more comprehensive understanding of the patient's health status. In particular, the introduction of a causal mask mechanism can not only prevent the leakage of future information but also dynamically adjust the historical dependence length according to the patient's status, thereby enhancing the system's response ability to emergencies. Finally, by setting an association weight threshold to trigger the alarm mechanism, real-time perception and intelligent judgment of the patient's abnormal behaviors (such as falls and violent struggles) and sudden changes in physiological indicators are achieved, providing a reliable basis for timely clinical intervention and improving the efficiency and safety of emergency care. Description of the Drawings

[0016] Figure 1 It is the overall method flowchart of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Example 1: Please refer to Figure 1 As shown, this example provides a target recognition method in the emergency care unit based on the Internet of Things, including the following steps: S1. Collect the video stream in the emergency care unit through a camera at a fixed frame rate (≥30fps). Use a high-definition network camera, deploy multiple monitoring nodes inside the emergency care unit, set the fixed frame rate (≥30fps) to ensure the continuity and accuracy of action capture, and transmit the video data to the edge computing device or cloud server in real time through the local area network; S2. Preprocess the collected video stream (including image enhancement of the original video stream, such as denoising, contrast adjustment, and light compensation), and extract image frames (each frame is input into the YOLOv7 model for analysis as an independent image); S3. Use the YOLOv7 model to perform real-time target detection on each frame in the image frames, identify the target patient, associate the targets in consecutive frames of the video stream through the DeepSORT algorithm, and combine the Kalman filter to track the position of the target patient; In this example, the YOLOv7 model consists of an optimized backbone network (CSPDarknet), a cross-frame attention module, a feature pyramid network, a path aggregation network, a detection head, and an anchor box mechanism, aiming to achieve efficient and accurate target detection; Using the YOLOv7 model to perform real-time target detection on each frame in the image frames and identify the target patient includes the following steps: S3.1. Input the image frame into the YOLOv7 model for forward propagation, and output the position, confidence, and class score of each predicted box of the target patient; Among them, this solution is a feature enhancement solution jointly optimized by spatio-temporal attention and occlusion density, that is: the cross-frame attention mechanism uses temporal information to enhance the current frame features, solves the problem that single-frame detection is sensitive to instantaneous occlusion, and combines the occlusion density of the depth sensor to guide feature fusion; the feature enhancement solution jointly optimized by spatio-temporal attention and occlusion density effectively solves problems such as dense targets, frequent occlusions, and complex semantics. This solution integrates a local occlusion density perception mechanism and a global spatio-temporal attention mechanism, can dynamically perceive the occlusion intensity between targets, and perform specific compensation on the occluded area during the feature extraction process to enhance the expression ability of the target boundary and detail features. At the same time, the cross-frame spatio-temporal attention strengthens the identity consistency modeling of the target between consecutive frames, enabling the model to have strong robustness and re-identification ability for short-term occlusions and targets reappearing after occlusion. In addition, this solution can also effectively model the behavioral semantics and dynamic change features of the target, helping to distinguish the micro-action differences between medical staff and patients, and improving the accuracy of abnormal state recognition. Overall, this feature enhancement mechanism significantly improves the stability, accuracy, and practicality of the multi-target recognition and tracking system in complex medical scenarios; Input the image frame into the YOLOv7 model for forward propagation, and output the position, confidence, and class score of each predicted bounding box for the target patient, including the following steps: S3.11. In the Backbone of the YOLOv7 model, for each residual block, during forward propagation, retain the feature maps corresponding to the past 3 frames (C3 / C4 / C5 layers), construct a feature cache pool (whenever a new image is processed, add the feature maps corresponding to each selected layer to the corresponding cache pool), perform a dot product operation on the feature map of the current frame and the historical frame feature maps to obtain a similarity matrix (this operation is to find the correlation or consistency between feature maps of different frames), and use attention weights to perform weighted summation on the historical features to enhance the key feature representation of moving targets in the current frame. Specifically: In each residual block of the YOLOv7 backbone network (CSPDarknet), introduce a cross-frame attention module (CFAM); S3.12. Obtain the depth information in the emergency care unit through a depth camera that works synchronously with the RGB camera, calculate the occlusion degree of the local area based on the depth map (based on neighborhood depth differences or gradient changes). For high-occlusion areas (interference caused by multiple medical staff gathering), retain low-level detail features to maintain positioning accuracy, while for low-occlusion areas, emphasize high-semantic-level features to enhance classification ability; S3.13. Use the PANet structure of the YOLOv7 model for multi-scale feature fusion (P3, P4, P5) to enhance the detection performance of small and large targets. Input the multi-scale fused features into the detection head of the YOLOv7 model (the detection head is a collection of one or more network layers that receive feature maps from the backbone network (Backbone) and the feature pyramid network (FPN)), output a prediction tensor containing bounding box offsets, object confidences, and class scores, use the anchor box mechanism to decode the offsets into the actual box positions in the image coordinate space, and calculate the final confidence value by combining the object confidence and class score in order to obtain high-quality and accurate prediction bounding boxes; S3.2. Use non-maximum suppression to remove overlapping prediction bounding boxes and filter out targets with a confidence higher than the threshold a; S3.3. Among all the prediction bounding boxes retained after screening, further analyze the class information of each prediction bounding box, identify the target patient, and output the final detection result list; Among them, among all the prediction bounding boxes retained after screening, further analyze the class information of each prediction bounding box to identify the target patient, including the following steps: S3.31. Traverse each prediction bounding box and extract the class information; S3.32. Use a predefined class mapping table to map the class ID to a semantic label; S3.33. Screen out the prediction boxes of the category of patients, set a conditional judgment, and only keep the prediction boxes of the category of patients; S3.34. Number the identified target patients, assign a unique ID to each identified patient for subsequent tracking and behavior analysis; S3.35. Output the list of identified target patients as the final detection result list.

[0019] Furthermore, associate the targets in consecutive frames of the video stream through the DeepSORT algorithm and combine Kalman filtering to track the positions of target patients, including the following steps: S3.4. Input the coordinates of the detection boxes output by YOLOv7 into the DeepSORT model. For each target detected by the YOLOv7 model in each frame, crop the bounding box area of the target and use the ReID sub-network (improved MobileNet or ResNet-18) to extract the appearance ReID embedding feature vector of the target. This vector has good discrimination and is used to measure the appearance similarity between different targets; S3.5. Initialize a Kalman filter for each detected target to predict its state vector in the next frame; S3.6. Use the state estimation obtained by Kalman filtering prediction and combine the cosine distance between the appearance ReID embedding feature vectors to construct a four-dimensional spatio-temporal-semantic collaborative cost matrix for target matching; Among them, this solution breaks through the traditional two-dimensional (appearance + motion) cost calculation and introduces the device connection status and spatial verification; in complex and high-density emergency care scenarios, it is easy for targets to be occluded, crossed, aggregated, etc. The traditional two-dimensional (appearance + motion) matching mechanism is difficult to handle these challenges. By constructing a four-dimensional spatio-temporal-semantic collaborative cost matrix and introducing more dimensional information (device status, spatial verification, etc.), the accuracy of target association is significantly improved. The four-dimensional spatio-temporal-semantic collaborative cost matrix solves the problem of ID jumps caused by target occlusion and aggregation in the emergency scenario by fusing four dimensions: spatio-temporal (Kalman prediction), appearance (ReID), device status (semantics), and spatial verification (RFID); Use the state estimation obtained by Kalman filtering prediction and combine the cosine distance between the ReID embedding vectors to construct a four-dimensional spatio-temporal-semantic collaborative cost matrix, including the following steps: S3.61. Generate a hybrid ReID feature by combining the texture features of the patient's hospital gown and the device connection status (whether there is an infusion tube or an electrocardiogram monitor), calculate the cosine similarity between the current frame and the targets in the historical trajectory, and set the initial weight to b (0.5 in this embodiment); S3.62. Calculate the Mahalanobis distance between targets based on the Kalman filter prediction results (the state vector of the Kalman filter includes the position and velocity of the target); S3.63. Check whether the targets detected in the current frame carry specific medical devices (infusion tubes), and compare with the device connection status in the historical trajectory. If the same device is detected in both the current frame and the historical trajectory, increase the confidence of this matching pair, with the weight +0.2; S3.64. Combine the RFID wristband positioning signal. If the Euclidean distance between the center of the detection box and the RFID coordinates is less than 0.5 meters, increase the confidence of this matching pair (if the condition is met, the confidence is multiplied by 1.2); S3.65. When it is detected that the acceleration of the patient in consecutive frames is greater than the threshold x (i.e., the acceleration when the target patient falls), reduce the appearance weight to 0.3 and increase the motion weight to 0.5, and preferentially rely on motion continuity to complete target association in emergency situations; S3.66. Synthesize the results of the above steps to construct a four-dimensional cost matrix. The rows in the cost matrix represent each detection box in the current frame, and the columns in the cost matrix represent each tracking trajectory in the previous frame; Among them, each element represents the total cost between a pair of candidate matches (i.e., a detection box in the current frame and a tracking trajectory in the previous frame); S3.7. Based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph (such as blocking associations that violate clinical operation logic) to optimize the accuracy of target ID assignment (if the matching is successful, retain the matching target ID; if a detection box is not matched, initialize it as a new target and assign a new ID; if a trajectory is not matched for consecutive N frames, mark it as "Lost" or "Deleted"); Furthermore, this solution converts clinical operation logic into matching rules through a medical behavior causal graph, solving the problem of insufficient understanding of medical scenarios by general algorithms (target aggregation caused by emergency first aid). By means of a hierarchical bidirectional mechanism, it separates the matching paths of emergency goals and regular goals to ensure zero latency in tracking critical patients. In the emergency care scenario, the behaviors and positions of patients, medical staff, and equipment follow clinical operation logic (patients are usually located in the bed area, and medical staff need to approach the equipment cabinet before operating the equipment). If it is detected that the predicted position of patient A suddenly appears next to the equipment cabinet (distance from the bed > 2 meters), and medical staff B is operating the equipment, then the matching between patient A and the equipment cabinet area is blocked to avoid mis-association. If the vital signs of patient C are abnormal (sudden drop in blood oxygen), and there is medical staff D nearby carrying a defibrillator, then patient C and medical staff D are preferentially associated. By pre-filtering to exclude obviously unreasonable candidate pairs, the number of matching pairs to be processed by the Hungarian algorithm is reduced, accelerating the matching process and reducing resource consumption (especially in high-density target scenarios). In high-dynamic scenarios such as first aid (multiple people surrounding a patient), traditional algorithms are prone to ID jumps due to target aggregation. The causal graph maintains the ID continuity of key targets (patients) through clinical rule constraints. The cost matrix relying only on appearance (ReID) and motion (Kalman filter) may produce logically conflicting matches (such as a patient "teleporting" instantaneously to an unreasonable area). In the pre-filtering stage, candidate pairs violating medical behavior rules are directly eliminated, generating a cost matrix cleaned by clinical logic to improve the input quality of the Hungarian algorithm. Based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-temporal - semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph, including the following steps: S3.71. Construct a medical behavior causal graph for the emergency care unit, and real-time analyze the action intentions of medical staff. If it is detected that a medical staff is holding a syringe (visual recognition) and moving towards the bed (trajectory prediction), then preferentially match the association between this medical staff and the patient in the corresponding bed. If the vital signs of the patient suddenly change (drop in blood oxygen), then force the retention of its ID and increase the matching priority. Define nodes (target type, action, equipment, area) and edges (causal relationship, dependency relationship): The causal graph of this embodiment is: "Medical staff holds a syringe → approaches the bed → matches the patient in that bed"; "The defibrillator is unlocked → detects abnormal electrocardiogram monitoring → preferentially matches nearby patients"; "Patient's blood oxygen drops → visual features become blurred → force the retention of ID and increase the matching weight"; S3.72. For target patients carrying first aid equipment (defibrillator) or with abnormal vital signs, construct an independent sub-cost matrix, and preferentially complete high-weight matching to ensure zero-latency tracking of critical targets; S3.73. Based on the conventional Hungarian algorithm, a detection frame of the current frame and a tracking trajectory of the previous frame are jointly verified in terms of spatial position, equipment status and vital signs (only logical matching pairs (patient-bed binding, medical staff-equipment operation permission matching) are allowed to enter the cost matrix calculation, and associations that violate clinical rules (patients "teleport" to non-bed areas, unauthorized personnel contacting high-risk equipment) are blocked). Specifically: if the predicted position of patient A conflicts with the topological relationship of the bed (moving to the next bed instantly), the abnormal match is blocked; if the movement direction of medical staff B conflicts with the equipment access logic (walking to the exit but not returning the equipment), manual review is triggered; S3.74, dynamically set the matching threshold according to the area crowding (calculated by the depth sensor), relax the appearance similarity threshold (0.7→0.5) in high-density areas (>3 people / ㎡), focus on motion continuity, and enable a strict threshold (0.8) in low-density areas to avoid long-distance misassociation; S3.75, output the final matching result with priority control; S3.8. Output the tracking result with the target patient ID.

[0020] S4, analyzing the action sequence of the target patient in the video stream through 3D CNN, and judging abnormalities based on the physiological data of the target patient; In this embodiment, the action sequence of the target patient in the video stream is analyzed by 3D CNN, and the abnormality is determined in combination with the physiological data of the target patient, including the following steps: S4.1. According to the tracking ID of the target patient, extract the 3D space-time cube of the continuous time window (16 frames / 2 seconds in this embodiment) from the video stream, and synchronously collect the physiological data (heart rate, blood pressure, blood oxygen) of the target patient in the corresponding time period, and standardize the video stream; S4.2, using the 3D space-time cube as the input of the 3D CNN model, extracting the video features in the spatial and temporal dimensions through the convolutional layer of the 3D CNN model, and outputting the target patient action feature vector, which represents the behavior pattern of the target patient in the time window; S4.3. Use continuous wavelet transform (CWT) to convert the target patient's physiological data (heart rate, blood pressure, blood oxygen) into a time-frequency graph, use the time-frequency graph as the input of the 2D CNN model, extract frequency domain features, and introduce causal masks to ensure that the prediction relies only on historical information to avoid future data leakage. If the patient is in a stable state (normal heart rate), use a standard causal mask; if a sudden abnormality is detected (sudden increase in heart rate), temporarily relax the mask restriction to allow the model to refer to earlier historical data (features of the past 5 minutes) to enhance the ability to respond to emergencies; S4.4. Calculate the joint feature vector of the target patient's motion features and frequency domain features using the cross-attention mechanism (wherein, the motion features and frequency domain features are unified into the same dimensional space), and perform anomaly classification based on the joint feature vector through an MLP classifier (specifically: input the fused joint feature vector into a multi-layer perceptron (MLP), and after non-linear transformation, output the prediction probabilities of various abnormal states, so as to realize the recognition of abnormal events), and trigger the alarm mechanism; Among them, calculating the joint feature vector of the target patient's motion features and frequency domain features using the cross-attention mechanism includes the following steps: S4.41. Use the cosine similarity to measure the cosine value of the angle between the target patient's motion feature vector and frequency domain feature vector to calculate the similarity score, which ranges from [-1, 1]. The larger the value, the more similar the two vectors are; S4.42. Apply the Softmax function to the similarity score to convert the similarity score into a probability distribution, so that each score falls within the range of [0, 1], and the sum of all scores is 1; S4.43. According to the similarity score (the weight after Softmax), perform weighted summation on the motion feature vector and frequency domain feature vector to form a new joint feature vector. The finally obtained joint feature vector contains the comprehensive information of the motion features and frequency domain features.

[0021] This embodiment is for fall detection of patients in the emergency intensive care unit: The motion features are: analyze the spatio-temporal cube of 16 consecutive frames (about 0.5 seconds) through 3D CNN to capture the continuous motion patterns of the patient's torso suddenly tilting (hip joint angle change > 60°), lower limb instability (knee joint acceleration > 3m / s²), and the head quickly approaching the ground. The spatio-temporal features extracted by the 3D convolutional kernel will be cosine similarity matched with the preset "fall template" (the limb movement frequency during the fall process is 2 - 5Hz) (> 0.85 triggers an alarm); Physiological data collaboration: Synchronously collect the heart rate data at the moment of falling (through an electrocardiogram monitor), detect that the heart rate suddenly rises by more than 30 bpm within 6 seconds, and combine it with the instantaneous fluctuation of blood pressure (systolic blood pressure drops by more than 20 mmHg within the same time period or in the few minutes immediately following). Convert the electrocardiogram signal into a time-frequency diagram through wavelet transform, and perform cross-attention calculation on the frequency domain features (sudden increase in high-frequency energy) extracted by 2D CNN and the motion features. If the correlation weight exceeds the threshold (0.7), it is determined as a fall event; Among them, when the patient is blocked by medical staff, the system fuses the unoccluded features of the first 3 frames through the Cross-Frame Attention Mechanism (CFAM) to maintain tracking continuity. If the patient's position suddenly deviates from the hospital bed area (moves to beside the equipment cabinet), the medical causal graph will block the misconnection between "fall" and "normal walking", and an alarm will be triggered only when there is an abnormal acceleration in the hospital bed area bound to the patient ID.

[0022] Embodiment 2: This embodiment provides an in-urgency-care-unit target recognition system based on the Internet of Things, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the above-mentioned in-urgency-care-unit target recognition method based on the Internet of Things.

[0023] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. An object recognition method in an emergency care unit based on the Internet of Things, characterized in that It includes the following steps: S1. Collect the video stream in the emergency intensive care unit at a fixed frame rate through a camera; S2. Preprocess the collected video stream and extract image frames; S3. Use the YOLOv7 model to perform real-time object detection on each frame in the image frames, identify the target patient, associate the targets in consecutive frames of the video stream through the DeepSORT algorithm, and combine the Kalman filter to track the position of the target patient. During the process of tracking the position of the target patient, construct a four-dimensional spatio-temporal-semantic collaborative cost matrix for optimization; S4. Analyze the action sequence of the target patient in the video stream through 3D CNN and judge abnormalities in combination with the physiological data of the target patient.

2. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 1, wherein: In the above S3, using the YOLOv7 model to perform real-time object detection on each frame in the image frames and identify the target patient includes the following steps: S3.

1. Input the image frame into the YOLOv7 model for forward propagation, and output the position, confidence, and class score of each prediction box of the target patient; S3.

2. Use non-maximum suppression to remove overlapping prediction boxes and screen out the targets with a confidence higher than the threshold a; S3.

3. Among all the prediction boxes retained after screening, further analyze the class information of each prediction box, identify the target patient, and output the final detection result list.

3. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 2, wherein: In the above S3.1, inputting the image frame into the YOLOv7 model for forward propagation and outputting the position, confidence, and class score of each prediction box of the target patient includes the following steps: S3.

11. In the Backbone of the YOLOv7 model, for each residual block, retain the feature maps corresponding to the past 3 frames during forward propagation, construct a feature cache pool, perform dot product operations on the feature map of the current frame and the feature maps of historical frames to obtain a similarity matrix, and use attention weights to perform weighted summation on historical features; S3.

12. Obtain the depth information in the emergency intensive care unit through a depth camera and calculate the occlusion degree of the local area according to the depth map; S3.

13. Use the PANet structure of the YOLOv7 model for multi-scale feature fusion, input the multi-scale fused features into the detection head of the YOLOv7 model, output a prediction tensor containing the bounding box offset, target confidence, and class score, use the anchor box mechanism to decode the offset into the actual box position in the image coordinate space, and calculate the final confidence value in combination with the target confidence and class score.

4. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 2, wherein: In the above S3.3, among all the prediction boxes retained after screening, further analyze the class information of each prediction box and identify the target patient includes the following steps: S3.

31. Traverse each prediction box and extract the class information; S3.

32. Use a predefined class mapping table to map the class ID to a semantic label; S3.

33. Screen out the prediction boxes with the class of patient; S3.

34. Number the identified target patients and assign a unique ID to each identified patient; S3.

35. Output the list of identified target patients as the final detection result list.

5. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 2, characterized in that: In S3, the DeepSORT algorithm is used to associate the targets in consecutive frames of the video stream, and the Kalman filter is combined to track the position of the target patient, including the following steps: S3.

4. Input the detection box coordinates output by YOLOv7 into the DeepSORT model. For each target detected by the YOLOv7 model in each frame, crop the bounding box area of the target, and use the ReID sub-network to extract the appearance ReID embedding feature vector of the target; S3.

5. Initialize a Kalman filter for each detected target to predict its state vector in the next frame; S3.

6. Using the state estimation obtained by Kalman filter prediction, combined with the cosine distance between the appearance ReID embedding feature vectors, construct a four-dimensional spatio-temporal-semantic collaborative cost matrix; S3.

7. Based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-temporal-semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph; S3.

8. Output the tracking results with the target patient ID.

6. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 5, wherein: In S3.6, using the state estimation obtained by Kalman filter prediction, combined with the cosine distance between the ReID embedding vectors, construct a four-dimensional spatio-temporal-semantic collaborative cost matrix, including the following steps: S3.

61. Generate a hybrid ReID feature by combining the texture features of the patient's gown and the device connection status, calculate the cosine similarity between the current frame and the targets in the historical trajectory, and set the initial weight to b; S3.

62. Calculate the Mahalanobis distance between the targets through the Kalman filter prediction results; S3.

63. Check whether the targets detected in the current frame carry specific medical devices, and compare them with the device connection status in the historical trajectory. If the same devices are detected in both the current frame and the historical trajectory, increase the confidence of this matching pair; S3.

64. Combining the RFID wristband positioning signal, if the Euclidean distance between the center of the detection box and the RFID coordinate is less than 0.5 meters, increase the confidence of this matching pair; S3.

65. When it is detected that the acceleration of the patient in consecutive frames is greater than the threshold x, reduce the appearance weight and increase the motion weight; S3.

66. Synthesize the results of the above steps to construct a four-dimensional cost matrix. The rows in the cost matrix represent each detection box in the current frame, and the columns in the cost matrix represent each tracking trajectory in the previous frame.

7. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 5, characterized in that: In S3.7, based on the Hungarian algorithm, perform minimum weight matching on the four-dimensional spatio-temporal-semantic collaborative cost matrix, associate the detection results in the current frame and the previous frame of the video stream, and pre-filter the matching candidate pairs through the medical behavior causal graph, including the following steps: S3.

71. Construct a medical behavior causal graph for the emergency intensive care unit to real-time analyze the action intentions of medical staff; S3.

72. For the target patients carrying first aid equipment, construct an independent sub-cost matrix; S3.

73. On the basis of the conventional Hungarian algorithm, perform joint verification of the spatial position, device status and vital signs of a detection box in the current frame and a tracking trajectory in the previous frame; S3.

74. Dynamically set the matching threshold according to the regional congestion degree; S3.

75. Output the final matching result with priority control.

8. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 1, characterized in that: In S4, the action sequence of the target patient in the video stream is analyzed by 3D CNN, and combined with the physiological data of the target patient to judge abnormalities, including the following steps: S4.

1. According to the tracking ID of the target patient, extract the 3D spatio-temporal cube of the continuous time window from the video stream, and synchronously collect the physiological data of the target patient in the corresponding time period; S4.

2. Use the 3D spatio-temporal cube as the input of the 3D CNN model, extract the video features in the spatial and temporal dimensions through the convolutional layer of the 3D CNN model, and output the action feature vector of the target patient; S4.

3. Use continuous wavelet transform to convert the physiological data of the target patient into a time-frequency diagram, use the time-frequency diagram as the input of the 2D CNN model, extract the frequency domain features, and introduce a causal mask; S4.

4. Use the cross-attention mechanism to calculate the joint feature vector of the action feature and the frequency domain feature of the target patient, and perform anomaly classification through the MLP classifier based on the joint feature vector to trigger the alarm mechanism.

9. The method for target recognition in an emergency care unit based on the Internet of Things according to claim 8, wherein: In S4.4, using the cross-attention mechanism to calculate the joint feature vector of the action feature and the frequency domain feature of the target patient includes the following steps: S4.

41. Use cosine similarity to measure the cosine value of the angle between the action feature vector and the frequency domain feature vector of the target patient to calculate the similarity score; S4.

42. Apply the Softmax function to the similarity score to convert the similarity score into a probability distribution; S4.

43. According to the similarity score, perform weighted summation on the action feature vector and the frequency domain feature vector to form a new joint feature vector.

10. An object recognition system in an emergency care unit based on the Internet of Things, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes a computer program to implement the steps of the target recognition method in the emergency care unit based on the Internet of Things according to any one of claims 1-9.

Citation Information

Patent Citations

  • Remote medical monitoring system and method

    CN112086209A

  • Dense object multi-target tracking method based on Deepsort

    CN117058193A

  • Target detection and tracking system based on improved YOLOv7 and DeepSORT

    CN117423031A

  • Deep learning-based sheep rumination behavior identification method and system

    CN119625779A

  • KR20240065816A

Cited By

  • Unmanned aerial vehicle real-time detection tracking method and device based on YOLOv8 and dynamic ROI

    CN121280954A

  • Clinical first-aid equipment information acquisition management method and system

    CN121483535A

  • A method and system for collecting and managing information on clinical emergency medical equipment

    CN121483535B

  • Remote monitoring and health management system for falling risk of stroke patient

    CN121839142A

  • Remote monitoring and health management system for fall risk in stroke patients

    CN121839142B