Fall identification method, device and system and storage medium

By deploying the posture detection model on the camera end and deploying the multimodal large model on the cloud, the problems of high false alarm rate and excessive load in the cloud are solved, and efficient and accurate fall recognition is achieved.

CN120337137APending Publication Date: 2025-07-18SHENZHEN XIAOPAI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510419803.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

Smart Images

  • Figure CN120337137A_ABST
    Figure CN120337137A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a fall identification method, device and system and a storage medium, and is used for improving fall identification accuracy and real-time performance. The method comprises the steps that a camera end detects the posture of a figure in a shot picture in real time based on a trained posture detection model so as to recognize the confidence coefficient of the posture category of the figure; when the confidence coefficient of the posture category is a falling category or a sitting category is greater than a preset confidence coefficient, sending a service request to a cloud end, so that the cloud end inputs a target image frame corresponding to the posture category and an associated cue word into a trained multi-modal large model to re-judge the falling behavior of the person; and when the camera end receives an indication message that the cloud end confirms the falling behavior of the person, alarm information is sent out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a fall recognition method, device, system, and storage medium. Background Art

[0002] The recognition of fall behavior is a core research direction in the field of elderly care, and it is also a key and difficult problem in technical implementation. With the increasing aging of the global population, accidental falls of the elderly have become an important factor affecting their health and quality of life. Therefore, how to accurately and efficiently detect and recognize fall behavior and timely send out warning or help signals to reduce the risks brought by accidental accidents is a key problem that the current intelligent care system urgently needs to solve.

[0003] Currently, the fall detection methods based on wearable devices on the market mainly rely on sensors such as gyroscopes, accelerometers, and positioning devices to judge whether a fall occurs by monitoring the changes in human body postures and motion states. However, such methods have obvious limitations, including strong dependence on devices, inconvenient wearing, poor user compliance, etc., and it is difficult to meet the actual application requirements. Compared with wearable devices, there are also image detection technologies based on single-end models such as YOLO that can achieve non-contact and continuous behavior monitoring.

[0004] However, in actual application scenarios, the visual features extracted by the YOLO-based detection model in situations such as sitting down, bending over, and lying down have certain similarities with falls, resulting in the model being prone to false alarms. Moreover, since fall behavior requires real-time alarms, the solution of simply performing behavior discrimination based on a large cloud model is likely to cause an excessive load on the server side and reduce the real-time performance. Summary of the Invention

[0005] This application relates to the field of artificial intelligence technology, and particularly to a fall recognition method, device, system, and storage medium to solve the technical problems that traditional fall behavior recognition is prone to false alarms and the server side has an excessive load.

[0006] A fall recognition method based on the fusion of object detection and multi-modal large model includes: The camera end, based on the trained pose detection model, real-time detects the pose of the person in the captured image to identify the confidence level of the pose category of the person; When the confidence level of the pose category being the fall category or the sitting category is greater than the preset confidence level, a service request is sent to the cloud, so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person; When the camera end receives the indication message from the cloud confirming the fall behavior of the person, an alarm message is sent.

[0007] Further, after the camera terminal receives the indication message from the cloud to confirm the person's falling behavior, the method further includes: The camera terminal sends a notification message to the terminal to cause the terminal to perform a pop-up reminder; Wherein, when the number of times that the confidence level of the posture category being the falling category or the sitting category detected by the camera terminal within the detection window is greater than the preset confidence level exceeds the first preset threshold, sending a service request to the multimodal large model or sending a notification message to the terminal is prohibited until the number of times that the camera terminal detects the standing category within the detection window exceeds the second preset threshold, then the function of sending a service request to the multimodal large model is restored to wait for the recognition of the next falling state.

[0008] Further, the multimodal large model is trained in the following manner: a. Construct a question-and-answer data set and an unlabeled third falling behavior data set, wherein the question-and-answer data set includes picture data and its corresponding question-and-answer text, video clip data and its corresponding question-and-answer text; b. Fine-tune the pre-trained vision-language model VLM based on the question-and-answer data set; c. Use the fine-tuned pre-trained vision-language model VLM to process the image data in the third falling behavior data set to generate a preliminary prediction result, and the preliminary prediction result includes a falling category and a text description of the falling scenario; d. Generate a falling data annotation result based on the preliminary prediction result; e. Add the falling data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM; Repeat steps c - e until the trained multimodal large model is obtained.

[0009] Further, adding the falling data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM includes: Adding the falling data annotation result that has been manually reviewed and corrected as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM.

[0010] Further, the posture detection model is trained in the following manner: Collect a first falling behavior data set and a second falling behavior data set, where the first falling behavior data set is a publicly available falling behavior data set, and the second falling behavior data set is a falling behavior data set obtained by recording falling behaviors in a target area from multiple different perspectives of cameras; After preprocessing the first fall behavior dataset and the second fall behavior dataset, label the image data of the preprocessed first fall behavior dataset and second fall behavior dataset to obtain a target fall behavior dataset. The label annotation includes object detection label annotation and behavior classification label annotation. Among them, the object detection label is used to annotate the position box of the fallen individual, and the classification labels include fall category, standing category, and sitting category; Send the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model.

[0011] Further, the step of sending the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model includes: Perform data augmentation on the target fall behavior dataset that has passed quality review and calibration. The data augmentation includes performing rotation, scaling, color jittering, blurring, and noise perturbation processing on the image data; Send the target fall behavior dataset after data augmentation into a neural network for training to obtain the trained pose detection model; Among them, the AdamW optimizer is used for fine-tuning during the training process.

[0012] Further, the neural network includes the YOLOv5 object detection network.

[0013] A fall recognition device includes: A detection module, configured to, based on the trained pose detection model, detect the pose of a person in a captured image in real time to identify the confidence level of the pose category of the person; A determination module, configured to, when the confidence level of the pose category being the fall category or the sitting category is greater than a preset confidence level, send a service request to the cloud, so that the cloud inputs the target image frame corresponding to the pose category and associated prompt words into the trained multimodal large model to re-determine the fall behavior of the person; A sending module, configured to, when receiving the indication message from the cloud confirming the fall behavior of the person, issue an alarm message.

[0014] A fall recognition system based on the fusion of object detection and multimodal large model. The fall recognition system includes a camera end and a cloud: The camera end is configured to, based on the trained pose detection model, detect the pose of a person in a captured image in real time to identify the confidence level of the pose category of the person; when the confidence level of the pose category being the fall category or the sitting category is greater than a preset confidence level, send a service request to the cloud; The cloud is configured to respond to the service request, input the target image frame corresponding to the posture category and the associated prompt words into the trained multi-modal large model to re-determine the falling behavior of the person, and send an indication message to the camera side after confirming the falling behavior of the person; The camera side is configured to send an alarm message when receiving the indication message from the cloud confirming the falling behavior of the person.

[0015] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps or functions of the falling recognition method as described in any one of the foregoing are implemented.

[0016] This application provides a falling recognition solution based on the fusion of object detection and multi-modal large model. The camera side, based on the trained posture detection model, real-time detects the posture of the person in the captured image to identify the confidence level of the posture category of the person; when the confidence level of the posture category being the falling category or the sitting category is greater than the preset confidence level, a service request is sent to the cloud, so that the cloud inputs the target image frame corresponding to the posture category and the associated prompt words into the trained multi-modal large model to re-determine the falling behavior of the person; when the camera side receives the indication message from the cloud confirming the falling behavior of the person, an alarm message is sent. It can be seen that this method combines the advantages of object detection and multi-modal semantic understanding. It can not only use the efficient object detection ability of the posture detection model to extract human body posture features, but also further perform image semantic-level understanding based on the associated prompt words through the multi-modal large model, and further report to the cloud large model for discrimination to decide whether to alarm, enhancing the perception ability of the falling environment, reducing the false detection caused by using a single detection algorithm, and reducing the false alarm problem caused by misjudgment of simple visual features; in addition, at the system level, deploying the falling posture detection model on the camera side for judgment first and deploying the multi-modal large model on the cloud can effectively avoid the technical problem that the solution of simply performing falling behavior discrimination based on the cloud is likely to cause excessive load on the cloud side. For the cloud, it is based on the posture judgment of the camera side, and the cloud is only called when the required category conditions are met, which can also reduce the high cost of calling the cloud. Whether it is cloud resources or bandwidth occupancy, it is reduced, and the real-time performance is also improved. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of a fall recognition method based on the fusion of object detection and multi-modal large model in an embodiment of the present application; Figure 2 It is a schematic system architecture diagram of a fall recognition system based on the fusion of object detection and multi-modal large model in an embodiment of the present application; Figure 3 It is a schematic diagram of a training process of a multi-modal large model in an embodiment of the present application; Figure 4 It is a schematic structural diagram of a fall recognition device based on the fusion of object detection and multi-modal large model in an embodiment of the present application; Figure 5 It is a schematic structural diagram of a computer device in an embodiment of the present application. Specific Embodiments

[0019] In order to make the technical problems, technical solutions and beneficial effects solved by the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] In order to overcome the problems and deficiencies existing in the related technical solutions, the embodiments of the present application provide a fall recognition solution suitable for the home camera scenario, which will be described separately below.

[0021] As Figure 1 shown, in one embodiment, a fall recognition method based on the fusion of object detection and multi-modal large model is provided, and the method includes the following steps: S101. The camera end, based on the trained pose detection model, real-time detects the pose of the person in the captured image to identify the confidence of the pose category of the person; S102. When the confidence of the pose category being the fall category or the sitting category is greater than the preset confidence, a service request is sent to the cloud, so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person; S103. When the camera end receives the indication message from the cloud confirming the fall behavior of the person, an alarm message is issued.

[0022] In this embodiment, the pose detection model and the multi-modal large model are pre-trained. The target detection algorithm based on the pose detection model is deployed on the camera side, and the further analysis and determination based on the multi-modal large model are deployed on the cloud. After the model training and optimization are completed, it is necessary to adapt to different hardware environments to enable the camera on the side to have real-time detection capabilities, while the cloud server undertakes complex large model inference tasks to achieve efficient fall detection, recognition and analysis. Among them, the pose detection model deployed on the camera side needs to be converted into a model format supported by the board on the camera side. For example, it can be seen that the original.pt model file of the trained pose detection model is converted into the onnx (Open Neural Network Exchange) format, and then the onnx format pose detection model is converted into the axmodel format supported by the board on the camera side. In addition, for the multi-modal large model deployed on the cloud, the cloud can use a dedicated inference framework for optimizing the multi-modal large model to achieve inference acceleration. For example, as an explanatory note, the cloud can use the dedicated inference framework Vllm (FastLLM Inference Engine) for optimizing the multi-modal large model to achieve inference acceleration, which can greatly improve the inference throughput, reduce the memory occupancy of the cloud, and support batch parallel inference, and is suitable for deploying the trained multi-modal large model on the cloud server; the deployed multi-modal large model can be called using an interface method compatible with various Api formats for the camera side to call, and is also applicable to multiple access platforms at the same time.

[0023] In the stage of fall behavior recognition, in this embodiment, the camera side will, based on the trained pose detection model, real-time detect the pose of the person in the captured image to identify the confidence level of the pose category of the person and the corresponding detection box. When the confidence level of the pose category being the fall category or the sitting category is greater than the preset confidence level, a service request is sent to the cloud so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person.

[0024] For example, the camera side real-time detects the pose of the person in the image, gives the confidence levels of the three categories of standing / falling / sitting and the corresponding detection boxes. When the confidence level of the person falling or sitting in the detected image is greater than 80%, the target image frame corresponding to the pose category of the person is reported to the cloud for further confirmation, and the associated prompt words of this target image frame are: "Whether the person at , , , and , , , ... has fallen?" Among them, the coordinates , , , , , , , represent the coordinates of the detection box of the person in the target image frame. If the cloud further confirms the fall behavior based on the multimodal large model, the cloud will send an instruction message to confirm the fall behavior of the person to the camera side. When the camera side receives the instruction message from the cloud to confirm the fall behavior of the person, an alarm message will be sent to complete the detection and early warning of the fall behavior in the detection screen of the camera side.

[0025] It can be seen that in this embodiment, a fall recognition method based on the fusion of object detection and multimodal large model is provided, which combines the advantages of object detection and multimodal semantic understanding. It can not only use the efficient object detection ability of the pose detection model to extract human body pose features, but also further perform image semantic-level understanding based on associated prompt words through the multimodal large model, and further report to the cloud large model for discrimination to decide whether to give an alarm, enhancing the perception ability of the fall environment, reducing the false detection caused by using a single detection algorithm, and reducing the false alarm problem caused by misjudgment of pure visual features; In addition, at the system level, deploying a fall pose detection model on the camera side for judgment first and deploying a multimodal large model on the cloud can effectively avoid the technical problem that the scheme of judging the fall behavior based solely on the cloud is likely to cause excessive load on the cloud side. For the cloud, it is based on the pose judgment of the camera side, and the cloud is only called when the required category conditions are met, which can also reduce the high cost of calling the cloud. Whether it is cloud resources or bandwidth occupancy, it is reduced, so that the real-time performance is also improved.

[0026] In one embodiment, after the camera side receives the instruction message from the cloud to confirm the fall behavior of the person, the method further includes: S104. The camera side sends a notification message to the terminal to make the terminal perform a pop-up reminder; Among them, when the number of times that the confidence level of the pose category being the fall category or the sitting category detected by the camera side in the detection window is greater than the preset confidence level exceeds the first preset threshold, sending a service request to the multimodal large model or sending a notification message to the terminal is prohibited until the number of times that the camera side detects the standing category in the detection window exceeds the second preset threshold, and then the function of sending a service request to the multimodal large model is restored to wait for the recognition of the next fall state.

[0027] In this embodiment, when the camera end receives the indication message from the cloud to confirm the person's falling behavior, the camera end sends a notification message to the terminal to enable the terminal to pop up a reminder for multiple warnings, so that the person holding the terminal or managing the terminal can receive the falling reminder, such as the nursing home administrator, so that the falling behavior can be effectively monitored.

[0028] In addition, in this embodiment, in order to prevent the terminal from receiving repeated notifications, when the number of times that the confidence level of the posture category being the falling category or the sitting category detected by the camera end within the detection window exceeds the first preset threshold, the service request to the multimodal large model or the notification message to the terminal is prohibited until the number of times that the camera end detects the standing category within the detection window exceeds the second preset threshold, and then the function of sending the service request to the multimodal large model is restored to wait for the recognition of the next falling state.

[0029] The above first preset threshold and second preset threshold can be set according to experience and are not specifically limited. As an exemplary illustration, both the first preset threshold and the second preset threshold can be 10. Similarly, for example, the length of the detection window (TemporalWindow) can be 20. Taking this as an example, when the camera detects the falling or sitting category within this detection window and it is confirmed by the cloud to be a fall, if the camera end detects the "falling category" or "sitting category" more than 10 times within the detection window, the service of the multimodal large model in the cloud will no longer be requested or the terminal will not be notified until the camera end detects the "standing category" more than 10 times within the detection window, and then the function of calling the service of the multimodal large model in the cloud is restored to wait for the next falling state.

[0030] It can be seen that in this embodiment, due to the addition of the detection window, the load of the cloud large model service is also reduced, and the cost is greatly reduced compared with the scheme of simply using the large model for falling recognition; and it can prevent the terminal from receiving multiple notification messages repeatedly.

[0031] In one embodiment, as Figure 2 shown, the multimodal large model is trained in the following manner: a. Construct a question-and-answer data set and an unlabeled third falling behavior data set, where the question-and-answer data set includes picture data and its corresponding question-and-answer text, video clip data and its corresponding question-and-answer text; b. Fine-tune the pre-trained vision-language model VLM based on the question-and-answer data set; c. Use the fine-tuned pre-trained Vision-Language Model (VLM) to process the image data in the third fall behavior dataset to generate a preliminary prediction result, where the preliminary prediction result includes a fall category and a text description of the fall scenario; d. Generate a fall data annotation result based on the preliminary prediction result; e. Add the fall data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained Vision-Language Model (VLM); Repeat steps c - e until a trained multi-modal large model is obtained.

[0032] In one embodiment, adding the fall data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained Vision-Language Model (VLM) includes: Adding the fall data annotation result that has been manually reviewed and corrected as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained Vision-Language Model (VLM).

[0033] In this embodiment, question-and-answer data for the multi-modal large model will be constructed. To improve the multi-modal large model's understanding ability of fall behaviors, when constructing the multi-modal question-and-answer dataset, image, video, and text information will be combined, and based on the pre-trained large model, continue with fine-tuning training to obtain a multi-modal large model capable of discriminating fall events. Specifically, the question-and-answer dataset includes picture data and its corresponding question-and-answer text, video clip data and its corresponding question-and-answer text, that is: 1. For the question-and-answer pair data based on picture data, design multiple different question methods for the guests, such as: "Is there anyone falling in this picture?", "Is the person in the picture in an abnormal state?", "Which category does the behavior of the person in the picture belong to: standing / falling / sitting?", "Is the person at coordinates [x0, y0, x1, y1] falling?" to ask questions about the image data in the dataset.

[0034] 2. Temporal question-and-answer pairs based on videos. The publicly collected data or the multi-view dataset collected also contains video data. For this type of data, perform slicing processing to obtain multiple segments with a duration of, for example, 5s, and design multiple groups of question methods for videos, such as: "Is the person in the video walking normally or falling?", "At which second in the video does someone fall to the ground?". The above two types of data are both stored in the json data format, and the fields include the image / video storage path, the user's question, and the corresponding answer.

[0035] For the annotation of the answer data, this method adopts an iterative optimization method to continuously optimize the multi-modal large model by gradually fine-tuning the model and expanding high-quality training data, which mainly includes the following processes: Initial model training: A pre-trained vision-language model (VLM) with a scale of 7B can be selected as the basic model, and a part of the artificially annotated question-and-answer data set is used to fine-tune the pre-trained vision-language model VLM to enhance its understanding ability of the falling behavior. Model-assisted data expansion: That is, the fine-tuned VLM model is used to process the unannotated falling data (the third falling behavior data set), and preliminary prediction results are automatically generated. The annotation results of the falling data are added to the training set as new data, including text descriptions such as the type of fall and the situation of the fall. Manual review and correction: The annotation results of the falling data generated by the model are reviewed and corrected manually to ensure the data quality. Incremental fine-tuning training: The newly added training data is added to the training set, and the VLM model is incrementally fine-tuned (Incremental Fine-tuning) to continuously optimize its recognition ability. Iterative optimization: Continuously repeat to gradually expand the data scale and improve the recognition ability of the multi-modal for complex falling scenarios.

[0036] It can be seen that in this embodiment, a method for constructing question-and-answer data of a multi-modal large model and its corresponding training process are provided. In order to improve the understanding ability of the large model for the falling behavior, a multi-modal question-and-answer data set is constructed, which combines image, video and text information, and continues to fine-tune and train on the basis of the pre-trained large model to obtain a multi-modal large model that can judge falling events; in the fine-tuning process, through an iterative optimization method, by gradually fine-tuning the model, expanding high-quality training data, and combining manual review, the continuous optimization of the fall detection model is realized, providing an accurate recognition model for the determination of the falling behavior again in the cloud, and reducing the false positive rate.

[0037] In addition, in one embodiment, the fine-tuning process of the above-mentioned multi-modal large model with 7B parameter scale can be fine-tuned using the LoRA (Low-Rank Adaptation) method. LoRA can effectively fine-tune the pre-trained large model with limited computing resources by introducing low-rank matrices, avoiding the high cost of modifying the entire model parameters. The LoRA fine-tuning process can be divided into the following steps: 1. Select the fine-tuning layer, usually select the self-attention layer or the feed-forward neural network layer in the Transformer architecture.

[0038] 2. Introduce low-rank matrices. For each fine-tuning layer, introduce two low-rank matrices and to transform the output of the multi-modal large model. For the weight of each layer, the updated weight of Lora is: 3. Fix the weights of the pre-trained multi-modal large model. Only the and parameters are updated during the LoRA training process. The weights of the pre-trained multi-modal large model network are frozen for adjustment, avoiding large-scale computational overhead.

[0039] 4. During the training process, the CrossEntropyLoss cross-entropy loss is used to calculate the loss between the tokenized label data and the model's reply token data, and the low-rank matrix parameters are updated.

[0040] 5. Merge the weights. Combining the low-rank matrix added by LoRA with the parameters of the pre-trained multi-modal large model can achieve similar performance to the full-scale fine-tuned model.

[0041] In this embodiment, for the fine-tuning of the multi-modal large model, by introducing a low-rank matrix, it is possible to effectively fine-tune the pre-trained multi-modal large model with limited computing resources, avoiding the high cost of modifying the entire model parameters.

[0042] In one embodiment, the pose detection model is trained as follows: Collect the first fall behavior dataset and the second fall behavior dataset. The first fall behavior dataset is a publicly available fall behavior dataset, and the second fall behavior dataset is a fall behavior dataset obtained by recording fall behaviors in the target area using cameras from multiple different perspectives; After preprocessing the first fall behavior dataset and the second fall behavior dataset, label the image data of the preprocessed first fall behavior dataset and the second fall behavior dataset to obtain the target fall behavior dataset. The label annotation includes object detection label annotation and behavior classification label annotation. Among them, the object detection label is used to label the position box of the falling individual, and the classification labels include fall category, standing category, and sitting category; Send the target fall behavior dataset that has passed quality review and calibration into the neural network for training to obtain the trained pose detection model.

[0043] In this embodiment, publicly available first fall behavior datasets are collected and organized. As an explanatory note, the first fall behavior datasets are publicly available fall behavior datasets, including but not limited to the following categories: the detection datasets Fall-Down-Det-v1, Fall-Down-Det-v2, and Lei2i Fall detection dataset for fall behavior, the Fall-Down-Cls-v1 and Fall-Down-Cls-v2 datasets dedicated to fall behavior classification, and the image data related to fall categories in the general action recognition dataset. Additionally, a second fall behavior dataset is obtained by recording fall behavior in the target area using cameras with multiple different perspectives. For example, 7 cameras with different perspectives are arranged in an indoor scene to construct a multi-camera scene for collecting fall videos. The scene covers different heights, angles, and spatial layouts to ensure comprehensive coverage of the target area and minimize occlusion problems. During the data collection process, the cameras synchronously record fall behavior from multiple perspectives, capturing fall events under different human postures, movement trajectories, and environmental factors to improve the diversity and generalization ability of the data.

[0044] Then, the label data for the detection model is constructed. For the collected multi-camera fall videos and public datasets, LabelImg can be used for data cleaning, duplicate removal, format standardization, and label annotation according to the requirements of the fall detection task. The annotation methods include but are not limited to: object detection labels (Bounding Box), which are used to label the position boxes (Bounding Box) of the fallen individuals, providing basic training data for the object detection task. Action classification labels (Action Labeling), which represent the classification of video frames or video segments, distinguishing between three behaviors: fall, stand, and sit, providing training data for the fall behavior classification model. After annotation, the data will be converted into the COCO standard format and undergo quality review to ensure the consistency and accuracy of data annotation.

[0045] In one embodiment, the step of feeding the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model includes: Performing data augmentation on the target fall behavior dataset that has passed quality review and calibration, where the data augmentation includes processing the image data by rotation, scaling, color jitter, blurring, and noise perturbation; Feeding the target fall behavior dataset after data augmentation into a neural network for training to obtain the trained pose detection model; Among them, the training process uses the AdamW optimizer for fine-tuning. In one embodiment, the neural network includes the YOLOv5 object detection network, or it can be other neural networks, and no specific limitation is made.

[0046] In this embodiment, the pose detection model can be fine-tuned. Based on YOLOV5s, performance fine-tuning is carried out. In order to improve the robustness of the model, the training data needs to be augmented before being fed into the model, including image augmentation operations such as rotation, scaling, color jitter, blurring, and noise perturbation. The AdamW optimizer is used in the fine-tuning process to improve training stability. In terms of the loss function design, for the classification task, Cross-Entropy Loss is used, and for the object detection task, IoU Loss + Focal Loss is used to handle the class imbalance problem.

[0047] In this embodiment, the pose detection model is trained by collecting the first fall behavior dataset and the second fall behavior dataset. The first fall behavior dataset is a publicly available fall behavior dataset, and the second fall behavior dataset is a fall behavior dataset obtained by recording fall behaviors in the target area from multiple different perspectives of cameras. In particular, for existing fall detection algorithms that cannot understand scene information, users need to manually delimit the detection area to exclude scenes such as living room sofas or bedroom beds to eliminate false detections of non-fall behaviors such as sleeping. Through the above-mentioned fall annotation data in the embodiments of the present application, the recognition in the above scenarios can be effectively improved, and the false positive rate can be reduced.

[0048] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0049] In one embodiment, as Figure 3 shown, the embodiments of the present application also provide a fall recognition system. The fall recognition system includes a camera end and a cloud end, where: The camera end is used to, based on the trained pose detection model, detect the pose of the person in the captured image in real time to identify the confidence level of the pose category of the person; when the confidence level of the pose category being the fall category or the sitting category is greater than the preset confidence level, a service request is sent to the cloud end; The cloud end is used to, in response to the service request, input the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person, and send an indication message to the camera end after confirming the fall behavior of the person; The camera end is used to send an alarm message when receiving the indication message from the cloud to confirm the falling behavior of the person.

[0050] It can be seen that in this embodiment, a fall recognition system based on the fusion of object detection and multimodal large models is provided, which combines the advantages of object detection and multimodal semantic understanding. It can not only use the efficient object detection ability of the pose detection model to extract human pose features, but also further understand the image at the semantic level through the multimodal large model based on associated prompt words, and then report it to the cloud large model for discrimination to decide whether to give an alarm, enhancing the perception ability of the falling environment, reducing the false detection caused by using a single detection algorithm, and reducing the false alarm problem caused by misjudgment of pure visual features. In addition, at the system level, the fall pose detection model is deployed on the camera side for preliminary judgment, and the multimodal large model is deployed on the cloud. This can effectively avoid the technical problem that the scheme of judging the fall behavior based solely on the cloud is prone to cause excessive load on the cloud side. For the cloud, it is based on the pose judgment of the camera end, and the cloud is only called when the required category conditions are met, which can also reduce the high cost of calling the cloud. Whether it is cloud resources or bandwidth occupancy, it is reduced, improving the real-time performance.

[0051] For the specific limitations of the cloud or the camera end in the fall recognition system, reference can be made to the relevant limitations of the cloud or the camera end in the fall recognition method described above, which will not be elaborated here.

[0052] In one embodiment, a fall recognition device is provided. The fall recognition device corresponds one-to-one to the fall recognition method in the above embodiment. As Figure 4 shown, the fall recognition device includes a detection module 101, a determination module 102, and a sending module 103. The detailed description of each functional module is as follows: The detection module 101 is used to detect the pose of the person in the captured image in real time based on the trained pose detection model to identify the confidence level of the pose category of the person. The determination module 102 is used to send a service request to the cloud when the confidence level of the pose category being the fall category or the sitting category is greater than the preset confidence level, so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multimodal large model to re-determine the fall behavior of the person. The sending module 103 is used to send an alarm message when receiving the indication message from the cloud to confirm the fall behavior of the person.

[0053] In one embodiment, the sending module 103 is further configured to: after the camera terminal receives the indication message from the cloud confirming the person's falling behavior, send a notification message to the terminal to enable the terminal to perform a pop-up reminder; wherein, when the number of times the confidence level of the posture category being the falling category or the sitting category in the detection window is greater than the preset confidence level exceeds the first preset threshold, sending a service request to the multimodal large model or sending a notification message to the terminal is prohibited until the number of times the camera terminal detects the standing category in the detection window exceeds the second preset threshold, and then the function of sending a service request to the multimodal large model is restored to wait for the recognition of the next falling state.

[0054] In one embodiment, the multimodal large model is trained in the following manner: a. Construct a question-and-answer data set and an unlabeled third falling behavior data set, wherein the question-and-answer data set includes picture data and its corresponding question-and-answer text, and video clip data and its corresponding question-and-answer text; b. Fine-tune the pre-trained vision-language model VLM based on the question-and-answer data set; c. Use the fine-tuned pre-trained vision-language model VLM to process the image data in the third falling behavior data set to generate a preliminary prediction result, and the preliminary prediction result includes a falling category and a text description of the falling scenario; d. Generate a falling data annotation result based on the preliminary prediction result; e. Add the falling data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM; Repeat steps c - e until the trained multimodal large model is obtained.

[0055] In one embodiment, adding the falling data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM includes: Adding the falling data annotation result that has been manually reviewed and corrected as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM.

[0056] In one embodiment, the posture detection model is trained in the following manner: Collect a first falling behavior data set and a second falling behavior data set, where the first falling behavior data set is a publicly available falling behavior data set, and the second falling behavior data set is a falling behavior data set obtained by recording the falling behavior in the target area from multiple different perspectives of cameras; After preprocessing the first fall behavior dataset and the second fall behavior dataset, label the image data of the preprocessed first fall behavior dataset and the second fall behavior dataset to obtain a target fall behavior dataset. The label annotation includes object detection label annotation and behavior classification label annotation. Among them, the object detection label is used to annotate the position box of the falling individual, and the classification labels include fall categories, standing categories, and sitting categories; Send the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model.

[0057] In one embodiment, the step of sending the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model includes: Perform data augmentation on the target fall behavior dataset that has passed quality review and calibration. The data augmentation includes rotating, scaling, color jittering, blurring, and noise perturbation processing on the image data; Send the target fall behavior dataset after data augmentation into a neural network for training to obtain the trained pose detection model; Among them, the AdamW optimizer is used for fine-tuning during the training process.

[0058] In one embodiment, the neural network includes the YOLOv5 object detection network.

[0059] It can be seen that in this embodiment, a fall recognition device based on the fusion of object detection and multimodal large models is provided, which reduces the false detection caused by using a single detection algorithm and reduces the false alarm problem caused by misjudgment of pure visual features; in addition, at the system level, the fall pose detection model is deployed on the camera side for preliminary judgment, and the multimodal large model is deployed on the cloud. This can effectively avoid the technical problem that the scheme of judging fall behavior based solely on the cloud is prone to cause excessive cloud load. For the cloud, it is based on the pose judgment of the camera side, and the cloud is only called when the required category conditions are met, which can also reduce the high cost of calling the cloud. Whether it is cloud resources or bandwidth occupancy, both are reduced, and the real-time performance is also improved.

[0060] For the specific limitations of the fall recognition device, reference can be made to the limitations of the fall recognition method in the above text, which will not be elaborated here. Each module in the above fall recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0061] In one embodiment, as Figure 5 shown, a computer device is provided, which can be a cloud or a camera end, including a network interface, a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps or functions of the cloud or the camera end in the fall recognition method in the above embodiment. To avoid repetition, it will not be elaborated here.

[0062] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps or functions of the cloud or the camera end in the fall recognition method in the above embodiment. To avoid repetition, it will not be elaborated here The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A fall recognition method based on the fusion of object detection and multi-modal large models, characterized in that Including: The camera end, based on the trained pose detection model, detects the pose of the person in the captured image in real time to identify the confidence level of the pose category of the person. When the confidence level of the pose category being the fall category or the sitting category is greater than the preset confidence level, a service request is sent to the cloud, so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multimodal large model to re-determine the fall behavior of the person. When the camera end receives the indication message from the cloud confirming the fall behavior of the person, an alarm message is issued.

2. The fall recognition method according to claim 1, wherein After the camera end receives the indication message from the cloud confirming the fall behavior of the person, the method further includes: The camera end sends a notification message to the terminal to enable the terminal to perform a pop-up reminder. Wherein, when the number of times that the camera end detects that the confidence level of the pose category being the fall category or the sitting category in the detection window is greater than the preset confidence level exceeds the first preset threshold, sending a service request to the multimodal large model or sending a notification message to the terminal is prohibited until the number of times that the camera end detects the standing category in the detection window exceeds the second preset threshold, and then the function of sending a service request to the multimodal large model is restored to wait for the next identification of the fall state.

3. The fall recognition method according to claim 1, wherein The multimodal large model is trained in the following manner: a. Construct a question-and-answer data set and an unlabeled third fall behavior data set, wherein the question-and-answer data set includes picture data and its corresponding question-and-answer text, video clip data and its corresponding question-and-answer text. b. Fine-tune the pre-trained vision-language model VLM based on the question-and-answer data set. c. Use the fine-tuned pre-trained vision-language model VLM to process the image data in the third fall behavior data set to generate a preliminary prediction result, and the preliminary prediction result includes a fall category and a text description of the fall situation. d. Generate a fall data annotation result based on the preliminary prediction result. e. Add the fall data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM. Repeat steps c-e until the trained multimodal large model is obtained.

4. The fall recognition method according to claim 3, characterized in that The adding the fall data annotation result as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM includes: Adding the fall data annotation result that has been manually reviewed and corrected as new data to the training set to perform incremental fine-tuning on the fine-tuned pre-trained vision-language model VLM.

5. The fall recognition method according to any one of claims 1-4, characterized in that, The pose detection model is trained in the following manner: Collect a first fall behavior data set and a second fall behavior data set, where the first fall behavior data set is a publicly available fall behavior data set, and the second fall behavior data set is a fall behavior data set obtained by recording the fall behavior of a target area from multiple different perspectives of cameras. After preprocessing the first fall behavior dataset and the second fall behavior dataset, label the image data of the preprocessed first fall behavior dataset and second fall behavior dataset to obtain a target fall behavior dataset. The label annotation includes object detection label annotation and behavior classification label annotation. Among them, the object detection label is used to label the position box of the fallen individual, and the classification labels include fall categories, standing categories, and sitting categories; Send the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model.

6. The fall recognition method according to claim 5, characterized in that, The step of sending the target fall behavior dataset that has passed quality review and calibration into a neural network for training to obtain the trained pose detection model includes: Perform data augmentation on the target fall behavior dataset that has passed quality review and calibration. The data augmentation includes performing rotation, scaling, color jitter, blurring, and noise perturbation processing on the image data; Send the target fall behavior dataset after data augmentation into a neural network for training to obtain the trained pose detection model; Among them, the AdamW optimizer is used for fine-tuning during the training process.

7. The fall recognition method according to claim 5, wherein The neural network includes the YOLOv5 object detection network.

8. A fall recognition device, characterized in that, It includes: A detection module for real-time detecting the pose of a person in the captured image based on the trained pose detection model to identify the confidence level of the pose category of the person; A determination module for, when the confidence level of the pose category being a fall category or a sitting category is greater than a preset confidence level, sending a service request to the cloud, so that the cloud inputs the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person; A sending module for, when receiving the instruction message from the cloud confirming the fall behavior of the person, issuing an alarm message.

9. A fall recognition system based on the fusion of object detection and multimodal large models, characterized in that, The fall recognition system includes a camera end and a cloud: The camera end is used for real-time detecting the pose of a person in the captured image based on the trained pose detection model to identify the confidence level of the pose category of the person; When the confidence level of the pose category being a fall category or a sitting category is greater than a preset confidence level, send a service request to the cloud; The cloud is used for responding to the service request, inputting the target image frame corresponding to the pose category and the associated prompt words into the trained multi-modal large model to re-determine the fall behavior of the person, and sending an instruction message to the camera end after confirming the fall behavior of the person; The camera end is used for, when receiving the instruction message from the cloud confirming the fall behavior of the person, issuing an alarm message.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps or functions of the fall recognition method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-modal model-based thrown object detection method and system

    CN122024186A