Fall detection method and system based on multi-modal large model, terminal and storage medium

Through the two-stage detection process of the multimodal large model, the initial inspection model quickly screens suspicious frames, and the re-inspection model conducts in-depth analysis, which solves the problems of low efficiency and low accuracy in existing technologies and realizes efficient and accurate fall detection.

CN120656098APending Publication Date: 2025-09-16GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ) +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510512556.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing fall detection methods are inefficient and inaccurate, lack targeted pre-screening mechanisms, and cannot effectively understand dynamic changes in consecutive video frames.

Method used

A two-stage detection process based on a multimodal large model is adopted. The initial inspection model quickly screens suspicious video frames, and the re-inspection model conducts in-depth analysis. Combining natural language processing and computer vision technology, we build the initial and re-inspection models for falls.

Benefits of technology

The efficiency and accuracy of fall detection are significantly improved, the false alarm rate is reduced, and the ability to understand the dynamic process of fall events is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656098A_ABST
    Figure CN120656098A_ABST
Patent Text Reader

Abstract

The invention discloses a tumble detection method and system based on a multi-mode large model, a terminal and a storage medium, and the method comprises the steps: obtaining real-time video information, inputting the real-time video information to a tumble initial detection model, and obtaining a suspicious video frame; obtaining text prompt information, inputting the text prompt information and the suspicious video frame into a tumble recheck model, and outputting a behavior analysis report; and obtaining a tumble detection result according to the behavior analysis report. Two-stage detection of the tumble behavior is realized by constructing the tumble initial detection model and the tumble redetection model, the tumble initial detection model can quickly screen out the possible tumble behavior, the tumble redetection model can deeply analyze the possible tumble behavior, and then the tumble detection result is output. The fall behavior detection efficiency and the fall detection result output accuracy are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a fall detection method, system, terminal, and computer-readable storage medium based on a multimodal large model. Background Art

[0002] For elderly or ill individuals living alone and without supervision, the ability to detect unusual activity patterns early is crucial. Within the context of existing intelligent assistance applications and environments, the recognition of physical activity and fall detection are considered essential features. Due to the significant impact of falls on health and healthcare costs, there is growing interest in automated fall detection methods.

[0003] However, existing fall detection methods generally adopt a single-stage detection process, which requires a comprehensive analysis of all video frames and lacks a targeted pre-screening mechanism. In addition, existing technologies rely on static image analysis and cannot effectively understand the dynamic changes in continuous video frames, resulting in low fall detection efficiency and inaccurate detection results.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a fall detection method, system, terminal and computer-readable storage medium based on a multimodal large model, aiming to solve the problems that fall detection methods in the existing technology generally adopt a single-stage detection process, require comprehensive analysis of all video frames, and lack a targeted pre-screening mechanism. In addition, the existing technology relies on static image analysis and cannot effectively understand the dynamic changes in continuous video frames, resulting in low detection efficiency of fall detection and low accuracy of detection results.

[0006] To achieve the above objectives, the present invention provides a fall detection method based on a multimodal large model, the fall detection method based on the multimodal large model comprising the following steps:

[0007] Acquire real-time video information, and input the real-time video information into a fall initial detection model to obtain a suspicious video frame;

[0008] Obtaining text prompt information, and inputting the text prompt information and the suspicious video frame into a fall re-examination model, and outputting a behavior analysis report;

[0009] A fall detection result is obtained according to the behavior analysis report.

[0010] Optionally, the multimodal large model-based fall detection method, wherein the step of acquiring real-time video information and inputting the real-time video information into a fall initial detection model to obtain a suspicious video frame, specifically includes:

[0011] Obtaining labeled fall images and labeled non-fall images;

[0012] Determine a first pre-trained model, and perform model training on the first pre-trained model based on the labeled fall image and the labeled non-fall image to obtain a fall initial detection model;

[0013] Acquiring real-time video information and inputting the real-time video information into the fall initial detection model;

[0014] The real-time video information is processed by using the fall initial detection model to identify suspicious video frames to obtain suspicious video frames.

[0015] Optionally, the fall detection method based on a multimodal large model, wherein the step of performing suspicious video frame identification processing on the real-time video information using the fall initial detection model to obtain the suspicious video frame, specifically includes:

[0016] Performing image standardization and size adjustment on the real-time video information using the fall initial detection model to obtain formatted image data;

[0017] Performing feature extraction processing on the formatted image data to obtain local features of the image;

[0018] The local features of the image are input into the fully convolutional network in the fall initial detection model to output suspicious key frames.

[0019] Optionally, the multimodal large model-based fall detection method, wherein the acquiring of text prompt information, inputting the text prompt information and the suspicious video frame into a fall re-examination model, and outputting a behavior analysis report, further comprises:

[0020] Obtaining an open source dataset and a self-built dataset, and preprocessing the open source dataset and the self-built dataset to obtain a training set and an evaluation set;

[0021] Determine a second pre-trained model, and use the LoRA method to fine-tune the second pre-trained model according to the training set to obtain an initial fall re-examination model;

[0022] The cross-validation method is used to perform model optimization processing on the initial fall re-examination model according to the evaluation set to obtain a fall re-examination model.

[0023] Optionally, in the fall detection method based on a multimodal large model, the preprocessing of the open source dataset and the self-built dataset to obtain a training set and an evaluation set specifically includes:

[0024] Performing semantic extraction processing and video frame annotation processing on the open source dataset and the self-built dataset to obtain a key frame dataset;

[0025] Performing data filtering processing on the key frame data set to obtain a filtered data set, wherein the filtered data set includes fall data and daily data;

[0026] Acquire preset structured data, and perform video content matching processing and annotation content matching processing on the preset structured data and the screening data set to obtain a sample data set;

[0027] The sample data set is divided into a training set and an evaluation set.

[0028] Optionally, the multimodal large model-based fall detection method, wherein the acquiring of text prompt information, inputting the text prompt information and the suspicious video frame into a fall re-examination model, and outputting a behavior analysis report, specifically includes:

[0029] Obtaining text prompt information, and inputting the text prompt information and the suspicious video frame into the fall review model;

[0030] The text prompt information and the suspicious video frame are coded and dimensionally aligned using the fall re-examination model to obtain a behavior analysis report.

[0031] Optionally, the fall detection method based on a multimodal large model, wherein the encoding and dimension alignment of the text prompt information and the suspicious video frame by the fall re-examination model to obtain a behavior analysis report, specifically includes:

[0032] Performing block processing and encoding processing on the suspicious video frame using the fall re-examination model to obtain a video encoding result;

[0033] Performing text encoding processing on the text prompt information through the fall re-examination model to obtain a text encoding result;

[0034] Performing dimension alignment processing on the video encoding result and the text encoding result to obtain a target encoding result;

[0035] The target coding result is subjected to structured generation processing to obtain a behavior analysis report.

[0036] In addition, to achieve the above-mentioned object, the present invention further provides a fall detection system based on a multimodal large model, wherein the fall detection system based on the multimodal large model includes:

[0037] A fall detection module is used to obtain real-time video information and input the real-time video information into a fall detection model to obtain suspicious video frames;

[0038] A fall recheck module is used to obtain text prompt information, input the text prompt information and the suspicious video frame into a fall recheck model, and output a behavior analysis report;

[0039] A fall detection result output module is used to obtain a fall detection result according to the behavior analysis report.

[0040] In the present invention, real-time video information is acquired and input into a fall initial detection model to obtain a suspicious video frame; text prompt information is acquired and input into a fall re-detection model to output a behavior analysis report; and a fall detection result is obtained based on the behavior analysis report. The present invention implements a two-stage detection of fall behavior by constructing a fall initial detection model and a fall re-detection model. The fall initial detection model can quickly screen out possible fall behaviors, while the fall re-detection model can perform an in-depth analysis of possible fall behaviors and then output a fall detection result, effectively improving the efficiency of fall behavior detection and the accuracy of the fall detection result output. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of a preferred embodiment of the fall detection method based on a multimodal large model of the present invention;

[0042] Figure 2 1 is a schematic diagram of the overall flow of a preferred embodiment of the fall detection method based on a multimodal large model of the present invention;

[0043] Figure 3 2 is a schematic diagram of the processing process of the initial fall detection model of a preferred embodiment of the fall detection method based on a multimodal large model of the present invention;

[0044] Figure 4 1 is a schematic diagram of a processing flow for generating a data set in a preferred embodiment of a fall detection method based on a multimodal large model of the present invention;

[0045] Figure 5 1 is a schematic diagram of a data set preparation process of a preferred embodiment of a fall detection method based on a multimodal large model of the present invention;

[0046] Figure 6 3 is a schematic diagram of fine-tuning the model of the second pre-trained model of the preferred embodiment of the fall detection method based on the multimodal large model of the present invention;

[0047] Figure 7 2 is a schematic diagram of a specific processing flow of a fall recheck model of a preferred embodiment of the fall detection method based on a multimodal large model of the present invention;

[0048] Figure 8 Schematic diagram of the processing of text prompt information and suspicious video frames by the fall recheck model of the preferred embodiment of the fall detection method based on the multimodal large model of the present invention;

[0049] Figure 9 2 is a schematic diagram of fall detection result recognition according to a preferred embodiment of the fall detection method based on a multimodal large model of the present invention;

[0050] Figure 10 1 is a structural diagram of a preferred embodiment of a fall detection system based on a multimodal large model of the present invention;

[0051] Figure 11 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0053] Traditional fall detection systems have the following major shortcomings: 1. Low detection efficiency: Using a single-stage detection process, there is a lack of a targeted pre-screening mechanism. The system needs to conduct a comprehensive analysis of all video frames, regardless of whether they contain fall behavior, which is inefficient. 2. Limited data comprehension capabilities: Relying on static image analysis, it is unable to effectively understand the dynamic changes in continuous video frames, making it difficult for the system to capture the continuity and complexity of fall behavior, especially when multiple stages and movement changes are involved. 3. High false alarm rate: Single-shot detection by a single model is susceptible to background interference and misjudgment of non-fall actions, resulting in a high false alarm rate.

[0054] To address the above-mentioned issues, the present invention provides a fall detection method based on a multimodal large language model (MLLM). Specific objectives include: 1. Improving detection efficiency: By introducing a two-stage detection process consisting of initial inspection and re-inspection, the first stage uses a traditional fall detection model to quickly screen out possible falls. The second stage utilizes the multimodal large language model to conduct in-depth analysis of suspicious video frames, effectively reducing unnecessary computation and improving overall detection efficiency. 2. Reducing false alarm rates: The initial screening in the first stage eliminates images that clearly indicate non-fall behavior, reducing false alarms. The second stage, re-inspection using the multimodal large language model, further confirms falls. This dual guarantee ensures the accuracy of detection results, significantly reducing the false alarm rate. 3. Understanding video content: By analyzing motion changes in video content and utilizing the multimodal large language model to understand dynamic information in consecutive video frames, the method enhances the ability to understand and discern fall events. 4. Innovating data processing: Innovating in data construction and processing, the method optimizes the construction of video datasets, improves the learning and recognition capabilities of the multimodal large language model, and enhances system performance.

[0055] The fall detection method based on the multimodal large model described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the fall detection method based on the multimodal large model includes the following steps:

[0056] Step S10: Acquire real-time video information, and input the real-time video information into a fall initial detection model to obtain a suspicious video frame.

[0057] The technologies collected in the present invention include: 1. Natural Language Processing: using large language models for conversation analysis and report generation. 2. Computer Vision: using image recognition and analysis to detect abnormal changes in human posture from single-frame images, providing preliminary candidate fall events for subsequent multimodal analysis. 3. Multimodal Artificial Intelligence: integrating multiple data modalities such as text, images, and videos, through deep learning and large model analysis, to comprehensively understand and analyze suspicious fall video frames and their contexts. 4. Fall Detection System: a technical system that detects human fall events by combining various large model technologies to ensure accurate and timely fall detection.

[0058] like Figure 2As shown, the present invention proposes a fall detection system based on the multimodal large model MLLM technology. Through the two stages of initial inspection + re-inspection, it realizes the process from video information collection to fall behavior detection. The system workflow is as follows: 1. Collect open source and self-built video data sets, use special data annotation methods to construct, process and generate video-behavior analysis data sets. 2. Use the processed data set to fine-tune and train the multimodal large model as a fall re-inspection model. 3. The camera monitors in real time and collects video information as the input of the fall initial inspection model based on traditional fall detection to detect possible fall behaviors in a single picture. 4. After the suspicious video frame is passed into the fall re-inspection model, the video content of the 30 frames before and after the key frame is automatically extracted as the input of the multimodal large model, combined with the prompt information, and a corresponding behavior analysis report is generated. 5. According to the analysis report, determine whether the behavior in the video is a fall and record the results.

[0059] The present invention can give full play to the video understanding and text generation capabilities of multimodal large model technology, greatly improving the efficiency and accuracy of fall detection while reducing false alarms, providing an innovative solution for fall behavior detection.

[0060] Specifically, labeled fall images and labeled non-fall images are obtained; a first pre-trained model is determined, and the first pre-trained model is trained based on the labeled fall images and the labeled non-fall images to obtain a fall initial detection model. Real-time video information is obtained and input into the fall initial detection model; the real-time video information is subjected to image normalization and resizing processing by the fall initial detection model to obtain formatted image data; feature extraction processing is performed on the formatted image data to obtain local image features; the local image features are input into the fully convolutional network in the fall initial detection model to output suspicious key frames.

[0061] like Figure 3 As shown, the processing process of the initial fall detection model includes: 1. Receiving real-time video: obtaining video streams from common monitoring equipment such as cameras, and extracting key frames as input of the fall detection model. 2. Detection model training: In the initial stage, a large number of labeled fall and non-fall images are used to train the YOLO model (i.e., the first pre-training model in the present invention) to learn to distinguish the features of these two states, so that it can recognize fall behaviors in a single image. 3. Initial fall detection: The improved YOLO model (i.e., the initial fall detection model in the present invention) is used for preliminary fall detection, and the video behavior is preliminarily classified according to the image picture. 4. Extracting suspicious video frames: The video frames identified as suspicious fall behaviors by the initial fall detection model, as well as the 30 frames before and after the video frame are extracted as input to the fall re-detection model.

[0062] Introduction to the first pre-training model: The YOLO (You Only Look Once) model is an algorithm model for image recognition and classification tasks. Within the YOLO system, a series of pre-processing operations will first be performed on the input image, including: image standardization (image standardization is image normalization, for example, normalizing the pixel values ​​0 to 255 to 0 to 1), size adjustment (size adjustment, that is, image resolution adjustment, for example, adjusting a 1080*1920 image to 720*640), etc., to convert the original image into image data in a unified format. Then, the present invention will perform feature extraction on these image data and identify local features in the image through a convolutional neural network (CNN) (extracting features by scanning the image with a convolution kernel).

[0063] YOLO then feeds the extracted features into a fully convolutional network (FCN), which is capable of learning the complex mapping between image features and object categories. The core of the YOLO model lies in its single-shot detection mechanism, which treats the object detection task as a regression problem, predicting bounding boxes and class probabilities directly on the image.

[0064] The trained YOLO-based fall detection model has efficient fall recognition capabilities in real-time video streams, quickly and accurately detecting fall events in various complex environments, and providing accurate preliminary detection results for subsequent multimodal large model analysis.

[0065] It can be understood that the initial fall detection model is the first behavior detection module in the fall detection system. Its main function is to extract key frames based on real-time video content, initially detect the behavior of key frames, output suspicious video key frames, and output 30 frames of video content before and after the frame (for subsequent re-inspection).

[0066] Among them, a single frame of suspected fall is detected first, and then the video clips of 30 frames before and after the single frame are further reviewed to further determine whether there is a fall. Because the current frame is likely to have fallen, the 30 frames before and after are most likely to have fallen (falling is a process of several seconds, and these few seconds need to be extracted (corresponding to the 30 frames before and after the video key frame)).

[0067] Step S20: Obtain text prompt information, input the text prompt information and the suspicious video frame into a fall re-examination model, and output a behavior analysis report.

[0068] Specifically, an open source data set and a self-built data set are obtained, and semantic extraction processing and video frame annotation processing are performed on the open source data set and the self-built data set to obtain a key frame data set; data filtering processing is performed on the key frame data set to obtain a screening data set, wherein the screening data set includes fall data and daily data; preset structured data is obtained, and video content matching processing and annotation content matching processing are performed on the preset structured data and the screening data set to obtain a sample data set; data set division processing is performed on the sample data set to obtain a training set and an evaluation set.

[0069] like Figure 4 As shown in the figure, the dataset generation process includes: 1. Collecting action video datasets: We extensively collect open source action video datasets and self-built datasets, covering a variety of common action postures, including falls. 2. Invoking the dataset processing process to generate a dataset: We use innovative data structures and data processing pipelines to generate high-quality and robust datasets for subsequent model training.

[0070] The specific processing flow of the dataset generation is as follows: This paper collects a large amount of high-quality and diverse video sample data, which plays a key role in the subsequent model fine-tuning and improving the model's ability to detect fall behavior. The following is the specific process of data collection and dataset construction, such as Figure 5As shown: 1. Data collection: Collect video datasets containing various actions, including not only falling behaviors, but also other common actions such as walking, sitting, standing, bending over, attempted falls, etc., to avoid overfitting the model to specific scenarios. At the same time, taking into account the diversity of non-behavioral actions, involving scenes (mainly indoors and a small amount of outdoor), people (different genders, ages, clothing and body shapes, etc.), camera angles: multiple angles and heights (looking down, looking straight, etc.) and ambient light conditions (including bright, dim, daylight and night, etc.). In order to enhance the model's ability to understand and detect falling behaviors, we also focused on collecting negative samples, that is, videos of daily activities that do not involve falling behaviors (such as: sitting down quickly, lying on the sofa, and squatting quickly, etc.), and in order to cope with the various camera placements in the home, random cropping and rotation are adopted to reduce false positives. 2. Video frame annotation and extraction: The CLIP model is used to perform semantic extraction on the collected video frames (the CLIP model performs semantic extraction and analysis on image frames to determine whether there are people in the image), annotate the picture content, and select video frames containing people as the key frame dataset (video frames containing people in the annotated pictures are retained, and those without people are deleted to save computing resources). 3. Data filtering: Based on the characteristics of the model usage scenario and the semantic information extracted from the video, some outdoor and overly blurry and dark scenes are filtered out, and data is cleaned and divided into fall data and other daily behavior data to improve the data quality of the fine-tuning dataset. 4. Video information structuring: The filtered video slices are paired with hand-crafted structured data (structured data refers to a video clip with four labels: overall description, whether there is a fall behavior, behavior category, and behavior analysis; while pairing refers to pairing one video with one label, and combining them to form a real training data) to achieve accurate matching of video content and annotated content, preparing for subsequent fine-tuning tasks. 5. Dataset Division: The dataset was divided into an 80% training set and a 20% evaluation set. Taking into account the proportion of falls in simulated reality, the proportion of falls in the evaluation set was adjusted to 5%-10%. 6. Dataset Storage: After data processing is complete, the training and evaluation sets are stored separately for easy subsequent use.

[0071] Dataset processing is the starting point for the production of a fall detection system. Its main function is to collect a large number of diverse behavioral action video clips, generate video keyframe data and corresponding structured data, and serve as the basis for subsequent fine-tuning training.

[0072] Determine a second pre-trained model, and use the LoRA method to fine-tune the second pre-trained model according to the training set to obtain an initial fall re-examination model; use the cross-validation method to optimize the initial fall re-examination model according to the evaluation set to obtain a fall re-examination model.

[0073] like Figure 6 The figure shows the detailed process for fine-tuning the second pre-trained model, including: 1. Fine-tuning: Extracting the training set from the dataset and fine-tuning it for fall behavior detection. 2. Evaluation and Verification: Regularly selecting the evaluation set data during the fine-tuning process to evaluate the model and check the results of the model fine-tuning.

[0074] Introduction to the second pre-trained model (the second pre-trained model in the present invention preferably adopts the pre-trained MiniCPMV model): 1. Fine-tuning settings: According to the characteristics of the fall detection task, design a suitable fine-tuning scheme, select hyperparameters such as the number of fine-tuning layers, learning rate, Batch-Size, and number of training rounds, and determine the training goal (the goal is to enhance the multi-graph understanding ability of the model) and the loss function (the cross-entropy loss function is preferably used in the present invention). 2. Model fine-tuning: Load the pre-trained MiniCPMV model and the prepared video clip data into the fine-tuning script using the LoRA fine-tuning technology, and start the fine-tuning process. Among them, LoRA (Low-Rank Adaptation) technology is an efficient model fine-tuning method that fine-tunes the model behavior by introducing small, low-rank matrices in the key layers of the model without making major modifications to the entire model structure. The model learns the overall description, whether it falls, behavior category, and behavior analysis in the video data in a supervised manner, and fine-tunes the original parameters to adapt to the fall detection task. 3. Evaluation and Verification: During the fine-tuning process, the generation effect of the model is regularly evaluated. Cross-validation and other methods are used to evaluate the performance of the model in fall detection. The fine-tuning strategy is adjusted according to the evaluation results to optimize the model (the process of LoRA fine-tuning is as follows: first, all parameter layers of the pre-trained model are frozen, and then the QKV layer in the Transformer model is replaced with Q+LoRA, K+LoRA, and V+LoRA. QKV remains frozen and not trained, and only the LoRA layer is trained. This way, the entire fine-tuning process is efficient and takes up very little computing power). In the LoRA fine-tuning process, there are many adjustable parameters. According to the fine-tuning results, the parameters can be readjusted and compared for optimization. 4. Model saving: After the fine-tuning process is completed, the fine-tuned model parameters are saved for subsequent use and deployment.

[0075] The core purpose of model fine-tuning is to fine-tune the MiniCPMV model to improve its understanding and recognition of fall behaviors, strengthening its recognition capabilities for subsequent fall re-detection models. Furthermore, these video clips use a standardized format and description method to ensure the consistency of the extracted video information in the final output.

[0076] Obtain text prompt information, and input the text prompt information and the suspicious video frame into the fall review model; perform block processing and encoding processing on the suspicious video frame through the fall review model to obtain a video encoding result; perform text encoding processing on the text prompt information through the fall review model to obtain a text encoding result; perform dimension alignment processing on the video encoding result and the text encoding result to obtain a target encoding result; perform structured generation processing on the target encoding result to obtain a behavior analysis report.

[0077] like Figure 7 As shown in the figure, the specific processing flow of the fall re-examination model is as follows: 1. Collect suspicious video frames: Based on the suspicious video frames detected by the initial fall detection model, extract 30 frames before and after the video frame as part of the input of the re-examination model. 2. Text input: By inputting text instructions related to fall detection (for example, the text instructions can be: "Please judge whether there is a fall behavior in this video?" or "Please describe what happened in the video?"), the fall re-examination model based on the multimodal large model is prompted to identify and detect the specified fall behavior. 3. Re-examination of fall behavior: The collected video frames and text are input into the re-examination model at the same time, and the video clip behavior detection results are obtained through computational reasoning. 4. Behavior report output: The re-examination model generates a behavior analysis report based on the input. The format is the structured data in the previous dataset generation, which serves as the behavior analysis report result.

[0078] Among them, the results of the behavioral analysis report include: 1. The overall description of the video; 2. Whether a fall occurred (yes or no); 3. The specific behavior category (falling, sitting down quickly, squatting, lying down); 4. Behavioral analysis (what characteristics are used to judge the above behaviors).

[0079] Algorithm Model Introduction: MiniCPM-V (the second pre-trained model in this invention) is a large, on-device multimodal model for image, text, and video understanding. Built on SigLip-400M and Qwen2-7B, this model possesses video understanding and contextual learning and reasoning capabilities. After fine-tuning training, this model learns more about video understanding of falls and strengthens its ability to detect falls, forming a multimodal model specifically for fall detection.

[0080] like Figure 8As shown in the figure, the fall re-examination model processes text prompt information and suspicious video frames as follows: 1. Video encoding: Take the continuous video frames in the video clip as input, divide the continuous images into blocks and encode them, and perform feature extraction and feature fusion through the encoding layer in the model. 2. Text encoding: Use text encoding for the text prompt content, align it with the feature vector dimension obtained by video encoding (for example, if the dimension of the video is 1024 dimensions, then the dimension of the text must be 1024 dimensions. If the dimension of the text is 512 dimensions, then you have to find a way to turn it into 1024 dimensions), and embed the video encoding result into the text encoding. 3. Large language model generation: The large language model generates a structured video clip behavior analysis report based on the input of the encoding result.

[0081] The fall recheck model is a key module in the fall detection system. Its main function is to output a formatted behavior analysis report based on suspicious video clips and text prompt information, thereby realizing the identification of fall behavior.

[0082] Step S30: Obtain a fall detection result according to the behavior analysis report.

[0083] like Figure 9 As shown in FIG, the process of result output is as follows: after two rounds of fall behavior detection, a behavior analysis report is output, and based on this, it is determined whether the video clip content contains fall behavior, and a final judgment is made to obtain the result.

[0084] Beneficial effects of the present invention:

[0085] 1. Significantly improve response speed: Automatically detect suspicious falls through artificial intelligence, achieve end-to-end automated processing from real-time monitoring to rapid response, and shorten the detection cycle from minutes to seconds. 2. Significantly reduce false alarm rate: Utilize a two-stage detection process, with the first stage quickly screening suspicious frames and the second stage conducting in-depth analysis and confirmation, replacing the traditional single detection method, significantly reducing false alarms and improving detection accuracy. 3. Enhance the depth of video understanding: Through in-depth analysis of video sequences using multimodal large models, it surpasses the detection limitations of traditional single-frame images, can accurately capture the dynamic process of fall events, and stimulate deeper video understanding capabilities. 3. Stable output of high-accuracy detection results: The system is trained on a large amount of high-quality data to ensure that the generated detection results have a high level of accuracy and stability, reducing reliance on manual review.

[0086] In summary, this invention provides a novel solution to the problems of low efficiency, limited data comprehension, and high false alarm rates in traditional fall detection technology. By combining the advanced YOLO network with a large multimodal model, this invention enables efficient and accurate detection of falls in video streams, significantly improving the accuracy and reliability of fall detection.

[0087] The data construction and processing part of the present invention is the core of the entire system of the present invention. Through the careful construction and processing of video frames, the present invention can better capture the dynamic process of the fall event, thereby providing richer and more accurate input data for the multimodal large model. This innovation in data processing flow not only improves the performance of the system, but also provides new ideas for the development of fall detection technology. In addition, the present invention also innovatively introduces the "initial screening + re-inspection" detection mechanism to form an efficient and accurate detection process. In the initial screening stage, the YOLO network can quickly screen out possible fall behaviors, and in the re-inspection stage, the multimodal large model conducts in-depth analysis of these suspicious frames to ensure the high accuracy of the detection results. This mechanism not only improves the detection efficiency, but also ensures the high quality and consistency of the generated content.

[0088] The present invention adopts a modular architecture design, and each module interacts with data through a standardized interface. This decoupling design allows each module to independently perform algorithm upgrades and performance optimization, and the entire system has the ability to continuously evolve and self-improve. At the same time, this architecture also facilitates the introduction of more types of detection models in the future and expands the detection capabilities of the system. The focus of the present invention is on the innovation of fall detection procedures and data processing. By deeply integrating artificial intelligence technology with fall detection, a new fall detection paradigm has been created. This paradigm not only improves the efficiency and accuracy of detection, but also provides a new data processing and construction process for the fall detection model, and provides more possibilities for research and application in related fields. Through the application of the present invention, false alarms and missed alarms can be effectively reduced, the reliability of fall detection can be improved, and safer monitoring protection can be provided for the elderly and people who need special care.

[0089] Furthermore, if Figure 10 As shown, based on the above-mentioned fall detection method based on the multimodal large model, the present invention also provides a fall detection system based on the multimodal large model, wherein the fall detection system based on the multimodal large model includes:

[0090] A fall detection module 51 is used to obtain real-time video information and input the real-time video information into a fall detection model to obtain suspicious video frames;

[0091] A fall recheck module 52 is configured to obtain text prompt information, input the text prompt information and the suspicious video frame into a fall recheck model, and output a behavior analysis report;

[0092] The fall detection result output module 53 is configured to obtain a fall detection result according to the behavior analysis report.

[0093] Furthermore, if Figure 11As shown, based on the above-mentioned fall detection method and system based on the multimodal large model, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 11 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0094] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a fall detection program 40 based on a multimodal large model is stored on the memory 20, and the fall detection program 40 based on the multimodal large model can be executed by the processor 10, thereby realizing the fall detection method based on the multimodal large model in the present application.

[0095] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program code or process data stored in the memory 20, such as executing the multimodal large model-based fall detection method.

[0096] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0097] In one embodiment, when the processor 10 executes the multimodal large model-based fall detection program 40 in the memory 20 , the steps of the multimodal large model-based fall detection method described above are implemented.

[0098] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a fall detection program based on a multimodal large model, and when the fall detection program based on the multimodal large model is executed by a processor, the steps of the fall detection method based on the multimodal large model as described above are implemented.

[0099] In summary, the present invention provides a fall detection method, system, and terminal based on a multimodal large model. The method comprises: obtaining real-time video information, and inputting the real-time video information into a fall initial detection model to obtain a suspicious video frame; obtaining text prompt information, and inputting the text prompt information and the suspicious video frame into a fall re-detection model to output a behavior analysis report; and obtaining a fall detection result based on the behavior analysis report. The present invention implements a two-stage detection of fall behavior by constructing a fall initial detection model and a fall re-detection model. The fall initial detection model can quickly screen out possible fall behaviors, and the fall re-detection model can perform an in-depth analysis of possible fall behaviors, and then output a fall detection result, effectively improving the detection efficiency of fall behaviors and the accuracy of the fall detection result output.

[0100] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0101] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0102] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A fall detection method based on a multimodal large model, characterized in that: The fall detection method based on the multimodal large model includes: Acquire real-time video information, and input the real-time video information into a fall initial detection model to obtain a suspicious video frame; Obtaining text prompt information, and inputting the text prompt information and the suspicious video frame into a fall re-examination model, and outputting a behavior analysis report; A fall detection result is obtained according to the behavior analysis report.

2. The fall detection method based on a multimodal large model according to claim 1, characterized in that: The acquiring of real-time video information and inputting the real-time video information into the fall initial detection model to obtain a suspicious video frame specifically includes: Obtaining labeled fall images and labeled non-fall images; Determine a first pre-trained model, and perform model training on the first pre-trained model based on the labeled fall image and the labeled non-fall image to obtain a fall initial detection model; Acquiring real-time video information and inputting the real-time video information into the fall initial detection model; The real-time video information is processed by using the fall initial detection model to identify suspicious video frames to obtain suspicious video frames.

3. The fall detection method based on a multimodal large model according to claim 2, characterized in that: The process of performing suspicious video frame identification processing on the real-time video information by using the fall initial detection model to obtain the suspicious video frame specifically includes: Performing image standardization and size adjustment on the real-time video information using the fall initial detection model to obtain formatted image data; Performing feature extraction processing on the formatted image data to obtain local features of the image; The local features of the image are input into the fully convolutional network in the fall initial detection model to output suspicious key frames.

4. The fall detection method based on a multimodal large model according to claim 1, characterized in that: The step of obtaining text prompt information, inputting the text prompt information and the suspicious video frame into a fall re-examination model, and outputting a behavior analysis report may also include: Obtaining an open source dataset and a self-built dataset, and preprocessing the open source dataset and the self-built dataset to obtain a training set and an evaluation set; Determine a second pre-trained model, and use the LoRA method to fine-tune the second pre-trained model according to the training set to obtain an initial fall re-examination model; The cross-validation method is used to perform model optimization processing on the initial fall re-examination model according to the evaluation set to obtain a fall re-examination model.

5. The fall detection method based on a multimodal large model according to claim 4, characterized in that: The preprocessing of the open source dataset and the self-built dataset to obtain a training set and an evaluation set specifically includes: Performing semantic extraction processing and video frame annotation processing on the open source dataset and the self-built dataset to obtain a key frame dataset; Performing data filtering processing on the key frame data set to obtain a filtered data set, wherein the filtered data set includes fall data and daily data; Acquire preset structured data, and perform video content matching processing and annotation content matching processing on the preset structured data and the screening data set to obtain a sample data set; The sample data set is divided into a training set and an evaluation set.

6. The fall detection method based on a multimodal large model according to claim 1, characterized in that: The obtaining of text prompt information, inputting the text prompt information and the suspicious video frame into the fall re-examination model, and outputting a behavior analysis report specifically includes: Obtaining text prompt information, and inputting the text prompt information and the suspicious video frame into the fall review model; The text prompt information and the suspicious video frame are coded and dimensionally aligned using the fall re-examination model to obtain a behavior analysis report.

7. The fall detection method based on a multimodal large model according to claim 6, characterized in that: The encoding and dimension alignment of the text prompt information and the suspicious video frame by the fall re-examination model to obtain a behavior analysis report specifically includes: Performing block processing and encoding processing on the suspicious video frame using the fall re-examination model to obtain a video encoding result; Performing text encoding processing on the text prompt information through the fall re-examination model to obtain a text encoding result; Performing dimension alignment processing on the video encoding result and the text encoding result to obtain a target encoding result; The target coding result is subjected to structured generation processing to obtain a behavior analysis report.

8. A fall detection system based on a multimodal large model, characterized in that: The fall detection system based on the multimodal large model includes: A fall detection module is used to obtain real-time video information and input the real-time video information into a fall detection model to obtain suspicious video frames; A fall recheck module is used to obtain text prompt information, input the text prompt information and the suspicious video frame into a fall recheck model, and output a behavior analysis report; A fall detection result output module is used to obtain a fall detection result according to the behavior analysis report.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a fall detection program based on a multimodal large model stored in the memory and executable on the processor. When the fall detection program based on the multimodal large model is executed by the processor, the steps of the fall detection method based on the multimodal large model according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a fall detection program based on a multimodal large model. When the fall detection program based on the multimodal large model is executed by a processor, the steps of the fall detection method based on a multimodal large model according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Ward monitoring method and system based on multi-modal large model and edge calculation, terminal and storage medium

    CN121053587A