Fall detection method and device, electronic equipment and storage medium

By improving the convolutional layer of the YOLO model, and using image recognition technology to achieve efficient and accurate fall detection, the problems of inefficiency and strong device dependence in traditional methods are solved, and are suitable for intelligent monitoring and public safety fields.

CN120340068APending Publication Date: 2025-07-18GUANGZHOU GAOKE COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424650.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, fall detection methods rely on manual diagnosis and are inefficient, wearable sensor devices are highly dependent, lack real-time and costly.

Method used

The convolution layer based on spatially separable convolution and reproducible parameterized convolution blocks are used to replace the YOLO model, and the object detection model is constructed, and fall detection is achieved through image recognition.

Benefits of technology

It improves the recognition accuracy of fall detection, reduces the system deployment cost, enhances the robustness of the model in different scenarios, and is suitable for environments with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340068A_ABST
    Figure CN120340068A_ABST
Patent Text Reader

Abstract

The invention discloses a fall detection method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a personnel state image set; personnel state images in the personnel state image set are marked with personnel state categories; performing detection training on a preset target detection model by taking the personnel state image set as a data set to obtain a personnel state detection model; wherein the target detection model is constructed based on space separable convolution and a reparameterized convolution block replacing a convolution layer in a YOLO model; performing state detection on the acquired image of the to-be-detected person by using the person state detection model to obtain a target state category; and determining whether the to-be-detected person falls down based on the target state category. According to the method, the standard convolution layer of the YOLO model is replaced by the space separable convolution and the re-parameterized convolution block, the multi-scale feature extraction capability of the model is enhanced, fall detection can be efficiently and accurately realized, and the method can be widely applied to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a fall detection method, device, electronic device, and storage medium. Background Art

[0002] For some specific groups of people (such as the elderly, children, or factory workers), in order to achieve effective care, it is necessary to detect the falls of relevant people. However, the traditional manual diagnosis method highly depends on the manual detection of caregivers, which is obviously inefficient. In related technologies, there are also fall detection methods that achieve fall detection through wearable sensors or manual monitoring, but these methods have problems such as strong device dependence, insufficient real-time performance, and high costs. Summary of the Invention

[0003] The main purpose of the embodiments of the present invention is to propose a fall detection method, device, electronic device, and storage medium, in order to solve at least one problem of the prior art. The present invention can efficiently and accurately achieve fall detection.

[0004] To achieve the above object, on the one hand, an embodiment of the present invention proposes a fall detection method, and the method includes:

[0005] Obtain a set of human state images; the human state images in the set of human state images are labeled with human state categories;

[0006] Use the set of human state images as a data set to detect and train a preset target detection model to obtain a human state detection model;

[0007] Wherein, the target detection model is constructed by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks;

[0008] Use the human state detection model to perform state detection on the captured image of the person to be detected to obtain a target state category;

[0009] Determine whether the person to be detected has fallen based on the target state category.

[0010] In some embodiments, obtaining the set of human state images includes the following steps:

[0011] Obtain human state images of various people in different scenarios;

[0012] In response to the labeling instruction of the target object, use a bounding box to select the position of the person included in the human state image and label the human state category of the corresponding person in the bounding box;

[0013] Wherein, the human state categories include a fall category and a non-fall category, and the non-fall category includes at least one of a standing category or a squatting category.

[0014] In some embodiments, taking the set of personnel status images as a dataset to detect and train a preset target detection model to obtain a personnel status detection model, includes the following steps:

[0015] Construct a target detection model by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks;

[0016] Input the personnel status images in the set of personnel status images into the target detection model for processing to obtain the status prediction categories;

[0017] Based on the status prediction categories, combine with the personnel status categories to construct a loss function, and iteratively converge and adjust the model parameters of the target detection model based on the loss function to obtain the personnel status detection model.

[0018] In some embodiments, the YOLO model uses YOLOv8; constructing a target detection model by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks, includes the following steps:

[0019] Replace the convolutional layers in the backbone network of the YOLO model with spatially separable convolutions and reparameterizable convolution blocks;

[0020] Replace the convolutional layers in the neck structure of the YOLO model with spatially separable convolutions.

[0021] In some embodiments, replacing the convolutional layers in the backbone network of the YOLO model with spatially separable convolutions and reparameterizable convolution blocks, includes the following steps:

[0022] Replace the first two convolutional layers in the backbone network of the YOLO model with spatially separable convolutions;

[0023] Replace the convolutional layers in the deep part of the backbone network of the YOLO model with reparameterizable convolution blocks.

[0024] In some embodiments, the method further includes the following steps:

[0025] Construct a target detection model using a YOLO model with a target quantization architecture in response to the target detection requirements;

[0026] Among them, the target detection requirements include real-time requirements and accuracy requirements.

[0027] In some embodiments, constructing a target detection model using a YOLO model with a target quantization architecture in response to the target detection requirements, includes the following steps:

[0028] Obtain the first requirement weight for real-time requirements and the second requirement weight for accuracy requirements;

[0029] Determine the requirement value of the target detection requirement according to the ratio of the first requirement weight to the second requirement weight;

[0030] Determine the architecture size of the target quantization architecture according to the requirement value mapping;

[0031] Wherein, the numerical size of the requirement value is negatively correlated with the architecture size of the target quantization architecture.

[0032] To achieve the above object, another aspect of the embodiments of the present invention proposes a fall detection device, the device includes:

[0033] A first module, configured to obtain a set of personnel status images; the personnel status images in the set of personnel status images are labeled with personnel status categories;

[0034] A second module, configured to use the set of personnel status images as a data set to detect and train a preset target detection model to obtain a personnel status detection model;

[0035] Wherein, the target detection model is constructed by replacing the convolutional layer in the YOLO model with a spatially separable convolution and a reparameterizable convolution block;

[0036] A third module, configured to use the personnel status detection model to perform status detection on the collected images of the person to be detected to obtain a target status category;

[0037] A fourth module, configured to determine whether the person to be detected has fallen based on the target status category.

[0038] In some embodiments, the device further includes:

[0039] A fifth module, configured to construct a target detection model using the YOLO model of the target quantization architecture in response to the target detection requirement;

[0040] Wherein, the target detection requirement includes a real-time requirement and an accuracy requirement.

[0041] To achieve the above object, another aspect of the embodiments of the present invention proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0042] To achieve the above object, another aspect of the embodiments of the present invention proposes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0043] The embodiments of the present invention have at least the following beneficial effects: The present invention provides a fall detection method, device, electronic device, and storage medium. The solution obtains a set of personnel status images; the personnel status images in the set of personnel status images are labeled with personnel status categories; the set of personnel status images is used as a data set to detect and train a preset target detection model to obtain a personnel status detection model; among them, the target detection model is constructed by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks; the personnel status detection model is used to detect the status of the collected images of the person to be detected to obtain the target status category; based on the target status category, it is determined whether the person to be detected has fallen. By replacing the standard convolutional layers of the YOLO model with spatially separable convolutions and reparameterizable convolution blocks, the present invention enhances the model's ability to extract multi-scale features, thereby improving the recognition accuracy of personnel status (such as falling, standing, etc.); the optimized target detection model can better adapt to different scenarios and improve the robustness of personnel status detection. Compared with traditional sensor solutions, the present invention can achieve high-precision fall detection only with ordinary cameras and computing devices, reducing the system deployment cost and enhancing the applicability. The present invention can achieve efficient and accurate personnel status detection in an environment with limited computing resources, providing reliable technical support for fields such as intelligent guardianship and public safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flowchart of the fall detection method provided by the embodiments of the present invention;

[0045] Figure 2 is a schematic diagram of the original YOLOv8 network structure provided by the embodiments of the present invention;

[0046] Figure 3 is a schematic diagram of the network result of the target detection model provided by the embodiments of the present invention;

[0047] Figure 4 is a schematic diagram of the structure of the RepVGGBlock module provided by the embodiments of the present invention;

[0048] Figure 5 is a schematic diagram of the structure of the fall detection device provided by the embodiments of the present invention;

[0049] Figure 6 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0051] It can be understood that the terms "first", "second", etc. used in the present invention can be used in the present invention to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information can also be referred to as the second information. Similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0052] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present invention, at least one includes one, two, or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0053] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the present invention are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0054] The fall detection method provided by the embodiments of the present invention relates to the technical field of data processing. The fall detection method provided by the embodiments of the present invention can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the fall detection method, etc., but is not limited to the above forms.

[0055] The present invention can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0056] Figure 1 is an optional flowchart of the fall detection method provided by an embodiment of the present invention, Figure 1 The method in may include but is not limited to steps S100 to S400.

[0057] S100. Obtain a set of human state images;

[0058] Among them, the human state images in the set of human state images are labeled with human state categories;

[0059] It should be noted that in some embodiments, step S100 may include the following steps: obtain human state images of various types of people in different scenarios; in response to a labeling instruction of a target object, use a bounding box to select the position of the person in the human state image and label the human state category of the corresponding person in the bounding box; among them, the human state category includes a fall category and a non-fall category, and the non-fall category includes at least one of a standing category or a squatting category.

[0060] Exemplarily, in some specific embodiments, images of people falling in different scenarios can be collected, and then a human fall data set can be labeled. Specifically, in the form of manual labeling, by labeling bounding boxes, two categories of standing and / or squatting and falling can be labeled. In addition, various posture categories (such as bending over, half-squatting, etc.) during the process from squatting to standing can also be added.

[0061] S200. Use the set of human state images as a data set to detect and train a preset target detection model to obtain a human state detection model;

[0062] Among them, the target detection model is constructed by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks;

[0063] It should be noted that, in some embodiments, step S200 may include the following steps: constructing a target detection model by replacing the convolution layer in the YOLO model based on spatially separable convolution and reparameterizable convolution blocks; inputting the personnel status images in the personnel status image set into the target detection model for processing to obtain the status prediction category; constructing a loss function based on the status prediction category combined with the personnel status category, and iteratively converging and adjusting the model parameters of the target detection model based on the loss function to obtain the personnel status detection model. Among them, in some specific application scenarios, the personnel status image set can be divided into a training set, a test set, and a verification set for model training, test fine-tuning, and effect verification.

[0064] For example, in some specific implementations, taking YOLOv8 as an example, the training of the personnel status detection model based on YOLOv8 can be implemented as follows:

[0065] 1. Data preparation stage:

[0066] ① Construction of personnel status image set:

[0067] Collection scenario: images of personnel operations (such as standing, bending over, raising hands, falling to the ground, etc.) acquired by factory operation monitoring cameras.

[0068] Data labeling: Use the LabelImg tool to label the bounding boxes and status categories of people in the image (for example: falling, squatting, and standing).

[0069] Dataset division: training set (70%): 7,000 annotated images; validation set (10%): 1,000 images (final performance evaluation); test set (20%): 2,000 images (for hyperparameter tuning). The training set can be divided according to the actual amount of data, and the default ratio is training: validation: test = 7:1:2. Generally, the training set has the most data, followed by the test set, and finally the validation set.

[0070] 2. Model training and optimization:

[0071] Target detection model selection: Use YOLOv8 pre-trained model.

[0072] The loss function can be the loss function that comes with YOLO.

[0073] Training parameters: initial learning rate 0.01, cosine annealing schedule; Batch Size = 32, 300 iterations, early stopping mechanism (termination if the mAP of the test set does not improve for 10 consecutive rounds).

[0074] 3. Verification and testing:

[0075] Testing stage: The mAP@0.5 metric of the test set can be monitored, and the data augmentation strategy can be adjusted (such as adding random rotation and lighting perturbation).

[0076] Verification result: If the mAP@0.5 of the validation set reaches the preset threshold, the detection requirements are met.

[0077] Among them, in some embodiments, the YOLO model adopts YOLOv8; a target detection model is constructed by replacing the convolutional layers in the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks, which may include the following steps: replacing the convolutional layers in the backbone network of the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks; replacing the convolutional layers in the neck structure of the YOLO model with depthwise separable convolutions.

[0078] Among them, in some embodiments, replacing the convolutional layers in the backbone network of the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks may include the following steps: replacing the first two convolutional layers in the backbone network of the YOLO model with depthwise separable convolutions; replacing the convolutional layers in the deep part of the backbone network of the YOLO model with reparameterizable convolution blocks.

[0079] Exemplarily, in some specific embodiments, the present invention mainly improves the YOLOv8 algorithm to optimize automatic fall detection. Among them, the original YOLOv8 network structure is as Figure 2 shown, and the network structure optimization implemented in the embodiments of the present invention is as follows. The optimized network structure is as Figure 3 shown:

[0080] 1. The present invention replaces the first two convolutional layers of the back-bone (backbone network) with depthwise separable convolutions (SSCONV), which improves the convolution efficiency.

[0081] 2. Replace two convolutions in the neck part with SSCONV.

[0082] 3. In the deep part of the back-bone, use the RepVGGBlock (i.e., reparameterizable convolution block) module to replace the convolutional layer, which is beneficial to feature extraction. Among them, the structure of the RepVGGBlock is as Figure 4 shown. Its main function is to be faster and more memory-saving during the stacking of multiple modules, and to extract features more effectively in the deep backbone network. Its structure is mainly composed of three branches: a 3x3 convolution branch, a 1x1 convolution branch, and an identity mapping branch (corresponding to the 3 branches from left to right in Figure 4 order).

[0083] For the convenience of explaining the specific composition of the network structure, the following is for Figure 2 , Figure 3 and Figure 4The proprietary names of the components of the network structure in [0] are explained as follows: Input represents the input of the image; CONV represents the convolutional layer; SSCONV represents the spatially separable convolution; C2f is a component in the YOLO algorithm, mainly used for feature extraction and feature fusion; RepVGGBlock represents the reparameterizable convolutional block; SPPF represents the fast spatial pyramid pooling layer; Up-Sample represents upsampling; Concat represents concatenation; Detect represents making prediction operations on the extracted feature map; BN represents the batch normalization layer; Activation represents the activation function.

[0084] It should also be noted that in some embodiments, the method may further include the following steps: constructing a target detection model using a YOLO model with a target quantization architecture in response to the target detection requirements; wherein, the target detection requirements include real-time requirements and accuracy requirements.

[0085] Exemplarily, in some specific embodiments, specifically, according to the emphasis on different requirements for fall detection in different application scenarios, different quantization architecture versions can be selected as the model application.

[0086] Wherein, in some embodiments, constructing a target detection model using a YOLO model with a target quantization architecture in response to the target detection requirements may include the following steps: obtaining the first requirement weight of the real-time requirement and the second requirement weight of the accuracy requirement; determining the requirement value of the target detection requirement according to the ratio of the first requirement weight to the second requirement weight; determining the architecture size of the target quantization architecture according to the mapping of the requirement value; wherein, the numerical size of the requirement value is negatively correlated with the architecture size of the target quantization architecture.

[0087] Exemplarily, in some specific embodiments, YOLOv8 has five different sizes (the quantization architecture is gradually improved from lightweight) of models, namely YOLOv8n, YOLOv8s, YOLOm, YOLOv8l, and YOLOv8x. Different-sized models can be used according to different usage scenario requirements. Specifically, generally speaking, the larger the model, the slower the inference speed and the higher the required video memory. If the hardware device requires deployment on edge devices, choose YOLOv8n or YOLOv8s. If the hardware resource performance is good and real-time performance is not required, and only detection accuracy is required, a larger (such as YOLOv8l or YOLOv8x) model can be selected according to the hardware. The five preset parameters can evaluate whether the accuracy and detection speed of the model can meet the application scenario during the training process. If the detection accuracy cannot be met during the testing process, data needs to be urgently collected and optimized in the real testing scenario.

[0088] In some specific application scenarios, for example, if the demand value falls within the range of 0.9 to 1.1 (the intermediate range), it can be considered that the requirements for real-time performance and accuracy in the detection demand are similar. Under balance, YOLOm can be adopted; if the demand value is less than 0.9 (the small range), it can be considered that the requirement for accuracy in the detection demand is relatively high, and YOLOv8l or YOLOv8x can be adopted. Further, a threshold node (such as 0.2) can be further set to distinguish the application conditions of YOLOv8l and YOLOv8x. If the demand value is less than 0.2 (the extremely small range), YOLOv8x is adopted, otherwise YOLOv8l is adopted; if the demand value is greater than 1.1 (the large range), it can be considered that the requirement for real-time performance in the detection demand is relatively high, and YOLOv8n or YOLOv8s can be adopted. Further, a threshold node (such as 5) can be further set to distinguish the application conditions of YOLOv8n and YOLOv8s. If the demand value is greater than 5, YOLOv8n is adopted, otherwise YOLOv8s is adopted.

[0089] S300. Use the personnel status detection model to perform status detection on the captured image of the person to be detected, and obtain the target status category.

[0090] Exemplarily, in some specific embodiments, the captured image of the person to be detected is input into the personnel status detection model trained through the foregoing steps to process and obtain the target status category of the relevant person to be detected.

[0091] S400. Determine whether the person to be detected has fallen based on the target status category.

[0092] Exemplarily, in some specific embodiments, if the target status category is the fall category, it can be confirmed that the person to be detected has fallen. If the target status category is the non-fall category (such as standing, squatting, etc.), it can be confirmed that the person to be detected has not fallen.

[0093] To explain the principle of the technical solution of the present invention in detail, the overall process of the present invention will be described below in conjunction with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.

[0094] First of all, it should be noted that for the design of existing large-scale automatic fall detection systems, the traditional manual diagnosis method is obviously inefficient and cannot meet the needs of current fall detection methods.

[0095] With the continuous development of the economy and computer technology, artificial intelligence technology has been widely developed. The solution of combining artificial intelligence technology with computers has always received extensive attention. The existing deep convolutional neural network has been greatly developed. In the field of image recognition, many scholars have designed different network architectures to solve image processing problems in different fields. Their goals are all to achieve automated image recognition faster and better, thereby replacing manual labor and improving the diagnosis efficiency and accuracy.

[0096] In order to improve the efficiency of automatically detecting people falling, the present invention provides a method for detecting falls by improving YOLOv8. As Figure 2 、 Figure 3 and Figure 4 , the present invention mainly solves the following problems:

[0097] 1. The present invention mainly improves the YOLOv8 algorithm to optimize the automatic recognition of fall detection.

[0098] 2. The present invention replaces the first two layers of convolution in the back-bone with spatially separable convolution (SSCONV), which improves the convolution efficiency.

[0099] 3. Replace two convolutions with SSCONV in the neck part.

[0100] 4. Replace the convolution layer with the RepVGGBlock module in the deep part of the back-bone, which is beneficial to feature extraction.

[0101] The technical solution of the present invention can be mainly realized as follows:

[0102] S1. Dataset preparation stage:

[0103] Collect images of people falling in different scenarios and label the dataset of people falling.

[0104] S2. Model training stage:

[0105] Divide the dataset into training set, test set, and validation set for training. The training parameters can be adjusted according to the actual training equipment situation. A better model is obtained.

[0106] S3. Model application:

[0107] YOLOv8 has five different-sized models: YOLOv8n, YOLOv8s, YOLOm, YOLOv8l, and YOLOv8x. Different-sized models can be used according to different usage scenario requirements.

[0108] Specifically, the parameters in YOLOv8 are optional. There are five different parameters in the model file, which are:

[0109] [depth, width, max_channels],

[0110] n: [0.33, 0.25, 1024],

[0111] s: [0.33, 0.50, 1024],

[0112] m: [0.67, 0.75, 768],

[0113] l: [1.00, 1.00, 512],

[0114] x: [1.00, 1.25, 512].

[0115] Generally speaking, for the five different models, the larger the model, the slower the inference speed and the higher the required video memory. If the hardware device requires deployment on edge devices, choose YOLOv8n or YOLOv8s. If the hardware resource performance is good and real-time performance is not required, and only detection accuracy is required, a larger model can be selected according to the hardware.

[0116] For the five preset parameters, the accuracy and detection speed of the model can be evaluated during the training process to see if they can meet the application scenario. If the detection accuracy cannot be met during the testing process, data needs to be urgently collected and optimized in the real test scenario.

[0117] In summary, in order to improve the efficiency of automatically detecting people's falls, the present invention provides an improved YOLOv8 fall detection method. Compared with the prior art, the present invention has at least the following beneficial effects:

[0118] Improve detection accuracy: By replacing the standard convolutional layer of the YOLO model with depthwise separable convolution and reparameterizable convolution blocks, the model's ability to extract multi-scale features is enhanced, thereby improving the recognition accuracy of human states (such as falls, standing, etc.).

[0119] Reduce computational complexity: Depthwise separable convolution can reduce the number of parameters and computational volume, and reparameterizable convolution blocks use different structures during training and inference phases to optimize computational efficiency, making the model more lightweight while maintaining high performance and suitable for edge device deployment.

[0120] Enhance generalization ability: The optimized object detection model can better adapt to complex scenarios such as different lighting, occlusion, and poses, improving the robustness of human state detection.

[0121] Real-time optimization: The improved model has higher computational efficiency during the inference phase, can meet the needs of real-time monitoring and fall detection, and is suitable for application scenarios with high real-time requirements such as elderly care monitoring and intelligent security.

[0122] Reduce hardware dependence: Compared with traditional sensor solutions, this technical solution only requires an ordinary camera and a computing device to achieve high-precision fall detection, reducing the system deployment cost and improving applicability.

[0123] Through the above technical improvements, the present invention can achieve efficient and accurate detection of personnel status in an environment with limited computing resources, providing reliable technical support for fields such as intelligent guardianship and public safety.

[0124] On the other hand, as Figure 5 shown, the embodiment of the present invention also provides a fall detection device 900, which may include:

[0125] A first module 901, configured to obtain a set of personnel status images; the personnel status images in the set of personnel status images are labeled with personnel status categories;

[0126] A second module 902, configured to use the set of personnel status images as a data set to detect and train a preset target detection model to obtain a personnel status detection model;

[0127] Wherein, the target detection model is constructed by replacing the convolutional layers in the YOLO model with spatially separable convolutions and reparameterizable convolution blocks;

[0128] A third module 903, configured to use the personnel status detection model to perform status detection on the captured images of the person to be detected to obtain a target status category;

[0129] A fourth module 904, configured to determine whether the person to be detected has fallen based on the target status category.

[0130] In some embodiments, the device may further include:

[0131] A fifth module, configured to construct a target detection model using the YOLO model with a target quantization architecture in response to a target detection requirement;

[0132] Wherein, the target detection requirements include real-time requirements and accuracy requirements.

[0133] The content of the method embodiment of the present invention is applicable to the device embodiment of the present invention. The functions specifically implemented by the device embodiment of the present invention are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method.

[0134] The embodiment of the present invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above fall detection method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0135] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0136] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device 1000 according to another embodiment. The electronic device 1000 includes:

[0137] A processor 1001, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;

[0138] A memory 1002, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the fall detection method of the embodiments of the present invention;

[0139] An input / output interface 1003, which is used to implement information input and output;

[0140] A communication interface 1004, which is used to implement communication interaction between this device and other devices, and can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0141] A bus 1005, which transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);

[0142] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.

[0143] The embodiments of the present invention also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above fall detection method is implemented.

[0144] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0145] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0146] The fall detection method, fall detection device, electronic device, and storage medium provided by the embodiments of the present invention obtain a set of human state images; the human state images in the set of human state images are labeled with human state categories; the set of human state images is used as a data set to detect and train a preset target detection model to obtain a human state detection model; wherein, the target detection model is constructed by replacing the convolutional layers in the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks; the human state detection model is used to perform state detection on the captured images of the person to be detected to obtain the target state category; based on the target state category, it is determined whether the person to be detected has fallen. By replacing the standard convolutional layers of the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks, the present invention enhances the model's ability to extract multi-scale features, thereby improving the recognition accuracy of human states (such as falling, standing, etc.); the optimized target detection model can better adapt to different scenarios and improve the robustness of human state detection. Compared with traditional sensor solutions, the present invention only requires an ordinary camera and a computing device to achieve high-precision fall detection, reducing the system deployment cost and improving the applicability. The present invention can achieve efficient and accurate human state detection in an environment with limited computing resources, providing reliable technical support for fields such as intelligent guardianship and public safety.

[0147] The embodiments described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation to the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0148] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0149] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0151] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0152] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0153] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or units can be in electrical, mechanical or other forms.

[0154] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.

[0157] The preferred embodiments of the embodiments of the present invention have been described above with reference to the drawings, but this does not limit the scope of rights of the embodiments of the present invention. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of rights of the embodiments of the present invention.

Claims

1. A fall detection method, characterized in that, The method includes the following steps: Obtain a set of human state images; the human state images in the set of human state images are labeled with human state categories; Use the set of human state images as a data set to detect and train a preset target detection model to obtain a human state detection model; Among them, the target detection model is constructed by replacing the convolutional layers in the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks; Use the human state detection model to perform state detection on the captured images of the person to be detected to obtain the target state category; Determine whether the person to be detected has fallen based on the target state category.

2. The fall detection method according to claim 1, characterized in that The obtaining of the set of human state images includes the following steps: Obtain the human state images of various people in different scenarios; In response to the annotation instruction of the target object, use a bounding box to select the position of the person in the human state image and label the human state category of the corresponding person in the bounding box; Among them, the human state categories include a fall category and a non-fall category, and the non-fall category includes at least one of a standing category or a squatting category.

3. The fall detection method according to claim 1, characterized in that The using the set of human state images as a data set to detect and train a preset target detection model to obtain a human state detection model includes the following steps: Construct the target detection model by replacing the convolutional layers in the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks; Input the human state images in the set of human state images into the target detection model for processing to obtain the state prediction category; Construct a loss function based on the state prediction category and the human state category, and iteratively converge and adjust the model parameters of the target detection model based on the loss function to obtain the human state detection model.

4. The fall detection method according to claim 3, wherein The YOLO model uses YOLOv8; the constructing of the target detection model by replacing the convolutional layers in the YOLO model with depthwise separable convolutions and reparameterizable convolution blocks includes the following steps: Use depthwise separable convolutions and reparameterizable convolution blocks to replace the convolutional layers in the backbone network of the YOLO model; Replace the convolutional layers in the neck structure of the YOLO model with depthwise separable convolutions.

5. The fall detection method according to claim 4, wherein The using depthwise separable convolutions and reparameterizable convolution blocks to replace the convolutional layers in the backbone network of the YOLO model includes the following steps: Replace the first two convolutional layers in the backbone network of the YOLO model with depthwise separable convolutions; Replace the convolutional layers in the deep part of the backbone network of the YOLO model with reparameterizable convolution blocks.

6. The fall detection method according to claim 1, wherein The method further includes the following steps: Construct the target detection model using the YOLO model with a target quantization architecture in response to the target detection requirement; Among them, the target detection requirements include real-time requirements and accuracy requirements.

7. The fall detection method according to claim 6, wherein The constructing the target detection model using the YOLO model with a target quantization architecture in response to the target detection requirement includes the following steps: Obtain the first requirement weight for the real-time requirement and the second requirement weight for the accuracy requirement; Determine the requirement value of the target detection requirement according to the ratio of the first requirement weight to the second requirement weight; Determine the architecture size of the target quantization architecture according to the mapping of the requirement value; Wherein, the numerical size of the requirement value is negatively correlated with the architecture size of the target quantization architecture.

8. A fall detection device, characterized in that, The device includes: A first module, configured to obtain a set of human state images; the human state images in the set of human state images are labeled with human state categories; A second module, configured to use the set of human state images as a data set to perform detection training on a preset target detection model to obtain a human state detection model; Wherein, the target detection model is constructed by replacing the convolutional layer in the YOLO model with a spatially separable convolution and a reparameterizable convolution block; A third module, configured to perform state detection on the captured image of the person to be detected by using the human state detection model to obtain a target state category; A fourth module, configured to determine whether the person to be detected has fallen based on the target state category.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.