A fall detection method and system in a multi-person scenario based on visual intelligence
By identifying and separating individual detection images in multi-person scenarios, combining human grid reconstruction model and fall detection model, the problem of low fall detection accuracy in multi-person scenarios is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411007525.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-25
AI Technical Summary
The existing fall detection technology has low accuracy in multi-person scenarios and is difficult to adapt to complex backgrounds and occlusion situations, especially when identifying and capturing key points of dynamic changes, the accuracy rate decreases.
By identifying different individuals in the image to be detected, the individual detection images are isolated, and the grid reconstruction model is used to perform grid reconstruction. Finally, the grid reconstruction image is input into the preset fall detection model to achieve fall detection in multiple scenarios.
It improves the accuracy of fall detection in multi-person scenarios, can more effectively identify and handle complex backgrounds and occlusion situations, and enhances the robustness of detection and feasibility of practical applications.
Smart Images

Figure CN119007283B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of behavior detection, and in particular, to a fall detection method and system in a multi-person scenario based on visual intelligence. Background Art
[0002] As China enters an aging society, the problem of elderly care has become increasingly severe. The physical function indicators of the elderly decline, and their mobility decreases. In particular, the lack of balance, reaction ability, and coordination ability may cause accidental falls. When an elderly person falls, if timely assistance is not obtained, they may even die. Therefore, fall detection for the elderly in families or other environments is a very meaningful research issue in the fields of computer vision and machine learning.
[0003] In the field of fall detection technology, current methods mainly rely on deep learning and computer vision to automatically identify fall events through video analysis. These technical solutions have their own focuses, including using deep learning for human key point detection and combining with support vector machines for classification, adopting multi-modal spatio-temporal skeletal dynamics feature fusion, and behavior classification based on the BAT-GCN model. These methods can effectively identify and classify fall behaviors to a certain extent, but these methods are mainly applied in simple scenarios and single-person situations.
[0004] Currently, relevant technologies are not very adaptable to complex backgrounds and occlusion situations. Especially in multi-person scenarios, the accuracy of key point detection and the ability to capture dynamic changes are limited, resulting in a decrease in the accuracy of fall detection. Secondly, although some methods attempt to improve the detection performance by fusing spatio-temporal features, it is still difficult to completely solve the problems of occlusion and multi-person recognition. These limitations reduce the feasibility and robustness of these technologies in practical applications.
[0005] Therefore, how to achieve fall detection in multi-person scenarios in public places such as parks, sports fields, and senior activity centers is a problem that needs to be solved currently. Summary of the Invention
[0006] In order to improve the accuracy of fall detection in multi-person scenarios, the present application provides a fall detection method and system in a multi-person scenario based on visual intelligence.
[0007] In the first aspect of the present application, a fall detection method in a multi-person scenario based on visual intelligence is provided. The method includes:
[0008] Obtain an image to be detected, where the image to be detected represents a video and / or an image that needs to be subjected to fall detection;
[0009] Identify different individuals in the image to be detected, determine the individual detection image, and the individual detection image includes an individual to be detected; according to the human body mesh reconstruction model, perform mesh reconstruction on the individual detection image to obtain the mesh reconstruction image;
[0010] Input the mesh reconstruction image into a preset fall detection model to obtain the detection result corresponding to the image to be detected, and the detection result is used to reflect whether the individual to be detected in the individual detection image has fallen.
[0011] As can be seen from the above technical solutions, by identifying different individuals in the image to be detected, separating multiple individuals to be detected in the image to be detected to obtain the individual detection image, then respectively performing mesh reconstruction on the individual detection image to obtain the mesh reconstruction image, and then obtaining the detection result corresponding to the image to be detected according to the fall detection model, fall detection in a multi-person scenario is realized, and at the same time, the accuracy of fall detection in a multi-person scenario is improved.
[0012] In a possible implementation manner, the image to be detected is obtained through the following method:
[0013] Obtain the initial image, and the initial image represents the original video and / or image;
[0014] Sample the video according to the frame rate of the video to obtain the intermediate image;
[0015] Adjust the format, encoding, resolution, and frame rate of the intermediate image to obtain the image to be detected.
[0016] In a possible implementation manner, the human body mesh reconstruction model is determined through the following method:
[0017] Obtain the image training data set, and the image training data set includes multiple videos and / or images;
[0018] Sample the video according to the frame rate of each video to obtain the portrait image data;
[0019] Train a preset reconstruction model according to the portrait image data to obtain the human body mesh reconstruction model, and the human body mesh reconstruction model is used to convert the portrait image data into a two-dimensional projection image.
[0020] In a possible implementation manner, training a preset reconstruction model according to the portrait image data to obtain the human body mesh reconstruction model includes:
[0021] Perform convolution operation on the portrait image data according to the image encoder to obtain the human body feature map;
[0022] Determine the prediction parameters of the human body mesh reconstruction model according to the feature vector of the human body feature map and the initial parameters, and the initial parameters are used to reflect the average situation of the portrait image data;
[0023] Based on the parameter regressor and the predicted parameters, obtain the loss function of the parameter regressor;
[0024] Determine the human body mesh reconstruction model according to the loss function of the parameter regressor.
[0025] In a possible implementation, the loss function of the parameter regressor is determined in the following way:
[0026]
[0027] where L reg represents the loss function of the parameter regressor, J represents the 3D joints in the portrait image data, represents the true value of the 3D joints, K represents the 2D key points after projecting J into the image coordinate system; represents the true value of the 2D key points, Θ represents the predicted parameters, represents the true value of the initial parameters, ‖·‖ represents the squared L2 norm, represents the error of the 2D key points, represents the error of the 3D joints, represents the error of the predicted parameters, λ 2d represents the weight coefficient of the error of the 2D key points, λ 3d represents the weight coefficient of the error of the 3D joints, λ para represents the weight coefficient of the error of the predicted parameters.
[0028] In a possible implementation, the fall detection model is determined in the following way:
[0029] Obtain a fall image training data set, where the fall image training data set includes multiple images;
[0030] Input the fall image training data set into the human body mesh reconstruction model to obtain an image reconstruction data set;
[0031] Optimize the pre-trained vision-language model according to the image reconstruction data set to obtain a fall detection model.
[0032] In a possible implementation, the pre-trained vision-language model includes an image encoder and a text encoder. The image encoder is used to convert the input image into a feature vector, and the text encoder is used to convert natural language text into a feature vector;
[0033] Optimizing the pre-trained vision-language model according to the image reconstruction data set to obtain a fall detection model includes:
[0034] Input the image reconstruction data in each image reconstruction data set into a deep learning model to obtain an image feature vector;
[0035] Segment the classification information of the image reconstruction data in the reconstruction dataset to obtain text embedding vectors;
[0036] Calculate the cosine similarity between the image feature vector and the text embedding vector;
[0037] Adjust the pre-trained vision-language model according to the cross-entropy objective function and the cosine similarity to obtain a fall detection model.
[0038] In a possible implementation, the cross-entropy objective function is determined in the following way:
[0039]
[0040] where, v i represents the image feature vector of the i-th image reconstruction data, t i represents the text embedding vector of the i-th image reconstruction data, τ represents the temperature coefficient, sim(·) represents the cosine similarity, exp(·) represents the natural exponential function, and log(·) represents the natural logarithm function.
[0041] In a possible implementation, the human body mesh reconstruction model is the SMPL model.
[0042] In the second aspect of this application, a fall detection system in a multi-person scenario based on visual intelligence is provided. The system includes:
[0043] A data acquisition module for acquiring the image to be detected, where the image to be detected represents a video and / or an image that needs to be subjected to fall detection; an individual recognition module for recognizing different individuals in the image to be detected and determining an individual detection image, where the individual detection image includes an individual to be detected;
[0044] A human body reconstruction module for performing mesh reconstruction on the individual detection image according to the human body mesh reconstruction model to obtain a mesh reconstruction image;
[0045] A fall detection module for inputting the mesh reconstruction image into a preset fall detection model to obtain a detection result corresponding to the image to be detected, where the detection result is used to reflect whether the individual to be detected in the individual detection image has fallen.
[0046] In summary, this application includes at least one beneficial technical effect:
[0047] By recognizing different individuals in the image to be detected, separating multiple individuals to be detected in the image to be detected to obtain an individual detection image, then respectively performing mesh reconstruction on the individual detection image to obtain a mesh reconstruction image, and then obtaining a detection result corresponding to the image to be detected according to the fall detection model, fall detection in a multi-person scenario is realized, and at the same time, the accuracy of fall detection in a multi-person scenario is improved. Brief Description of the Drawings
[0048] Figure 1 is a schematic flowchart of the fall detection method in a multi-person scenario based on visual intelligence provided by this application.
[0049] Figure 2 is a schematic structural diagram of the fall detection system in a multi-person scenario based on visual intelligence provided by this application.
[0050] Figure 3 is a schematic structural diagram of the electronic device provided by this application.
[0051] In the figure, 201 is the data acquisition module; 202 is the individual recognition module; 203 is the human body reconstruction module; 204 is the fall detection module; 301 is the CPU; 302 is the ROM; 303 is the RAM; 304 is the I / O interface; 305 is the input part; 306 is the output part; 307 is the storage part; 308 is the communication part; 309 is the driver; 310 is the removable medium. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.
[0053] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after unless otherwise specified.
[0054] Fall detection aims to automatically identify human fall behaviors through various methods and technologies to reduce the injuries caused by falls, especially for the elderly population. This technology has broad application potential in scenarios such as outdoor multi-person sports activities and elderly care centers. In outdoor sports venues, such as parks or sports fields, it can accurately identify and separate the actions of each participant, detect fall events, and provide instant information for emergency medical response. In elderly care centers, this technology can be used to continuously monitor the activities of residents, timely detect fall risks, and thus take preventive measures to ensure the safety and health of the elderly.
[0055] Currently, there are mainly three existing fall detection methods, namely fall detection based on wearable devices, fall detection based on depth cameras, and fall detection based on ordinary cameras. Among them, the method based on wearable devices must be carried at all times, which brings great inconvenience to users and has little practical application value; the method based on depth cameras is difficult to promote in practice due to its high cost; while the method based on ordinary cameras is cheap and convenient to use, but has high requirements for algorithms.
[0056] The following further describes the embodiments of the present application in detail with reference to the accompanying drawings of the specification.
[0057] The embodiments of the present application provide a fall detection method in a multi-person scenario based on visual intelligence. The main processes of the above method are described as follows.
[0058] As Figure 1 shown:
[0059] Step S101: Obtain the image to be detected.
[0060] Specifically, the image to be detected represents the video and / or image that needs to be subjected to fall detection. Obtain the initial image, and the initial image represents the original video and / or image, that is, the video or picture that has not been processed at all. According to the frame rate of the above video, sample the above video to obtain an intermediate image. Adjust the format, encoding, resolution, and frame rate of the above intermediate image to obtain the image to be detected.
[0061] Adjusting the video to a unified format, encoding, resolution, and frame rate is to adapt to the human body mesh reconstruction model and avoid problems with the obtained mesh reconstruction image due to format, encoding, resolution, and frame rate issues, which may affect the subsequent detection results.
[0062] Step S102: Identify different individuals in the image to be detected and determine the individual detection image.
[0063] Specifically, the individual detection image includes one individual to be detected. It can be understood that in the embodiments provided by the present application, the YOLOv8 algorithm is used to identify different individuals in the image to be detected. The YOLOv8 algorithm is an object recognition model that can effectively detect multiple human bodies and can accurately identify and determine the specific position of the human body in the image in real time.
[0064] In a specific example, the YOLOv8 algorithm is used to perform human body detection on the image to be detected to obtain the regional information of different individuals in the image to be detected. Combining the regional information, the image to be detected is divided into regions to obtain the individual detection image.
[0065] Step S103: According to the human body mesh reconstruction model, perform mesh reconstruction on the individual detection image to obtain the mesh reconstruction image.
[0066] The above-mentioned human body mesh reconstruction model is determined in the following manner:
[0067] An image training data set is obtained, and the above-mentioned image training data set includes a plurality of videos and / or images. According to the frame rate of each of the above-mentioned videos, the above-mentioned videos are sampled to obtain portrait image data. The preset reconstruction model is trained according to the above-mentioned portrait image data to obtain a human body mesh reconstruction model, and the above-mentioned human body mesh reconstruction model is used to convert portrait image data into a two-dimensional projection image.
[0068] In a specific implementation manner, the above-mentioned image training data set is derived from publicly available data sets, such as Human3.6M, MPI-INF-3DHP, etc. The above-mentioned image training data set contains two types of data, images and videos, and both types of data contain the true values of 3D joints, 2D key points, and SMPL parameters. In other implementation manners, other publicly available or non-public data sets can also be used, and this is not limited.
[0069] For the videos in the image training data set, a frame rate downsampling operation is performed to remove redundant frames. The above-mentioned sampling operation can be to sample one frame every n frames, or random sampling can be performed, or each frame can be analyzed to determine portrait image data. In other implementation manners, other methods can be used to sample the videos, and this is not limited.
[0070] The above reconstruction model is SMPL (Skinned Multi-Person Linear Model). SMPL is a computer graphics model for human pose estimation and shape reconstruction. SMPL accurately captures and renders individual shapes through shape and pose parameters. SMPL parameters are the parameters of the SMPL model, and the parameters of the SMPL model include shape parameters and pose parameters. The shape parameters are a shape vector composed of 10 scalar values, and each scalar value can be interpreted as the amount of expansion / contraction of the human body object along a certain direction. These parameters are used to describe the overall shape of the human body, and each parameter can vary independently to produce different body forms. The pose parameters are composed of a pose vector of 24×3 scalar values and are used to maintain the relative rotation of the joints with respect to their parameters. In the axis-angle rotation representation, each rotation is encoded as an arbitrary three-dimensional vector, allowing for flexible movement of the joints in space. In addition, the SMPL model also involves other parameters, such as action parameters, which represent the rotation relationship relative to the parent node, including facial key points, finger joint points, and key points on the limbs. Action parameters are used to capture finer actions and pose changes. The SMPL model combines shape and pose parameters and other related parameters to generate a highly realistic human mesh model. In the example provided in this application, the SMPL parameters include a 10-dimensional body shape control vector and a 3×24-dimensional pose control vector. The body shape control vector controls a person's height, weight, and build, and the pose control vector determines a person's skeletal pose. With these parameters, a human mesh with 6890 vertices is reconstructed.
[0071] The input of the human mesh reconstruction model is an image containing a real human body, and the output is a two-dimensional projection image of the human mesh model reconstructed using SMPL parameters. It can be understood that the above human mesh reconstruction model includes the rendering of the image. Input the image training data in the above image training dataset into the human mesh reconstruction model to obtain a sequence of rendered frames containing the SMPL human reconstruction results, that is, the rendered mesh reconstruction image. Since SMPL parameters can reconstruct a human mesh, and the human mesh is a three-dimensional mesh, it needs to be rendered into a two-dimensional image for subsequent operations. Render the human mesh according to the preset camera parameters. In this application, the two-dimensional image space obtained by rendering is selected for subsequent fall detection operations.
[0072] Further, perform a convolution operation on the above portrait image data according to the image encoder to obtain a human body feature map. Determine the prediction parameters of the above human mesh reconstruction model according to the feature vector of the above human body feature map and the initial parameters, and the above initial parameters are used to reflect the average situation of the above portrait image data. Obtain the loss function of the above parameter regressor according to the parameter regressor and the above prediction parameters. Determine the above human mesh reconstruction model according to the loss function of the above parameter regressor.
[0073] In a specific implementation, all the image training data in the above image training dataset are input into an image encoder based on ResNet-50 for convolution operations to obtain a human feature map. The eigenvalue of the point on the two-dimensional grid in the above human feature map is sampled at equal intervals, and is passed into a multi-layer perceptron for dimensionality reduction operations. Finally, the eigenvalues after dimensionality reduction of all points are concatenated to obtain the feature vector of the human feature map. The average value of the true values of the SMPL parameters of all the image training data in the above image training dataset is calculated as the initial SMPL parameter, that is, the initial parameter Θ0. Θ0 and the obtained feature vector are passed into a parameter regressor to obtain a residual value. The residual value is added to the initial parameter Θ0 to obtain Θ1, and Θ1 is used as the updated SMPL parameter; this process is repeated to obtain the final SMPL prediction parameter Θ, and the training of the parameter regressor is completed. In the implementation provided in this application, a loss threshold is set, and when the value of the loss function is less than the loss threshold, the training is stopped, and the training of the parameter regressor is completed to obtain a human mesh reconstruction model.
[0074] The loss function of the above parameter regressor is determined in the following manner:
[0075]
[0076] where L reg represents the loss function of the above parameter regressor, J represents the 3D joints in the above portrait image data, represents the true value of the above 3D joints, K represents the 2D key points after projecting J into the image coordinate system; represents the true value of the above 2D key points, Θ represents the above prediction parameter, represents the true value of the above initial parameter, ‖·‖ represents the squared L2 norm, represents the error of the above 2D key points, represents the error of the above 3D joints, represents the error of the above prediction parameter, λ 2d represents the weight coefficient of the error of the above 2D key points, λ 3d represents the weight coefficient of the error of the above 3D joints, λ para represents the weight coefficient of the error of the above prediction parameter.
[0077] Step S104: Input the mesh reconstruction image into a preset fall detection model to obtain the detection result corresponding to the image to be detected.
[0078] Specifically, the above detection result is used to reflect whether the individual to be detected in the above individual detection image has fallen. In the example provided in this application, the above fall detection model is a pre-trained vision-language model (Contrastive Language-Image Pre-Training, CLIP). CLIP is a multi-modal learning framework that has an image encoder and a text encoder, and understands and correlates visual concepts with natural language descriptions by jointly learning image and text representations.
[0079] Inputting the grid reconstruction images corresponding to different individuals to be detected into the fall detection model can obtain the similarity between the classification information corresponding to the individual to be detected and "<Occurred / Not Occurred> fall behavior". The one with a higher similarity is the detection result corresponding to the image to be detected. Finally, the individual to be detected corresponding to the grid reconstruction image with "fall behavior occurred" is obtained, thereby realizing fall detection in a multi-person scenario.
[0080] Furthermore, the above fall detection model is determined in the following manner:
[0081] Obtain a fall image training dataset, where the fall image training dataset includes multiple images. Input the fall image training dataset into the above human body grid reconstruction model to obtain an image reconstruction dataset. Optimize the pre-trained vision-language model according to the above image reconstruction dataset to obtain the above fall detection model.
[0082] Furthermore, the above pre-trained vision-language model includes an image encoder and a text encoder. The image encoder is used to convert the input image into a feature vector, and the text encoder is used to convert natural language text into a feature vector. Input the image reconstruction data in each of the above image reconstruction datasets into a deep learning model to obtain image feature vectors. Split the classification information of the image reconstruction data in the above image reconstruction dataset to obtain text embedding vectors. Calculate the cosine similarity between the above image feature vectors and the above text embedding vectors. Adjust the above pre-trained vision-language model according to the cross-entropy objective function and the above cosine similarity to obtain a fall detection model.
[0083] Furthermore, the above cross-entropy objective function is determined in the following manner:
[0084]
[0085] where, v i represents the image feature vector of the i-th above image reconstruction data, t i represents the text embedding vector of the i-th above image reconstruction data, τ represents the temperature coefficient, sim(·) represents the above cosine similarity, exp(·) represents the natural exponential function, and log(·) represents the natural logarithm function.
[0086] In a specific implementation, for the videos in the fall image training dataset, when processing a video with T frames, each frame is processed by the image encoder in CLIP, and the backbone of the image encoder is ResNet-50. Each frame of the image first passes through the initial convolutional layer of ResNet-50, and then passes through a series of residual blocks with skip connections. After passing through the last residual block, the resulting feature map passes through a global average pooling layer to average the spatial information within each feature channel, thereby converting the two-dimensional feature map into a one-dimensional feature vector. Finally, through a fully connected layer, the output feature vectors all have a fixed 512 dimensions. After completing this series of operations, T 512-dimensional frame-level feature vectors are obtained. The above series of residual blocks can be understood as the residual network of the deep learning framework, and the last residual block is the output of the residual network. Regarding the obtained T 512-dimensional frame-level feature vectors as a T×512 matrix, a temporal pooling operation is performed to calculate the average value of the T feature values for each feature dimension, obtaining a 512-dimensional video-level feature vector representing the entire video feature, that is, the image feature vector.
[0087] It can be understood that each fall image training data in the fall image training dataset includes, in addition to the image itself, classification information corresponding to the image, and the above classification information is used to reflect whether the corresponding above image has a fall. The classification information corresponding to each image is passed into a splitter to be split into multiple split items, and then the split items are passed into the text encoder in CLIP to be converted into a 512-dimensional text embedding vector. In the example provided in this application, the above classification information is a complete sentence. For example, a fall behavior occurs or a fall behavior does not occur, and the complete sentence needs to be split into single characters or words. The text encoder can only process the split content and cannot directly encode the complete sentence.
[0088] When there are N pieces of image reconstruction data in the image reconstruction dataset, the image reconstruction dataset is divided into multiple batches. In the example provided in this application, there are 32 pieces of the above image reconstruction data in each batch, and there are a total of K = N / 32 batches, and E large loops are performed. In each large loop, K batch loops are performed. In each batch loop, the 32 pieces of image reconstruction data in this batch are used to fine-tune with the objective function to complete an update of the model parameters.
[0089] The above fine-tuning needs to calculate the cosine similarity between the image feature vector and the text embedding vector, and maximize the cosine similarity to achieve the fine-tuning of the image encoder and the text encoder of CLIP. The maximization of the above cosine similarity can be understood as maximizing the cross-entropy objective function.
[0090] An embodiment of the present application provides a fall detection system in a multi-person scenario based on visual intelligence. Refer to Figure 2 , the fall detection system in a multi-person scenario based on visual intelligence includes:
[0091] A data acquisition module 201, configured to acquire an image to be detected, where the image to be detected represents a video and / or an image that needs to be subjected to fall detection;
[0092] An individual recognition module 202, configured to recognize different individuals in the image to be detected and determine an individual detection image, where the individual detection image includes an individual to be detected;
[0093] A human body reconstruction module 203, configured to perform grid reconstruction on the individual detection image according to a human body grid reconstruction model to obtain a grid reconstruction image;
[0094] A fall detection module 204, configured to input the grid reconstruction image into a preset fall detection model to obtain a detection result corresponding to the image to be detected, where the detection result is used to reflect whether the individual to be detected in the individual detection image has fallen.
[0095] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0096] An embodiment of the present application discloses an electronic device. Refer to Figure 3 , the electronic device includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 307 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus. An input / output (I / O) interface 304 is also connected to the bus.
[0097] The following components are connected to the I / O interface 304: an input section 305 including a keyboard, a mouse, etc.; an output section 306 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 307 including a hard disk, etc.; and a communication section 308 including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section 308 performs communication processing via a network such as the Internet. A drive 309 is also connected to the I / O interface 304 as required. A removable medium 310, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 309 as required so that a computer program read therefrom is installed into the storage section 307 as required.
[0098] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart Figure 1 can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a machine-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 308, and / or installed from the removable medium 310. When the computer program is executed by a central processing unit (CPU) 301, the above functions defined in the apparatus of the present application are executed.
[0099] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, radio frequency (RF), etc., or any suitable combination of the above.
[0100] The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the application involved in this application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing application concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions applied in this application.
Claims
1. A fall detection method in a multi-person scenario based on visual intelligence, characterized in that: include: Acquire an image to be detected, where the image to be detected represents a video on which fall detection is required; Identify different individuals in the image to be detected and determine an individual detection image, wherein the individual detection image includes an individual to be detected; perform mesh reconstruction on the individual detection image according to a human body mesh reconstruction model to obtain a mesh reconstruction image, wherein the human body mesh reconstruction model is used to convert portrait image data into a two-dimensional projection image of a three-dimensional human body mesh; Inputting the grid-reconstructed image into a preset fall detection model to obtain a detection result corresponding to the image to be detected, wherein the detection result is used to reflect whether the individual to be detected in the individual detection image has fallen; The fall detection model is determined in the following way: Acquire a fall image training data set, wherein the fall image training data set includes a plurality of images, and the images are video images; Inputting the fall image training data set into the human body mesh reconstruction model to obtain an image reconstruction data set; According to the image reconstruction data set, the pre-trained visual language model is optimized to obtain the fall detection model; The pre-trained visual language model includes an image encoder and a text encoder, wherein the image encoder is used to convert an input image into a feature vector, and the text encoder is used to convert a natural language text into a feature vector; The method of optimizing the pre-trained visual language model according to the image reconstruction data set to obtain the fall detection model includes: Inputting the image reconstruction data in each of the image reconstruction data sets into a deep learning model to obtain an image feature vector; segmenting the classification information of the image reconstruction data in the image reconstruction data sets to obtain a text embedding vector; Calculating the cosine similarity between the image feature vector and the text embedding vector; According to the cross entropy objective function and the cosine similarity, adjusting the pre-trained visual language model to obtain a fall detection model; The cross entropy objective function is determined in the following way: Wherein, vi represents the image feature vector of the i-th image reconstruction data, ti represents the text embedding vector of the i-th image reconstruction data, τ represents the temperature coefficient, sim(·) represents the cosine similarity, exp(·) represents the natural exponential function, and log(·) represents the natural logarithm function.
2. The fall detection method in a multi-person scenario based on visual intelligence according to claim 1 is characterized in that: The image to be detected is obtained by the following method: Acquire an initial image, wherein the initial image represents an original video and / or image; Sampling the video according to the frame rate of the video to obtain an intermediate image; The format, encoding, resolution and frame rate of the intermediate image are adjusted to obtain the image to be detected.
3. The fall detection method in a multi-person scenario based on visual intelligence according to claim 1 is characterized in that: The human body mesh reconstruction model is determined by: Acquire an image training data set, wherein the image training data set includes a plurality of videos and / or images; Sampling the videos according to the frame rate of each video to obtain portrait image data; A preset reconstruction model is trained according to the portrait image data to obtain a human body mesh reconstruction model, and the human body mesh reconstruction model is used to convert the portrait image data into a two-dimensional projection image.
4. The fall detection method in a multi-person scenario based on visual intelligence according to claim 3 is characterized in that: The step of training a preset reconstruction model according to the portrait image data to obtain a human body mesh reconstruction model comprises: According to the image encoder, a convolution operation is performed on the portrait image data to obtain a human body feature map; Determining prediction parameters of the human body mesh reconstruction model according to the feature vector and initial parameters of the human body feature map, wherein the initial parameters are used to reflect the average situation of the portrait image data; Obtaining a loss function of the parameter regressor according to the parameter regressor and the prediction parameters; The human body mesh reconstruction model is determined according to the loss function of the parameter regressor.
5. The fall detection method in a multi-person scenario based on visual intelligence according to claim 4 is characterized in that: The loss function of the parameter regressor is determined as follows: Wherein, Lreg represents the loss function of the parameter regressor, J represents the 3D joint in the portrait image data, represents the true value of the 3D joint, and K represents the 2D key point after J is projected into the image coordinate system; represents the true value of the 2D key point, Θ represents the prediction parameter, represents the true value of the initial parameter, ‖·‖ represents the squared L2 norm, represents the error of the 2D key point, represents the error of the 3D joint, represents the error of the prediction parameter, λ2d represents the weight coefficient of the error of the 2D key point, λ3d represents the weight coefficient of the error of the 3D joint, and λpara represents the weight coefficient of the error of the prediction parameter.
6. The fall detection method in a multi-person scenario based on visual intelligence according to claim 1, characterized in that: The human body mesh reconstruction model is an SMPL model.
7. A fall detection system in a multi-person scenario based on visual intelligence, characterized in that: include: A data acquisition module, used to acquire an image to be detected, wherein the image to be detected represents a video for which fall detection is required; An individual recognition module, used to identify different individuals in the image to be detected and determine an individual detection image, wherein the individual detection image includes an individual to be detected; A human body reconstruction module, used to perform grid reconstruction on the individual detection image according to the human body grid reconstruction model to obtain a grid reconstructed image; A fall detection module is used to input the grid reconstructed image into a preset fall detection model to obtain a detection result corresponding to the image to be detected, wherein the detection result is used to reflect whether the individual to be detected in the individual detection image falls; wherein the fall detection model is determined by: Acquire a fall image training data set, where the fall image training data set includes a plurality of images; Inputting the fall image training data set into the human body mesh reconstruction model to obtain an image reconstruction data set; According to the image reconstruction data set, the pre-trained visual language model is optimized to obtain the fall detection model; The pre-trained visual language model includes an image encoder and a text encoder, wherein the image encoder is used to convert an input image into a feature vector, and the text encoder is used to convert a natural language text into a feature vector; The method of optimizing the pre-trained visual language model according to the image reconstruction data set to obtain the fall detection model includes: Inputting the image reconstruction data in each of the image reconstruction data sets into a deep learning model to obtain an image feature vector; segmenting the classification information of the image reconstruction data in the image reconstruction data sets to obtain a text embedding vector; Calculating the cosine similarity between the image feature vector and the text embedding vector; According to the cross entropy objective function and the cosine similarity, the pre-trained visual language model is adjusted to obtain a fall detection model; the cross entropy objective function is determined by: Wherein, vi represents the image feature vector of the i-th image reconstruction data, ti represents the text embedding vector of the i-th image reconstruction data, τ represents the temperature coefficient, sim(·) represents the cosine similarity, exp(·) represents the natural exponential function, and log(·) represents the natural logarithm function.
Citation Information
Patent Citations
Indoor tumble detection method and system based on scene recognition
CN113743339A
Human body tumble detection method and feature extraction model acquisition method and device
CN115482494A
Three-dimensional human body reconstruction method, training method and related device
CN116704130A
Multi-target pedestrian potential safety hazard behavior comprehensive identification method
CN117994609A