A fall behavior recognition method, device, equipment and storage medium

By performing target recognition and tracking on videos and combining current and historical recognition results, the accuracy and continuity of fall behavior recognition are improved, solving the problem of insufficient recognition accuracy in existing technologies.

CN117197751BActive Publication Date: 2026-01-09SUZHOU KEDA TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311256196.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2026-01-09
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

Existing image and video-based fall detection technologies have shortcomings in terms of accuracy and continuity, especially in crowded and complex situations where they are prone to false positives and false negatives, resulting in poor recognition performance.

Method used

By performing target recognition and tracking on the video to be detected, the behavior image sequence of the target to be identified is obtained. Combined with a pre-trained fall behavior recognition model, the current and historical recognition results are used to determine whether the target has fallen, thereby enhancing continuity and accuracy.

Benefits of technology

This improves the accuracy of fall behavior recognition, reduces the false detection rate, and ensures the correctness and reliability of subsequent processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197751B_ABST
    Figure CN117197751B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a fall behavior recognition method, device and equipment and a storage medium. The method comprises: performing target recognition on a video to be detected to determine at least one target to be recognized; performing target tracking on each target to be recognized to determine a behavior image sequence corresponding to each target to be recognized; when the current fall behavior recognition is non-first-time recognition, determining a last-recognized behavior image set and a current behavior image set to be recognized according to the behavior image sequence, and determining a current behavior image sequence according to the last-recognized behavior image set and the current behavior image set to be recognized; inputting the current behavior image sequence into a pre-trained fall behavior recognition model to determine a current fall behavior recognition result, and determining a target fall behavior recognition result of the target to be recognized according to the current fall behavior recognition result and a historical fall behavior recognition result. The accuracy of determining the state of personnel falling is improved, and the false detection rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision recognition, and particularly relates to a fall behavior recognition method and device, equipment and a storage medium. BACKGROUND

[0002] With the development and progress of artificial intelligence technology, target recognition technology has been used in many fields. In addition to identifying targets, target recognition technology can also identify target behaviors, and further evolves into behavior recognition algorithms. Behavior recognition algorithms play an important role in the field of computer vision and artificial intelligence, and have a wide range of applications. Its purpose is to analyze and process input images, videos, sensor data, etc., to identify and infer human behavior and actions, and it is widely used in intelligent monitoring, traffic management, intelligent transportation systems, health monitoring and elderly care, human-computer interaction and virtual reality, video analysis and content management, and unmanned driving and robotics.

[0003] In actual use, the recognition of human fall behavior has stronger demand in the safety field. Current fall detection technology based on images and videos has been widely used in fields such as life, transportation and security.

[0004] However, fall detection technology based only on images uses only the spatial dimension information of images and does not use time dimension information, and the single feature dimension leads to incomplete recognition information and low recognition accuracy. Current fall detection technology based on videos often performs single-process target recognition and fall behavior recognition on a piece of completed video data. In the case of personnel concentration, complex actions, and occlusion that cannot continuously collect the motion process of the same object, there are problems such as easy mis-detection and easy missed detection, which greatly reduces the recognition effect, and leads to low accuracy of personnel fall state determined directly from the video and poor actual usability. SUMMARY

[0005] The present application provides a fall behavior recognition method, device, equipment and storage medium, which improves the continuity of the acquisition of human behavior images, ensures the continuity of the behavior of the personnel in the recognition process, fully considers the influence of the historical behavior state on the determination of the task fall behavior, and improves the accuracy of the determination of the personnel fall state.

[0006] In a first aspect, the embodiments of the present application provide a fall behavior recognition method, comprising:

[0007] acquiring a to-be-detected video, and performing target recognition on the to-be-detected video to determine at least one to-be-recognized target;

[0008] tracking each to-be-identified target in the to-be-detected video to determine a behavior image sequence corresponding to each to-be-identified target;

[0009] For each behavior image sequence, when the current fall behavior recognition is non-first-time recognition, determining a last-recognized behavior image set and a current to-be-identified behavior image set according to the behavior image sequence, and determining the current behavior image sequence according to the last-recognized behavior image set and the current to-be-identified behavior image set;

[0010] inputting the current behavior image sequence into the pre-trained fall behavior recognition model to determine a current fall behavior recognition result, and determining a target fall behavior recognition result of the to-be-identified target according to the current fall behavior recognition result and a historical fall behavior recognition result.

[0011] In a second aspect, an embodiment of the present application further provides a fall behavior recognition device, comprising:

[0012] a target determination module configured to acquire a to-be-detected video, and perform target recognition on the to-be-detected video to determine at least one to-be-identified target;

[0013] an image sequence determination module configured to track each to-be-identified target in the to-be-detected video to determine a behavior image sequence corresponding to each to-be-identified target;

[0014] a current sequence determination module configured to, for each behavior image sequence, when the current fall behavior recognition is non-first-time recognition, determine a last-recognized behavior image set and a current to-be-identified behavior image set according to the behavior image sequence, and determine the current behavior image sequence according to the last-recognized behavior image set and the current to-be-identified behavior image set;

[0015] an identification result determination module configured to input the current behavior image sequence into the pre-trained fall behavior recognition model to determine a current fall behavior recognition result, and determine a target fall behavior recognition result of the to-be-identified target according to the current fall behavior recognition result and a historical fall behavior recognition result.

[0016] In a third aspect, an embodiment of the present application further provides a fall behavior recognition device, comprising:

[0017] at least one processor; and

[0018] a memory in communication with the at least one processor; wherein

[0019] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the fall behavior recognition method provided by the embodiments of the present application.

[0020] In a fourth aspect, the embodiments of the present application further provide a storage medium containing computer executable instructions for executing the fall behavior recognition method provided by the embodiments of the present application when executed by a computer processor.

[0021] The embodiments of the present application provide a fall behavior recognition method, device, equipment and storage medium. The method comprises the following steps: obtaining a to-be-detected video, performing target recognition on the to-be-detected video, and determining at least one to-be-recognized target; performing target tracking on each to-be-recognized target in the to-be-detected video, and determining a behavior image sequence corresponding to each to-be-recognized target; for each behavior image sequence, when the current fall behavior recognition is not the first time, determining a last-recognized behavior image set and a current to-be-recognized behavior image set according to the behavior image sequence, and determining a current behavior image sequence according to the last-recognized behavior image set and the current to-be-recognized behavior image set; inputting the current behavior image sequence into a pre-trained fall behavior recognition model, determining a current fall behavior recognition result, and determining a target fall behavior recognition result of the to-be-recognized target according to the current fall behavior recognition result and a historical fall behavior recognition result. By using the above technical solution, when the to-be-detected video is obtained, the to-be-detected video is preferentially subjected to target recognition and target tracking, and a plurality of behavior image sequences of the to-be-recognized targets which need to be subjected to fall behavior recognition are obtained. Then, for each behavior image sequence, when the fall behavior recognition is not the first time, part of the behavior images which have not been recognized are taken as the current to-be-recognized behavior image set, and the behavior images which have been subjected to fall behavior recognition are determined as the last-recognized behavior image set. The current behavior image sequence is obtained by combining the two sets of images, and is input into the pre-trained fall behavior recognition model for current fall behavior recognition. Then, the target fall behavior recognition result of whether the to-be-recognized target corresponding to the behavior image sequence falls is determined by combining the current fall behavior recognition result and the historical fall behavior recognition result of the last recognition. Since the current behavior image sequence contains both recognized and unrecognized images, the fall behavior recognition model can fully consider the information in the last recognition in one recognition, thereby enhancing the continuity of two consecutive recognitions, and making the fall behavior recognition model more accurate for current fall behavior recognition. When determining the target fall behavior recognition result of whether the target falls, the historical fall behavior recognition result with the highest correlation with the current fall recognition result is taken as the basis for judging whether the target falls, thereby improving the accuracy of determining the state of the personnel falling, reducing the false detection rate, and enabling the subsequent processing according to the target fall behavior recognition result to be performed correctly.

[0022] It is to be understood that the embodiments described herein are merely meant to identify key or important features of the embodiments of the present application, and are not intended to limit the scope of the present application. Other features of the present application will become readily apparent from the description set forth below. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0024] Figure 1 A flow chart of a fall behavior recognition method provided by an embodiment of the present application;

[0025] Figure 2 A flow chart of a fall behavior recognition method provided by another embodiment of the present application;

[0026] Figure 3 A flow chart of an example of combining a set of historical behavior images to be merged and a set of current behavior images to be recognized in time sequence to determine a current behavior image sequence provided by an embodiment of the present application;

[0027] Figure 4 A structural schematic diagram of a fall behavior recognition device provided by an embodiment of the present application;

[0028] Figure 5 A structural schematic diagram of a fall behavior recognition device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of protection of the present application.

[0030] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and the above-described accompanying drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to such a process, method, product, or device.

[0031] Figure 1 A flowchart of a fall behavior identification method provided for an embodiment of the present application, the embodiment of the present application can be applicable to the case of identifying the possible fall behavior of personnel in the acquired video. The method can be executed by a fall behavior identification device, which can be realized by software and / or hardware, and can be configured in a fall behavior identification apparatus. Optionally, the fall behavior identification apparatus can be an electronic device, which can be a notebook, a desktop computer, a smart tablet, etc. The embodiment of the present application does not limit this.

[0032] As shown in Figure 1 , the fall behavior identification method provided by the embodiment of the present application specifically includes the following steps:

[0033] S101, acquiring a to-be-detected video, and performing target identification on the to-be-detected video to determine at least one to-be-identified target.

[0034] In the embodiment, the to-be-detected video can be specifically understood as a video captured by a camera device arranged in a scene to be monitored. The to-be-identified target can be specifically understood as an object in the to-be-detected video that has a possible fall behavior. Generally, the to-be-identified target is a person in the to-be-detected video.

[0035] Specifically, the device is set up in the scene to be monitored, and after the setup is completed, the video in the monitoring area is collected by the camera device arranged in the scene, and the collected video is determined as the to-be-detected video. The target identification algorithm relying on computer vision technology is used to detect each person target in the monitoring scene, that is, the to-be-detected video is detected by the target identification algorithm with the person as the target, and each person detected is determined as the to-be-identified target.

[0036] Optionally, the target recognition can be performed by a pre-trained target recognition model, which can be a YOLOv5 model based on deep learning, or other models with target recognition function, and the embodiments of the present application do not limit this. YOLOv5 has a lower computational complexity while maintaining a high detection accuracy, and is suitable for real-time target detection in a computing resource limited scenario. In the embodiments of the present application, YOLOv5 model can be used for target recognition to reduce the data resource requirement in the fall behavior recognition process, so as to efficiently and accurately identify the to-be-recognized target for subsequent processing.

[0037] S102, tracking each to-be-recognized target in the to-be-detected video to determine a behavior image sequence corresponding to each to-be-recognized target.

[0038] In the embodiments, the behavior image sequence can be understood as an image sequence obtained by arranging images containing behaviors of to-be-recognized targets in time sequence, which are obtained by capturing the to-be-detected video in real time.

[0039] Specifically, after target recognition, the position of each to-be-recognized target in each frame of the to-be-detected video can be determined and labeled by a bounding box. For each to-be-recognized target, the target tracking algorithm can be used to associate the regions related to the to-be-recognized target in different frames, thereby realizing the tracking of the to-be-recognized target in the to-be-detected video. The region of the to-be-recognized target can be extracted in each to-be-detected video frame in which the to-be-recognized target is recognized, and the behavior image of the to-be-recognized target corresponding to the capture time of the to-be-detected video frame can be obtained. Then, all the behavior images of the to-be-recognized targets obtained are arranged according to the capture time, and the behavior image sequence corresponding to the to-be-recognized target can be obtained.

[0040] Exemplarily, when the fall behavior is first identified, the behavior image sequence can include a plurality of frames of behavior images obtained after target identification is performed on the to-be-identified target, starting from the starting frame of the to-be-detected video, and the number of images is the number required to meet the input requirement of the fall behavior identification model; when the fall behavior is not identified for the first time, the behavior image sequence can include a plurality of frames of behavior images used for the last fall behavior identification, and a plurality of frames of new behavior images obtained due to real-time acquisition and real-time target identification of the to-be-detected video from the time when the plurality of frames of behavior images used for the last fall behavior identification are input into the fall behavior identification model to the current time, which will be used for the current fall behavior identification; or, when the fall behavior is not identified for the first time, the behavior image sequence can include a plurality of frames of historical behavior images obtained by frame extraction on the plurality of frames of behavior images used for the last fall behavior identification, and a plurality of frames of new behavior images obtained due to real-time acquisition and real-time target identification of the to-be-detected video from the time when the plurality of frames of behavior images used for the last fall behavior identification are input into the fall behavior identification model to the current time, which will be used for the current fall behavior identification; in this case, the number of images in the determined behavior image sequence will be kept to the number required to meet the input requirement of the fall behavior identification model, thereby reducing the data storage space required to maintain the behavior image sequence. The above-mentioned determination manners of the behavior image sequence are optional implementation manners exemplified by the embodiments of the present application, and can be adaptively selected according to actual application requirements, which are not limited by the embodiments of the present application.

[0041] Optionally, multi-target tracking can be performed on each to-be-identified target in the to-be-detected video, and the behavior image sequence belonging to the same to-be-identified target is determined according to the tracking result.

[0042] Exemplarily, different to-be-identified targets can be tracked simultaneously based on the determined plurality of to-be-identified targets, for example, the DeepSORT (Deep Simple Online and Realtime Tracking) algorithm can be used to realize real-time tracking of multiple targets in a monitoring scene, and when the algorithm is applied to the embodiments of the present application, real-time tracking of the plurality of to-be-identified targets in the to-be-detected video is realized, and the behavior image sequence corresponding to each to-be-identified target is obtained.

[0043] In the embodiments of the present application, the multi-target tracking algorithm is used to track different to-be-identified targets simultaneously, and the plurality of to-be-identified targets identified in the to-be-detected video are tracked uniformly in one processing, thereby improving the data calculation efficiency of the target identification process.

[0044] S103, for each behavior image sequence, when the current fall behavior recognition is not the first time, determining a last-recognized behavior image set and a current-to-be-recognized behavior image set according to the behavior image sequence, and determining the current behavior image sequence according to the last-recognized behavior image set and the current-to-be-recognized behavior image set.

[0045] In the embodiment, the last-recognized behavior image set can be specifically understood as a set of behavior images in the behavior image sequence used for recognition in the last time when the fall behavior recognition is performed. The current-to-be-recognized behavior image set can be specifically understood as a set of behavior images in the behavior image sequence that need to be continuously recognized according to the time sequence at the current time. The current behavior image sequence can be specifically understood as an image sequence used for the fall behavior recognition at the current time.

[0046] Specifically, for each behavior image sequence, that is, for each to-be-recognized target, when the fall behavior recognition operation corresponding to the current time is not the first time, the set of behavior images used for the last fall behavior recognition in the behavior image sequence can be determined as the last-recognized image set. After the last-recognized image set is determined, the behavior images in the behavior image sequence that have not been used for the fall behavior recognition can be determined, a plurality of behavior images that are adjacent to the latest behavior image in the last-recognized image set can be determined as the behavior images that need to be currently recognized, a set of the behavior images can be determined as the current-to-be-recognized behavior image set, and the behavior images in the last-recognized behavior image set and the current-to-be-recognized behavior image set can be combined according to the time sequence to obtain the current behavior image sequence.

[0047] In the embodiment, the last-recognized behavior image and the current-to-be-recognized behavior image are combined to generate the current behavior image sequence used for the current behavior recognition, so that the current behavior image sequence contains both the information that has not been recognized and the historical information that has been recognized once, the historical behavior can be well used as a basis for judging the fall at the current time, and the accuracy of the subsequent fall behavior determination according to the current behavior image sequence is improved.

[0048] S104, inputting the current behavior image sequence into the pre-trained fall behavior recognition model, determining a current fall behavior recognition result, and determining a target fall behavior recognition result of the to-be-recognized target according to the current fall behavior recognition result and a historical fall behavior recognition result.

[0049] In the embodiment, the fall behavior recognition model can be specifically understood as a neural network model that performs feature extraction and action classification on the input image sequence, and outputs a behavior recognition result of the classified continuous action of the person in the image sequence. Optionally, the TSM (Temporal Shift Module) behavior recognition model can be used as the fall behavior recognition model in the embodiment, and other models with behavior recognition capability can also be used as the fall behavior recognition model, which is not limited in the embodiment. The current fall behavior recognition result can be specifically understood as an identification result of the current behavior image sequence obtained by classifying the behavior of the current behavior image sequence, which is used to represent whether the target to be identified has a fall behavior. The historical fall behavior recognition result can be specifically understood as an identification result obtained by a fall behavior recognition at a previous time.

[0050] Specifically, the current behavior image sequence is input into the pre-trained fall behavior recognition model, and the fall behavior recognition model is used to perform feature extraction, action correlation and action classification on each behavior image in the current behavior image sequence, and output a current fall behavior recognition result of whether the target to be identified falls at the current time. Then, the current fall behavior recognition result and the historical fall behavior recognition result are combined to determine whether the target to be identified is indeed in a falling state at the current time, and the determined result is used as the target fall behavior recognition result of the target to be identified.

[0051] The technical scheme of the embodiment is as follows: a video to be detected is acquired, target recognition is performed on the video to be detected, and at least one target to be recognized is determined; target tracking is performed on each target to be recognized in the video to be detected, and a behavior image sequence corresponding to each target to be recognized is determined; for each behavior image sequence, when the current fall behavior recognition is non-first-time recognition, a last-recognized behavior image set and a current behavior image set to be recognized are determined according to the behavior image sequence, and the current behavior image sequence is determined according to the last-recognized behavior image set and the current behavior image set to be recognized; the current behavior image sequence is input into a pre-trained fall behavior recognition model, a current fall behavior recognition result is determined, and a target fall behavior recognition result of the target to be recognized is determined according to the current fall behavior recognition result and a historical fall behavior recognition result. By using the above technical scheme, when the video to be detected is acquired, target recognition and target tracking are preferentially performed on the video to be detected, a plurality of behavior image sequences of the targets to be recognized that need to be recognized are obtained, and then for each behavior image sequence, when the fall behavior is non-first-time recognition, part of the behavior images that have not been recognized are taken as the current behavior image set to be recognized, and the behavior images that have been recognized are taken as the last-recognized behavior image set, the current behavior image sequence for current fall behavior recognition is obtained by combining the two sets of images, the current behavior image sequence is input into the pre-trained fall behavior recognition model for current fall behavior recognition, and then the target fall behavior recognition result of whether the target to be recognized falls is determined according to the current fall behavior recognition result and the historical fall behavior recognition result of the last recognition. Since the current behavior image sequence contains both recognized and unrecognized images, the fall behavior recognition model can fully consider the information in the last recognition in one recognition, the continuity of two consecutive recognitions is enhanced, and then the fall behavior recognition model is more accurate for current fall behavior recognition, and when the target fall behavior recognition result of whether the target falls is determined, the historical fall behavior recognition result with the highest correlation with the current fall recognition result is taken as the basis for judging whether the target falls, the accuracy of determining the state of the person falling is improved, the false detection rate is reduced, and then the subsequent processing according to the target fall behavior recognition result can be correctly performed.

[0052] Figure 2A flowchart of a fall behavior recognition method provided for another embodiment of the application, the embodiment of the application is further optimized on the basis of the above-mentioned optional technical solutions, the set of historical behavior images to be merged obtained by frame extraction from the last recognized behavior image set extracted from the behavior image sequence is combined with the current recognized behavior image set extracted from the behavior image sequence according to the time sequence to obtain a current behavior image sequence containing both historical behavior information and current expected detection, and then the current fall behavior recognition considering the motion state of the recognized target before the current recognition can be performed by using the current behavior image sequence and the pre-trained fall behavior recognition model, and the target fall behavior recognition result of the recognized target is determined as a fall only when two fall behaviors are continuously recognized, avoiding the reduction of recognition result accuracy caused by single recognition misjudgment, improving the accuracy of determining the personnel fall state, and reducing the false detection rate. At the same time, by frame extraction and size conversion on the obtained original detection video before target recognition, and then format conversion on the original video frame after pre-conversion, the data amount required during format conversion is reduced, and the efficiency of fall behavior recognition is improved.

[0053] As shown in Figure 2 , the fall behavior recognition method provided by the embodiment of the application specifically includes the following steps:

[0054] S201, acquiring a detection video, frame extracting the detection video, and determining a set of original video frames.

[0055] In the embodiment, the set of original video frames can be understood as a set of video frames arranged in time sequence and directly extracted from a plurality of video frames constituting the detection video.

[0056] Specifically, after acquiring the detection video of the specified area, the software with frame extraction function is used to perform frame extraction processing on the detection video, so as to reduce the number of video frames needing target recognition, determine each video frame obtained by frame extraction as the original video frame of the detection video, and determine the set of original video frames constituted by arranging the original video frames in time sequence.

[0057] Exemplarily, the detection video can be frame extracted by ffmpeg, and the corresponding size image can be obtained according to the shooting resolution of the actual camera device. In the embodiment of the application, the size of 1920*1080 is taken as an example. Since the scheme of the application aims to recognize the motion behavior of the detection target in the detection video, the continuity of the motion after frame extraction needs to be considered, so the frame extraction interval should not be too large. The detection video can be frame extracted by selecting frame extraction mode to obtain the set of original video frames.

[0058] S202, scaling, random translation, rotation and scale transformation are performed on each original video frame in the original video frame set, so that each transformed original video frame meets the input requirements of the pre-trained target recognition model.

[0059] In this embodiment, the target recognition model can be specifically understood as a neural network model that performs feature extraction, candidate region determination, classification regression and other processes on the input image, and outputs the region information and confidence information of the desired recognized target in the image.

[0060] Specifically, since the original video frames in the to-be-detected video captured by the camera equipment often have high resolution and large size, the target recognition model often does not need a large image input to obtain a correct recognition result, in order to generate an image for inputting the target recognition model for target recognition, the original video frame needs to be preprocessed such as scaling, random translation, rotation and scale transformation, so as to obtain an original video frame meeting the input requirements of the pre-trained target recognition model.

[0061] S203, format conversion is performed on each transformed original video frame to determine the to-be-detected video frame set.

[0062] Specifically, since the image data format of the original video frame often depends on the camera equipment that collects it, and the pre-trained target recognition model often needs the image input to be in a pre-set fixed grid image data format, after completing the size transformation of the original video frame, format conversion is further performed on each transformed original video frame, so that the original video frame after completing the size transformation and format conversion is a to-be-detected video frame that can be directly input into the target recognition model, and then each to-be-detected video frame is sorted according to the collection order of the original video frame to obtain the to-be-detected video frame set.

[0063] It should be noted that when the image data format of the original video frame is NV12 and the image data format required by the target recognition model is RGB format, the conventional method is to first convert the format of the original video frame to RGB format, and then perform scale transformation on the RGB format. In this embodiment, the original video frame in NV12 format is first scaled, and then the scaled image is converted to RGB format, so that the calculation amount when converting to RGB is smaller, and when performing large data amount operation, the overall operation amount can be reduced.

[0064] For example, the random translation, rotation and scale transformation in this embodiment can be implemented in the following manner:

[0065] Random translation is as follows:

[0066]

[0067] wherein (x, y, l) represents the matrix before image translation, (x', y', l) represents the matrix after image translation, d x and d y are the pixel amounts of image translation on the x-axis and y-axis respectively.

[0068] The scale transformation is as follows:

[0069]

[0070] wherein (x, y, l) represents the matrix before image size transformation, (x", y", l) represents the matrix after image size transformation, s x and s y are the scale transformation factors of image on the x-axis and y-axis respectively.

[0071] The rotation is as follows:

[0072]

[0073] wherein (x, y, l) represents the matrix before image rotation, (x"', y"', l) represents the matrix after image rotation, and is the rotation angle.

[0074] In the above example, if the image input size requirement of the target recognition model is 640*640, after random translation, rotation and scale transformation, the size of the original video frame can be converted from 1920*1080 to 640*640 to meet the input requirement of the target recognition model.

[0075] S204, input the set of video frames to be detected into the target recognition model, and determine at least one target to be recognized according to the output result of the target recognition model.

[0076] Specifically, the set of video frames to be detected is input into the target recognition model, and the target recognition model performs target recognition on each video frame to be detected to obtain the recognition frame in which a human image may exist in each video frame to be detected. The targets with similar features in each recognition frame are determined as the same target to be recognized, so that at least one target to be recognized in the set of video frames to be detected can be determined according to the output result of the target recognition model.

[0077] S205, target tracking is performed on each target to be recognized in the video to be detected, and a behavior image sequence corresponding to each target to be recognized is determined.

[0078] Specifically, since each target to be identified is identified in the video to be detected, the target recognition model can output the bounding box position and class information of each target, and then when tracking different targets to be identified, a feature vector can be extracted from the bounding box position corresponding to each target to be identified, which can capture the semantic and visual information of the target to be identified, and then the similarity between the feature vectors can be used to associate the targets to be identified in the current frame and the previous frame. For the associated targets to be identified, the target state is predicted, updated and managed, that is, for each target to be identified in the video to be detected, the video frame where the target to be identified is located and the corresponding behavior image in the video frame are determined, and the obtained behavior images are associated in time sequence to obtain the corresponding behavior image sequence.

[0079] S206, for each behavior image sequence, determine whether the current fall behavior recognition for the behavior image sequence is the first recognition, if yes, execute S207; if no, execute S208.

[0080] Specifically, for each behavior image sequence, it is determined whether the fall behavior recognition for the behavior image sequence at the current time, i.e., for the target to be identified corresponding to the behavior image sequence, is the first time. The way of constructing the image sequence for fall behavior recognition is different between the first recognition and the non-first recognition. In the first recognition, step S207 is executed; in the non-first recognition, step S208 is executed.

[0081] S207, the sequence of the first frame behavior image in the behavior image sequence as the starting point and the continuous preset number of frame behavior images is determined as the current behavior image sequence, and S212 is executed.

[0082] In this embodiment, the preset number can be understood as the number of images in the image sequence input to the fall behavior recognition model according to the actual training of the fall behavior recognition model and the actual needs.

[0083] Specifically, when the current fall behavior recognition of the behavior image sequence is the first recognition, it can be considered that all behavior images in the behavior image sequence have not been input to the fall behavior recognition model for fall behavior recognition. At this time, the sequence of the first frame behavior image as the starting point and the previous preset number of frame behavior images containing the first frame can be directly determined as the current behavior image sequence corresponding to the fall behavior recognition at the current time, and S212 is further executed.

[0084] S208, extracting the behavior image last input to the fall behavior recognition model from the behavior image sequence to determine the last identified behavior image set.

[0085] Specifically, when the current fall behavior recognition of the behavior image sequence is not the first recognition, it can be considered that there is at least a group of behavior images in the behavior image sequence which have been input into the fall behavior recognition model, at this time, the behavior images in the behavior image sequence which were input into the fall behavior recognition model last time can be extracted, it can be considered that the historical behavior information closely related to the current fall behavior recognition is contained therein, and the set composed of the extracted behavior images is determined as the last recognition behavior image set.

[0086] In S209, the last recognition behavior image set is frame-extracted to determine a to-be-merged historical behavior image set.

[0087] Specifically, according to the actual demand, the proportion of the required historical behavior information in the current behavior image sequence when the current fall behavior recognition is performed is determined, the number of to-be-merged historical behavior images to be extracted from the last recognition behavior image set is determined, that is, the preset number is multiplied by the proportion to determine the number of images in the to-be-merged historical behavior image set, after the number of images is determined, the last recognition behavior image set can be uniformly frame-extracted based on the number of images, the images obtained by frame-extraction are determined as to-be-merged historical behavior images, and the set composed of the to-be-merged historical behavior images arranged in the order of occurrence is determined as the to-be-merged historical behavior image set.

[0088] For example, the number of images in the to-be-merged historical behavior image set can be half of the preset number in S207, that is, the historical behavior information in the current behavior image sequence and the current to-be-recognized image information each accounts for half, at this time, the last recognition behavior image set can be frame-extracted every other frame to obtain the to-be-merged historical behavior image set.

[0089] Optionally, when the last recognition behavior image set is frame-extracted, the last recognition behavior image set is uniformly frame-extracted starting from a non-first frame position. Since the frame-extraction of the last recognition behavior image set is non-first frame extraction each time, after the frame-extraction of the last recognition behavior image set is performed for multiple times, the behavior images far from the current time for fall behavior recognition image frame will no longer be retained in the to-be-merged historical behavior image set, so that in the current behavior image sequence composed of the to-be-merged historical behavior image set and the current to-be-recognized behavior image set, the historical behavior can be fully retained, and the influence of the behavior images at a remote time on the current fall behavior recognition can be gradually reduced.

[0090] For example, assuming that 10 frames are required to input the current behavior image sequence into the fall behavior recognition model, the current set of images to be recognized during the first recognition can be represented as [1,2,3,4,5,6,7,8,9,10]. After completing one fall behavior recognition, this current set of images to be recognized will be used as the previous set of images for the next recognition. If the previous set of images to be recognized is extracted using the first frame extraction method, and combined with the current set of images to be recognized to construct the current behavior image sequence, then the corresponding image frames in the second constructed current behavior image sequence can be represented as [1,3,5,7,9,11,12,13,14,15], and the corresponding image frames in the third constructed current behavior image sequence can be represented as...

[0091] [1,5,9,12,14,16,17,18,19,20], and so on, it can be seen that the first frame of the behavior image furthest from the current moment is always in the current behavior image sequence. However, the first frame of the behavior image no longer affects the current behavior recognition. If it is always included in the current behavior image sequence, it will reduce the accuracy of the fall behavior recognition result based on the current behavior image sequence. However, if a non-first frame extraction method is used to extract frames from the previous recognized behavior image set and combine it with the current behavior image set to be recognized to construct the current behavior image sequence, then the corresponding image frames in the second constructed current behavior image sequence can be represented as [2,4,6,8,10,11,12,13,14,15], and the corresponding image frames in the third constructed current behavior image sequence can be represented as...

[0092] [4,8,11,13,15,16,17,18,19,20], and so on, it can be seen that the behavior image farthest from the current moment in the current behavior image sequence will gradually approach the behavior image at the current moment as time goes by, thus improving the accuracy of the fall behavior recognition result based on the current behavior image sequence.

[0093] S210. Extract the behavior images located after the previous set of recognized behavior images from the behavior image sequence in chronological order, and determine the current set of behavior images to be recognized.

[0094] The ratio of the number of images in the current set of images to be identified to the number of images in the set of historical images to be merged is a preset ratio.

[0095] In this embodiment, the preset ratio can be understood as the ratio of the number of current behavior images to be identified to the number of historical behavior images to be merged, which is set in advance according to the actual situation. In other words, it can be understood as the ratio between the current image and the historical image in the current behavior image sequence.

[0096] Specifically, the latest identified behavior image in the previous identified behavior image set is determined in the behavior image sequence, and a behavior image after the latest identified behavior image is taken as a starting point. According to a preset proportion of the current to-be-identified behavior image and the historical behavior information in the current behavior image sequence, the number of the current to-be-identified behavior image is determined, and then the same number of behavior images in the behavior image sequence are extracted in time sequence to form a set, which is determined as the current to-be-identified behavior image set. Optionally, the preset proportion can be one, that is, the number of the current to-be-identified behavior image is consistent with the number of the to-be-merged historical behavior image.

[0097] S211, combining the to-be-merged historical behavior image set and the current to-be-identified behavior image set in time sequence to determine the current behavior image sequence.

[0098] Specifically, since the behavior image sequence is generated in time sequence, the collection time of each to-be-merged historical behavior image and each current to-be-identified behavior image extracted from the behavior image sequence should be included. The to-be-merged historical behavior image set and the current to-be-identified behavior image set are sequentially combined according to the collection time of each behavior image, and after the combined image sequence is processed to adapt to the fall behavior recognition model, the current behavior image sequence for current time fall behavior recognition is determined.

[0099] Optionally, Figure 3 A flowchart example of combining the to-be-merged historical behavior image set and the current to-be-identified behavior image set in time sequence to determine the current behavior image sequence is provided for an embodiment of the application, as shown in FIG. 8, which specifically includes the following steps: Figure 3

[0100] S2111, combining the to-be-merged historical behavior image set and the current to-be-identified behavior image set in time sequence to determine the current to-be-processed behavior image sequence.

[0101] Specifically, the to-be-merged historical behavior image set and the current to-be-identified behavior image set are sequentially combined according to the collection time of each behavior image, and the combined image sequence is determined as the current to-be-processed behavior image sequence.

[0102] S2112, scaling, random translation, rotation and scale transformation are performed on each current to-be-processed behavior image in the current to-be-processed behavior image sequence, so that the size of each current to-be-processed behavior image after transformation meets the input requirements of the fall behavior recognition model.

[0103] ​Specifically, each current behavior image in the current behavior image sequence is scaled, randomly translated, rotated and scaled according to the processing manner in S202, so that the size of each current behavior image after processing is the input size required by the fall behavior recognition model.

[0104] According to the above example, since each current behavior image in the current behavior image sequence is an action image extracted in a video frame after target tracking, the sizes of the extracted action images are different due to different action behaviors of the target to be recognized in different frames. When inputting the fall behavior recognition model, each image input thereto needs to be of a uniform size. If the image input size required by the fall behavior recognition model is 224*224, the size of each current behavior image will be converted to 224*224 after processing in S2112, so as to meet the processing requirement of the model.

[0105] S2113, the set of each transformed current behavior image is determined as the current behavior image sequence.

[0106] S212, the current behavior image sequence is input into the pre-trained fall behavior recognition model to determine a current fall behavior recognition result.

[0107] S213, it is determined whether the current fall behavior recognition result and the historical fall behavior recognition result are both falls. If yes, S214 is performed; if no, S215 is performed.

[0108] Specifically, by determining whether the current fall behavior recognition result and the historical fall behavior recognition result are both falls, if yes, it is considered that the target to be recognized is in a fall state through the fall behavior recognition model twice in succession, and S214 is performed at this time; if no, it is considered that the current fall behavior recognition result is not a fall, or the last fall state recognition result for the target to be recognized is not a fall, that is, the target to be recognized has not been recognized to fall twice in succession, and it is not determined whether the recognition result of the fall behavior recognition model is misrecognition, and S215 is performed at this time.

[0109] S214, the target fall behavior recognition result of the target to be recognized is determined as a fall.

[0110] Optionally, after determining that the target fall behavior recognition result of the target to be recognized is a fall, corresponding alarm or other subsequent processing can be performed according to the target fall behavior recognition result, and at the same time, the fall behavior recognition for the behavior image sequence corresponding to the target to be recognized can be ended, thereby reducing the amount of data to be processed.

[0111] S215, the current fall behavior recognition result is determined as a new historical fall behavior recognition result, and S208 is returned to be performed.

[0112] Specifically, after the current fall behavior recognition result is obtained, if the current fall behavior recognition result and the historical fall behavior recognition result cannot be used to determine whether the to-be-recognized target is in a falling state, it can be temporarily considered that the to-be-recognized target is not in a falling state at the current detection moment. At this time, the current fall behavior recognition result is saved as the historical fall behavior recognition result, and the execution of S208 is returned to determine the current behavior image sequence input into the fall behavior recognition model as the last recognized behavior image set, and the next fall behavior recognition of the to-be-recognized target is performed. After the behavior image sequence corresponding to the to-be-recognized target has been completely input into the fall behavior recognition model for processing, if two consecutive fall recognition results are not recognized, it can be considered that the target fall behavior recognition result of the to-be-recognized target is determined as not falling.

[0113] It can be understood that for the to-be-recognized target, the behavior image sequence corresponding to the to-be-recognized target can be updated according to the real-time collected video to be detected. Since the current behavior image sequence corresponding to the current fall behavior recognition result is frame-extracted and reserved as the basis for the next behavior recognition when the target fall behavior recognition result cannot be directly determined according to the current fall behavior recognition result, the historical behavior can be fully reserved and the influence of the video frame at a remote moment on the current behavior recognition can be gradually reduced, thereby improving the accuracy of each fall behavior recognition of the to-be-recognized target.

[0114] The technical scheme of the embodiment, by frame-extracting the last recognized behavior image set extracted from the behavior image sequence to obtain the to-be-merged historical behavior image set, and combining the current recognized behavior image set extracted from the behavior image sequence according to the time sequence to obtain the current behavior image sequence containing historical behavior information and current expected detection, and then using the current behavior image sequence and the pre-trained fall behavior recognition model to perform the current fall behavior recognition considering the motion state of the recognized target before the current recognition, the target fall behavior recognition result of the to-be-recognized target is determined as falling only when two consecutive fall behaviors are recognized, thereby avoiding the reduction of recognition result accuracy caused by single recognition misjudgment, improving the accuracy of determination of the falling state of the personnel, and reducing the false detection rate. At the same time, by frame-extracting and size-converting the obtained original video to be detected before target recognition, and then converting the format of the original video frame after the pre-conversion, the data amount required during format conversion is reduced, and the efficiency of fall behavior recognition is improved.

[0115] Figure 4 A structural schematic diagram of a fall behavior recognition device provided by an embodiment of the present application is shown in FIG. 1. Figure 4 As shown in FIG. 1, the fall behavior recognition device includes a target determination module 31, an image sequence determination module 32, a current sequence determination module 33, and a recognition result determination module 34.

[0116] The target determination module 31 is configured to acquire a video to be detected and perform target recognition on the video to be detected to determine at least one target to be recognized. The image sequence determination module 32 is configured to perform target tracking on each target to be recognized in the video to be detected to determine a behavior image sequence corresponding to each target to be recognized. The current sequence determination module 33 is configured to, for each behavior image sequence, when a current fall behavior recognition is a non-first recognition, determine a last recognized behavior image set and a current to-be-recognized behavior image set according to the behavior image sequence, and determine a current behavior image sequence according to the last recognized behavior image set and the current to-be-recognized behavior image set. The recognition result determination module 34 is configured to input the current behavior image sequence into a pre-trained fall behavior recognition model, determine a current fall behavior recognition result, and determine a target fall behavior recognition result of the target to be recognized according to the current fall behavior recognition result and a historical fall behavior recognition result.

[0117] The technical scheme of the embodiment of the present application, when the video to be detected is acquired, preferentially performs target recognition and target tracking on the video to be detected to obtain a plurality of behavior image sequences of the targets to be recognized, which need to be recognized, and then, for each behavior image sequence, when a fall behavior recognition is a non-first recognition, part of the behavior images that have not been recognized are taken as a current to-be-recognized behavior image set, and the behavior images that have been recognized are taken as a last recognized behavior image set, the current behavior image sequence for the current fall behavior recognition is obtained by combining the two sets of images, the current behavior image sequence is input into a pre-trained fall behavior recognition model to recognize the current fall behavior, and then, the target fall behavior recognition result of whether the target to be recognized falls is determined by combining the current fall behavior recognition result and a last recognized historical fall behavior recognition result. Since the current behavior image sequence contains both recognized and unrecognized images, the fall behavior recognition model can fully consider the information in the last recognition in one recognition, the continuity of the two consecutive recognitions is enhanced, and then the fall behavior recognition model is more accurate for the current fall behavior recognition, and when the target fall behavior recognition result of whether the target falls is determined, the historical fall behavior recognition result with the highest correlation with the current fall recognition result is taken as the basis for judging whether the target falls, the accuracy of determining the state of the person falling is improved, the false detection rate is reduced, and then the subsequent processing according to the target fall behavior recognition result can be correctly performed.

[0118] Optionally, the current sequence determination module 33 comprises:

[0119] The last image set determination unit is configured to extract behavior images input into the fall behavior recognition model last time from the behavior image sequence, and determine a last recognized behavior image set.

[0120] The to-be-merged image set determination unit is configured to extract frames from the last recognized behavior image set, and determine a to-be-merged historical behavior image set.

[0121] The current image set determination unit is configured to extract behavior images after the last recognized behavior image set from the behavior image sequence in chronological order, and determine a current to-be-recognized behavior image set. A ratio of a number of images in the current to-be-recognized behavior image set to a number of images in the to-be-merged historical behavior image set is a preset proportion.

[0122] The current image sequence determination unit is configured to combine the to-be-merged historical behavior image set and the current to-be-recognized behavior image set in chronological order, and determine a current behavior image sequence.

[0123] Optionally, the current image sequence determination unit is specifically configured to:

[0124] combine the to-be-merged historical behavior image set and the current to-be-recognized behavior image set in chronological order, and determine a current to-be-processed behavior image sequence.

[0125] scale, randomly translate, rotate, and scale transform each current to-be-processed behavior image in the current to-be-processed behavior image sequence, so that sizes of the transformed current to-be-processed behavior images meet input requirements of the fall behavior recognition model.

[0126] determine a set of the transformed current to-be-processed behavior images as the current behavior image sequence.

[0127] Optionally, the recognition result determination module 34 is specifically configured to:

[0128] if the current fall behavior recognition result and the historical fall behavior recognition result are both falls, determine a target fall behavior recognition result of the to-be-recognized target as a fall;

[0129] otherwise, determine the current fall behavior recognition result as a new historical fall behavior recognition result, and return to execute the steps of determining the last recognized behavior image set and the current to-be-recognized behavior image set according to the behavior image sequence.

[0130] Optionally, the current sequence determination module 33 is further configured to:

[0131] for each behavior image sequence, when the current fall behavior recognition is a first recognition, determine a sequence of a preset number of behavior images starting from a first row behavior image in the behavior image sequence as the current behavior image sequence.

[0132] Optionally, the target determination module 31 comprises:

[0133] a video frame extraction unit configured to extract frames from the video to be detected to determine a set of original video frames;

[0134] a size transformation unit configured to scale, randomly translate, rotate and scale transform each original video frame in the set of original video frames so that each transformed original video frame conforms to the input requirement of the pre-trained target recognition model;

[0135] a format conversion unit configured to convert the format of each transformed original video frame to determine a set of video frames to be detected;

[0136] a target determination unit configured to input the set of video frames to be detected into the target recognition model and determine at least one target to be recognized according to the output result of the target recognition model.

[0137] Optionally, the image sequence determination module 32 is specifically configured to:

[0138] perform multi-target tracking on each target to be recognized in the video to be detected and determine a behavior image sequence belonging to the same target to be recognized according to the tracking result.

[0139] The fall behavior recognition device provided by the embodiments of the present application can execute the fall behavior recognition method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0140] Figure 5 A structural schematic diagram of a fall behavior recognition device provided by an embodiment of the present application. The fall behavior recognition device 40 can be an electronic device, which is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0141] As Figure 5As shown, the fall behavior recognition device 40 includes at least one processor 41, and a memory, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., communicatively connected to the at least one processor 41, where the memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 42 or loaded from the storage unit 48 into the random access memory (RAM) 43. Various programs and data required for the operation of the fall behavior recognition device 40 can also be stored in the RAM 43. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0142] Various components in the fall behavior recognition device 40 are connected to the I / O interface 45, including an input unit 46, such as a keyboard, a mouse, etc., an output unit 47, such as various types of displays, speakers, etc., a storage unit 48, such as a magnetic disk, an optical disk, etc., and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the fall behavior recognition device 40 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0143] The processor 41 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 performs various methods and processes described above, such as the fall behavior recognition method.

[0144] In some embodiments, the fall behavior recognition method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the fall behavior recognition device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded onto the RAM 43 and executed by the processor 41, one or more steps of the fall behavior recognition method described above can be performed. Alternatively, in other embodiments, the processor 41 can be configured to perform the fall behavior recognition method by any other appropriate means, such as by means of firmware.

[0145] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0146] Computer programs used to implement the processes of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program

[0147] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0149] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0150] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0151] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0152] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure. Any alternatives, modifications, equivalents, and the like of all of the above described devices, systems, methods, etc. are intended to be encompassed by the present disclosure.

Claims

1. A fall behavior recognition method, characterized by, The method comprises: acquiring a to-be-detected video, and performing target recognition on the to-be-detected video to determine at least one to-be-recognized target; performing target tracking on each to-be-recognized target in the to-be-detected video to determine a behavior image sequence corresponding to each to-be-recognized target; for each behavior image sequence, when a current fall behavior recognition is not a first recognition, determining a last-recognized behavior image set and a current to-be-recognized behavior image set according to the behavior image sequence, and determining a current behavior image sequence according to the last-recognized behavior image set and the current to-be-recognized behavior image set; inputting the current behavior image sequence into a pre-trained fall behavior recognition model to determine a current fall behavior recognition result, and determining a target fall behavior recognition result of the to-be-recognized target according to the current fall behavior recognition result and a historical fall behavior recognition result; wherein the historical fall behavior recognition result is a recognition result obtained by a fall behavior recognition at a time point before a current time point; wherein the determination of the target fall behavior recognition result of the to-be-recognized target according to the current fall behavior recognition result and the historical fall behavior recognition result comprises: if the current fall behavior recognition result and the historical fall behavior recognition result are both falls, determining the target fall behavior recognition result of the to-be-recognized target as a fall; otherwise, determining the current fall behavior recognition result as a new historical fall behavior recognition result, and returning to perform the step of determining the last-recognized behavior image set and the current to-be-recognized behavior image set according to the behavior image sequence.

2. The method of claim 1, wherein, The determination of the last-recognized behavior image set and the current to-be-recognized behavior image set according to the behavior image sequence, and the determination of the current behavior image sequence according to the last-recognized behavior image set and the current to-be-recognized behavior image set, comprise: extracting a behavior image last input into the fall behavior recognition model from the behavior image sequence to determine a last-recognized behavior image set; frame extraction is performed on the last-recognized behavior image set to determine a to-be-merged historical behavior image set; extracting behavior images located after the last-recognized behavior image set from the behavior image sequence in chronological order to determine a current to-be-recognized behavior image set; wherein a ratio of a number of images in the current to-be-recognized behavior image set to a number of images in the to-be-merged historical behavior image set is a preset proportion; combining the to-be-merged historical behavior image set and the current to-be-recognized behavior image set in chronological order to determine a current behavior image sequence.

3. The method of claim 2, wherein, The combination of the to-be-merged historical behavior image set and the current to-be-recognized behavior image set in chronological order to determine a current behavior image sequence comprises: combining the to-be-merged historical behavior image set and the current to-be-recognized behavior image set in chronological order to determine a current to-be-processed behavior image sequence; scaling, random translation, rotation and scale transformation are performed on each current to-be-processed behavior image in the current to-be-processed behavior image sequence, so that the size of each current to-be-processed behavior image after transformation meets the input requirement of the fall behavior recognition model; The set of each transformed current to-be-processed behavior image is determined as a current behavior image sequence.

4. The method of claim 1, wherein, After the behavior image sequences corresponding to each of the to-be-identified targets are determined, the method further includes: For each of the behavior image sequences, when the current fall behavior is identified for the first time, a sequence composed of a first frame behavior image in the behavior image sequence and a preset number of continuous frame behavior images is determined as a current behavior image sequence.

5. The method according to any one of claims 1 to 4, characterized in that, The target identification on the to-be-detected video includes: Frame extraction is performed on the to-be-detected video to determine a set of original video frames; Each original video frame in the set of original video frames is scaled, randomly translated, rotated, and dimensionally transformed so that each transformed original video frame meets the input requirements of a pre-trained target identification model; Each transformed original video frame is format-converted to determine a set of to-be-detected video frames; The set of to-be-detected video frames is input into the target identification model, and at least one to-be-identified target is determined according to an output result of the target identification model.

6. The method according to any one of claims 1-4, characterized in that, The target tracking on each of the to-be-identified targets in the to-be-detected video includes: Multi-target tracking is performed on each of the to-be-identified targets in the to-be-detected video, and behavior image sequences belonging to the same to-be-identified target are determined according to a tracking result. 7.A fall behavior recognition apparatus, characterized by, The method includes: A target determination module is configured to acquire a to-be-detected video and perform target identification on the to-be-detected video to determine at least one to-be-identified target; An image sequence determination module is configured to perform target tracking on each of the to-be-identified targets in the to-be-detected video to determine behavior image sequences corresponding to each of the to-be-identified targets; A current sequence determination module is configured to, for each of the behavior image sequences, when the current fall behavior is identified for the first time, determine a last-identified behavior image set and a current to-be-identified behavior image set according to the behavior image sequence, and determine a current behavior image sequence according to the last-identified behavior image set and the current to-be-identified behavior image set; An identification result determination module is configured to input the current behavior image sequence into a pre-trained fall behavior identification model to determine a current fall behavior identification result, and determine a target fall behavior identification result of the to-be-identified target according to the current fall behavior identification result and a historical fall behavior identification result; The historical fall behavior identification result is an identification result obtained by a fall behavior identification at a time point before a current time point; The determination of the target fall behavior identification result of the to-be-identified target according to the current fall behavior identification result and the historical fall behavior identification result includes: If both the current fall behavior identification result and the historical fall behavior identification result are falls, the target fall behavior identification result of the to-be-identified target is determined as a fall; Otherwise, the current fall behavior identification result is determined as a new historical fall behavior identification result, and the step of determining a last-identified behavior image set and a current to-be-identified behavior image set according to the behavior image sequence is performed again.

8. A fall behavior recognition apparatus, characterized by, The method includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the fall behavior identification method of any one of claims 1-6.

9. A storage medium containing computer-executable instructions, wherein: the computer executable instructions, when executed by a computer processor, are used to perform the fall behavior identification method of any one of claims 1-6.

Citation Information

Patent Citations

  • Human body fall detection alarm device based on multiple sensors

    CN102800170A

  • Fall risk prediction method, system and device

    CN110379131A