Behavior recognition method, electronic device and readable storage medium

By extracting the human skeleton information and continuous sequences in the real-time video stream, and using the Softmax classifier for behavior recognition, the problem of the inability to accurately identify real-time continuous video streams in the prior art is solved, and efficient and accurate behavior recognition is achieved.

CN114373219BActive Publication Date: 2025-05-13CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011097409.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-14
Publication Date
2025-05-13
Estimated Expiration
2040-10-14

AI Technical Summary

Technical Problem

The prior art cannot accurately identify real-time continuous video streams containing multiple actions and backgrounds, and there are problems such as high labor costs and difficult to guarantee the quality of identification.

Method used

By obtaining the image frames collected in real time, extracting human skeleton information, and obtaining two consecutive skeleton sequences, the Softmax classifier is used to determine whether trusted recognition results can be obtained and accurate behavior recognition results can be output.

Benefits of technology

Real-time accurate identification of continuous video streams containing multiple actions and backgrounds is achieved, and is suitable for behavior recognition scenarios that require emergency processing, reducing labor costs and improving identification quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114373219B_ABST
    Figure CN114373219B_ABST
Patent Text Reader

Abstract

The present invention provides a behavior recognition method, an electronic device and a readable storage medium, belonging to the field of intelligent recognition technology, the method comprising: obtaining image frames collected in real time; extracting human skeleton information in the image frames; obtaining a first human skeleton sequence and a second human skeleton sequence, the first human skeleton sequence and the second human skeleton sequence both including human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and differs from the current image frame by d image frames, d is less than L; according to the first human skeleton sequence and the second human skeleton sequence, determining whether a credible recognition result can be obtained. The present invention can perform real-time and accurate behavior recognition on a continuous video stream containing multiple actions and backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent recognition technology, and in particular to a behavior recognition method, an electronic device and a readable storage medium. Background Art

[0002] Behavior recognition is a basic technology in the fields of intelligent monitoring and human-computer interaction. Through deep learning and other algorithm technologies, the video data collected by the camera is analyzed to perceive the behavior of people in the video scene. With the development of 5G and the continuous advancement of edge intelligent applications in the Internet of Things, the application value of behavior recognition technology will continue to expand, and its application fields will be very wide, such as elderly care, security, education, shopping, etc.

[0003] Currently, most behavior recognition is done by monitoring personnel observing surveillance videos. This behavior recognition method has problems such as high labor costs, inability to concentrate for a long time, and difficulty in ensuring recognition quality. With the development of deep learning, many behavior recognition algorithms have emerged. However, these behavior recognition algorithms are mostly used in offline non-real-time scenarios and cannot accurately recognize behaviors in real-time continuous video streams containing multiple actions and backgrounds. Summary of the invention

[0004] In view of this, the present invention provides a behavior recognition method, an electronic device and a readable storage medium, which are used to solve the problem that it is currently impossible to accurately recognize behaviors in a real-time continuous video stream containing multiple actions and backgrounds.

[0005] In order to solve the above technical problems, in a first aspect, the present invention provides a behavior recognition method, comprising:

[0006] Get the image frames collected in real time;

[0007] Extracting human skeleton information in the image frame;

[0008] Acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and differs from the current image frame by d image frames, d and L are positive integers and d is less than L;

[0009] Based on the first human skeleton sequence and the second human skeleton sequence, determine whether a credible recognition result can be obtained, and when it is determined that a credible recognition result can be obtained, output the credible recognition result.

[0010] Optionally, the step of acquiring the first human skeleton sequence and the second human skeleton sequence includes:

[0011] The human skeleton information extracted from the image frames starting from the second image frame is sequentially added to a preset queue of length L until the preset queue is full, and the L human skeleton information in the preset queue is obtained as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs;

[0012] Delete the earliest added d human skeleton information in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

[0013] Optionally, the method further includes:

[0014] If the human skeleton information is not successfully extracted from the image frame, the existing human skeleton information in the preset queue is cleared.

[0015] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence includes:

[0016] Inputting the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0017] Inputting the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0018] It is determined whether the credible identification result can be obtained according to the first probability value and the second probability value.

[0019] Optionally, determining whether the credible identification result can be obtained according to the first probability value and the second probability value includes:

[0020] If the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions, then it is determined that the reliable identification result can be obtained.

[0021] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, includes:

[0022] Running a behavior recognition algorithm model, and having the behavior recognition algorithm model determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained;

[0023] The method further comprises: determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, and further comprising:

[0024] If human skeleton information is not successfully extracted from the image frame, the operation of the behavior recognition algorithm model is terminated.

[0025] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, includes:

[0026] If the second human skeleton sequence is obtained, the behavior recognition algorithm model is run;

[0027] The behavior recognition algorithm model determines whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained.

[0028] In a second aspect, the present invention further provides an electronic device, comprising:

[0029] An image acquisition module, used to acquire image frames collected in real time;

[0030] A skeleton data extraction module, used to extract human skeleton information in the image frame;

[0031] A skeleton sequence acquisition module is used to acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and there is a difference of d image frames between the first image frame and the current image frame, d and L are positive integers and d is less than L;

[0032] The recognition module is used to determine whether a reliable recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and output the reliable recognition result when it is determined that a reliable recognition result can be obtained.

[0033] Optionally, the skeleton sequence acquisition module includes:

[0034] A first acquisition unit is used to sequentially add the human skeleton information extracted from the image frames starting from the second image frame to a preset queue of length L until the preset queue is full, and acquire the L human skeleton information in the preset queue as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs;

[0035] The second acquisition unit is used to delete the earliest d human skeleton information added in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

[0036] Optionally, the electronic device further includes:

[0037] The queue clearing module is used to clear the existing human skeleton information in the preset queue if the human skeleton information is not successfully extracted from the image frame.

[0038] Optionally, the identification module includes:

[0039] A first classification unit is used to input the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0040] A second classification unit is used to input the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0041] A judgment unit is used to determine whether the credible recognition result can be obtained according to the first probability value and the second probability value.

[0042] Optionally, the judgment unit is used to determine that the trusted identification result can be obtained if the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions.

[0043] Optionally, the recognition module is used to run a behavior recognition algorithm model, and the behavior recognition algorithm model determines whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained;

[0044] The electronic device further comprises:

[0045] The termination module is used to terminate the operation of the behavior recognition algorithm model if human skeleton information is not successfully extracted from the image frame.

[0046] Optionally, the identification module includes:

[0047] an activation unit, configured to run a behavior recognition algorithm model if the second human skeleton sequence is acquired;

[0048] The recognition unit is used to determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence by the behavior recognition algorithm model, and output the credible recognition result when it is determined that a credible recognition result can be obtained.

[0049] In a third aspect, the present invention further provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, the steps in any one of the above-mentioned behavior recognition methods are implemented.

[0050] In a fourth aspect, the present invention further provides a readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in any of the above-mentioned behavior recognition methods.

[0051] The beneficial effects of the above technical solution of the present invention are as follows:

[0052] In the embodiment of the present invention, the image frames of a continuous video stream containing multiple actions and backgrounds can be segmented, and real-time and accurate recognition results can be given. Therefore, it can be applied to behavior recognition scenarios where the continuous video stream containing multiple actions and backgrounds collected by data acquisition devices such as cameras without segmentation needs to be recognized, and it can also be applied to behavior recognition scenarios where real-time and accurate recognition results need to be obtained, for example, it can be applied to some behavior recognition scenarios for handling emergency situations. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of a behavior recognition method in Embodiment 1 of the present invention;

[0054] Figure 2 A schematic diagram of a queue-based sliding window mechanism in an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of a queue-based sliding window action scoring in an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of a general real-time online behavior recognition process in an embodiment of the present invention;

[0057] Figure 5 This is a schematic diagram of the structure of an electronic device in Embodiment 2 of the present invention;

[0058] Figure 6 This is a schematic diagram of the structure of an electronic device in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.

[0060] In the related technologies, there are a lot of researches on algorithms in the direction of offline non-real-time behavior recognition. Depending on the processed data sets, there are also the following two offline non-real-time behavior recognition algorithms:

[0061] 1) Behavior recognition algorithms that recognize edited and only contain one segmented action sample. Currently, there are many models that use deep learning methods to learn the spatiotemporal features of optical flow, video or skeleton data for behavior recognition. Since the dataset used is a single segmented action, this type of behavior recognition algorithm is a classification model.

[0062] 2) Identify an action recognition algorithm that contains a background segment and multiple action samples. Since the processed data contains multiple unsegmented actions, research in this field mainly divides the problem into two sub-problems. First, a regression algorithm is used to segment the action start boundary, and then a classification algorithm is used for action recognition.

[0063] However, the above two offline non-real-time behavior recognition algorithms have certain shortcomings, so they cannot perform real-time online behavior recognition. Specifically:

[0064] The above single action behavior recognition algorithm can only recognize one action segment of the input at a time, so it cannot process continuous video segments containing multiple actions;

[0065] The above-mentioned multi-action behavior recognition algorithm cannot provide real-time online behavior recognition results because it needs to use the context information of the entire video.

[0066] For real-time online behavior recognition, not only can it not directly obtain the start of an action, but it can also not use the future data information of an unfinished action like offline behavior recognition algorithms. In related technologies, there is a method of using future video generation algorithms to generate future video frames to obtain additional information for behavior recognition algorithm research. However, this real-time behavior recognition algorithm requires the use of additional video generation algorithms, the model is complex, and it is impossible to provide accurate real-time behavior recognition evaluation criteria.

[0067] See also Figure 1 , Figure 1 A flow chart of a behavior recognition method provided in the first embodiment of the present invention includes the following steps:

[0068] Step 11: Acquire the real-time collected image frames;

[0069] Step 12: extracting human skeleton information in the image frame;

[0070] Specifically, a posture estimation algorithm may be used to preprocess image frames collected in real time to extract human skeleton information (or human skeleton data).

[0071] Step 13: obtaining a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and differs from the current image frame by d image frames, d and L are positive integers and d is less than L;

[0072] Specifically, the first image frame is before the current image frame, and the sequence number of the first image frame differs from that of the current image frame by d.

[0073] Step 14: Determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and output the credible recognition result when it is determined that a credible recognition result can be obtained.

[0074] In an embodiment of the present invention, a posture estimation algorithm is used to preprocess image frames collected in real time to extract human skeleton information (or skeleton data). The background segmentation can be naturally completed according to whether the human skeleton information is extracted from the collected video image frames, thereby eliminating the interference of such background segments without human activities in the real-time continuously collected video image frames to behavior recognition.

[0075] The behavior recognition method provided by the embodiment of the present invention can segment the image frames of a continuous video stream containing multiple actions and backgrounds, and give real-time and accurate recognition results. Therefore, it can be applied to the behavior recognition scenario where the continuous video stream containing multiple actions and backgrounds collected by a data acquisition device such as a camera without segmentation is needed to be recognized, and it can also be applied to the behavior recognition scenario where real-time and accurate recognition results need to be obtained, for example, it can be applied to some behavior recognition scenarios for handling emergency situations.

[0076] The above behavior recognition method is illustrated below with examples.

[0077] In one optional specific implementation manner, the obtaining of the first human skeleton sequence and the second human skeleton sequence includes:

[0078] The human skeleton information extracted from the image frames starting from the second image frame is sequentially added to a preset queue of length L until the preset queue is full, and the L human skeleton information in the preset queue is obtained as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs;

[0079] Delete the earliest added d human skeleton information in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

[0080] Through the study of behavior recognition algorithms, it is found that the most important factor affecting the accuracy of behavior recognition results is the partial key frames in the entire action. Therefore, whether all the processes of an action can be completely input into the behavior recognition algorithm model is not very important. Therefore, the embodiment of the present invention proposes a queue-based sliding window mechanism to store key frames, realize the segmentation of action candidate segments, and solve the problem that the current offline behavior recognition algorithm cannot segment the start of the action in the real-time video stream. In addition, since the image frames have been preprocessed by using the posture estimation algorithm, the human skeleton data has been extracted, and the background segments have been well segmented, the embodiment of the present invention can use a single action behavior recognition algorithm that recognizes one action at a time to perform online real-time behavior recognition. That is to say, after using the posture estimation algorithm to preprocess the video stream image frames collected in real time, extracting the human skeleton data, and realizing the segmentation of the action candidate segments based on the sliding window mechanism of the queue, the single action behavior recognition algorithm can be used to perform behavior recognition on the action candidate segments in real time, and the image frames of the continuous video stream containing multiple actions and backgrounds can be accurately recognized.

[0081] The above queue-based sliding window mechanism is described in detail below.

[0082] See also Figure 2 After using the posture estimation algorithm to preprocess the real-time collected video stream image frames, the real-time video stream can be regarded as a skeleton sequence of multiple continuous actions. Set the sliding window size to L, the window sliding step length d, and the window sliding step length d can be set according to the computing power of the behavior recognition algorithm model deployment device. The sliding window uses a queue to cache the action sequence (i.e., the human skeleton information sequence) with a total length of L frames from the current image frame. When the data in the queue is full, all the data in the queue is input into the behavior recognition algorithm model for recognition, and the earliest 1 to d frames of data in the queue are deleted to store the latest skeleton data. This queue-based sliding window mechanism realizes the segmentation of continuous action frames, while ensuring that the behavior recognition algorithm model currently processes the behavior data at the current moment.

[0083] The embodiment of the present invention proposes a queue-based sliding window mechanism, which ensures the real-time performance of the behavior recognition result and realizes the dynamic segmentation of continuous action frames.

[0084] Optionally, the method further includes:

[0085] If the human skeleton information is not successfully extracted from the image frame, the existing human skeleton information in the preset queue is cleared.

[0086] In an embodiment of the present invention, behavior recognition is required based on human skeleton information of continuous image frames. Therefore, if the human skeleton information is not successfully extracted from the latest image frame, the existing human skeleton information in the queue is cleared until the queue is filled with a human skeleton sequence of L continuous image frames. All human skeleton information in the queue is input into the behavior recognition algorithm model for recognition, and at the same time, the earliest 1 to d frame data in the queue are deleted to store the latest skeleton data.

[0087] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence includes:

[0088] Inputting the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0089] Inputting the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0090] It is determined whether the credible identification result can be obtained according to the first probability value and the second probability value.

[0091] The division of continuous action frames can be completed through a queue-based sliding window mechanism, but this division does not guarantee that the starting part of a key frame of an action is exactly divided out, and it may even include the beginning and end parts of two actions. Therefore, the embodiment of the present invention needs to determine whether the image frame corresponding to the previous human skeleton sequence happens to include the key frame part of a certain action based on two consecutive human skeleton sequences.

[0092] The Softmax classifier is the most commonly used classification method in neural networks. For an input X, it can determine which category it belongs to among N categories. For a single action recognition algorithm, it deals with a classification problem, so this type of behavior recognition algorithm usually adds a Softmax layer at the end of the network to complete the action classification. The human skeleton sequence is input to give a score of which of the N actions it belongs to. The higher the score, the greater the possibility that the input action frame belongs to a certain category. The highest score is considered to be the correct category of the input action frame. However, the score range is very wide, and it would be better if it could be converted into a probability. Softmax is a normalization function that converts the (-∞, +∞) score into a set of probabilities and makes their sum equal to 1.

[0093]

[0094] Among them, s i Represents the model's score for the input X in the i-th category.

[0095] against Figure 2 The real-time video stream containing human skeleton information shown in Figure 2 shows that as the sliding window moves, the scores given by the Softmax function are as follows: Figure 3 As shown. Figure 3 It can be seen that for a sliding window containing two actions, the action judgment will not be very certain, and the scores corresponding to all actions will not be very prominent, such as the score corresponding to window 1; as the sliding window moves, the key frame sequences it contains will gradually increase, and the scores of the corresponding actions will show a fluctuating upward trend, such as the scores corresponding to windows 2-4; when the window roughly includes the key frame part, its score will fluctuate and decrease as the window slides, such as the scores of windows 4-6.

[0096] In summary, the embodiment of the present invention sets the following threshold determination strategy.

[0097] In one optional specific implementation manner, determining whether the credible recognition result can be obtained according to the first probability value and the second probability value includes:

[0098] If the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions, it is determined that the credible recognition result can be obtained. The credible recognition result is the first action.

[0099] In an embodiment of the present invention, the second probability value is the probability value corresponding to the N preset actions obtained by inputting the first human skeleton sequence into the Softmax classifier, that is, the probability value corresponding to the image frame of the current window, and the first probability value is the probability value corresponding to the N preset actions obtained by inputting the second human skeleton sequence into the Softmax classifier, that is, the probability value corresponding to the image frame of the previous window.

[0100] Specifically, we can use k t Represents the highest score in the Softmax scoring result in the tth window:

[0101]

[0102] in, is the set threshold, when When , it indicates that the classification algorithm of behavior recognition cannot be very sure that the human skeleton sequence in the input video segment is a certain action. It may be the connection between two actions. The score of the behavior recognition result of the current window for a certain action is greater than And it is lower than the maximum score given by the previous window, indicating that the sliding window has passed the key frame part of the current action, and a reliable behavior recognition result can be given at this time.

[0103] In another optional specific implementation, if the maximum probability value among the second probability values ​​is greater than a first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, the difference between the maximum probability value among the first probability values ​​and the maximum probability value among the second probability values ​​is greater than a second preset value, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions, then it is determined that the reliable identification result can be obtained.

[0104] In order to avoid fluctuations in the recognition results of the Softmax function and affect the accuracy of the recognition results, a judgment condition is added in the embodiment of the present invention: the difference between the maximum probability value in the first probability values ​​and the maximum probability value in the second probability values ​​is greater than a second preset value.

[0105] Specifically, we can use k t Represents the highest score in the Softmax scoring result in the tth window:

[0106] Give reliable recognition results,

[0107] in, and is the set threshold, and The values ​​of are greater than 0 and less than 1. When , it indicates that the classification algorithm of behavior recognition cannot be very sure that the human skeleton sequence in the input video segment is a certain action. It may be the connection between two actions. The score of the behavior recognition result of the current window for a certain action is greater than and the decrease from the maximum score given in the previous window is greater than , it indicates that the sliding window has passed the key frame part of the current action, and a reliable behavior recognition result can be given at this time.

[0108] The threshold judgment strategy provided by the embodiment of the present invention can be used to distinguish whether the sliding window just contains the key frame part of a certain action, so as to provide a reliable behavior recognition result.

[0109] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, includes:

[0110] Running a behavior recognition algorithm model, and having the behavior recognition algorithm model determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained;

[0111] The method further comprises: determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, and further comprising:

[0112] If human skeleton information is not successfully extracted from the image frame, the operation of the behavior recognition algorithm model is terminated.

[0113] Specifically, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence includes:

[0114] Inputting the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0115] Inputting the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0116] It is determined whether the credible identification result can be obtained according to the first probability value and the second probability value.

[0117] Therefore, after obtaining the second human skeleton sequence, it is necessary to run the behavior recognition algorithm model to input the second human skeleton sequence into the behavior recognition algorithm model to obtain a first probability value, and then input the first human skeleton sequence into the behavior recognition algorithm model to obtain a second probability value. Afterwards, if the human skeleton information is not successfully extracted from the image frame, the operation of the behavior recognition algorithm model is terminated.

[0118] Optionally, determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, includes:

[0119] If the second human skeleton sequence is obtained, the behavior recognition algorithm model is run;

[0120] The behavior recognition algorithm model determines whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained.

[0121] In actual scenarios, there will be a large number of people in the current field of view who perform such meaningless actions, and there will also be many continuous background frames. In order to avoid the useless computing and storage overhead caused by these situations, an embodiment of the present invention provides a dynamic activation strategy for the form recognition algorithm model, which realizes the dynamic activation of the behavior recognition algorithm model by judging whether the current frame can extract skeleton data and the status of the cache queue.

[0122] Specifically, when no skeleton information can be extracted in the current frame, it means that there is no current behavior recognition task, and the behavior recognition algorithm model can be terminated. When skeleton information frames are continuously stored in the buffer area corresponding to the preset queue for a period of time, since the action of walking through the field of view is very short, it will not fill the buffer queue. Therefore, the behavior recognition algorithm model will be activated again only when the buffer queue is filled up cumulatively.

[0123] The following example illustrates the above behavior recognition method. Figure 4 , Figure 4 This is a general real-time online behavior recognition process diagram, which is mainly divided into the following four steps:

[0124] 1) By using the posture estimation algorithm to preprocess the image frames and extract the human skeleton information, the background segments can be segmented well and the interference of the background segments can be eliminated.

[0125] 2) Use a queue-based sliding window mechanism to store key frames. Specifically, after extracting human skeleton information for the current image frame, determine whether the human skeleton information is successfully extracted. If the human skeleton information is successfully extracted, store the extracted skeleton information in the cache queue. If the human skeleton information is not successfully extracted, determine whether there is skeleton information in the queue. If there is skeleton information in the queue, clear the queue. Determine whether the cache queue is full. If the cache queue is full, input all the data in the queue into the behavior recognition algorithm model for behavior recognition.

[0126] 3) Based on the action score given by the Softmax classifier and the prior knowledge brought by the sliding window mechanism, a threshold judgment strategy is designed to evaluate the single behavior recognition result and provide real-time and accurate behavior recognition results. Specifically, the behavior recognition algorithm model performs behavior recognition on the human skeleton information in the sliding window and judges the credibility of the behavior recognition result. If it is determined to be credible, the current action recognition result is output. In other words, the embodiment of the present invention proposes a confidence threshold strategy, combined with a queue-based sliding window mechanism, using the recognition results of multiple sliding windows, and controls the credibility of the behavior recognition result according to the threshold judgment.

[0127] 4) In the above process, in order to reduce the load of the system, the embodiment of the present invention also provides a dynamic activation strategy for the behavior recognition algorithm, which can dynamically release the computing and storage space when there is no behavior recognition task, avoiding invalid calculation and storage overhead of the behavior recognition algorithm model. Specifically, after extracting the human skeleton information for the current image frame, it is determined whether the human skeleton information is successfully extracted. If the human skeleton information is not successfully extracted, it is determined whether the behavior recognition algorithm model is running. If the behavior recognition algorithm model is already running, the behavior recognition algorithm model is terminated. When using a queue-based sliding window mechanism for key frame storage, if the cache queue is full, it is determined whether the behavior recognition algorithm model is activated. If not activated, the behavior recognition algorithm model is activated.

[0128] See also Figure 5 , Figure 5 : is a schematic diagram of the structure of an electronic device provided in Embodiment 2 of the present invention. The electronic device 50 includes:

[0129] An image acquisition module 51 is used to acquire image frames collected in real time;

[0130] A skeleton data extraction module 52, used to extract human skeleton information in the image frame;

[0131] The skeleton sequence acquisition module 53 is used to acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and there is a difference of d image frames between the first image frame and the current image frame, d and L are positive integers and d is less than L;

[0132] The recognition module 54 is used to determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and output the credible recognition result when it is determined that a credible recognition result can be obtained.

[0133] Optionally, the skeleton sequence acquisition module 53 includes:

[0134] A first acquisition unit is used to sequentially add the human skeleton information extracted from the image frames starting from the second image frame to a preset queue of length L until the preset queue is full, and acquire the L human skeleton information in the preset queue as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs;

[0135] The second acquisition unit is used to delete the earliest d human skeleton information added in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

[0136] Optionally, the electronic device further includes:

[0137] The queue clearing module is used to clear the existing human skeleton information in the preset queue if the human skeleton information is not successfully extracted from the image frame.

[0138] Optionally, the identification module 54 includes:

[0139] A first classification unit is used to input the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0140] A second classification unit is used to input the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0141] A judgment unit is used to determine whether the credible recognition result can be obtained according to the first probability value and the second probability value.

[0142] Optionally, the judgment unit is used to determine that the trusted identification result can be obtained if the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions.

[0143] Optionally, the recognition module 54 is used to run a behavior recognition algorithm model, and the behavior recognition algorithm model determines whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained;

[0144] The electronic device further comprises:

[0145] The termination module is used to terminate the operation of the behavior recognition algorithm model if human skeleton information is not successfully extracted from the image frame.

[0146] Optionally, the identification module 54 includes:

[0147] an activation unit, configured to run a behavior recognition algorithm model if the second human skeleton sequence is acquired;

[0148] The recognition unit is used to determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence by the behavior recognition algorithm model, and output the credible recognition result when it is determined that a credible recognition result can be obtained.

[0149] The embodiment of the present invention is a product embodiment corresponding to the above-mentioned method embodiment 1, so it will not be described in detail here. Please refer to the above-mentioned embodiment 1 for details.

[0150] See also Figure 6 , Figure 6 : is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. The electronic device 60 includes a processor 61, a memory 62, and a program stored in the memory 62 and executable on the processor 61; when the processor 61 executes the program, the following steps are implemented:

[0151] Get the image frames collected in real time;

[0152] Extracting human skeleton information in the image frame;

[0153] Acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and differs from the current image frame by d image frames, d and L are positive integers and d is less than L;

[0154] Based on the first human skeleton sequence and the second human skeleton sequence, determine whether a credible recognition result can be obtained, and when it is determined that a credible recognition result can be obtained, output the credible recognition result.

[0155] In the embodiment of the present invention, the image frames of a continuous video stream containing multiple actions and backgrounds can be segmented, and real-time and accurate recognition results can be given. Therefore, it can be applied to behavior recognition scenarios where the continuous video stream containing multiple actions and backgrounds collected by data acquisition devices such as cameras without segmentation needs to be recognized, and it can also be applied to behavior recognition scenarios where real-time and accurate recognition results need to be obtained, for example, it can be applied to some behavior recognition scenarios for handling emergency situations.

[0156] Optionally, the processor 61 may further implement the following steps when executing the program:

[0157] The obtaining of the first human skeleton sequence and the second human skeleton sequence comprises:

[0158] The human skeleton information extracted from the image frames starting from the second image frame is sequentially added to a preset queue of length L until the preset queue is full, and the L human skeleton information in the preset queue is obtained as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs;

[0159] Delete the earliest added d human skeleton information in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

[0160] Optionally, the processor 61 may further implement the following steps when executing the program:

[0161] If the human skeleton information is not successfully extracted from the image frame, the existing human skeleton information in the preset queue is cleared.

[0162] Optionally, the processor 61 may further implement the following steps when executing the program:

[0163] The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence includes:

[0164] Inputting the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1;

[0165] Inputting the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively;

[0166] It is determined whether the credible identification result can be obtained according to the first probability value and the second probability value.

[0167] Optionally, the processor 61 may further implement the following steps when executing the program:

[0168] The determining, according to the first probability value and the second probability value, whether the credible identification result can be obtained includes:

[0169] If the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions, then it is determined that the reliable identification result can be obtained.

[0170] Optionally, the processor 61 may further implement the following steps when executing the program:

[0171] The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, comprises:

[0172] Running a behavior recognition algorithm model, and having the behavior recognition algorithm model determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained;

[0173] The method further comprises: determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, and further comprising:

[0174] If human skeleton information is not successfully extracted from the image frame, the operation of the behavior recognition algorithm model is terminated.

[0175] Optionally, the processor 61 may further implement the following steps when executing the program:

[0176] The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, comprises:

[0177] If the second human skeleton sequence is obtained, the behavior recognition algorithm model is run;

[0178] The behavior recognition algorithm model determines whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained.

[0179] The specific working process of the embodiment of the present invention is consistent with that in the above method embodiment 1, so it will not be repeated here. Please refer to the description of the method steps in the above method embodiment 1 for details.

[0180] Embodiment 4 of the present invention provides a readable storage medium on which a program is stored, and when the program is executed by a processor, the steps in any one of the behavior recognition methods in the above embodiment 1 are implemented. For details, please refer to the description of the method steps in the above corresponding embodiment.

[0181] The above-mentioned readable storage medium includes computer-readable storage medium. Computer-readable storage media include permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0182] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A behavior recognition method, characterized in that: include: Get the image frames collected in real time; Extracting human skeleton information in the image frame; Acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and differs from the current image frame by d image frames, d and L are positive integers and d is less than L; Determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained; The obtaining of the first human skeleton sequence and the second human skeleton sequence comprises: The human skeleton information extracted from the image frames starting from the second image frame is sequentially added to a preset queue of length L until the preset queue is full, and the L human skeleton information in the preset queue is obtained as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs; Delete the earliest added d human skeleton information in the preset queue, continue to store the human skeleton information extracted from the image frame in chronological order until the preset queue is filled again, and obtain the latest L human skeleton information in the preset queue as the first human skeleton sequence.

2. The method according to claim 1, characterized in that: Also includes: If the human skeleton information is not successfully extracted from the image frame, the existing human skeleton information in the preset queue is cleared.

3. The method according to claim 1, characterized in that The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence includes: Inputting the second human skeleton sequence into a Softmax classifier to obtain first probability values ​​corresponding to N preset actions respectively; wherein N is a positive integer greater than 1; Inputting the first human skeleton sequence into the Softmax classifier to obtain second probability values ​​corresponding to the N preset actions respectively; It is determined whether the credible identification result can be obtained according to the first probability value and the second probability value.

4. The method according to claim 3, characterized in that The determining, according to the first probability value and the second probability value, whether the credible identification result can be obtained includes: If the maximum probability value among the second probability values ​​is greater than the first preset value, the maximum probability value among the first probability values ​​is greater than the maximum probability value among the second probability values, and the maximum probability value among the second probability values ​​and the maximum probability value among the first probability values ​​are both probability values ​​corresponding to the first action among the N preset actions, then it is determined that the reliable identification result can be obtained.

5. The method according to claim 1, characterized in that The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, comprises: Running a behavior recognition algorithm model, and having the behavior recognition algorithm model determine whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained; The method further comprises: determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, and further comprising: If human skeleton information is not successfully extracted from the image frame, the operation of the behavior recognition algorithm model is terminated.

6. The method according to claim 1, characterized in that The determining whether a credible recognition result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and outputting the credible recognition result when it is determined that a credible recognition result can be obtained, comprises: If the second human skeleton sequence is obtained, the behavior recognition algorithm model is run; The behavior recognition algorithm model determines whether a credible recognition result can be obtained based on the first human skeleton sequence and the second human skeleton sequence, and outputs the credible recognition result when it is determined that a credible recognition result can be obtained.

7. An electronic device, characterized in that: include: An image acquisition module, used to acquire image frames collected in real time; A skeleton data extraction module, used to extract human skeleton information in the image frame; A skeleton sequence acquisition module is used to acquire a first human skeleton sequence and a second human skeleton sequence, wherein the first human skeleton sequence and the second human skeleton sequence both include human skeleton information of L consecutive image frames, the last human skeleton information in the first human skeleton sequence is the human skeleton information of the current image frame, the last human skeleton information in the second human skeleton sequence is the human skeleton information of the first image frame, the first image frame is before the current image frame and there is a difference of d image frames between the first image frame and the current image frame, d and L are positive integers and d is less than L; an identification module, configured to determine whether a credible identification result can be obtained according to the first human skeleton sequence and the second human skeleton sequence, and output the credible identification result when it is determined that a credible identification result can be obtained; Among them, the obtaining of the first human skeleton sequence and the second human skeleton sequence includes: adding the human skeleton information extracted from the image frames starting from the second image frame to a preset queue of length L in sequence until the preset queue is full, and obtaining the L human skeleton information in the preset queue as the second human skeleton sequence; the second image frame is the image frame to which the first human skeleton information in the second human skeleton sequence belongs; deleting the earliest added d human skeleton information in the preset queue, and continuing to store the human skeleton information extracted from the image frames in chronological order until the preset queue is filled again, and obtaining the latest L human skeleton information in the preset queue as the first human skeleton sequence.

8. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that: When the processor executes the program, the steps in the behavior recognition method according to any one of claims 1 to 6 are implemented.

9. A readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the behavior recognition method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • System and method for processing lower-limb muscle sound signals for exoskeleton robots

    CN104666052A

  • Parallelized human body behavior identification method

    CN104899561A