Online behavior recognition method, device and equipment for typical actions in substation room

By acquiring and processing frame sequences in the substation room, and using random and continuous frame division strategies for action recognition, the problems of poor behavior recognition effect and poor universality in the prior art are solved, and accurate identification and classification of videos of different action lengths and density are achieved.

CN119131894BActive Publication Date: 2025-05-09EAST CHINA BRANCH OF STATE GRID CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411153849.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-05-09
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of indoor personnel in substations with poor recognition effect, poor universality and limited generalization capabilities, especially when dealing with video clips of different action lengths and densities.

Method used

By obtaining the frame sequence to be identified, pose detection is performed, and the frame sequence is divided into different sets of sub-sequences according to the random frame division strategy and the continuous frame division strategy, the action segments are identified and labeled, and finally, after integrating the recognition results, typical actions and recognition action categories are accurately positioned.

Benefits of technology

It realizes that video clips of different action lengths and density can be identified in the substation room without a large amount of data labeling and training, accurately locate typical actions and accurately detect action categories, improving the versatility and reliability of behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131894B_ABST
    Figure CN119131894B_ABST
Patent Text Reader

Abstract

The present application discloses an online behavior recognition method, device and equipment for typical actions in a substation room. The method includes: obtaining a frame sequence to be recognized; performing posture detection on the frame sequence to be recognized; dividing the frame sequence to be recognized with posture information detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy; respectively recognizing the action segments and the action labels of the action segments in the first subsequence set and the second subsequence set, and determining a first recognition result and a second recognition result of the frame sequence to be recognized; according to the first recognition result and the second recognition result, determining multiple target action segments in the frame sequence to be recognized, and the action category corresponding to each target action segment. The present application does not require a large amount of data to be annotated in advance, nor does it require a large amount of training in advance, and can accurately identify the segments where typical actions occur in different video segments in the substation room recognition scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video image processing, and in particular to an online behavior recognition method, device and equipment for typical actions in a substation room. Background Art

[0002] Human behavior recognition is an important research direction in the field of computer vision and artificial intelligence. It aims to achieve automatic recognition of specific behaviors by capturing and analyzing human movements.

[0003] In related technologies, pure action recognition networks usually process video clips. Their input requires only one action to be recognized. The deep network extracts the temporal and spatial information contained in the video frame, and then predicts the category of the action in the current video through the classification layer, but it cannot locate the location where the action occurs. Sequential action detection adds action location prediction on the basis of action category detection.

[0004] However, the temporal action detection based on the neural network model requires a large amount of data with temporal information annotation to train the neural network model in order to achieve good positioning and classification results, which will bring a lot of time loss. It is very difficult to obtain data sets for specific scenarios, and additional temporal annotation will cause more costs. Moreover, the neural network model trained on a specific data set usually only has good results in the same type of scenario. Therefore, the temporal action detection based on the neural network model has poor versatility and limited generalization ability.

[0005] For the behavior recognition of personnel in the substation room, real-time performance and high accuracy are required. In addition to the performance of the above-mentioned detection network, another important factor affecting the accuracy of behavior recognition is the quality of the input video. There are many typical actions in the substation room, such as making a phone call, opening the door of the electric cabinet, closing the door of the electric cabinet, crossing the isolation belt, bypassing the isolation belt, standing up, squatting, running, walking, crossing obstacles, and talking with two people. However, during the maintenance process, the staff does not always have specific actions that need to be detected, and there are a large number of irrelevant frame sequences. In the related technology, sliding and inputting a certain number of frame sequences one by one according to a fixed time length will greatly affect the effect of behavior recognition. In addition, the division of frame sequences also affects the effect of behavior recognition. The time length of each action is different, and the action speed of different people is also different. How to accurately judge the start and end of the action is also an important issue. Summary of the invention

[0006] In view of this, the present application provides an online behavior recognition method, device and equipment for typical actions indoors in a substation. It does not require a large amount of data to be labeled in advance, nor does it require a large amount of training in advance. It can recognize video clips of different action lengths and different action densities in the indoor recognition scenario of the substation, accurately locate the clips where typical actions occur, and accurately detect the category of the action.

[0007] According to one aspect of the present application, there is provided an online behavior recognition method for typical actions indoors in a substation, comprising: obtaining a frame sequence to be recognized; wherein the frame sequence to be recognized comprises a plurality of frame images to be recognized arranged in chronological order; performing posture detection on the frame sequence to be recognized; dividing the frame sequence to be recognized in which posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy, respectively; respectively identifying action segments in the first subsequence set and the second subsequence set and action labels of the action segments, and determining a first recognition result and a second recognition result of the frame sequence to be recognized; and determining a plurality of target action segments in the frame sequence to be recognized, and action categories corresponding to each of the target action segments, according to the first recognition result and the second recognition result.

[0008] According to another aspect of the present application, there is provided an online behavior recognition device for typical actions in a substation room, comprising:

[0009] An acquisition module, used for acquiring a frame sequence to be identified; wherein the frame sequence to be identified includes a plurality of frame images to be identified arranged in chronological order;

[0010] A detection module, used for performing posture detection on the frame sequence to be identified;

[0011] A division module, used for dividing the frame sequence to be identified in which the posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy respectively;

[0012] an identification module, used to identify the action segments and the action tags of the action segments in the first subsequence set and the second subsequence set respectively, and determine a first identification result and a second identification result of the frame sequence to be identified;

[0013] A determination module is used to determine a plurality of target action segments and action categories corresponding to each of the target action segments in the to-be-identified frame sequence according to the first recognition result and the second recognition result.

[0014] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the steps of the online behavior identification method for typical actions in a substation room are implemented.

[0015] By means of the above technical scheme, the present application provides an online behavior recognition method, device and equipment for typical actions in a substation room, wherein the frame sequence to be recognized in which posture information is detected is divided into different subsequence sets respectively through a random frame division strategy and a continuous frame division strategy, and the first recognition result and the second recognition result of the frame sequence to be recognized are determined by respectively identifying different subsequence sets. Then, the first recognition result and the second recognition result are integrated and mutually verified to reduce the possibility of misjudgment, thereby determining multiple target action segments in the frame sequence to be recognized, and the action category corresponding to each target action segment. The present application does not require a large amount of data to be annotated in advance, nor does it require a large amount of training in advance. It can recognize video segments of different action lengths and different action densities in the recognition scene of the substation room, accurately locate the segments where typical actions occur, and accurately detect the category of the action, thereby improving the versatility and reliability of behavior recognition.

[0016] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A schematic flow chart of an online behavior recognition method for typical actions in a substation provided by an embodiment of the present application is shown;

[0019] Figure 2 Another flow chart of the online behavior recognition method for typical actions in a substation provided by an embodiment of the present application is shown;

[0020] Figure 3 Another flow chart of the online behavior recognition method for typical actions in a substation provided by an embodiment of the present application is shown;

[0021] Figure 4 A structural block diagram of an online behavior recognition device for typical actions in a substation provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.

[0023] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.

[0024] Those skilled in the art will appreciate that, unless expressly stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "connected" to another element, it may be directly connected or connected to the other element, or there may be intermediate elements. In addition, the "connection" or "connection" used herein may include wireless connection or wireless fusion. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0025] The online behavior recognition method for typical actions in the substation provided in the embodiment of the present application can be applied to the terminal, can also be applied to the server, and can also be software running in the terminal or the server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application of the online behavior recognition method for typical actions in the substation, etc., but is not limited to the above forms.

[0026] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in a variety of different forms and should not be interpreted as being limited to the embodiments set forth herein. It should be understood that these embodiments are provided to make the disclosure of the present application thorough and complete, and to fully convey the concepts of these exemplary embodiments to those of ordinary skill in the art.

[0027] In this embodiment, an online behavior recognition method for typical actions in a substation room is provided. Figure 1 As shown, the method includes:

[0028] Step 101: Obtain a frame sequence to be identified.

[0029] The frame sequence to be identified includes a plurality of frame images to be identified arranged in time sequence, namely, frame videos.

[0030] In this embodiment, a certain number of continuous frame images to be recognized are cached according to the time sequence of generation of the frame images to be recognized as a frame sequence to be recognized, so as to facilitate subsequent centralized processing.

[0031] Among them, the number of cached frame images to be identified can be set according to the specific needs of the actual application scenario, so that videos of various lengths can be processed at one time, thereby improving the versatility of the present application.

[0032] It is worth mentioning that for online videos, real-time processing can be achieved through a loop method, that is, a certain number of frame images to be identified are cached each time as a frame sequence to be identified. After the current frame sequence to be identified is identified, the next frame sequence to be identified is cached for identification. The relationship between recognition accuracy and recognition time is balanced by setting parameters such as the number of subsequences, so that the recognition time of each frame sequence to be identified is close to the total time corresponding to the frame sequence to be identified, thereby achieving real-time processing.

[0033] Step 102: Perform posture detection on the frame sequence to be identified.

[0034] In this embodiment, posture detection is performed on the frame sequence to be identified to extract the posture information of the human body in the frame sequence to be identified, provide accurate input information for the preset action recognition model in the subsequent steps, and improve the reliability of recognition.

[0035] Exemplarily, human posture information of each frame image to be identified in a frame sequence to be identified, such as human key points such as joint positions, is extracted through a preset posture detection model, and a thermal image is formed based on the extracted human posture information, thereby determining the input information required by the preset action recognition model in subsequent steps.

[0036] It should be noted that the preset posture detection model is pre-trained and can detect the posture information of the input image, and this application does not impose any restrictions.

[0037] Step 103 , dividing the frame sequence to be identified in which the posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy respectively.

[0038] In this embodiment, the frame sequence to be identified in which gesture information is detected is divided into different subsequence sets according to different division strategies to capture actions occurring in different areas of the frame sequence to be identified and of different durations, thereby improving the comprehensiveness of identification.

[0039] Step 104 , identifying the action segments and the action tags of the action segments in the first subsequence set and the second subsequence set respectively, and determining a first recognition result and a second recognition result of the frame sequence to be recognized.

[0040] In this embodiment, according to different subsequence sets of the frame sequence to be identified, the behavioral features in the frame sequence to be identified are captured from different angles and levels, reducing the limitations of a single method, being applicable to actions of different types and lengths, and video clips of different action intensities, thereby enhancing the versatility of behavior recognition.

[0041] Step 105 , determining a plurality of target action segments and an action category corresponding to each target action segment in the frame sequence to be identified according to the first recognition result and the second recognition result.

[0042] In this embodiment, the first recognition result and the second recognition result are integrated so that the first recognition result and the second recognition result can verify and complement each other, thereby reducing deviations and making the final recognition result more reliable.

[0043] It is worth mentioning that in this embodiment, the first subsequence set is obtained through a random frame division strategy, and the first subsequence set is identified to obtain a first identification result, and the second subsequence set is obtained through a continuous frame division strategy, and the second subsequence set is identified to obtain a second identification result. The two identification processes are parallel, that is, they are carried out simultaneously, independently of each other, and do not interfere with each other, thereby improving the efficiency of behavior identification to meet the requirements of real-time online behavior identification in the indoor scene of the substation. In addition, this embodiment combines the first identification result and the second identification result obtained by the two processing processes to reduce the error of behavior identification to meet the requirements of accuracy of online behavior identification in the indoor scene of the substation.

[0044] Furthermore, as a refinement and extension of the specific implementation methods of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, in step 103, the frame sequence to be identified in which posture information is detected is divided into a first subsequence set according to a random frame division strategy, specifically including: randomly generating a first preset number of first target subsequences according to multiple different first preset lengths of the frame sequence to be identified in which posture information is detected.

[0045] In this example, based on the R-CNN (Region-based Convolutional Neural Networks) target detection method, a certain number and length of frame sequences are preset in the time dimension, and the frame sequence to be identified is randomly divided into multiple first target subsequences to more carefully analyze actions of different lengths in different regions in the frame sequence to be identified.

[0046] Exemplarily, a first length list including multiple different first preset lengths is obtained, a first preset number of first target subsequences are randomly generated from the frame sequence to be identified according to the first length list, and the first preset number of first target subsequences are used as the first subsequence set.

[0047] In actual application scenarios, since the length of each action is different and random, and the action speeds of different people are also inconsistent, the length of the action segment corresponding to each action is different, and the density of actions in different video segments is also different. This embodiment randomly samples the frame sequence to be identified through multiple different first preset lengths to cover actions of different lengths. At the same time, this embodiment adjusts the number of first target subsequences (first preset number) to adapt to video segments with different action intensities, thereby improving the versatility of behavior recognition.

[0048] The first preset length is determined according to the longest action and the shortest action in the actual application scenario. For example, in the behavior recognition scenario in the substation room, it is considered that the shortest action corresponds to 15 frames of images and the longest action corresponds to 90 frames of images. Therefore, within the range of 15-90 frames, sampling can be performed every 10 frames. For example, the first length list includes multiple different first preset lengths of 15 frames, 25 frames, 34 frames, 45 frames, 55 frames, 65 frames, 75 frames, 85 frames and 90 frames.

[0049] like Figure 2 As shown, in one embodiment, when a frame sequence to be identified in which posture information is detected is randomly generated into a first preset number of first target subsequences according to a plurality of different first preset lengths, action segments and action tags of the action segments in the first subsequence set are identified in step 104 to determine a first recognition result of the frame sequence to be identified, including:

[0050] Step 201 : Recognize the first subsequences in the first subsequence set respectively by using a preset action recognition model to determine the recognition result of the first subsequence.

[0051] The recognition result of the first subsequence includes the action label, the action start frame image, the action end frame image, and the confidence of the action label matched by the first subsequence.

[0052] Specifically, step 201, that is, identifying the first subsequences in the first subsequence set respectively through the preset action recognition model to determine the recognition result of the first subsequence, specifically includes: inputting the first target subsequence into the preset action recognition model to obtain the recognition result of the first target subsequence.

[0053] The recognition result of the first target subsequence includes the first action label matched by the first target subsequence, the first action start frame image, the first action end frame image, and the confidence of the first action label.

[0054] In this embodiment, each randomly divided first target subsequence is identified by a preset action recognition model, and possible actions in each first target subsequence are determined, as well as the starting position and the ending position of the action, thereby achieving action positioning and action classification.

[0055] It should be noted that the preset action recognition model is pre-trained based on the typical actions that need to be monitored in the substation room, and can predict the typical actions in the input video clip, locate the starting position and the ending position of the action in the input video clip, and output the probability of the action, that is, the confidence. The higher the confidence, the greater the probability of the action in the input video clip, and the more accurate the output recognition result. For example, in the behavior recognition scenario in the substation room, it is necessary to monitor 11 typical actions of the staff, namely: making a phone call, opening the electric cabinet door, closing the electric cabinet door, crossing the isolation belt, bypassing the isolation belt, standing up, squatting, running, walking, crossing obstacles, and talking with two people. After the preset action recognition model is trained based on these 11 actions, each typical action is corresponding to a classification head, and each classification head will output the confidence when the classification head corresponds to the typical action. The preset action recognition model determines the action label of the input video clip based on the typical action corresponding to the highest confidence among the 11 confidences, so as to mark which typical action exists in the input video clip.

[0056] Exemplarily, the preset action recognition model predicts a typical action (first action label) present in the first target subsequence, and determines the probability of the presence of the typical action in the first target subsequence (confidence of the first action label), and determines the starting position (first action starting frame image) and the ending position (first action ending frame image) of the typical action in the first target subsequence, and determines the first action starting frame number and the first action ending frame number according to the order of the first action starting frame image and the first action ending frame image in the frame sequence to be identified.

[0057] Step 202: According to the recognition result of the first subsequence, remove the first subsequence that matches the same action label to obtain the target subsequence.

[0058] In this embodiment, by removing the first subsequence with the same action label, redundant information can be reduced, making the target subsequence more concise and clear, which helps to improve the accuracy of action recognition. In addition, reducing the first subsequence matched to the same action label can reduce the amount of calculation during system operation and improve system performance and efficiency.

[0059] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, step 202, that is, according to the recognition result of the first subsequence, removing the first subsequence matched to the same action label to obtain the target subsequence, specifically includes: determining multiple first target subsequences matched to the same first action label as a screening set; sorting the multiple first target subsequences in the screening set according to the confidence level to determine the sequence order; determining the first target subsequence located before a preset position in the sequence order as a primary subsequence; determining the repetition degree between the comparison subsequence and the primary subsequence according to the sequence length; wherein the comparison subsequence is the first target subsequence other than the primary subsequence in the screening set; removing the comparison subsequences in the screening set whose repetition degree is greater than or equal to the preset repetition degree threshold to obtain the target subsequence.

[0060] In this example, the redundant and too close first target subsequences are removed according to the recognition result of the first target subsequence by using the non-maximum suppression method, thereby improving the accuracy and uniqueness of the recognition result. In addition, the primary subsequence with a higher confidence is used as the deduplication target, and the repetition between the comparison subsequence and the primary subsequence is referred to for targeted deduplication processing, thereby retaining as much original data as possible while reducing important and repeated data, which helps to reduce the error rate and facilitate accurate recognition results.

[0061] Exemplarily, first, determine the total set S containing all first target subsequence recognition results. Where S = [s1, s2, ..., s N ],s i (i=1,...,N) is the recognition result of any first target subsequence, and N is the total number of first target subsequences (first preset number). Next, multiple first target subsequences matching the same first action label are extracted from the total set S as the screening set For subsequent comparison and screening. s j (j=1,...,t) is any first target subsequence matched to the same first action label, and t is the number of first target subsequences matched to the same first action label.

[0062] It is understandable that in the total set S, there are multiple screening sets Each filter set Include multiple first target subsequences that match the same first action tag. For example, if the first action tags matched by the first target subsequence include "make a phone call", "open the electric cabinet door", "close the electric cabinet door", "cross the isolation belt", "bypass the isolation belt", "stand up", and "squat", then multiple first target subsequences that match the "make a phone call" action tag are used as a screening set. Similarly, multiple first target subsequences matching the action label of "open the electric cabinet door" are used as another screening set With this type of heap, we can get seven screening sets

[0063] Furthermore, for any filter set Filter collections based on confidence The first target subsequences in the screening set are sorted to determine the sequence order. Then, the first target subsequence located before the preset position in the sequence order is determined as the primary subsequence to determine the most representative and reliable first target subsequence with the same first action label. The first target subsequences other than the primary subsequence in the screening set are determined as the comparison subsequences.

[0064] The preset position is determined according to the confidence of the first action label obtained. For example, if the first few confidence values ​​in the sequence order are high and close, it means that the recognition results corresponding to these confidences are relatively accurate. Therefore, the first target subsequences corresponding to these confidences are determined as the primary subsequences, and the comparison subsequences are screened according to each primary subsequence. If the first confidence in the sequence order is much higher than other confidences, it means that the first confidence is the most reliable. Therefore, the first target subsequence corresponding to the first confidence is determined as the primary subsequence.

[0065] Specifically, first, determine the length of the repetition between the comparison subsequence and the primary subsequence, and determine the repetition degree of the comparison subsequence based on the ratio of the repetition length to the length of the primary subsequence. If the repetition degree of the comparison subsequence is greater than or equal to the preset repetition degree threshold, it means that the comparison subsequence and the primary subsequence overlap too much, and the recognition result of the comparison subsequence is not as accurate as the recognition result of the primary subsequence. Therefore, the comparison subsequence with a repetition degree greater than or equal to the preset repetition degree threshold is removed from the screening set. The first target subsequences are removed from the frame sequence to be identified, thereby reducing repeated recognition results and improving the accuracy and simplicity of the final output, so that the remaining first target subsequences can accurately reflect the action conditions of different regions in the frame sequence to be identified.

[0066] The preset repetition threshold is determined according to the recognition accuracy required in the actual application scenario. For example, the preset repetition is set to 50%, 70%, etc. to ensure a higher recognition accuracy.

[0067] Further, according to the removed filter set The first target subsequence in determines the target subsequence.

[0068] Step 203: determining the first action segment in the frame sequence to be identified according to the action start frame image and the action end frame image of the target subsequence.

[0069] In this embodiment, the first action segment with a typical action in the frame sequence to be identified is extracted through the action start frame image and the action end frame image, so as to accurately locate the range of the action and reduce interference factors.

[0070] Further, as a refinement and extension of the specific implementation method of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, step 203, that is, determining the first action segment in the frame sequence to be identified based on the action start frame image and the action end frame image of the target subsequence, specifically includes: determining the first action segment in the frame sequence to be identified based on the first action start frame image and the first action end frame image of the target subsequence.

[0071] Step 204: taking the action label corresponding to the first action segment and the confidence level of the action label as the first recognition result of the frame sequence to be recognized.

[0072] Further, as a refinement and extension of the specific implementation method of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, step 204, that is, taking the action label corresponding to the first action clip and the confidence of the action label as the first recognition result of the frame sequence to be identified, specifically includes: taking the first action label corresponding to the first action clip and the confidence of the first action label as the first recognition result of the frame sequence to be identified.

[0073] In this embodiment, the recognition results of the target subsequence are mapped to the frame sequence to be recognized, and the start and end of each recognized action are determined, so as to extract the fragments where typical actions occur in the frame sequence to be recognized, and at the same time determine the category of each action, thereby obtaining the first recognition result of the frame sequence to be recognized.

[0074] Exemplarily, according to the first action start frame image number and the first action end frame image number of the target subsequence obtained in step 201, the starting position and the ending position of the typical action corresponding to the target subsequence are determined in the frame sequence to be identified, and according to the area between the starting position and the ending position, the first action segment corresponding to the target subsequence is determined, and then the first action label matching the target subsequence is drawn on each frame image to be identified in the first action segment, and the confidence of the first action label is determined at the same time, and the frame sequence to be identified with the first action label is output as the first recognition result to intuitively display the action information identified in the frame sequence to be identified.

[0075] For example, through the recognition results of the target subsequence, three first action segments (first recognition results) are determined in the frame sequence to be recognized, which are: the 5th to 32nd frames (the first first action segment) are "making a phone call" action (first action label), and the probability or accuracy of the 5th to 32nd frames being the "making a phone call" action is 98% (confidence of the first action label); the 40th to 60th frames (the second first action segment) are "squatting" action (first action label), and the probability or accuracy of the 40th to 60th frames being the "squatting" action is 95% (confidence of the first action label); the 75th to 96th frames (the third first action segment) are "running" action (first action label), and the probability or accuracy of the 75th to 96th frames being the "running" action is 90% (confidence of the first action label).

[0076] In one embodiment, the online behavior recognition method of typical actions indoors in a substation of this embodiment also includes: if the same frame image to be identified in the frame sequence to be identified belongs to multiple first action fragments, the multiple first action fragments to which the frame image to be identified belongs are sorted according to the confidence level to obtain a fragment order; and the same frame image to be identified is deleted from other action fragments.

[0077] Among them, the other action clips are the first action clips except the first first action clip in the clip sequence.

[0078] In this embodiment, if the same frame image to be identified in the frame sequence to be identified belongs to multiple first action segments, it means that there is action overlap in the first action segments. Therefore, the first action segment to which the frame image to be identified most likely belongs is selected according to the confidence level, thereby reducing the repeated frame images to be identified in the first action segments, making the identification of the frame sequence to be identified more reliable.

[0079] Exemplarily, if the same frame image to be identified in the frame sequence to be identified belongs to multiple first action segments, that is, the same frame image to be identified has different first action labels, then the first action label with the highest confidence is selected to be drawn on the frame image to be identified, so that the frame image to be identified belongs to the first action segment with a more accurate recognition result.

[0080] It is worth mentioning that the more the number of first subsequences randomly divided into the frame sequence to be identified, the more accurate the positioning and classification of the action in the first recognition result obtained, but the increase in the number of first subsequences will increase the recognition time of the frame sequence to be identified. As shown in Table 1, as the number of first subsequences increases, the recognition time of the frame sequence to be identified also increases. Therefore, if you want to achieve online real-time output of accurate recognition results of the frame sequence to be identified, you need to balance the relationship between the accuracy of the recognition result and the recognition time by setting the number of first subsequences (first preset number), so that the recognition time is close to the total time corresponding to the frame sequence to be identified while ensuring accurate recognition. For example, as shown in Table 1, if the frame sequence to be identified is 20 seconds, you can choose to randomly divide the frame sequence to be identified into 50 first subsequences. At this time, the recognition time required is about 35 seconds, which is equivalent to a small amount of time exceeding the original video to obtain accurate recognition results.

[0081] Table 1

[0082] First subsequence quantity (first preset quantity) Identification time / s 1 (and the length of the first subsequence is fixed) 15.0361 50 35.5656 100 56.8242 300 156.6328

[0083] Further, as a refinement and extension of the specific implementation methods of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, in step 103, the frame sequence to be identified in which the posture information is detected is divided into a first subsequence set according to a random frame division strategy, specifically including: determining a reference frame image in the frame sequence to be identified in which the posture information is detected according to a preset interval length; based on the reference frame image, randomly generating a second preset number of second target subsequences according to a plurality of different second preset lengths of the frame sequence to be identified in which the posture information is detected, so that each second target subsequence can include the reference frame image.

[0084] In this embodiment, based on the YOLO (You Only Look Once) target detection method, multiple reference frame images are evenly selected from a frame sequence to be identified of a certain length, and multiple second target subsequences are randomly generated according to the reference frame images to ensure that the generated second target subsequences can cover all areas of the frame sequence to be identified, thereby avoiding the problem of missing actions when a small number of first target subsequences are randomly generated, and increasing the diversity and accuracy of behavior recognition.

[0085] Exemplarily, the number of regions that need to be evenly divided into the frame sequence to be identified is obtained, and the preset interval length is determined according to the length of the frame sequence to be identified and the number of regions. The frame sequence to be identified in which the posture information is detected is evenly divided into multiple regions according to the preset interval length, and a reference frame image is determined in each region. Then, a second length list including multiple different second preset lengths is obtained, and a second preset number of second target subsequences are randomly generated according to the second length list with the reference frame image as the base point, so that each second target subsequence can include the reference frame image.

[0086] For example, if the length of the frame sequence to be identified is 120 frames, if it is evenly divided into 6 sections, the preset interval length is 20 frames, and the multiple areas after even division are: Frame 1-Frame 20, Frame 21-Frame 40, Frame 41-Frame 60, Frame 61-Frame 80, Frame 81-Frame 100, Frame 101-Frame 120. Further, the five frame images to be identified, namely, Frame 20, Frame 40, Frame 60, Frame 80, and Frame 100, are determined as reference frame images. Taking the 20th frame image to be identified as the reference frame image as an example, each second target subsequence randomly generated based on the 20th frame image to be identified must include the 20th frame image to be identified.

[0087] It should be noted that the setting principle of the second preset length and the second preset number is the same as the setting principle of the first preset length and the first preset number in the above steps. That is, the second preset length is also determined according to the longest action and the shortest action in the actual application scenario, and by setting the second target subsequence number (the second preset number), the relationship between the accuracy of the recognition result and the recognition time can also be balanced. For example, by adjusting the number of the second target subsequences, a 20-second frame sequence to be recognized is divided into multiple second target subsequences for recognition, and the time to obtain the first recognition result is about 30 seconds.

[0088] In one embodiment, when a reference frame image is determined in a frame sequence to be identified in which gesture information is detected according to a preset interval length; and when a second preset number of second target subsequences are randomly generated according to the reference frame image from the frame sequence to be identified in which gesture information is detected according to a plurality of different second preset lengths, in step 104, action segments and action tags of the action segments in the first subsequence set are identified to determine a first recognition result of the frame sequence to be identified, the step includes:

[0089] The first subsequences in the first subsequence set are respectively identified by using a preset action recognition model to determine a recognition result of the first subsequence.

[0090] The recognition result of the first subsequence includes the action label, the action start frame image, the action end frame image, and the confidence of the action label matched by the first subsequence.

[0091] Furthermore, as a refinement and extension of the specific implementation methods of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, the first subsequence in the first subsequence set is recognized respectively by a preset action recognition model to determine the recognition result of the first subsequence, specifically including: inputting the second target subsequence into the preset action recognition model to obtain the recognition result of the second target subsequence.

[0092] The recognition result of the second target subsequence includes the second action label matched by the second target subsequence, the second action start frame image, the second action end frame image, and the confidence of the second action label.

[0093] In this embodiment, similar to the above step 201, each second target subsequence is identified by a preset action recognition model, typical actions that may exist in each second target subsequence are determined, and the starting position and the ending position of the action are determined, thereby achieving the positioning and classification of the action.

[0094] Exemplarily, the preset action recognition model predicts the typical action (second action label) present in the second target subsequence, as well as the probability of the presence of the typical action in the second target subsequence (confidence of the second action label), and determines the starting position (second action starting frame image) and ending position (second action ending frame image) of the typical action in the second target subsequence, and determines the second action starting frame number and the second action ending frame number based on the order of the second action starting frame image and the second action ending frame image in the frame sequence to be identified.

[0095] Further, after obtaining the recognition results of each second target subsequence, similarly to step 202, based on the recognition results of the second target subsequence, the second target subsequences that match the same second action label are removed to obtain the target subsequence. Then, similarly to step 203, the first action segment in the frame sequence to be identified is determined based on the second action start frame image and the second action end frame image of the target subsequence. And, similarly to step 204, the second action label corresponding to the first action segment and the confidence of the second action label are used as the first recognition result of the frame sequence to be identified, that is, through the target subsequence, multiple first action segments can be determined in sequence in the frame sequence to be identified, and the second action label corresponding to the first action segment and the confidence of the second action label are obtained at the same time.

[0096] Furthermore, as a refinement and extension of the specific implementation methods of the above-mentioned embodiment, in order to fully illustrate the specific implementation process of this embodiment, in step 103, the frame sequence to be identified in which posture information is detected is divided into a second subsequence set according to the continuous frame division strategy, specifically including: the frame sequence to be identified in which posture information is detected is divided into multiple second subsequences in sequence according to the third preset length.

[0097] In one embodiment, when the frame sequence to be identified in which the posture information is detected is divided into a plurality of second subsequences in sequence according to a third preset length, the action segments and the action tags of the action segments in the second subsequence set are identified in step 104 to determine the second recognition result of the frame sequence to be identified, specifically including: sorting the plurality of second subsequences according to the time when the second subsequences are generated; determining the target recognition sequence according to the first second subsequence of the generated timing; identifying the target recognition sequence through a preset action recognition model to determine the recognition result of the target recognition sequence; wherein the recognition result of the target recognition sequence includes the third action tag matched by the target recognition sequence, the third action end frame image , the confidence of the third action label; if the confidence of the third action label is greater than or equal to the first preset threshold and less than the second preset threshold, the target recognition sequence is superimposed according to the second subsequence adjacent to the target recognition sequence until the confidence of the third action label of the superimposed target recognition sequence is greater than or equal to the second preset threshold; if the confidence of the third action label is greater than or equal to the second preset threshold, the target recognition sequence is determined as a target frame sequence; according to the target frame sequence and the third action end frame image, the second action segment in the frame sequence to be recognized is determined; the third action label corresponding to the second action segment and the confidence of the third action label are used as the second recognition result of the frame sequence to be recognized.

[0098] In this embodiment, in combination with the situation in actual application scenarios where the occurrence of an action cannot be predicted in most cases and only when the action ends can be determined, unlike the first action segment which simultaneously determines the exact starting position and the exact ending position of the action, this embodiment makes a judgment based on the confidence level, gradually increases the number of frames for recognition, and integrates the information of adjacent frame sequences through superposition processing until the boundary of the action is clearly identified, thereby determining the exact ending position of the action after the action occurs, thereby improving the comprehensiveness of behavior recognition and making the recognition result more reliable.

[0099] Specifically, based on the judgment that the higher the action positioning accuracy of the frame sequence is, the higher the confidence level corresponding to the frame sequence is, the present embodiment gradually increases the number of frames for recognition until the confidence level reaches a second preset threshold value, and an accurate action end position is obtained. Similarly, multiple accurate action end positions are sequentially determined in the frame sequence to be recognized, and then multiple second action segments are sequentially determined in the frame sequence to be recognized according to the action end position, so as to achieve accurate positioning of the action end position.

[0100] Exemplarily, based on the concept of gradually superimposing a certain number of frame images with gesture information to be identified from the frame sequence to be identified in which gesture information is detected, the frame sequence to be identified in which gesture information is detected is divided into a plurality of second subsequences in sequence according to a third preset length, so that the second subsequences are gradually superimposed in the generation sequence for identification, so as to improve the efficiency of identification. The plurality of second subsequences are sorted according to the time when the second subsequences are generated, and the generation sequence of the second subsequences is determined.

[0101] The third preset length is based on the number of video frames superimposed each time in the actual application scenario. add Sure.

[0102] Furthermore, if there is a second subsequence that has not been input into the preset action recognition model for recognition, that is, the recognition of the frame sequence to be recognized has not been completed, the following steps in this embodiment are continued.

[0103] First, initialize a variable F to record the current cumulative number of recognition frames (the length of the target recognition sequence) temp is 0, that is, F temp =0, and then enters an inner loop of gradually superimposing the second subsequence according to the generation timing for identification.

[0104] Before each superposition of the second subsequence, it is necessary to compare the current cumulative number of recognition frames (the length of the target recognition sequence) F temp The preset maximum cache frame number F max The size of is used to control the length of the sequence input to the preset action recognition model each time to not exceed the set maximum value, thereby ensuring the recognition efficiency of the preset action recognition model and avoiding the recognition time of the preset action recognition model being too long.

[0105] If the current cumulative recognition frame number (the length of the target recognition sequence) F temp Less than the preset maximum cache frame number F max , then obtain the second subsequence according to the generation timing to superimpose the target recognition sequence and re-determine the target recognition sequence.

[0106] It should be noted that when the first recognition is performed, the target recognition sequence has not been determined. Therefore, the second subsequence is obtained according to the generation sequence to superimpose the target recognition sequence, and the target recognition sequence step is re-determined, which is manifested in taking the first second subsequence of the generation sequence as the target recognition sequence.

[0107] Furthermore, the target recognition sequence is input into a preset action recognition model for recognition, and the recognition result of the target recognition sequence is determined, wherein the recognition result of the target recognition sequence includes the third action label matched by the target recognition sequence, the third action end frame image, and the confidence of the third action label.

[0108] If the confidence of the third action label is less than the first preset threshold, it means that there is no typical action in the target recognition sequence and it is necessary to continue the recognition, that is, gradually superimpose the second subsequence in the target recognition sequence according to the generation sequence until the confidence corresponding to the superimposed target recognition sequence is greater than or equal to the second preset threshold.

[0109] It should be noted that if there is no typical action in the target recognition sequence, after the target recognition sequence is input into the preset action recognition model, the recognition result will still be output, but the confidence of the third action label in the output recognition result will be relatively low. In combination with actual application scenarios, not all specific actions that need to be detected exist at all times. Therefore, in the step of gradually superimposing the second subsequence in the target recognition sequence according to the generation sequence, if the confidence corresponding to the superimposed target recognition sequence is always less than the first preset threshold, it means that there is no action that needs to be detected in this part of the frame sequence to be recognized. Therefore, in the length F of the target recognition sequence, the confidence of the third action label in the output recognition result is relatively low. temp Greater than the preset maximum cache frame number F max When the length of the target recognition sequence F temp It is reset to 0, the cache is released, and the target recognition sequence is obtained by re-accumulating and superimposing in the unrecognized area of ​​the frame sequence to be recognized, so as to avoid wasting computing resources and improve recognition efficiency.

[0110] If the confidence of the third action label is greater than or equal to the first preset threshold and less than the second preset threshold, it means that there is a typical action in the target recognition sequence, but the typical action has not ended. At this time, in each frame image to be recognized included in the second subsequence of the target recognition sequence that is most recently superimposed, the third action label matching the target recognition sequence is drawn to indicate that there is an action corresponding to the third action label in the most recently superimposed second subsequence.

[0111] After the third label is drawn, the next inner loop is performed, that is, the current cumulative recognition frame number (the length of the target recognition sequence) is compared. temp The preset maximum cache frame number F max If the current cumulative recognition frame number (the length of the target recognition sequence) is F temp Less than the preset maximum cache frame number F max , then the second subsequence is obtained according to the generation time sequence to superimpose the target recognition sequence, and the target recognition sequence is re-determined until the confidence of the third action label of the superimposed target recognition sequence is greater than or equal to the second preset threshold, and the typical action is determined to be over. And the target frame sequence is determined according to the target recognition sequence, wherein the target frame sequence is a sequence with the third action label drawn in the target recognition sequence.

[0112] If the confidence of the third action label is greater than or equal to the second preset threshold, it means that the typical action in the target recognition sequence has ended. Then, the exact position where the action ends is determined according to the third action end frame image of the target recognition sequence, and the third action end frame number is determined according to the order of the third action end frame image in the frame sequence to be recognized. In this way, the second action fragment is determined in the frame sequence to be recognized according to the target frame sequence and the third action end frame image.

[0113] The first preset threshold and the second preset threshold are determined according to factors such as the preset action recognition model used and the recognition accuracy requirement, etc. For example, the first preset threshold is set to 0.6, and the second preset threshold is set to 0.7 or higher.

[0114] It can be understood that after the first second action segment is determined according to the above steps, the same as the above principle, starting from the second subsequence of the first generation sequence in the second subsequence that has not been recognized (not input into the preset action recognition model), the second subsequence is gradually superimposed according to the generation sequence for recognition until the exact position of the end of the second action (the second third action end frame image) is recognized, thereby determining the second second action segment. And so on, until all second subsequences are recognized, that is, the recognition of the frame sequence to be recognized is completed, and the exact positions of the end of multiple actions and multiple second action segments can be determined in the frame sequence to be recognized, and the third action label corresponding to the second action segment and the confidence of the third action label are used as the second recognition result of the frame sequence to be recognized.

[0115] For example, in the case of the first recognition, that is, when the first second subsequence of the generated sequence is used as the target recognition sequence among all the second subsequences, if the confidence corresponding to the target recognition sequence is greater than or equal to the first preset threshold and less than the second preset threshold, it means that there is a typical action in the first second subsequence of the generated sequence, but the action has not ended. Therefore, a third action label matching the target recognition sequence is drawn on each frame image to be recognized included in the first second subsequence of the generated sequence to indicate that the action exists in this area of ​​the frame sequence to be recognized.

[0116] In the case of the first recognition, if the confidence of the third action label corresponding to the target recognition sequence is less than the first preset threshold, it means that there is no typical action in the first second subsequence of the generated sequence, so it is necessary to superimpose the second second subsequence of the generated sequence to the first second subsequence of the generated sequence, and re-determine the first two second subsequences of the generated sequence as the target recognition sequence, and input the target recognition sequence into the preset action recognition model for recognition. Then, the relationship between the confidence corresponding to the target recognition sequence and the first preset threshold and the second preset threshold is determined to determine whether it is necessary to continue to superimpose the second subsequence.

[0117] If the confidence corresponding to the target recognition sequence is greater than or equal to the first preset threshold value and less than the second preset threshold value when the first two second subsequences of the generation sequence are used as the target recognition sequence, it means that there is a typical action in the latest superimposed second subsequence, so the third action label matched by the target recognition sequence is drawn in each frame image to be recognized included in the second second subsequence of the generation sequence to indicate that the action exists in this area of ​​the frame sequence to be recognized. However, since the confidence corresponding to the target recognition sequence is less than the second preset threshold value, it means that the action has not ended, and the second subsequence needs to be continued to be superimposed on the target recognition sequence according to the generation sequence until the confidence corresponding to the superimposed target recognition sequence is greater than or equal to the second preset threshold value, and it is determined that the action is ended.

[0118] For example, if the third second subsequence of the generated sequence is superimposed on the target recognition sequence, that is, the first three second subsequences of the generated sequence are used as the target recognition sequence, and the confidence of the third label corresponding to the target recognition sequence is greater than or equal to the second preset threshold, it means that the action is ended, and then the exact end position of the action is determined in the target recognition sequence (the third action end frame image), so as to determine the first second action segment in the frame sequence to be recognized according to the target frame sequence (the frame image to be recognized with the third action label drawn in the target recognition sequence) and the third action end frame image, and at the same time determine the confidence of the third action label corresponding to the second action segment.

[0119] After obtaining the first second action segment, the same principle as above is used, starting from the fourth second subsequence of the generated time sequence, the second subsequences are gradually superimposed for recognition until the second second action segment is determined. Similarly, multiple second action segments are sequentially determined in the frame sequence to be recognized, and the third action label and the confidence of the third action label corresponding to each second action segment are determined, thereby obtaining the second recognition result of the frame sequence to be recognized.

[0120] It is worth mentioning that in this embodiment, the three recognition processes of randomly dividing the frame sequence to be recognized into multiple first target subsequences, and recognizing the first target subsequences to determine the first recognition result, dividing the frame sequence to be recognized into multiple second target subsequences according to the reference frame image, and recognizing the second target subsequences to determine the first recognition result, and dividing the frame sequence to be recognized into multiple second subsequences in turn, and recognizing the second recognition result according to the superimposed second subsequences are also parallel, that is, they are carried out simultaneously, independently, and do not interfere with each other. This embodiment uses three parallel and different recognition processes to perform differentiated recognition on the frame sequence to be recognized, so that the advantages of different recognition processes complement each other, while improving the efficiency of behavior recognition, and increasing the comprehensiveness of the first recognition result obtained under the random division strategy.

[0121] like Figure 3As shown, in one embodiment, step 105, i.e., determining a plurality of target action segments and action categories corresponding to each target action segment in the frame sequence to be identified according to the first recognition result and the second recognition result, specifically includes the following steps:

[0122] Step 301: Acquire a plurality of different approximate segment conditions.

[0123] Step 302: Compare the first action segment in the first recognition result and the second action segment in the second recognition result with the approximate segment conditions respectively, and take the first action segment and the second action segment that meet the same approximate segment conditions as an approximate segment set.

[0124] In this embodiment, the first recognition result and the second recognition result are integrated to determine the first action segment and the second action segment that belong to the same area in the frame sequence to be identified, so that in the subsequent steps, the first action segment and the second action segment that belong to the same area are mutually verified to determine a more reliable recognition result.

[0125] Specifically, in the above three parallel and different recognition processes, the randomly generated first target subsequence can capture the typical actions existing in different areas of the frame sequence to be identified. Then, the first target subsequence is further selected according to the recognition result of the first target subsequence to select the first action segment that best matches the real result. The second target subsequence is randomly generated by the uniformly selected reference frame image, which can ensure that the typical actions existing at each time point will be captured. The first action segment is obtained according to the second target subsequence, which can make up for the problem of missing actions when the number of randomly generated first target subsequences is relatively small, and more evenly cover the various areas of the frame sequence to be identified, and then determine the accurate starting position and accurate ending position of each typical action in the frame sequence to be identified according to the first action segment. At the same time, the frame sequence to be identified is gradually identified according to the superimposed second subsequence, and the starting position of the action is not paid attention to, but the confidence is used as the judgment basis to determine the accurate ending position of each typical action in the frame sequence to be identified, thereby obtaining the second action segment. It is precisely because this embodiment uses three different recognition processes in parallel that it is able to perceive different action classifications in the same video clip, thereby determining whether the recognized actions are consistent through comparison and determining the final recognition result, thereby reducing the error of behavior recognition without the need for a lot of training and ensuring the accuracy of behavior recognition.

[0126] Exemplarily, multiple different approximate fragment conditions are obtained, each approximate fragment condition corresponds to a different area of ​​the frame sequence to be identified, and the first action fragment in the first recognition result and the second action fragment in the second recognition result are compared with the approximate fragment conditions respectively. The first action fragment and the second action fragment that meet the same approximate fragment conditions are taken as an approximate fragment set, thereby determining the first action fragment and the second action fragment corresponding to the same area of ​​the frame sequence to be identified.

[0127] It can be understood that there are multiple approximate segment sets, and each approximate segment set corresponds to a different area of ​​the frame sequence to be identified.

[0128] Step 303, if the first action label corresponding to the first action segment, the second action label and the third action label corresponding to the second action segment in the approximate segment set are the same, then determine the target action segment based on the first action segment and the second action segment, and determine the action category corresponding to the target action segment based on the first action label or the second action label or the third action label.

[0129] In this embodiment, since the first recognition result and the second recognition result of the frame sequence to be recognized are based on the same preset action recognition model, the action labels corresponding to the frame sequences in the same area of ​​the frame sequence to be recognized in the first recognition result and the second recognition result should be consistent. However, due to different strategies for dividing subsequences, the lengths of the recognized action segments will be inconsistent. This embodiment verifies the first recognition result and the second recognition result to accurately locate the segments where the action occurs in the frame sequence to be recognized, thereby improving the reliability and stability of behavior recognition.

[0130] Exemplarily, the approximate fragment set is represented as:

[0131]

[0132] Among them, label k1 is the first action label corresponding to the first action clip, is the starting frame number of the first action corresponding to the first action segment, label is the frame number of the first action ending corresponding to the first action segment; k2 is the second action label corresponding to the first action segment, is the starting frame number of the second action corresponding to the first action segment, label is the end frame number of the second action corresponding to the first action segment; k3 is the third action label corresponding to the second action fragment, is the starting frame number of the third action corresponding to the second action segment, It is the end frame number of the third action corresponding to the second action segment.

[0133] It should be noted that when obtaining multiple second action segments of the frame sequence to be identified according to the second subsequence set, only the termination of the action is concerned, and the exact starting position of the action is not output. For the convenience of representation, the starting frame number of the third action corresponding to the second action segment is The first frame image to be identified in the time sequence of the second action segment is determined in the sequence of frames to be identified, but the starting frame number of the third action is Does not participate in subsequent calculations.

[0134] Furthermore, if the first action label label in the approximate segment set k1 , second action label label k2 And the third action label label k3 If the sequence number of the first action start frame in the approximate segment set is consistent, it means that a reliable recognition result has been obtained. The second action start frame number The target action start frame number of the target action segment is determined by the average value of The second action ends the frame number The frame number of the third action end The target action segment is determined by the average value of the target action segment. Thus, the target action segment is determined in the frame sequence to be identified according to the target action start frame number and the target action end frame number, and the target action segment is determined according to the first action label label k1 Or the second action label label k2 Or the third action label label k3 Determine the action category corresponding to the target action segment to reduce the deviation of the recognition results caused by different subsequence division strategies and improve the positioning accuracy of the target action segment.

[0135] Among them, the target action fragment is expressed as:

[0136]

[0137] X out is the target action segment, is the action category corresponding to the target action clip, is the starting frame number of the target action, The frame number where the target action ends.

[0138] For example, the first action segment in the approximate segment set includes: the 6th frame (first action start frame sequence number) to the 30th frame (first action end frame sequence number) as the "door opening" action (first action label), and the 4th frame (second action start frame sequence number) to the 26th frame (second action end frame sequence number) as the "door opening" action (second action label), and the second action segment is the 8th frame (third action start frame sequence number) to the 40th frame (third action end frame sequence number) as the "door opening" action (third action label). At this time, the first action label, the second action label and the third action label are all the same, then according to the average value of the 6th frame and the 4th frame, the target action start frame sequence number - the 5th frame is determined, and according to the average value of the 30th frame, the 26th frame and the 40th frame, the target action end frame sequence number - the 32nd frame is determined, so that the 5th frame to the 32nd frame (target action segment) in the frame sequence to be identified is the "door opening" action (action category corresponding to the target action segment).

[0139] Step 304 , if any two of the first action label, the second action label and the third action label are the same, then a target action segment is determined according to the action segments corresponding to any two of them, and an action category corresponding to the target action segment is determined according to the action labels corresponding to any two of them.

[0140] Step 305, if the first action label, the second action label and the third action label are all different, the first action label, the second action label and the third action label are sorted according to the confidence level to determine the label order; the target action segment is determined according to the action segment corresponding to the first action label in the label order, and the action category of the target action segment is determined according to the first label in the label order.

[0141] In this example, if the action labels corresponding to the frame sequences in the same area of ​​the frame sequence to be identified in the first recognition result and the second recognition result are inconsistent, it means that an abnormality may have occurred during the recognition process, such as the first subsequence not covering the area, or identifying other actions, etc. This embodiment comprehensively compares the first recognition result and the second recognition result based on the confidence of the action label, so that this embodiment can still provide a relatively reliable recognition result under abnormal conditions.

[0142] Exemplarily, if any two of the first action label, the second action label and the third action label in the approximate segment set are the same, then the target action start frame number of the target action segment is determined according to the average value of the action start frame numbers of the action segments corresponding to the two identical action labels, and the target action end frame number of the target action segment is determined according to the average value of the action end frame numbers of the action segments corresponding to the two identical action labels, thereby determining the target action segment in the frame sequence to be identified according to the target action start frame number and the target action end frame number, and determining the action category corresponding to the target action segment according to the two identical action labels.

[0143] For example, the first action segment in the approximate segment set includes: the 6th frame (first action start frame sequence number) to the 30th frame (first action end frame sequence number) as the "door opening" action (first action label), and the 4th frame (second action start frame sequence number) to the 26th frame (second action end frame sequence number) as the "door opening" action (second action label), and the second action segment is the 8th frame (third action start frame sequence number) to the 40th frame (third action end frame sequence number) as the "door closing" action (third action label). At this time, the first action label is the same as the second action label, so the target action start frame sequence number - the 5th frame is determined according to the average value of the 6th frame and the 4th frame, and the target action end frame sequence number - the 28th frame is determined according to the average value of the 30th frame and the 26th frame, so that the 5th frame to the 28th frame in the frame sequence to be identified is determined as the "door opening" action.

[0144] Furthermore, if the first action label, the second action label and the third action label in the approximate segment set are all different, the action segment corresponding to the action label with the highest confidence is selected as the target action segment, and the category of the target action segment is determined according to the action label with the highest confidence to improve the accuracy of behavior recognition.

[0145] For example, the first action segment in the approximate segment set includes: the 50th frame (first action start frame sequence number) to the 70th frame (first action end frame sequence number) is the "door opening" action (first action label), and the 45th frame (second action start frame sequence number) to the 80th frame (second action end frame sequence number) is the "door closing" action (second action label), and the second action segment is the 50th frame (third action start frame sequence number) to the 65th frame (third action end frame sequence number) is the "standing" action (third action label). At this time, the first action label, the second action label and the third action label are all different, then the action label with the highest confidence is selected. If the "standing" action (third action label) has the highest confidence, it is determined that the 50th frame to the 65th frame in the frame sequence to be identified is the "standing" action.

[0146] In one embodiment, a long video containing multiple actions is used to simulate the input of an online video stream, and the actions in the long video are annotated in advance. The effects of the random frame division strategy and the continuous frame division strategy alone and in parallel are compared, and the indicator is the difference between the recognition results and the annotation results of the start and end of the action. Since the continuous frame division strategy only focuses on the termination of the action, the judgment of the starting position of the target action is determined by the starting position corresponding to the random frame division strategy, and the judgment of the termination position of the target action is determined jointly by the starting positions corresponding to the random frame division strategy and the continuous frame division strategy. As shown in Table 2, Table 2 shows the results of identifying a video stream of 30 frames / s. The results show that the effect of the random frame division strategy and the continuous frame division strategy in parallel, integrated and mutually verified exceeds the effect of a single strategy.

[0147] Table 2

[0148]

[0149] It should be noted that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0150] Furthermore, if Figure 4 As shown, as a specific implementation of the above-mentioned online behavior recognition method for typical actions in the substation room, an embodiment of the present application provides an online behavior recognition device 400 for typical actions in the substation room, and the online behavior recognition device 400 for typical actions in the substation room includes: an acquisition module 401, a detection module 402, a division module 403, an identification module 404 and a determination module 405.

[0151] The acquisition module 401 is used to acquire a frame sequence to be identified; wherein the frame sequence to be identified includes a plurality of frame images to be identified arranged in time sequence;

[0152] A detection module 402, used for performing posture detection on the frame sequence to be identified;

[0153] A division module 403 is used to divide the frame sequence to be identified in which the posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy respectively;

[0154] An identification module 404 is used to identify the action segments and action tags of the action segments in the first subsequence set and the second subsequence set, respectively, and determine a first identification result and a second identification result of the frame sequence to be identified;

[0155] The determination module 405 is used to determine a plurality of target action segments and an action category corresponding to each target action segment in the frame sequence to be identified according to the first recognition result and the second recognition result.

[0156] In one embodiment, the recognition module 404 is specifically used to recognize the first subsequence in the first subsequence set respectively through a preset action recognition model to determine the recognition result of the first subsequence; wherein the recognition result of the first subsequence includes the action label, action start frame image, action end frame image, and confidence of the action label matched by the first subsequence; according to the recognition result of the first subsequence, the first subsequence matched with the same action label is removed to obtain the target subsequence; according to the action start frame image and action end frame image of the target subsequence, the first action segment in the frame sequence to be recognized is determined; and the action label corresponding to the first action segment and the confidence of the action label are used as the first recognition result of the frame sequence to be recognized.

[0157] In one embodiment, the division module 403 is specifically configured to randomly generate a first preset number of first target subsequences from a frame sequence to be identified in which posture information is detected according to a plurality of different first preset lengths.

[0158] In one embodiment, the recognition module 404 is specifically used to input the first target subsequence into a preset action recognition model to obtain a recognition result of the first target subsequence; wherein the recognition result of the first target subsequence includes the first action label matched by the first target subsequence, the first action start frame image, the first action end frame image, and the confidence of the first action label.

[0159] In one embodiment, the division module 403 is specifically used to determine a reference frame image in a frame sequence to be identified in which posture information is detected according to a preset interval length; based on the reference frame image, the frame sequence to be identified in which posture information is detected is randomly generated according to a plurality of different second preset lengths, and a second preset number of second target subsequences are generated, so that each second target subsequence can include the reference frame image.

[0160] In one embodiment, the recognition module 404 is specifically used to input the second target subsequence into a preset action recognition model to obtain a recognition result of the second target subsequence; wherein the recognition result of the second target subsequence includes a second action label matched by the second target subsequence, a second action start frame image, a second action end frame image, and a confidence level of the second action label.

[0161] In one embodiment, the identification module 404 is specifically used to determine multiple first subsequences that match the same action label as a screening set; sort the multiple first subsequences in the screening set according to the confidence level to determine the sequence order; determine the first subsequence located before a preset position in the sequence order as a primary subsequence; determine the repetition degree between the comparison subsequence and the primary subsequence according to the sequence length; wherein the comparison subsequence is the first subsequence in the screening set except the primary subsequence; remove the comparison subsequences in the screening set whose repetition degree is greater than or equal to the preset repetition threshold to obtain the target subsequence.

[0162] In one embodiment, the identification module 404 is specifically used to sort the multiple first action segments to which the frame image to be identified belongs according to the confidence level to obtain a segment order if the same frame image to be identified in the frame sequence to be identified belongs to multiple first action segments; and delete the same frame image to be identified from other action segments; wherein the other action segments are the first action segments other than the first first action segment in the segment order.

[0163] In one embodiment, the division module 403 is specifically configured to divide the frame sequence to be identified in which the posture information is detected into a plurality of second subsequences in sequence according to a third preset length.

[0164] In one embodiment, the recognition module 404 is specifically used to sort multiple second subsequences according to the time when the second subsequence is generated; determine the target recognition sequence according to the first second subsequence of the generation time sequence; recognize the target recognition sequence through a preset action recognition model to determine the recognition result of the target recognition sequence; wherein the recognition result of the target recognition sequence includes a third action label matched by the target recognition sequence, a third action end frame image, and a confidence of the third action label; if the confidence of the third action label is greater than or equal to the first preset threshold and less than the second preset threshold, the target recognition sequence is superimposed according to the second subsequence adjacent to the target recognition sequence until the confidence of the third action label of the superimposed target recognition sequence is greater than or equal to the second preset threshold; if the confidence of the third action label is greater than or equal to the second preset threshold, the target recognition sequence is determined as a target frame sequence; according to the target frame sequence and the third action end frame image, a second action segment in the frame sequence to be recognized is determined; and the third action label corresponding to the second action segment and the confidence of the third action label are used as the second recognition result of the frame sequence to be recognized.

[0165] For the specific definition of the online behavior recognition device for typical actions in the substation room, please refer to the definition of the online behavior recognition method for typical actions in the substation room above, which will not be repeated here. Each module in the above-mentioned online behavior recognition device for typical actions in the substation room can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0166] Based on the above Figures 1 to 3 The method shown, and Figure 4In order to achieve the above-mentioned purpose, the embodiment of the present application further provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figures 1 to 3 The online behavior recognition method for typical actions in the substation room is shown.

[0167] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0168] Those skilled in the art will appreciate that the computer device structure provided in this embodiment does not limit the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0169] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages and saves the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to realize communication between the components inside the storage medium, and communication with other hardware and software in the physical device.

[0170] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or the embodiments of the present application can be implemented by hardware.

[0171] Those skilled in the art will appreciate that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily necessary for implementing the present application. Those skilled in the art will appreciate that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the description of the implementation scenario, or can be changed accordingly and located in one or more devices different from the present implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple submodules.

[0172] The above serial numbers of this application are only for description and do not represent the advantages and disadvantages of the implementation scenarios. The above disclosure is only a few specific implementation scenarios of this application, but this application is not limited to them, and any changes that can be thought of by technicians in this field should fall within the scope of protection of this application.

Claims

1. An online behavior recognition method for typical actions in a substation room, characterized in that: The method comprises: Acquire a frame sequence to be identified; wherein the frame sequence to be identified includes a plurality of frame images to be identified arranged in time sequence; Performing posture detection on the frame sequence to be identified; Dividing the frame sequence to be identified in which the posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy respectively; Respectively identifying the action segments and the action tags of the action segments in the first subsequence set and the second subsequence set, and determining a first recognition result and a second recognition result of the frame sequence to be recognized; Determining, according to the first recognition result and the second recognition result, a plurality of target action segments in the to-be-recognized frame sequence, and an action category corresponding to each of the target action segments; The step of dividing the frame sequence to be identified by detecting the posture information into a second subsequence set according to the continuous frame division strategy includes: Dividing the frame sequence to be identified in which the posture information is detected into a plurality of second subsequences in sequence according to a third preset length; The step of identifying the action segments and the action tags of the action segments in the second subsequence set and determining a second recognition result of the frame sequence to be recognized includes: sorting the plurality of second subsequences according to the time when the second subsequences were generated; Determine a target identification sequence according to the first of the second subsequences of the generated time sequence; The target recognition sequence is recognized by a preset action recognition model to determine a recognition result of the target recognition sequence; wherein the recognition result of the target recognition sequence includes a third action label matched by the target recognition sequence, a third action end frame image, and a confidence level of the third action label; If the confidence of the third action label is greater than or equal to the first preset threshold and less than the second preset threshold, the target recognition sequence is superimposed according to the second subsequence adjacent to the target recognition sequence until the confidence of the third action label of the superimposed target recognition sequence is greater than or equal to the second preset threshold; If the confidence of the third action label is greater than or equal to the second preset threshold, determining the target recognition sequence as a target frame sequence; Determining a second action segment in the to-be-identified frame sequence according to the target frame sequence and the third action end frame image; using a third action label corresponding to the second action segment and a confidence level of the third action label as a second recognition result of the frame sequence to be recognized; The step of determining a plurality of target action segments and action categories corresponding to each of the target action segments in the to-be-identified frame sequence according to the first recognition result and the second recognition result includes: Obtain multiple different approximate fragment conditions; Compare the first action segment in the first recognition result and the second action segment in the second recognition result with the approximate segment condition respectively, and take the first action segment and the second action segment that meet the same approximate segment condition as an approximate segment set; If the first action label, the second action label and the third action label corresponding to the first action segment in the approximate segment set are the same, the target action segment is determined according to the first action segment and the second action segment, and the action category corresponding to the target action segment is determined according to the first action label or the second action label or the third action label; If any two of the first action label, the second action label and the third action label are the same, determining the target action segment according to the action segments corresponding to the any two action labels, and determining the action category corresponding to the target action segment according to the action labels corresponding to the any two action labels; If the first action label, the second action label and the third action label are all different, the first action label, the second action label and the third action label are sorted according to the confidence levels of the first action label, the second action label and the third action label to determine the label order; The target action segment is determined according to the action segment corresponding to the first action tag in the tag sequence, and the action category of the target action segment is determined according to the first tag in the tag sequence.

2. The online behavior recognition method for typical actions in a substation room according to claim 1 is characterized in that: The step of identifying the action segments and the action tags of the action segments in the first subsequence set and determining a first recognition result of the frame sequence to be identified includes: Recognize the first subsequences in the first subsequence set respectively by using a preset action recognition model, and determine the recognition result of the first subsequence; wherein the recognition result of the first subsequence includes the action label matched by the first subsequence, the action start frame image, the action end frame image, and the confidence of the action label; According to the recognition result of the first subsequence, remove the first subsequence matched with the same action label to obtain a target subsequence; Determining a first action segment in the frame sequence to be identified according to the action start frame image and the action end frame image of the target subsequence; The action label corresponding to the first action segment and the confidence of the action label are used as a first recognition result of the frame sequence to be recognized.

3. The online behavior recognition method for typical actions in a substation room according to claim 2 is characterized in that: The step of dividing the frame sequence to be identified by detecting the posture information into a first subsequence set according to a random frame division strategy includes: Randomly generate a first preset number of first target subsequences from the frame sequence to be identified in which the posture information is detected according to a plurality of different first preset lengths; The identifying the first subsequences in the first subsequence set respectively by using a preset action recognition model to determine the recognition result of the first subsequence includes: The first target subsequence is input into the preset action recognition model to obtain a recognition result of the first target subsequence; wherein the recognition result of the first target subsequence includes a first action label matched by the first target subsequence, a first action start frame image, a first action end frame image, and a confidence level of the first action label.

4. The online behavior recognition method for typical actions in a substation room according to claim 2 is characterized in that: The step of dividing the frame sequence to be identified by detecting the posture information into a first subsequence set according to a random frame division strategy includes: Determining a reference frame image in the frame sequence to be identified in which the posture information is detected according to a preset interval length; According to the reference frame image, randomly generate a second preset number of second target subsequences from the frame sequence to be identified in which posture information is detected according to a plurality of different second preset lengths, so that each of the second target subsequences can include the reference frame image; The identifying the first subsequences in the first subsequence set respectively by using a preset action recognition model to determine the recognition result of the first subsequence includes: The second target subsequence is input into the preset action recognition model to obtain a recognition result of the second target subsequence; wherein the recognition result of the second target subsequence includes a second action label matched by the second target subsequence, a second action start frame image, a second action end frame image, and a confidence level of the second action label.

5. The online behavior recognition method for typical actions in a substation room according to claim 2 is characterized in that: The removing the first subsequences matched to the same action label according to the recognition result of the first subsequence to obtain the target subsequence includes: Determine a plurality of the first subsequences that match the same action label as a screening set; Sort the plurality of first subsequences in the screening set according to the confidence level to determine a sequence order; Determine the first subsequence located before a preset position in the sequence order as a primary subsequence; Determining the degree of repetition between the comparison subsequence and the primary subsequence according to the sequence length; wherein the comparison subsequence is the first subsequence in the screening set except the primary subsequence; The comparison subsequences whose repetition degree in the screening set is greater than or equal to a preset repetition degree threshold are removed to obtain the target subsequence.

6. The online behavior recognition method for typical actions in a substation room according to claim 2 is characterized in that: The method further comprises: If the same frame image to be identified in the sequence of frames to be identified belongs to multiple first action segments, sorting the multiple first action segments to which the frame image to be identified belongs according to the confidence level to obtain a segment order; The same frame image to be identified is deleted from other action segments; wherein the other action segments are the first action segments except the first first action segment in the segment sequence.

7. An online behavior recognition device for typical actions in a substation room, characterized in that: The device comprises: An acquisition module, used for acquiring a frame sequence to be identified; wherein the frame sequence to be identified includes a plurality of frame images to be identified arranged in chronological order; A detection module, used for performing posture detection on the frame sequence to be identified; A division module, used for dividing the frame sequence to be identified in which the posture information is detected into a first subsequence set and a second subsequence set according to a random frame division strategy and a continuous frame division strategy respectively; an identification module, used to identify the action segments and the action tags of the action segments in the first subsequence set and the second subsequence set respectively, and determine a first identification result and a second identification result of the frame sequence to be identified; a determination module, configured to determine a plurality of target action segments and action categories corresponding to each of the target action segments in the to-be-identified frame sequence according to the first recognition result and the second recognition result; The step of dividing the frame sequence to be identified by detecting the posture information into a second subsequence set according to the continuous frame division strategy includes: Dividing the frame sequence to be identified in which the posture information is detected into a plurality of second subsequences in sequence according to a third preset length; The step of identifying the action segments and the action tags of the action segments in the second subsequence set and determining a second recognition result of the frame sequence to be recognized includes: sorting the plurality of second subsequences according to the time when the second subsequences were generated; Determine a target identification sequence according to the first of the second subsequences of the generated time sequence; The target recognition sequence is recognized by a preset action recognition model to determine a recognition result of the target recognition sequence; wherein the recognition result of the target recognition sequence includes a third action label matched by the target recognition sequence, a third action end frame image, and a confidence level of the third action label; If the confidence of the third action label is greater than or equal to the first preset threshold and less than the second preset threshold, the target recognition sequence is superimposed according to the second subsequence adjacent to the target recognition sequence until the confidence of the third action label of the superimposed target recognition sequence is greater than or equal to the second preset threshold; If the confidence of the third action label is greater than or equal to the second preset threshold, determining the target recognition sequence as a target frame sequence; Determining a second action segment in the to-be-identified frame sequence according to the target frame sequence and the third action end frame image; using a third action label corresponding to the second action segment and a confidence level of the third action label as a second recognition result of the frame sequence to be recognized; The step of determining a plurality of target action segments and action categories corresponding to each of the target action segments in the to-be-identified frame sequence according to the first recognition result and the second recognition result includes: Obtain multiple different approximate fragment conditions; Compare the first action segment in the first recognition result and the second action segment in the second recognition result with the approximate segment condition respectively, and take the first action segment and the second action segment that meet the same approximate segment condition as an approximate segment set; If the first action label, the second action label and the third action label corresponding to the first action segment in the approximate segment set are the same, the target action segment is determined according to the first action segment and the second action segment, and the action category corresponding to the target action segment is determined according to the first action label or the second action label or the third action label; If any two of the first action label, the second action label and the third action label are the same, determining the target action segment according to the action segments corresponding to the any two action labels, and determining the action category corresponding to the target action segment according to the action labels corresponding to the any two action labels; If the first action label, the second action label and the third action label are all different, the first action label, the second action label and the third action label are sorted according to the confidence levels of the first action label, the second action label and the third action label to determine the label order; The target action segment is determined according to the action segment corresponding to the first action tag in the tag sequence, and the action category of the target action segment is determined according to the first tag in the tag sequence.

8. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the program, the online behavior recognition method for typical actions in a substation room according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video action detection method based on scale attention hole convolution network

    CN111611847A

  • Video action recognition method and device, electronic equipment and storage medium

    CN112115788A