Behavior recognition method and device
By acquiring and processing the pulse stream information captured by the UAV, segmenting and fusing features, and combining the pulse residual module and deterministic weight calculation, the ambiguity problem of target behavior recognition under high-speed UAV motion is solved, and the recognition accuracy and stability are improved.
Patent Information
- Application Number
- CN202510713862.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-05-30
AI Technical Summary
When a drone moves at high speed, the pixel information shifts due to the change in the position of the target object, resulting in image blur, which affects the accurate recognition of the target's behavioral characteristics. In particular, the information in the key areas cannot be accurately extracted, reducing the recognition accuracy.
By obtaining the light intensity change information in the pulse stream, dividing it into multiple sub-pulse streams, performing feature fusion and video frame restoration, and combining the pulse residual module and deterministic weight calculation, the target behavior can be identified.
It improves the accuracy and stability of drone target recognition in complex environments, enhances the ability to extract key features from blurred images, and improves the reliability of behavior recognition.
Smart Images

Figure CN120236217B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a behavior recognition method and device. Background Art
[0002] One way to identify target behavior features is to collect and analyze images of the target to estimate its movements and behaviors. In a scenario where a drone is identifying the behavior of a ground target, the drone uses its onboard imaging equipment to capture images of the ground target and then uses image processing algorithms or machine learning models to analyze the captured image sequences and identify the target's behavior category.
[0003] Existing target behavior recognition methods rely on clear image input. Due to the high-speed motion of drones, the target object's position on the imaging plane changes during the exposure time. This shift causes the pixel information to shift spatially, altering the target's visual appearance. For example, when identifying the behavioral characteristics of ground-moving targets (such as vehicles and pedestrians), previously clear contours, textures, and other features become blurred due to motion blur. This blurring significantly interferes with the accurate recognition of key areas of the target's behavioral characteristics. Key information, such as human body movements and vehicle turn signals, cannot be accurately extracted, resulting in a significant decrease in recognition accuracy. Summary of the Invention
[0004] In view of this, the present application proposes a behavior recognition method and device to solve the problem of misalignment of behavior recognition features that may occur when drones move at high speeds in related technologies.
[0005] The first embodiment of the present application provides a behavior recognition method, the method comprising:
[0006] Acquire a pulse stream of the object to be identified; the pulse stream includes information on light intensity changes of each pixel at different time points;
[0007] For any time scale among a plurality of predefined time scales, dividing the pulse stream into a plurality of sub-pulse streams according to the time scale;
[0008] Performing feature fusion on all sub-pulse streams corresponding to multiple time scales to obtain a pulse fusion feature; the pulse fusion feature includes light intensity change information of each pixel at multiple time scales;
[0009] Restoring the pulse fusion features to obtain multiple video frames;
[0010] For any video frame, calculating a certainty weight of the video frame based on the probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame;
[0011] The target behavior of the object to be identified is predicted based on the probability values of the multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
[0012] In an embodiment of the present application, calculating the deterministic weight of the video frame according to the probability values of the multiple behavior categories in the video frame and the total number of the multiple behavior categories includes:
[0013] For any video frame, calculating first probability values of multiple behavior categories of the object to be identified in the video frame;
[0014] determining second probability values of the plurality of behavior categories according to the first probability values of the plurality of behavior categories and a preset constant;
[0015] calculating a prediction uncertainty of the video frame based on a total number of the plurality of behavior categories and a sum of a plurality of second probability values;
[0016] A certainty weight of the video frame is calculated according to the prediction uncertainty of the video frame.
[0017] In the embodiment of the present application, after obtaining the plurality of video frames, the method further includes:
[0018] Identify blurry area features and clear area features in each video frame;
[0019] Comparing the clear area features of the plurality of video frames to determine the target position; the area features of the target position of a preset number of video frames among the plurality of video frames are clear area features;
[0020] Filtering a video frame to be processed from the multiple video frames; the region feature of the target position of the video frame to be processed is a fuzzy region feature;
[0021] The fuzzy region feature of the target position of the to-be-processed video frame is replaced with the clear region feature of the target position of the adjacent video frame.
[0022] In the embodiment of the present application, after obtaining the plurality of video frames, the method further includes:
[0023] For any video frame, calling multiple pulse residual modules to extract features of the video frame in sequence to obtain target image features of the video frame;
[0024] Determine probability values of multiple behavior categories of the object to be identified in the video frame according to the target image features.
[0025] In an embodiment of the present application, calling multiple pulse residual modules to sequentially extract features of the video frame to obtain target image features of the video frame includes:
[0026] When i is equal to 1, calling the i-th pulse residual module to perform feature extraction on the video frame to obtain the output image feature of the i-th pulse residual module;
[0027] When i is greater than 1, the i-th pulse residual module is called to extract the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module;
[0028] When i is equal to j, the output image feature of the j-th pulse residual module is used as the target image feature of the video frame.
[0029] In the embodiment of the present application, the i-th pulse residual module is called to perform feature extraction on the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module, including:
[0030] Perform continuous pulse convolution processing on the output image features of the i-1th pulse residual module to obtain the first intermediate feature;
[0031] Extracting features of the output image of the (i-1)th pulse residual module to obtain local features, and normalizing the local features to obtain second intermediate features;
[0032] The first intermediate feature and the second intermediate feature are feature fused, and the feature fusion result is residually spliced with the output image feature of the i-1th pulse residual module to obtain the output image feature of the i-th pulse residual module.
[0033] In an embodiment of the present application, the pulse stream is divided into a plurality of sub-pulse streams according to the time scale, including:
[0034] Determine an initial point according to the time scale; the initial point is 1 / 2 of the time scale;
[0035] Using the initial point as the splitting midpoint and the time scale as the splitting length, splitting the pulse stream into a sub-pulse stream corresponding to the initial point;
[0036] The next point is determined according to the initial point and 1 / 2 of the time scale, the next point is used as the splitting midpoint, the time scale is used as the splitting length, and a sub-pulse stream corresponding to the next point is split from the pulse stream.
[0037] In the embodiment of the present application, the final behavior category prediction probability is calculated based on the probability values of the multiple behavior categories in the multiple video frames and the deterministic weight of each video frame, including:
[0038] Filtering a plurality of target video frames from the plurality of video frames; the target video frames are video frames whose deterministic weights are greater than a preset weight threshold among the plurality of video frames;
[0039] A weighted average is performed on the probability values of multiple behavior categories in the plurality of target video frames to predict the target behavior of the object to be identified; wherein the weight of the weighted average is the deterministic weight of each target video frame.
[0040] In the embodiment of the present application, all sub-pulse streams corresponding to multiple time scales are subjected to feature fusion to obtain pulse fusion features, including:
[0041] Constructing a pulse stream sequence corresponding to the time scale according to the multiple sub-pulse streams, to obtain multiple pulse stream sequences corresponding one-to-one to the multiple time scales;
[0042] The plurality of pulse stream sequences are subjected to feature fusion to obtain a pulse fusion feature.
[0043] An embodiment of the second aspect of the present application provides a behavior recognition device, including:
[0044] A pulse stream acquisition module is used to acquire a pulse stream of an object to be identified; the pulse stream includes information on light intensity changes of each pixel at different time points;
[0045] a pulse stream splitting module, configured to split the pulse stream into a plurality of sub-pulse streams according to any one of a plurality of predefined time scales;
[0046] A feature fusion module is used to fuse the features of all sub-pulse streams corresponding to multiple time scales to obtain a pulse fusion feature; the pulse fusion feature includes the light intensity change information of each pixel at multiple time scales;
[0047] A video frame restoration module, configured to restore the pulse fusion features to obtain multiple video frames;
[0048] a certainty weight calculation module, configured to calculate, for any video frame, a certainty weight of the video frame based on probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame;
[0049] The target behavior prediction module is used to predict the target behavior of the object to be identified based on the probability values of multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
[0050] An embodiment of the third aspect of the present application provides an electronic device, which includes a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the behavior recognition method described in the first aspect above by executing the computer instructions.
[0051] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the behavior recognition method described in the first aspect above.
[0052] Additional aspects and advantages of the present application will be given in part in the description below and in part will become apparent from the description below or learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0054] Figure 1 A flow chart of a behavior recognition method provided by an embodiment of the present application is shown;
[0055] Figure 2 A schematic diagram of the structure of a pulse neural network provided in one embodiment of the present application is shown;
[0056] Figure 3 A schematic diagram of the structure of a behavior recognition device provided in one embodiment of the present application is shown;
[0057] Figure 4 A schematic structural diagram of an electronic device provided in one embodiment of the present application is shown;
[0058] Figure 5 A schematic diagram of a storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0060] It should be noted that, unless otherwise specified, the technical or scientific terms used in this application should have the common meanings understood by those skilled in the art to which this application belongs.
[0061] According to an embodiment of the present application, an embodiment of a behavior recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0062] In this embodiment, a behavior recognition method is provided. Figure 1 is a flow chart of a behavior recognition method according to an embodiment of the present application. Figure 1 As shown, the process includes the following steps:
[0063] Step S101: Acquire a pulse stream of an object to be identified.
[0064] In an embodiment of the present application, the pulse stream contains information on the light intensity changes of each pixel at different time points. The pulse stream is obtained by photographing the object to be identified by a pulse camera, such as a pulse stream obtained by photographing road objects by a drone equipped with a pulse camera in high-speed motion.
[0065] Step S102 : for any time scale among a plurality of predefined time scales, dividing the pulse stream into a plurality of sub-pulse streams according to the time scale.
[0066] Specifically, the size of the time scale can be customized and is not specifically limited here.
[0067] In practical applications, the blur level of each frame of a video image is different, and the scene may not always be blurred, which makes the scheme of using images with the same blur level as input invalid. To this end, the pulse stream input is classified into long and short time series sources to drive the model to perform cross-level fusion based on the image frames with different blur in the sequence. Assume that the pulse stream output by the pulse camera is And it is a binary domain. This paper proposes to cut the pulse sub-flows of different time scales to obtain the characteristics of different time sequence pulse flows.
[0068] In some specific embodiments, the above step S102 includes steps S1021 to S1023:
[0069] Step S1021: determining an initial point according to the time scale.
[0070] Specifically, the initial point is 1 / 2 of the time scale. For example, when the time scale is 1, the initial point is 0.5; and when the time scale is 2, the initial point is 1.
[0071] Step S1022 , using the initial point as the segmentation midpoint and the time scale as the segmentation length, to segment the pulse stream into a sub-pulse stream corresponding to the initial point.
[0072] Specifically, for example, when 0.5 is used as the segmentation midpoint and 1 is used as the segmentation length, the sub-pulse stream corresponding to the initial point position is a pulse stream in the time range of 0-1.
[0073] More specifically, the above steps S1021 to S1022 can be understood by the following formula:
[0074]
[0075] Among them, when the initial point When it is 0.5, the time scale N When it is 1, the sub-pulse flow corresponding to the initial point S(N) =[0, 1].
[0076] Step S1023, determining the next point according to the initial point and 1 / 2 of the time scale, taking the next point as the splitting midpoint and the time scale as the splitting length, and splitting the sub-pulse stream corresponding to the next point from the pulse stream.
[0077] Specifically, the above example is used to illustrate step S1023: when the initial point is 0.5, time scale N When it is 1, the next bit = + N / 2=0.5+0.5=1; the current position When 1, the time scale N When it is 1, the sub-pulse flow corresponding to the next bit S(N) =[0.5, 1.5].
[0078] In the embodiment of the present application, the pulse stream can be divided into multiple sub-pulse streams according to the time scale through the above method, for example: [0, 1], [0.5, 1.5], [1, 2], [1.5, 2.5], ...;
[0079] Similarly, by dividing the pulse stream according to multiple time scales in the above manner, multiple sub-pulse streams corresponding to each time scale can be obtained, as shown in the following formula:
[0080]
[0081] in, i Indicates the sequence number of different time scales, S(N i ) Represents the sub-pulse stream obtained by segmentation at different time scales, for example:
[0082]
[0083] in, Represents different time scales, Indicates the maximum time scale.
[0084] For example: when the time scale is 1, [0, 1], [0.5, 1.5], [1, 2], [1.5, 2.5], ....
[0085] When the time scale is 2, [0, 2], [1, 3], [2, 4], [3, 5], ....
[0086] When the time scale is K, [0, K], [K / 2, 3K / 2], [K, 2K], [3K / 2, 5K / 2], ....
[0087] In some specific embodiments, after step S102, the method further includes:
[0088] A pulse stream sequence corresponding to the time scale is constructed according to the multiple sub-pulse streams, so as to obtain multiple pulse stream sequences corresponding to the multiple time scales one by one.
[0089] Step S103 , performing feature fusion on all sub-pulse streams corresponding to multiple time scales to obtain pulse fusion features.
[0090] The pulse fusion feature includes light intensity change information of each pixel at multiple time scales.
[0091] In some specific embodiments, the plurality of pulse stream sequences may be subjected to feature fusion to obtain pulse fusion features in the following manner:
[0092] First, the pulse stream sequences at various time scales are input into a multi-layer temporal convolutional network (TCN) to model the dynamic characteristics of pulses of different durations. TCN consists of multiple layers of one-dimensional convolutions, with the convolution kernel size and expansion rate increasing layer by layer, expanding the receptive field layer by layer, enabling it to capture dependencies over longer time spans while maintaining the temporal order.
[0093] Then, the pulse dynamic features of multiple time scales are averaged and fused to obtain a unified temporal feature representation, namely the pulse fusion feature;
[0094] In the embodiment of the present application, the above steps can refer to the following formula:
[0095]
[0096] in, K represents the total number of time scales, T represents the pulse fusion feature, TCN Represents a multi-level temporal convolutional network, S(N i ) Indicates the i The pulse stream sequence corresponding to a time scale.
[0097] Step S104: restoring the pulse fusion features to obtain multiple video frames.
[0098] In an embodiment of the present application, the pulse fusion features can be restored to obtain a video frame through a decoder composed of upsampling, convolution, and ReLU functions.
[0099] In some specific embodiments, after step S104, the method further includes:
[0100] Step a1: for any video frame, call multiple pulse residual modules to extract features of the video frame in sequence to obtain target image features of the video frame.
[0101] In an embodiment of the present application, feature refinement and noise suppression can be achieved through a pulse neural network including multiple pulse residual modules, and feature joint denoising can be performed using a pulse neural network: its core is the pulse residual module, each pulse residual module captures and removes hidden noise existing at different times of continuous pulse streams.
[0102] In some specific embodiments, the above step a1 includes steps a11 to a13:
[0103] Step a11: when i is equal to 1, calling the i-th pulse residual module to perform feature extraction on the video frame to obtain output image features of the i-th pulse residual module;
[0104] Step a12: when i is greater than 1, calling the i-th pulse residual module to extract the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module;
[0105] Step a13: when i is equal to j, the output image feature of the j-th pulse residual module is used as the target image feature of the video frame.
[0106] In some specific embodiments, the above step a12 further includes steps a121 to a123:
[0107] Step a121: Perform continuous pulse convolution processing on the output image features of the i-1th pulse residual module to obtain a first intermediate feature.
[0108] Specifically, continuous pulse convolution processing refers to performing two pulse convolution operations consecutively (e.g. Figure 2 in SCU → SCU ”), in this way, the spatial structure extraction capability is enhanced and the local enhanced representation is obtained, namely the first intermediate feature, as shown in the following formula:
[0109]
[0110] Among them, n-1 is used to refer to the previous pulse residual module, that is, i-1; n represents the current pulse residual module, represents the output image feature of the above i-1th pulse residual module, SCU Represents pulse convolution processing.
[0111] In some embodiments, pulse convolution processing SCU It is mainly used to simulate membrane potential dynamics, convert pulses into time feature maps, and use leaky integrate and discharge (LIF) neurons to balance biological characteristics and computational complexity. The expression of LIF is: Here, the input of the spiking neural network is the time series representation of the pulse stream, and the output is the cumulative membrane potential expression of the spiking neuron at each time step, as shown in the following formula:
[0112]
[0113] in, The neurons in time and status The membrane potential at Indicates the firing state (1 if the neuron is firing; 0 otherwise, Represents the update of neuronal membrane potential. Generally speaking, neuronal membrane potential The membrane potential at the previous time step , input current and time constant If the membrane potential exceeds the threshold , the neuron will emit a pulse, the pulse state Set to 1, the membrane potential will be reset to Otherwise, the membrane potential will continue to update and gradually decay.
[0114] Step a122: extract the output image features of the (i-1)th pulse residual module to obtain local features, and normalize the local features to obtain second intermediate features.
[0115] Specifically, feature extraction is performed to obtain local features, and local features are normalized, for example Figure 2 As shown in “Conv→tdBN” in
[15] and the following formula, this approach can pay more attention to temporal consistency and normalization stability.
[0116]
[0117] in, Conv represents feature extraction, tdBN Indicates normalization processing.
[0118] Step a123: perform feature fusion on the first intermediate feature and the second intermediate feature, and perform residual splicing on the feature fusion result and the output image feature of the i-1th pulse residual module to obtain the output image feature of the i-th pulse residual module.
[0119] Specifically, for example Figure 2 As shown in the following formula: The first intermediate feature and the second intermediate feature are processed by multi-head attention through MCU. MAU combines spatial and temporal attention mechanisms to effectively focus on relevant signals and reduce noise interference. Residual learning ensures that low-level features are not lost and are effectively propagated through skip connections. The output of each time step is .
[0120]
[0121] Step a2: determining the probability values of multiple behavior categories of the object to be identified in the video frame according to the target image features.
[0122] In this step, the pooling operation is applied to the target image features to obtain the frame-level vector, and the category prediction score is output through the Linear layer. , that is, the probability values of multiple behavior categories of the object to be identified in the video frame. For example, the probability that the behavior category of the object to be identified in the video frame is running is 80%, and the probability of walking is 20%.
[0123] In some specific embodiments, after obtaining the plurality of video frames, the method further includes steps b1 to b4:
[0124] Step b1: Identify blurry area features and clear area features in each video frame.
[0125] Step b2: comparing the clear area features of the multiple video frames to determine the target position.
[0126] Among them, the regional features of the target position of a preset number of video frames in the multiple video frames are clear regional features. For example, there are a total of 100 pictures involving vehicle A, and the license plate of vehicle A is clearly visible in more than 60 of the pictures, then the license plate position can be used as the target position.
[0127] Step b3: Filtering out the video frames to be processed from the multiple video frames.
[0128] The region feature of the target position of the video frame to be processed is a fuzzy region feature. For example, pictures with unclear license plate positions are selected from 100 pictures, and these pictures are used as the video frames to be processed.
[0129] Step b4: replacing the blurred region features of the target position of the to-be-processed video frame with the clear region features of the target position of the adjacent video frame.
[0130] Specifically, during the replacement process, the video frame that is closest in time to the video frame to be processed and whose target position is a clear area feature can be first screened out; then the clear area feature of the target position in the video frame will be used to replace the fuzzy area feature of the target position in the video frame to be processed.
[0131] In steps b1-b4 above, this application considers the situation where pixels in different regions of the same frame have different degrees of blur. Therefore, the impact of blur on pulse fluctuations is reduced by leveraging the pixel distribution of adjacent moments. Therefore, blur correction is performed within the sub-pulse stream by modeling the spatial similarity of features at each moment and performing forward and backward shared alignment. This collaboration enables the use of context at different scales to better remove possible blur in the pulse stream and better restore behavioral characteristics.
[0132] Step S105 , for any video frame, calculate the certainty weight of the video frame according to the probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame.
[0133] In some specific embodiments, the above step S105 further includes steps S1051 to S1054:
[0134] Step S1051 : for any video frame, calculating first probability values of multiple behavior categories of the object to be identified in the video frame.
[0135] In this embodiment of the present application, the prediction score of each video frame can be extracted by the feature extraction network, linear layer and ReLU , the prediction score is a probability distribution including the first probability values of multiple behavior categories of the object to be identified in the video frame, for example: .
[0136] Step S1052 : determining second probability values of the plurality of behavior categories according to the first probability values of the plurality of behavior categories and a preset constant.
[0137] Specifically, Gumbel-Softmax is used to perform differentiable sampling on the first probability values of multiple behavior categories, and is converted into the second probability values of multiple behavior categories (such as the parameters of the Dirichlet distribution) through linear mapping and ReLU activation as shown in the following formula:
[0138]
[0139] in, Represents the predicted score of the object to be identified in the video frame, that is, the first probability value of multiple behavior categories, ; and by using a linear layer and a linear rectification (ReLU) layer to transform the probability Connecting to Dirichlet distribution Parameters.
[0140] Step S1053 : Calculate the prediction uncertainty of the video frame according to the total number of the multiple behavior categories and the sum of the multiple second probability values.
[0141] Specifically, prediction uncertainty refers to the reliability or trustworthiness of the prediction result for the behavior category of the object to be identified in the video frame. High uncertainty means that the prediction result for the behavior category of a particular video frame is uncertain and may have a large error. Conversely, low uncertainty indicates that the prediction result for the behavior category of a particular video frame is relatively confident and reliable.
[0142] More specifically, the prediction uncertainty of a video frame is calculated as follows:
[0143]
[0144] in, represents the Dirichlet intensity, K is the number of behavioral categories, represents the forecast uncertainty, and is inversely proportional, and Indicates in i The probability distribution of all categories observed in the visual frame, that is, the first probability value of the above multiple behavior categories, Indicated by The parameters of the converted Dirichlet distribution are the second probability values of multiple behavior categories. In short, evidence is obtained and the first probability values of multiple behavior categories are linked to the second probability values of multiple behavior categories to form an evidence-based class probability model. During the inference process, the Dempster synthesis rule is used to synthesize the evidence sources of different visual frames to calculate the prediction uncertainty of each frame. . Among them, the prediction uncertainty of any video frame is If it is higher than a certain threshold, it indicates that the behavioral appearance features observed in the video frame are unreliable.
[0145] Step S1054 : Calculate the certainty weight of the video frame according to the prediction uncertainty of the video frame.
[0146] Specifically, higher forecast uncertainty This means that the behavior appearance features observed in this video frame are unreliable and should be assigned a lower weight value. , as shown in the following formula:
[0147]
[0148] In an embodiment of the present application, the certainty weight of a video frame is used to represent the reliability or credibility of the prediction result of the behavior category of the object to be identified in the video frame. When the certainty weight of a video frame is low, the video frame can be ignored in the subsequent process of judging the target behavior of the object to be identified. Conversely, when the certainty weight of a video frame is high, the video frame can be given priority in the subsequent process of judging the target behavior of the object to be identified.
[0149] Step S106 : predicting the target behavior of the object to be identified based on the probability values of the multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
[0150] In some specific embodiments, the above step S106 includes steps S1061 and S1062:
[0151] Step S1061 : Filter out a number of target video frames from the multiple video frames.
[0152] Specifically, the target video frame refers to a video frame among the multiple video frames whose deterministic weight is greater than a preset weight threshold. The preset weight threshold can be set according to actual conditions and is not specifically limited here.
[0153] Step S1062 , performing weighted averaging on the probability values of multiple behavior categories in the target video frames to predict the target behavior of the object to be identified; wherein the weight of the weighted averaging is the deterministic weight of each target video frame.
[0154] The above steps S1061 and S1062 are described with an example:
[0155] First, the probability values of multiple behavior categories of the object to be identified in each target video frame are determined (for example, the probability of walking is 0.4, the probability of running is 0.3, and the probability of jumping is 0.3).
[0156] Next, we perform a weighted average of these behavior probability values, where the weights are the deterministic weights of each target video frame. For example, if there are three target video frames with deterministic weights of 0.8, 0.9, and 0.75, and their corresponding behavior probability values are [0.4, 0.3, 0.3], [0.5, 0.2, 0.3], and [0.3, 0.4, 0.3], the weighted average is calculated as follows:
[0157]
[0158]
[0159]
[0160] ≈[0.406,0.294,0.3]
[0161] Finally, based on the weighted average probability values, the behavior with the highest probability is selected as the prediction result. In this example, walking has the highest probability (approximately 0.406), so we can determine that the behavior category of the object to be identified is walking.
[0162] This embodiment introduces a pulse camera modality to assist in addressing the misalignment of behavioral features during high-speed motion, supplementing potentially lost behavioral information. This method divides the pulse stream into multiple substreams along the time series, and then uses a pulse multi-source element network to perform batch reconstruction. Furthermore, based on the complete sequence representation, a Dirichlet distribution is introduced to redistribute weights, using the impact of blur on the behavioral representation as a constraint.
[0163] The embodiments of the present application can enhance key features in blurred images, improve target recognition capabilities, and enable drones to have stronger target perception capabilities in complex environments.
[0164] Corresponding to the implementation of the above behavior recognition method, the embodiment of the present application also provides a behavior recognition device for executing the behavior recognition method described in the above embodiment. Figure 3As shown, the behavior recognition device includes:
[0165] A pulse stream acquisition module is used to acquire a pulse stream of an object to be identified; the pulse stream includes information on light intensity changes of each pixel at different time points;
[0166] a pulse stream splitting module, configured to split the pulse stream into a plurality of sub-pulse streams according to any one of a plurality of predefined time scales;
[0167] A feature fusion module is used to fuse the features of all sub-pulse streams corresponding to multiple time scales to obtain a pulse fusion feature; the pulse fusion feature includes the light intensity change information of each pixel at multiple time scales;
[0168] A video frame restoration module, configured to restore the pulse fusion features to obtain multiple video frames;
[0169] a certainty weight calculation module, configured to calculate, for any video frame, a certainty weight of the video frame based on probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame;
[0170] The target behavior prediction module is used to predict the target behavior of the object to be identified based on the probability values of multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
[0171] Optionally, the deterministic weight calculation module is also used to: calculate, for any video frame, the first probability values of multiple behavior categories of the object to be identified in the video frame; determine the second probability values of the multiple behavior categories based on the first probability values of the multiple behavior categories and preset constants; calculate the prediction uncertainty of the video frame based on the total number of the multiple behavior categories and the sum of the multiple second probability values; and calculate the deterministic weight of the video frame based on the prediction uncertainty of the video frame.
[0172] Optionally, the device also includes: a video frame processing module, used to identify blurred area features and clear area features in each video frame; compare the clear area features of the multiple video frames to determine the target position; the area features of the target position of a preset number of video frames among the multiple video frames are clear area features; screen out the video frames to be processed from the multiple video frames; the area features of the target position of the video frames to be processed are blurred area features; replace the blurred area features of the target position of the video frames to be processed with the clear area features of the target position of the adjacent video frames.
[0173] Optionally, the device further comprises:
[0174] A feature extraction module is used to call multiple pulse residual modules to extract features of any video frame in sequence to obtain target image features of the video frame;
[0175] The probability value determination module is used to determine the probability values of multiple behavior categories of the object to be identified in the video frame according to the target image features.
[0176] Optionally, the feature extraction module is also used to, when i is equal to 1, call the i-th pulse residual module to perform feature extraction on the video frame to obtain the output image features of the i-th pulse residual module; when i is greater than 1, call the i-th pulse residual module to perform feature extraction on the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module; when i is equal to j, use the output image features of the j-th pulse residual module as the target image features of the video frame.
[0177] Optionally, the feature extraction module is also used to perform continuous pulse convolution processing on the output image features of the i-1th pulse residual module to obtain a first intermediate feature; perform feature extraction on the output image features of the i-1th pulse residual module to obtain a local feature, and normalize the local feature to obtain a second intermediate feature; perform feature fusion on the first intermediate feature and the second intermediate feature, and perform residual splicing on the feature fusion result and the output image feature of the i-1th pulse residual module to obtain the output image feature of the i-th pulse residual module.
[0178] Optionally, the pulse stream segmentation module is also used to determine the initial point position according to the time scale; the initial point position is 1 / 2 of the time scale; the initial point position is used as the segmentation midpoint, the time scale is used as the segmentation length, and a sub-pulse stream corresponding to the initial point position is segmented from the pulse stream; the next point position is determined according to the initial point position and 1 / 2 of the time scale, the next point position is used as the segmentation midpoint, the time scale is used as the segmentation length, and a sub-pulse stream corresponding to the next point position is segmented from the pulse stream.
[0179] Optionally, the target behavior prediction module is further used to filter out several target video frames from the multiple video frames; the target video frames refer to video frames in the multiple video frames whose deterministic weights are greater than a preset weight threshold; the probability values of multiple behavior categories in the several target video frames are weighted averaged to predict the target behavior of the object to be identified; wherein the weight of the weighted average is the deterministic weight of each target video frame.
[0180] The behavior recognition device provided in the above-mentioned embodiment of the present application and the behavior recognition method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0181] The present application also provides an electronic device to perform the above behavior recognition method. Figure 4 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 4 As shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402 and a communication interface 403, and the processor 400, the communication interface 403 and the memory 401 are connected via the bus 402; the memory 401 stores a computer program that can be run on the processor 400, and when the processor 400 runs the computer program, it executes the behavior recognition method provided in any of the aforementioned embodiments of the present application.
[0182] Memory 401 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. Communication between the system network element and at least one other network element is achieved through at least one communication interface 403 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.
[0183] Bus 402 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. Memory 401 is used to store programs, and processor 400 executes the programs upon receiving execution instructions. The behavior recognition method disclosed in any of the aforementioned embodiments may be applied to processor 400 or implemented by processor 400.
[0184] The processor 400 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 400 or by software instructions. The above processor 400 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 401 , and the processor 400 reads the information in the memory 401 and completes the steps of the above method in combination with its hardware.
[0185] The electronic device provided in the embodiment of the present application and the behavior recognition method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.
[0186] The present application also provides a computer-readable storage medium corresponding to the behavior recognition method provided in the above embodiment. Figure 5 The computer-readable storage medium shown is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the behavior recognition method provided by any of the aforementioned embodiments is executed.
[0187] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0188] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the behavior recognition method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0189] It should be noted that:
[0190] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0191] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following schematic diagram: the claimed application requires more features than the features expressly recited in each claim. Rather, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0192] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination.
[0193] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A behavior recognition method, characterized in that: The method comprises: Acquire a pulse stream of the object to be identified; the pulse stream includes information on light intensity changes of each pixel at different time points; For any time scale among a plurality of predefined time scales, dividing the pulse stream into a plurality of sub-pulse streams according to the time scale; Performing feature fusion on all sub-pulse streams corresponding to multiple time scales to obtain a pulse fusion feature; the pulse fusion feature includes light intensity change information of each pixel at multiple time scales; Restoring the pulse fusion features to obtain multiple video frames; For any video frame, calculating a certainty weight of the video frame based on the probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame; The target behavior of the object to be identified is predicted based on the probability values of the multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
2. The method according to claim 1, characterized in that Calculating the deterministic weight of the video frame according to the probability values of the multiple behavior categories in the video frame and the total number of the multiple behavior categories includes: For any video frame, calculating first probability values of multiple behavior categories of the object to be identified in the video frame; determining second probability values of the plurality of behavior categories according to the first probability values of the plurality of behavior categories and a preset constant; calculating a prediction uncertainty of the video frame based on a total number of the plurality of behavior categories and a sum of a plurality of second probability values; A certainty weight of the video frame is calculated according to the prediction uncertainty of the video frame.
3. The method according to claim 1 or 2, characterized in that After obtaining the plurality of video frames, the method further includes: Identify blurry area features and clear area features in each video frame; Comparing the clear area features of the plurality of video frames to determine the target position; the area features of the target position of a preset number of video frames among the plurality of video frames are clear area features; Filtering a video frame to be processed from the multiple video frames; the region feature of the target position of the video frame to be processed is a fuzzy region feature; The fuzzy region feature of the target position of the to-be-processed video frame is replaced with the clear region feature of the target position of the adjacent video frame.
4. The method according to claim 1, wherein After obtaining the plurality of video frames, the method further includes: For any video frame, calling multiple pulse residual modules to extract features of the video frame in sequence to obtain target image features of the video frame; Determine probability values of multiple behavior categories of the object to be identified in the video frame according to the target image features.
5. The method according to claim 4, characterized in that Calling multiple pulse residual modules to sequentially extract features from the video frame to obtain target image features of the video frame, including: When i is equal to 1, calling the i-th pulse residual module to perform feature extraction on the video frame to obtain the output image feature of the i-th pulse residual module; When i is greater than 1, the i-th pulse residual module is called to extract the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module; When i is equal to j, the output image feature of the j-th pulse residual module is used as the target image feature of the video frame.
6. The method according to claim 5, characterized in that Call the i-th pulse residual module to extract the output image features of the i-1-th pulse residual module to obtain the output image features of the i-th pulse residual module, including: Perform continuous pulse convolution processing on the output image features of the i-1th pulse residual module to obtain the first intermediate feature; Extracting features of the output image of the (i-1)th pulse residual module to obtain local features, and normalizing the local features to obtain second intermediate features; The first intermediate feature and the second intermediate feature are feature fused, and the feature fusion result is residually spliced with the output image feature of the i-1th pulse residual module to obtain the output image feature of the i-th pulse residual module.
7. The method according to claim 1 or 2, characterized in that The pulse stream is divided into a plurality of sub-pulse streams according to the time scale, comprising: Determine an initial point according to the time scale; the initial point is 1 / 2 of the time scale; Using the initial point as the splitting midpoint and the time scale as the splitting length, splitting the pulse stream into a sub-pulse stream corresponding to the initial point; The next point is determined according to the initial point and 1 / 2 of the time scale, the next point is used as the splitting midpoint, the time scale is used as the splitting length, and a sub-pulse stream corresponding to the next point is split from the pulse stream.
8. The method according to claim 1 or 2, characterized in that Calculating a final behavior category prediction probability based on the probability values of the multiple behavior categories in the multiple video frames and the deterministic weight of each video frame includes: Filtering a plurality of target video frames from the plurality of video frames; the target video frames are video frames whose deterministic weights are greater than a preset weight threshold among the plurality of video frames; A weighted average is performed on the probability values of multiple behavior categories in the plurality of target video frames to predict the target behavior of the object to be identified; wherein the weight of the weighted average is the deterministic weight of each target video frame.
9. The method according to claim 1 or 2, characterized in that The pulse fusion features are obtained by fusion of all sub-pulse streams corresponding to multiple time scales, including: Constructing a pulse stream sequence corresponding to the time scale according to the multiple sub-pulse streams, to obtain multiple pulse stream sequences corresponding one-to-one to the multiple time scales; The plurality of pulse stream sequences are subjected to feature fusion to obtain a pulse fusion feature.
10. A behavior recognition device, characterized in that: The device comprises: A pulse stream acquisition module is used to acquire a pulse stream of an object to be identified; the pulse stream includes information on light intensity changes of each pixel at different time points; a pulse stream splitting module, configured to split the pulse stream into a plurality of sub-pulse streams according to any one of a plurality of predefined time scales; A feature fusion module is used to fuse the features of all sub-pulse streams corresponding to multiple time scales to obtain a pulse fusion feature; the pulse fusion feature includes the light intensity change information of each pixel at multiple time scales; A video frame restoration module, configured to restore the pulse fusion features to obtain multiple video frames; a certainty weight calculation module, configured to calculate, for any video frame, a certainty weight of the video frame based on probability values of multiple behavior categories in the video frame and the total number of the multiple behavior categories; the certainty weight represents the certainty of the multiple behavior categories in the video frame; The target behavior prediction module is used to predict the target behavior of the object to be identified based on the probability values of multiple behavior categories in the multiple video frames and the deterministic weight of each video frame.
Citation Information
Patent Citations
High-speed thrown object detection method and system based on cooperation of event camera and visual camera
CN112800860A
Super-resolution improvement method for motion image fusing pulse data and frame data
CN114782492A