Feature processing method, action localization method, device and apparatus

By performing temporal dimension partitioning and fusion processing on the feature sequences of the target video, the problem of inaccurate feature representation of image or video data is solved, achieving more efficient information representation and more accurate analysis results.

CN114463662BActive Publication Date: 2025-11-11ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111580387.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-11-11
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

In existing technologies, the features extracted from image or video data are difficult to accurately represent the original information, resulting in low accuracy of the analysis results.

Method used

By determining the target feature sequence based on the features of each video frame in the target video, dividing it according to the temporal dimension, performing a first fusion process along the temporal direction, and finally performing a second fusion process, the video features of the target video are obtained.

Benefits of technology

It improves the ability of target video features to express information, better expresses the temporal and semantic information between video frames, and enhances the accuracy of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463662B_ABST
    Figure CN114463662B_ABST
Patent Text Reader

Abstract

This application discloses a feature processing method, an action localization method, a device, and an apparatus. The method includes: determining a target feature sequence of the target video based on the features of each video frame in the target video; dividing the target feature sequence according to a temporal dimension to obtain multiple sub-feature sequences, wherein each sub-feature sequence includes features from at least one video frame; performing a first fusion process on each sub-feature sequence along the temporal direction to obtain intermediate features corresponding to each sub-feature sequence; and performing a second fusion process on the multiple intermediate features to obtain the video features of the target video. This approach can improve the expressive power of features on information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a feature processing method, an action localization method, a computer device, and a storage device. Background Technology

[0002] With the rapid development of information technologies such as mobile internet, cloud computing, and the Internet of Things, more and more data is being generated. Various fields possess massive and diverse data resources, which contain immense value.

[0003] The value of data is typically extracted through analysis and mining. For example, in fields such as video surveillance and education, features can be extracted from image or video data, analyzed, and thus yield analytical results. However, currently, the features extracted from image or video data often fail to represent the original information of the data, resulting in low accuracy of the analytical results. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a feature processing method, motion localization method, device, and apparatus that can improve the ability of features to express information.

[0005] To address the aforementioned issues, the first aspect of this application provides a feature processing method, comprising: determining a target feature sequence of the target video based on the features of each video frame in the target video; dividing the target feature sequence according to a temporal dimension to obtain multiple sub-feature sequences, wherein each sub-feature sequence includes features of at least one video frame; performing a first fusion processing on each sub-feature sequence in the multiple sub-feature sequences along the temporal direction to obtain intermediate features corresponding to each sub-feature sequence; and performing a second fusion processing on the multiple intermediate features to obtain the video features of the target video.

[0006] To address the aforementioned issues, a second aspect of this application provides an action localization method. This method includes: determining the features of an action to be detected in a target video as a target feature sequence; processing the target feature sequence using the aforementioned feature processing method to obtain video features of the target video; and determining video frame localization information of the action to be detected in the target video based on the video features of the target video; wherein the video frame localization information represents video frames in the target video that contain the action to be detected.

[0007] To address the aforementioned problems, a third aspect of this application provides a computer device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement any step of the aforementioned feature processing method and / or motion positioning method.

[0008] To address the aforementioned problems, a fourth aspect of this application provides a storage device that stores program data executable by a processor, the program data being used to implement any step of the aforementioned feature processing method and / or motion localization method.

[0009] The above scheme determines the target feature sequence of the target video based on the features of each video frame in the target video; it divides the target feature sequence according to the temporal dimension to obtain multiple sub-feature sequences. Since the sub-features of different temporal dimensions are different spatial information of the target feature sequence, the features can be described from the temporal dimension of the video frames of the target video. Then, along the temporal direction, each sub-feature sequence in the multiple sub-feature sequences is subjected to a first fusion process to obtain the intermediate features corresponding to each sub-feature sequence. The multiple intermediate features are subjected to a second fusion process to obtain the video features of the target video. This scheme can make full use of the dependencies and semantic information between the video frames in the temporal sequence of the target video. Compared with the target feature sequence, the video features of the target video can better express the information of different video frames in the temporal sequence of the target video, thus improving the expressive ability of the target video features to express information. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:

[0011] Figure 1 This is a flowchart illustrating an embodiment of the feature processing method of this application;

[0012] Figure 2 This application Figure 1 A schematic diagram illustrating an embodiment of step S12;

[0013] Figure 3 This application Figure 1 A flowchart illustrating an embodiment of step S13;

[0014] Figure 4 This is an example schematic diagram of an embodiment of the feature processing method of this application;

[0015] Figure 5 This is a flowchart illustrating an embodiment of the motion positioning method of this application;

[0016] Figure 6 This application Figure 5 A flowchart illustrating an embodiment of step S23;

[0017] Figure 7 This application Figure 5 A flowchart illustrating another embodiment of step S23;

[0018] Figure 8 This is a schematic diagram of the structure of an embodiment of the motion localization network of this application;

[0019] Figure 9 This is a schematic diagram of the structure of an embodiment of the feature processing device of this application;

[0020] Figure 10 This is a schematic diagram of the structure of an embodiment of the motion positioning device of this application;

[0021] Figure 11 This is a schematic diagram of the structure of an embodiment of the computer device of this application;

[0022] Figure 12 This is a schematic diagram of the structure of an embodiment of the storage device of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0025] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0026] This application provides the following embodiments, and each embodiment is described in detail below.

[0027] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the feature processing method of this application. This embodiment may include the following steps:

[0028] S11: Determine the target feature sequence of the target video based on the features of each video frame in the target video.

[0029] The features of each video frame are extracted sequentially from the target video according to their temporal order to determine the target feature sequence of the target video. These features can be color features or optical flow features, etc.

[0030] In some implementations, features from n adjacent video frames in the target video are extracted at time t1 to obtain a target feature, and features from n adjacent video frames in the target video are extracted at time t2 to obtain another target feature. Here, time t2 is after time t1, n is an integer greater than or equal to 1, and the n video frames at time t2 may or may not overlap with the n video frames at time t1. This process is repeated to obtain a target feature sequence for the target video. In some application scenarios, multiple target features can be arranged in chronological order to obtain a target feature sequence.

[0031] In some implementations, a sliding window approach can be used to extract the target feature sequence of video frames in the target video. For example, the video frames of the target video can be used as a window, and several groups of video frames can be selected sequentially from the target video in a sliding window manner. Each group of video frames includes one or more video frames.

[0032] In some implementations, the sliding window is configured with a step size and a size, which represent the number of video frames. This allows multiple groups of video frames to be selected from the target video according to the step size; a group of video frames can include one or more video frames. Color features or optical flow features can then be extracted from each group of video frames to obtain a target feature sequence from the target video. The target feature sequence is ordered according to the chronological order of the video frames from which the color features or optical flow features were extracted.

[0033] In some implementations, the target feature sequence is the overall feature of the video frames of the target video, which may or may not contain the action to be detected. The target feature sequence can be obtained using networks or models with feature extraction capabilities, such as the TSN model.

[0034] S12: Divide the target feature sequence according to the temporal dimension to obtain multiple sub-feature sequences, wherein each sub-feature sequence includes features of at least one video frame.

[0035] The target feature sequence can include a temporal dimension and a feature dimension. The temporal dimension is related to the length of the target video; the longer the target video, the more video frames there are, resulting in more extracted features and a longer temporal dimension. The feature dimension represents the content within a video frame.

[0036] The target feature sequence can be divided according to the temporal dimension to obtain multiple sub-feature sequences, where each sub-feature sequence includes features from at least one video frame. In some implementations, the temporal dimension can represent the time point information for extracting the target feature sequence, recording the time point information for each time the features of n consecutive video frames are extracted in a sliding window manner. In the temporal dimension, the earliest time of the video frame corresponding to the feature in the (i+1)th sub-feature sequence (i.e., the time point corresponding to the first video frame in the n consecutive video frames corresponding to the (i+1)th sub-feature sequence) is no earlier than the earliest time of the video frame corresponding to the feature in the ith sub-feature sequence (i.e., the time point corresponding to the first video frame in the n consecutive video frames corresponding to the ith sub-feature sequence).

[0037] In some implementations, a sub-feature sequence may include features of at least one or more video frames; that is, a sub-feature sequence may include features of one or more video frames. In the above example, a sub-feature sequence may include features of n consecutive video frames. Consecutive or adjacent sub-feature sequences may include features of partially identical video frames, or they may include features of completely different video frames. This application does not impose any limitations on this.

[0038] In some implementations, as an example, please refer to Figure 2 If the target feature sequence extracted from a video is Where t represents the time dimension and c represents the feature dimension. Target feature sequence It can be represented in the form of a matrix.

[0039] Here, the temporal dimension t can represent the length of the target video, and the feature dimension c can represent the content in the video frame. For example, obtaining the target feature sequence... At time t1, features of n consecutive video frames in the target video can be extracted to obtain a target feature T1. At time t2, features of n consecutive video frames in the target video can be extracted to obtain a target feature T2. Here, time t2 is after time t1, n is an integer greater than or equal to 1, and the n video frames at time t2 may or may not overlap with the n video frames at time t1.

[0040] The target feature T1 extracted at time t1 can be placed in the target feature sequence. The first column of the matrix, i.e., the first column of the matrix (1×c), can be used to represent the features of n consecutive video frames. The target feature T2 extracted at time t2 is placed in the target feature sequence. The second column of the matrix. From this, the target feature sequence can be obtained. The length of the feature dimension 'c' can be determined based on the specific application scenario to define the target feature sequence. Size.

[0041] Additionally, the target feature sequence F can be divided into t sub-feature sequences F1, F2, ..., Fc along the temporal dimension. t That is, the time series matrix F is divided into multiple sub-feature sequences according to the columns of the matrix. Each column can be considered as a sub-feature sequence. For example, the target feature T1 extracted at time t1 is divided into a sub-feature sequence, and the target feature T2 extracted at time t2 is divided into a sub-feature sequence. The multiple sub-feature sequences F1, F2, ..., F3 obtained after the division are: t Sort according to time sequence.

[0042] S13: Perform the first fusion process on each sub-feature sequence in the multiple sub-feature sequences along the time sequence direction to obtain the intermediate features corresponding to each sub-feature sequence.

[0043] The first fusion process can be performed on each sub-feature sequence in the multiple sub-feature sequences along the time sequence direction. For example, the first fusion process can be performed on each sub-feature sequence in the multiple sub-feature sequences along the first time sequence direction and the second time sequence direction respectively, so as to obtain the intermediate features corresponding to each sub-feature sequence.

[0044] In some implementations, the first fusion process may include convolution and superposition. For example, the i-th sub-sequence feature may be convolved, and then the result of the convolution may be superimposed on the (i+1)-th sub-feature sequence. Then, the (i+1)-th sub-feature sequence after superposition may be convolved again. The first fusion process may be performed on each sub-feature sequence sequentially along the temporal direction in this manner.

[0045] In some implementations, a first fusion process is performed on multiple sub-feature sequences along a first temporal direction to obtain fused features corresponding to each sub-feature sequence in the first temporal direction; and a first fusion process is performed on each sub-feature sequence along a second temporal direction to obtain fused features corresponding to each sub-feature sequence in the second temporal direction. Thus, multiple intermediate features can be obtained by combining multiple fused features from the first temporal direction and multiple fused features from the second temporal direction.

[0046] In some implementations, a first fusion process can be performed on each sub-feature sequence of the multiple sub-feature sequences along a first temporal direction to obtain fused features corresponding to each sub-feature sequence. Then, a first fusion process can be performed on the multiple fused features along a second temporal direction to obtain intermediate features corresponding to the multiple fused features.

[0047] In some implementations, the temporal direction can be determined according to the extension direction of the temporal dimension. For example, the first temporal direction is forward, which is the forward extension direction of the temporal dimension, and the second temporal direction is backward, which is the reverse extension direction of the temporal dimension. For example, the order of the first temporal directions corresponding to multiple sub-feature sequences is F1, F2, ..., F... t The sorting of the second temporal direction corresponding to multiple sub-feature sequences is F. t F1, F2, ..., F1. This application uses these as examples to illustrate embodiments, but the application is not limited thereto.

[0048] In other embodiments, the first timing direction is reverse and the second timing direction is forward. The first timing direction and the second timing method are relative, and this application is not limited thereto.

[0049] S14: Perform a second fusion process on multiple intermediate features to obtain the video features of the target video.

[0050] A second fusion process is performed using the intermediate features corresponding to each of the multiple sub-feature sequences to obtain the video features of the target video after processing the target feature sequence.

[0051] In some implementations, the temporal dimensions of the video features of the target video and the target feature sequence may or may not be the same. Different fusion processes can be applied to the two cases. The following describes the case where the temporal dimensions of the video features of the target video and the target feature sequence are the same.

[0052] In some implementations, multiple intermediate features can be fused along a temporal dimension to obtain the video features of the target video corresponding to the target feature sequence. The video features of the target video obtained in this process have the same temporal dimension as the target feature sequence. Therefore, after performing a second fusion process on the intermediate features corresponding to each sub-feature sequence, the order of the intermediate features is consistent with the order of each sub-feature sequence, that is, the corresponding temporal dimension is consistent.

[0053] In some implementations, the second fusion process may include concatenation along a temporal dimension. That is, the intermediate features may be concatenated into a matrix of the same size as the matrix of the target feature sequence, following the form of sub-feature sequences divided into matrices. For example, the intermediate features corresponding to each sub-feature sequence may be arranged sequentially along a temporal dimension, with the intermediate features corresponding to the i-th sub-sequence and the intermediate features corresponding to the (i+1)-th sub-sequence arranged adjacently, thereby concatenating all intermediate features into the video features of the target video.

[0054] After the second fusion process, the intermediate features corresponding to adjacent target feature sequences are adjacent.

[0055] In this embodiment, a target feature sequence of the target video is determined based on the features of each video frame in the target video. The target feature sequence is divided according to the temporal dimension to obtain multiple sub-feature sequences. Since the sub-features of different temporal dimensions are different spatial information of the target feature sequence, the features can be described from the temporal dimension of the video frames of the target video. Then, along the temporal direction, each sub-feature sequence in the multiple sub-feature sequences is subjected to a first fusion process to obtain the intermediate features corresponding to each sub-feature sequence. The multiple intermediate features are subjected to a second fusion process to obtain the video features of the target video. This can make full use of the dependencies and semantic information between the video frames in the temporal sequence of the target video. Compared with the target feature sequence, the video features of the target video can better express the information of different video frames in the temporal sequence of the target video, thus improving the expressive ability of the target video features to express information.

[0056] In some embodiments, please refer to Figure 3 Step S13 above, which involves performing a first fusion process on each sub-feature sequence in the multiple sub-feature sequences along the temporal direction to obtain the intermediate features corresponding to each sub-feature sequence, may include the following steps:

[0057] In step S13, the first fusion process is performed on each sub-feature sequence of the multiple sub-feature sequences along the first time sequence direction to obtain the fused features corresponding to the multiple sub-feature sequences. This may include the following steps S131 to S133. Specifically:

[0058] S131: Sequentially use each sub-feature sequence as the first target sub-feature according to the first time sequence direction.

[0059] In some implementations, if the first temporal direction is positive, then the order of the first temporal directions corresponding to the multiple sub-feature sequences is F1, F2, ..., F... t Therefore, each sub-feature sequence F1, F2, ..., F can be sequentially arranged according to the first temporal direction of the time dimension. tThis application uses this as an example to illustrate an embodiment, but it is not limited thereto.

[0060] In other implementations, if the first temporal direction is reversed, then the order of the first temporal directions corresponding to the multiple sub-feature sequences is F. t F1, F2, ..., F1. Therefore, each sub-feature sequence F can be sequentially arranged according to the first temporal direction of the time dimension. t F1, F2, ..., F1 are used as the first target sub-features.

[0061] S132: Based on the first target sub-feature and the first reference sub-feature, obtain the first feature to be processed, wherein the first reference sub-feature is the sub-feature sequence preceding the first target sub-feature in the first temporal direction.

[0062] In some implementations, if the first target sub-feature is the first sub-feature sequence in the first temporal direction, then the first target sub-feature is taken as the first feature to be processed. That is, when the first temporal direction is positive, if the first sub-feature sequence F1 ordered according to the time dimension is the first target sub-feature, then the first sub-feature sequence F1 can be taken as the first feature to be processed.

[0063] In some implementations, if the first target sub-feature is not the first sub-feature sequence in the first temporal direction, the fused feature corresponding to the first target sub-feature and the first reference sub-feature is superimposed to obtain the first feature to be processed. For example, when the first target sub-feature is sub-feature sequence F2, since the preceding sub-feature sequence F1 of the first target sub-feature F2 is also a sub-feature sequence, convolution can be performed on sub-feature sequence F1 to obtain the fused feature corresponding to sub-feature sequence F1. The first target sub-feature F2 is then superimposed with the first reference sub-feature to obtain the first feature to be processed corresponding to the first target sub-feature F2.

[0064] In some implementations, the fused feature corresponding to the preceding sub-feature sequence in the first temporal direction is obtained by performing steps S131 to S133. In this process, the preceding sub-feature sequence in the first temporal direction in this description is used as the first target sub-feature.

[0065] S133: Use the first convolution kernel to perform convolution processing on the first feature to be processed to obtain the fused feature corresponding to the first target sub-feature.

[0066] In some implementations, when performing the step of dividing the target feature sequence according to the time dimension to obtain multiple sub-features, a first convolutional kernel can be generated.

[0067] The first convolution kernel can be used to perform convolution processing on the first feature to be processed to obtain the fused feature corresponding to the first target sub-feature.

[0068] In some implementations, if the first target sub-feature is the first sub-feature sequence F1 in the first temporal direction, the sub-feature sequence F1 is convolved using the first convolution kernel to obtain the fused feature corresponding to the first target sub-feature F1.

[0069] In some implementations, if the first target sub-feature is a sub-feature sequence F2, since the previous sub-feature sequence of the sub-feature sequence F2 in the first temporal direction is a sub-feature sequence F1, the first feature to be processed is obtained by superimposing the fusion feature corresponding to the first target sub-feature F2 and the sub-feature sequence F1. The first convolution kernel can be used to perform convolution processing on the first feature to be processed to obtain the fusion feature corresponding to the first target sub-feature F2.

[0070] By performing steps S131 to S133 above, the fused features corresponding to each sub-feature sequence in the first temporal direction can be obtained.

[0071] In step S13, the first fusion process is performed on each of the multiple fusion features along the second time sequence direction to obtain multiple intermediate features corresponding to the multiple fusion features. This may include the following steps S134 to S136, as follows:

[0072] S134: Sequentially take the fused features corresponding to each sub-feature sequence as the second target sub-features according to the second time sequence direction.

[0073] In some implementations, the second temporal direction is reversed, and the order of the first temporal directions corresponding to the multiple sub-feature sequences is F. t F1, F2, ..., F1. Therefore, each sub-feature sequence F can be sequentially arranged according to the second temporal direction of the time dimension. t The fusion features corresponding to F1, F2, ..., F1 are used as the second target sub-features. This application uses this as an example to illustrate the embodiment, but this application is not limited thereto.

[0074] In other implementations, if the second temporal direction is positive, then the order of the second temporal directions corresponding to the multiple sub-features is F1, F2, ..., F... t Therefore, following the second temporal direction of the time dimension, each sub-feature sequence F1, F2, ..., F can be sequentially arranged. t The corresponding fusion features are used as the second target sub-features.

[0075] In some embodiments, each second target sub-feature referred to in this embodiment is a fused feature corresponding to each sub-feature sequence obtained in steps S131 to S133 above, which is used as a sub-feature sequence after stacking. Steps S134 to S136 are then performed using each sub-feature sequence after stacking. That is, the fused features corresponding to each of the above-obtained sub-feature sequences are used as the second target sub-features in sequence according to the second temporal direction. The following description uses this embodiment as an example.

[0076] S135: Based on the second target sub-feature and the second reference sub-feature, obtain the second feature to be processed, wherein the second reference sub-feature is the previous fusion feature of the second target sub-feature in the second temporal direction.

[0077] In some implementations, if the second target sub-feature is the fused feature corresponding to the first sub-feature sequence in the second temporal direction, i.e., the first fused feature, then the second target sub-feature is used as the second feature to be processed. When the second temporal direction is reversed, the first sub-feature sequence F in reverse temporal order is... t When the corresponding fusion feature is the second target sub-feature, then the first sub-feature sequence F can be... t The corresponding fusion feature is used as the second feature to be processed.

[0078] In some implementations, if the second target sub-feature is not the fused feature corresponding to the first sub-feature sequence in the second temporal direction, i.e., it is not the first fused feature, then the second target sub-feature is superimposed with the intermediate feature corresponding to the second reference sub-feature to obtain the second feature to be processed. For example, the second target sub-feature is the sub-feature sequence F. t-1 At that time, due to the sub-feature sequence F t-1 The preceding sub-feature sequence in the second temporal direction is F t The second target sub-feature F can be used t-1 The sub-feature is sequence F t The corresponding intermediate features are superimposed to obtain the second feature to be processed.

[0079] In some implementations, the sub-feature sequence F t-1 The corresponding intermediate features are obtained by executing steps S134 to S136. During this process, the sub-feature sequence F... t-1 As the second target sub-feature.

[0080] S136: Use the second convolution kernel to perform convolution processing on the second feature to be processed, and obtain the intermediate feature corresponding to the second target sub-feature.

[0081] In some implementations, after performing step S133 above, a second convolution kernel can be generated.

[0082] The second convolution kernel is used to perform convolution processing on the second feature to be processed, resulting in intermediate features corresponding to the second target sub-feature. For example, the second target sub-feature is the sub-feature sequence F. t When considering the corresponding fusion features, a second convolution kernel can be used to convolve the fusion features to obtain the intermediate features corresponding to the second target sub-feature. For example, the second target sub-feature is the sub-feature sequence F. t-1 When fusion features are involved, due to the second temporal direction, the sub-feature sequence F t-1 The previous sub-feature sequence is F t Then the second target sub-feature F can be... t The corresponding fusion features and sub-feature sequences F t The corresponding intermediate features are superimposed to obtain the intermediate features corresponding to the second target sub-feature.

[0083] In some implementations, the first convolution kernel of the first convolution process is different from the second convolution kernel of the second convolution process, and the first convolution kernel and the second convolution kernel may have the same or different sizes.

[0084] By performing steps S134 to S136 above, intermediate features corresponding to each second target sub-feature in the second temporal direction can be obtained.

[0085] Please see Figure 4 The following description uses an example to illustrate steps S131 to S136. Assume the target feature sequence extracted from the target video is... Where t represents the temporal dimension and c represents the feature dimension, the target feature sequence F can be divided into t sub-feature sequences F1, F2, ..., Fc along the temporal dimension. t A 1×k first convolution kernel K1 is generated, where k is a positive odd number. The parameters of the first convolution kernel are initialized using random numbers that follow a normal distribution. The value of k can be selected according to the actual application scenario, and this application does not impose any restrictions on it. This process can be referred to the specific implementation process of the above embodiments, and will not be repeated here.

[0086] By performing a first convolution on the first sub-feature sequence F1 in the first temporal direction using the first convolution kernel K1, the fused feature corresponding to the sub-feature sequence F1 is obtained. The fusion features corresponding to the sub-feature sequence F1 The sub-feature sequence F2 is superimposed to obtain the superimposed sub-feature sequence F2, which is also the first feature C2 to be processed corresponding to the sub-feature sequence F2. Then, the superimposed sub-feature sequence F2 (the first feature C2 to be processed) is convolved with the first convolution kernel K1 to obtain the fused feature corresponding to the sub-feature sequence F2. Each sub-feature sequence is processed sequentially in this manner until the sub-feature sequence F is obtained. t-1 Corresponding fusion features Superimposed on sub-feature sequence F t The first convolution kernel is used to perform convolution processing on the superimposed result (the first feature to be processed) to obtain the sub-feature sequence F. t Corresponding fusion features This process yields the fused features corresponding to each sub-feature sequence. Through this process, the temporal information of the target video frames is transferred from the first temporal direction.

[0087] Then, another second convolution kernel K2 is generated. The parameters of the second convolution kernel K2 can be generated in the same way as the first convolution kernel K1. Specifically, the second convolution kernel K2 can be generated with the same size 1×k as the first convolution kernel K1, or it can be generated with a different size 1×m than the first convolution kernel K1. The first convolution kernel K1 and the second convolution kernel K2 can be selected according to the specific application scenario, and this application does not impose any restrictions.

[0088] Using the second convolution kernel K2 to process the sub-feature sequence F in the second temporal direction t Corresponding fusion features Perform a second convolution process to obtain the sub-feature sequence F. t Corresponding intermediate features Intermediate features Add to fusion features (Superimposed sub-feature sequences F) t-1 The second feature N2 to be processed is obtained from the first feature, and then the second convolution kernel K2 is used to process the sub-feature sequence F. t-1 The corresponding second feature N2 to be processed is subjected to a second convolution process to obtain the sub-feature sequence F. t-1 Corresponding intermediate features The above operations are performed sequentially until the intermediate feature corresponding to F2 is superimposed onto the superimposed sub-feature sequence F1 (fusion feature). Then, using a second convolution kernel, we perform convolution processing to obtain the intermediate features corresponding to the sub-feature sequence F1. This process completes the transmission of temporal information of the target video frames from the second temporal direction. In other words, it completes the first fusion process on each sub-feature sequence, obtaining the intermediate features corresponding to each sub-feature sequence.

[0089] Finally, the obtained t features (intermediate features) can be subjected to a second fusion process according to the temporal dimension, and concatenated into a video feature F with the same dimension as the target feature sequence F. p That is, the video features F pAs a video feature of the target video, it integrates the temporal information of the target video and the semantic information between video frames.

[0090] The fused target features fully utilize the temporal information of the target video, as well as the semantic information between video frames, enabling the video features of the target video to more accurately express the information of the target video, thereby improving the ability of video features to express the information of the target video and video frames.

[0091] In some embodiments, the above feature processing method can also be applied to feature processing in application scenarios where multiple video frames are related or connected, so as to use the video features of the obtained target video for analysis.

[0092] In some embodiments, the target feature sequence described above can be obtained by performing feature detection on the action to be detected, and thus can be applied to any scenario where the action to be detected needs to be identified, so that the action to be detected can be identified based on the target features.

[0093] The identification of the action to be detected can be, in this case, identifying whether the action to be detected exists in an image or image sequence, classifying the action to be detected, or locating the action to be detected, i.e., determining the image (video frame) in the image sequence (video) that contains the action to be detected.

[0094] Taking the application of the above feature processing method to an action localization scenario as an example, this application will explain it below.

[0095] Please see Figure 5 , Figure 5 This is a flowchart illustrating an embodiment of the motion localization method of this application. This embodiment may include the following steps:

[0096] S21: Determine the features of the action to be detected in the target video as the target feature sequence.

[0097] The target video may contain actions (one or more types), or it may not contain actions. All actions contained in the target video are considered as actions to be detected; therefore, the action localization method in this application aims to obtain the localization information of each action present in the target video. The target feature sequence contains semantic and temporal information between various video frames in the target video. Furthermore, the target feature sequence may include RGB features (color features) and / or optical flow features of one or more video frames.

[0098] The target feature sequence can be obtained using a model or network with feature extraction capabilities. For example, the target feature sequence of a target video can be obtained using the TSN model. This application will illustrate this by taking RGB features as an example.

[0099] In some embodiments, the features can be preprocessed by scaling the target feature sequence using methods such as linear interpolation to a fixed length L, so as to obtain a preprocessed target feature sequence.

[0100] The target feature sequence is input into the action localization network for processing, so that the action localization network can locate the target action in the target video based on the target feature sequence. That is, the action localization network can perform the following steps: processing the target feature sequence to obtain target features in step S22, and locating the target features to obtain the video frame localization information of the action to be detected in the target video in step S23.

[0101] S22: Use feature processing methods to process the target feature sequence to obtain the video features of the target video.

[0102] The feature processing method in this step is the feature processing method described in the above embodiments. The specific implementation process can be referred to the specific implementation process of the above embodiments, and will not be repeated here.

[0103] By using feature processing methods to process the target feature sequence, the resulting video features of the target video can better express the semantic information between each video frame in the target video, as well as the temporal information between each video frame.

[0104] S23: Based on the video features of the target video, determine the video frame location information of the action to be detected in the target video; wherein, the video frame location information is used to represent the video frames in the target video that contain the action to be detected.

[0105] The video features of the target video obtained from the above steps can be used to locate the video frames of the action to be detected in the target video, so as to determine the starting video frame, ending video frame, evaluation value of the action, etc.

[0106] In this embodiment, the target video is located by extracting the target feature sequence from the target video and using the video features of the target video obtained by the above feature processing method. The video features can better express the semantic and temporal information between each video frame in the target video, and thus the obtained location information is more accurate.

[0107] In some embodiments, please refer to Figure 6 Step S23 above, which determines the video frame location information of the action to be detected in the target video based on the target features, may include the following steps:

[0108] S231: Based on the video features of the target video, predict the first estimated value of each candidate video frame in the target video as the starting video frame, and the second estimated value of each candidate video frame as the ending video frame.

[0109] The target video corresponds to a pre-defined set of candidate video frames, which includes multiple candidate video frames. Each candidate video frame in the set is represented by an identifier, which can be a frame number or a corresponding time. Multiple candidate video frames in the set can form multiple video segments, and the starting frame of a valid video segment should be earlier than the ending frame.

[0110] Video frame location information refers to the location information of video segments. It can include the location information of all video segments that can be formed by multiple candidate video frames in the candidate video frame set, or it can include only the location information of valid video segments, or it can include only the location information of some video segments among the valid video segments. The following embodiments will be described using the example of only including the location information of valid video segments (hereinafter referred to as target video segments).

[0111] The reference video frame may include a start video frame and / or an end video frame. The start video frame is the first video frame of the video segment in which the action to be detected occurs, and the end video frame is the last video frame of the video segment in which the action to be detected occurs. The estimated value represents the probability that a candidate video frame is the reference video frame.

[0112] If the reference video frame includes a start video frame and an end video frame, and the estimated value includes a first estimated value and a second estimated value, then based on the video features of the target video, the first estimated value for each candidate video frame as the start video frame and the second estimated value for each candidate video frame as the end video frame can be predicted.

[0113] Among them, the video features of the target video can be directly predicted to obtain the first and second estimates of each candidate video frame.

[0114] In some implementations, to improve prediction accuracy, the video features of the target video can be determined as a new target feature sequence. The new target feature sequence is then processed using the aforementioned feature processing method to obtain new video features of the target video. The new video features of the target video are then predicted to obtain a first estimate and a second estimate for each candidate video frame.

[0115] In some implementations, the video features of the target video can be subjected to a first convolution process and a second convolution process respectively to obtain a first target feature and a second target feature. The first target feature and the second target feature are then determined as new target feature sequences, and the new target feature sequences are processed using the above feature processing method to obtain new first target features and second target features. The new first target features and second target features are then predicted to obtain a first estimate and a second estimate for each candidate video frame.

[0116] S232: Select a subset of candidate video frames based on the first and second estimates of each candidate video frame.

[0117] If the reference video frame only includes the start video frame, then a portion of the candidate start video frames are selected based on the first estimate of each candidate video frame; if the reference video frame only includes the end video frame, then a portion of the candidate end video frames are selected based on the second estimate of each candidate video frame; if the reference video frame includes both the start and end video frames, then a portion of the candidate start and candidate end video frames are selected based on the first and second estimates of each candidate video frame.

[0118] In some implementations, candidate video frames whose first estimated value satisfies a first condition are determined as candidate start video frames; and candidate video frames whose second estimated value satisfies a second condition are determined as candidate end video frames.

[0119] In some implementations, the first condition may be a first estimated value greater than or equal to a first estimated threshold, or it may be that the first estimated value of all candidate video frames is arranged in descending order and has a higher sequence number. Correspondingly, the second condition may be a second estimated value greater than or equal to a second estimated threshold, or it may be that the second estimated value of all candidate video frames is arranged in descending order and has a higher sequence number.

[0120] In some implementations, the first condition can be that the first estimated value is an extreme point, and the second condition is that the second estimated value is an extreme point. Under this condition, the first estimated values ​​of all starting video frames and the second estimated values ​​of all ending video frames can be traversed to obtain and mark all extreme points, and the video frames corresponding to the extreme points can be used as a subset of candidate video frames. That is, the index values ​​of all extreme points in the starting video frame can be used as the starting point of all prediction results, and the index values ​​of all extreme points in the ending video frame can be used as the ending point of all prediction results.

[0121] S233: Based on the selected candidate video frames, obtain the video frame location information.

[0122] The video frame positioning information referred to in this embodiment is the positioning information of the target video segment.

[0123] If the reference video frames only include the starting video frame, then the target video segment is a video segment that starts with the candidate starting video frame and ends with the last video frame of the target video. The location information of the target video segment may include a first estimate of the starting video frame.

[0124] If the reference video frames only include the ending video frame, then the target video segment is a video segment that starts with the first video frame of the target video and ends with a candidate ending video frame. The location information of the target video segment may include a second estimate of the ending video frame.

[0125] If the reference video frames include a start video frame and an end video frame, then at least one candidate start video frame and one candidate end video frame corresponding to the target video segment can be selected from the selected candidate start video frames and candidate end video frames. Each target video segment corresponds to one candidate start video frame and one candidate end video frame, and the candidate start video frame corresponding to the target video segment is temporally earlier than the candidate end video frame corresponding to the target video segment. Based on at least one candidate start video frame and candidate end video frame corresponding to the target video segment, video frame positioning information is determined.

[0126] In some implementations, each target video segment corresponds to a candidate start video frame and a candidate end video frame, that is, the index value corresponding to the extreme point of the end video frame is greater than the index value corresponding to the extreme point of the start video frame.

[0127] Valid video segments among those formed by the selected candidate start and end video frames can be considered as video segments that meet the positioning conditions and are used as target video segments. The positioning information of the target video segment is determined based on its start and end video frames. The positioning information of the target video segment may include a first estimated value of the start video frame and a second estimated value of the end video frame.

[0128] In some embodiments, the video localization information includes a first estimate of the starting video frame, a second estimate of the ending video frame, and an evaluation value indicating the presence of an action to be detected. This will be explained below.

[0129] In some embodiments, please refer to Figure 7 In step S23 above, obtaining video frame location information based on the selected candidate video frames may further include the following steps:

[0130] S234: Based on the video features of the target video, obtain the regional features of each target video segment.

[0131] After the start and end video frames that satisfy the positioning conditions of each of the target video segments, the steps in this embodiment can be performed.

[0132] The regional features corresponding to the target video segment are the features of the video frames included in the target video segment in the video features of the target video.

[0133] S235: Identify each target video segment based on the corresponding regional features to obtain the action evaluation value of each target video segment.

[0134] In some implementations, the action evaluation value includes a first evaluation value and / or a second evaluation value. The first evaluation value indicates whether a detectable action exists, and can describe the probability that the detectable action exists in the target video segment. The second evaluation value indicates the regressivity of the detectable action, and can represent the continuity of the detectable action in the target video segment, or the continuity of video frames in the target video segment where the detectable action occurs.

[0135] In some of the above methods, the action evaluation value includes a first evaluation value indicating whether there is an action to be detected. If the first evaluation value is greater than a first threshold, the target video segment is determined as the final target video segment.

[0136] In some of the above methods, the action evaluation value includes a second evaluation value representing the regressivity of the action to be detected. Target video segments with a second evaluation value greater than a second threshold are then determined as the final target video segments.

[0137] In some of the above methods, if the action evaluation value includes a first evaluation value and a second evaluation value, then the target video segment whose first evaluation value is greater than a first threshold and whose second evaluation value is greater than a second threshold is determined as the final target video segment.

[0138] In some of the above methods, after obtaining the video frame location information of the action to be detected in the target video, the method may further include: determining the video frame corresponding to the video frame location information in the target video based on the video frame location information, that is, based on the first estimated value of the starting video frame, the second estimated value of the ending video frame, and the evaluation value of the existence of the action to be detected. The video frame corresponding to the video frame location information is the video frame included in the target video segment. Therefore, based on the video frame corresponding to the video frame location information, the action to be detected is identified, the action to be detected contained in the target video segment is detected, and the action to be detected is classified. For example, the classification could be actions such as "raising a hand," "waving," or "making a heart shape." This application is not limited to this.

[0139] In some implementations, if there may be more than one target video segment corresponding to the same type of action to be detected, the following situation may occur:

[0140] Scenario 1: The same type of action to be detected may correspond to multiple duplicate location information. That is, multiple target video segments obtained do indeed contain the action to be detected, but the length of some target video segments obtained from location is shorter than the actual time segment length of the target video segments, meaning that the existing action to be detected is incomplete.

[0141] Scenario 2: Due to insufficient accuracy of the first / second estimate in the prediction, some target video segments may not actually contain the action to be detected.

[0142] Therefore, in step S233, the target time segment can be filtered based on the action evaluation value. The filtering method includes, but is not limited to, non-maximum suppression algorithms. The filtered target time segment is then used as the final target video segment, and the video frame positioning information obtained in step S233 is used as the positioning information for the final target video segment. Specifically, a target video segment with a first evaluation value greater than a first threshold can be determined as the final target video segment. Alternatively, a target video segment with a second evaluation value greater than a second threshold can be determined as the final target video segment. Or, a target video segment with both a first evaluation value greater than the first threshold and a second evaluation value greater than the second threshold can be determined as the final target video segment.

[0143] Please see Figure 8 As an example, the action localization method described above in this application can be implemented using an action localization network. The target feature sequence detected from the target video can be input into this action localization network to obtain video frame localization information. Specifically:

[0144] The action localization network may include a base encoding network, a start prediction network, an end prediction network, and an action prediction network. The base encoding network, start prediction network, and end prediction network may each include a target feature sequence fusion module, which can be used to perform any step of the aforementioned feature processing method.

[0145] The target feature sequence is extracted from the target video V. The target feature sequence is then scaled using methods such as linear interpolation to a fixed length L, resulting in a preprocessed target feature sequence.

[0146] The target feature sequence extracted from the target video is input into the action localization network.

[0147] The target feature sequence is processed using a basic coding network to obtain the video features of the target video. Specifically, the target feature sequence can be processed sequentially using the convolutional layer (Conv1d, k: 3, out: 256, ReLU) of the basic coding network, the target feature sequence fusion module (i.e., the feature processing method), and the convolutional layer (Conv1d, k: 3, out: 256, ReLU) to obtain the video features. Here, Conv1d represents a one-dimensional convolution, k represents the size of the convolutional kernel, out represents the dimension of the output features of the convolutional layer, and ReLU represents a non-linear activation function. Convolution processing and activation operations can be used to extract effective information from the features of the target video.

[0148] Next, the video features of the target video are input into the start prediction network, end prediction network, and action prediction network for processing.

[0149] The video features are processed using a start prediction network. Specifically, the start prediction network's convolutional layer (Conv1d, k: 3, out: 256, ReLU), target feature sequence fusion module (i.e., feature processing method), and convolutional layer (Conv1d, k: 3, out: 1, Sigmoid) are used sequentially to process the video features, yielding the first estimated values ​​(Start scores) for each candidate video frame as the start video frame. Here, Sigmoid represents the Sigmoid activation function. The first estimated values ​​(Start scores) of the start video frames can be represented by the matrix S∈R. 1×L express.

[0150] In some implementations, when the target feature sequence fusion module processes video features, the convolution result of the convolutional layer (Conv1d, k: 3, out: 256, Relu) on the video features can be used as a new target feature sequence to obtain new video features.

[0151] The video features are processed using an end prediction network. Specifically, the video features can be processed sequentially using the convolutional layer (Conv1d, k: 3, out: 256, ReLU), the target feature sequence fusion module (i.e., the feature processing method), and the convolutional layer (Conv1d, k: 3, out: 1, Sigmoid) of the end prediction network. This yields a second estimate (End scores) for each candidate video frame as the end video frame. This second estimate (End scores) can be represented by the matrix E∈R. 1×L express.

[0152] Then, the video features are processed using an action prediction network. Specifically, the video features are processed sequentially using the following convolutional layers: (Conv1d, k: 1, out: 256, ReLU), boundary matching module, convolutional layer (Conv3d, k: ns, out: 512, ReLU), convolutional layer (Conv2d, k: 1, out: 128, ReLU), convolutional layer (Conv2d, k: 3, out: 128, ReLU), convolutional layer (Conv2d, k: 3, out: 128, ReLU), and convolutional layer (Conv2d, k: 1, out: 2, Sigmoid). This yields the first and second evaluation values ​​(Existence & Reg) for the action to be detected. Here, Conv2d represents a two-dimensional convolution, and Conv3d represents a three-dimensional convolution. ns represents the number of samples in the boundary matching module.

[0153] In this embodiment, the first evaluation value indicates whether there is an action to be detected, which can be represented by the matrix EX∈R. L×L The second evaluation value represents the regression of the action to be detected, which can be represented by the matrix C∈R. L×L express.

[0154] In some implementations, the boundary matching module in the action prediction network described above can be a Boundary Matching Layer, that is, it adopts the idea of ​​a Boundary-Matching Network (BMN). The boundary matching module can convert the extracted feature sequence into a BM confidence map.

[0155] In some implementations, the boundary matching module may designate candidate video frames whose first estimated value is greater than a first estimated threshold as candidate start video frames, and candidate video frames whose second estimated value is greater than a second estimated threshold as candidate end video frames. The candidate start video frames and candidate end video frames constitute at least one target video segment.

[0156] The boundary matching module can iterate through the first estimated value S∈R of the initial video frame. 1×L The second estimated value E∈R of the end video frame 1×L The extreme points are identified and marked. The indices of all extreme points in S are used as the starting points for all predictions, and the indices of all extreme points in E are used as the ending points for all predictions. All starting points in S are matched with all ending points in E, meaning the ending point index in E must be greater than the starting point index in S. Candidate starting and ending video frames that meet the conditions are combined to form at least one target video segment.

[0157] Through the above motion localization network, the video localization information can be obtained as follows:

[0158] .

[0159] in, This indicates the number of actions to be detected in the target video. These represent the starting video frame, ending video frame, first estimate, and second estimate of the nth action to be detected, respectively.

[0160] In some implementations, the first estimate of the video location information can be... Multiply by the second estimate To obtain the confidence level of the action to be detected.

[0161] In some implementations, after obtaining the confidence scores of all detected actions in the target video, duplicate action sequences can be removed using a non-maximum suppression algorithm.

[0162] In some implementations, the video localization information following the aforementioned repetitive action sequence can be input into an action classification network to classify each action to be detected. This action classification network is pre-trained. The processing of the target video (i.e., segments containing action) by the classification network during training is consistent with the usage process. Furthermore, the action segments are labeled with annotation information used to identify the category to which the action belongs within the segment.

[0163] In some embodiments, the action localization network described above can be trained. During training, the processing of sample videos by the action localization network is consistent with the processing of target videos in the usage process described above. The specific training method is as follows:

[0164] The sample videos contain annotation information, and the annotation information corresponding to the sample videos is as follows:

[0165] .

[0166] Where, N g v represents the number of actual actions present in the sample video. n s and v n e c represents the start and end video frame numbers of the nth action. n This indicates the category to which the nth action belongs. The annotation information depends on the task of the action localization network. If, during use, the action localization network is used to predict the estimated values ​​of the start and end frames of the target video segment, then the annotation information includes the estimated values ​​of each video frame in the sample video; if, during use, the action localization network is also used to predict evaluation values, then the annotation information also includes the evaluation values ​​of each video frame in the sample video. The format of the annotation information is consistent with the format of the localization information predicted by the action localization network.

[0167] Extract the target feature sequence from the sample video, then scale the features to a fixed length L, and annotate the information. The process also involves scaling, whereby the scaled target feature sequence is input into the action localization network for processing, predicting a first estimate of the starting video frame, a second estimate of the ending video frame, a first evaluation value, and a second evaluation value. This processing procedure can be specifically referred to in the above embodiments, and will not be elaborated upon here.

[0168] When training the action localization network, the initial learning rate is between 0.001 and 0.1, and the learning rate decay strategy is to multiply by 0.1 every 10 epochs. The loss function consists of four parts, namely the first estimate prediction loss. Second estimate predicts loss First assessment value predicts loss Second evaluation value predicts loss The prediction loss of the action localization network can be expressed as:

[0169]

[0170] in, - The weights used to adjust each loss can be selected based on the actual situation during training.

[0171] In conjunction with the above embodiments, this application also provides a feature processing apparatus. Please refer to... Figure 9 , Figure 9 This is a schematic diagram of the structure of an embodiment of the feature processing device of this application. The feature processing device 30 includes an acquisition module 31, a division module 32, a first fusion module 33, and a second fusion module 34. The acquisition module 31, the division module 32, the first fusion module 33, and the second fusion module 34 are connected.

[0172] The acquisition module 31 is used to determine the target feature sequence of the target video based on the features of each video frame in the target video.

[0173] The segmentation module 32 is used to segment the target feature sequence according to the temporal dimension to obtain multiple sub-feature sequences, wherein each sub-feature sequence includes features of at least one video frame.

[0174] The first fusion module 33 is used to perform a first fusion process on each of the multiple sub-feature sequences along the time sequence direction to obtain the intermediate features corresponding to each sub-feature sequence.

[0175] The second fusion module 34 is used to perform a second fusion process on multiple intermediate features to obtain the video features of the target video.

[0176] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.

[0177] In conjunction with the above embodiments, this application also provides a motion positioning device. Please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of an embodiment of the motion positioning device of this application. The motion positioning device 40 includes a feature module 41, a target module 42, and a positioning module 43. The feature module 41, the target module 42, and the positioning module 43 are connected.

[0178] The feature module 41 is used to determine the features of the action to be detected in the target video as the target feature sequence.

[0179] The target module 42 is used to process the target feature sequence using feature processing methods to obtain the video features of the target video.

[0180] The positioning module 43 is used to determine the video frame positioning information of the action to be detected in the target video based on the video features of the target video; wherein, the video frame positioning information is used to indicate the video frame in the target video that contains the action to be detected.

[0181] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.

[0182] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 50 includes a memory 51 and a processor 52, wherein the memory 51 and the processor 52 are coupled to each other. The memory 51 stores program data, and the processor 52 is used to execute the program data to implement the steps in any of the embodiments of the feature processing method and / or action positioning method described above.

[0183] In this embodiment, processor 52 can also be referred to as CPU (Central Processing Unit). Processor 52 may be an integrated circuit chip with signal processing capabilities. Processor 52 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 52 can be any conventional processor.

[0184] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.

[0185] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a storage device. Please refer to [link to relevant documentation]. Figure 12 , Figure 12 This is a schematic diagram of a storage device according to an embodiment of the present application. The storage device 60 stores program data 61 that can be executed by a processor. The program data 61 can be executed by the processor to implement the steps of any embodiment of the feature processing method and / or motion positioning method described above.

[0186] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.

[0187] In this embodiment, the storage device 60 can be a medium that can store program data 61, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Alternatively, it can be a server that stores the program data 61. The server can send the stored program data 61 to other devices for execution, or it can run the stored program data 61 itself.

[0188] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0189] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0190] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0191] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage device, which is a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application.

[0192] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0193] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A feature processing method, characterized in that, The method includes: Based on the features of each video frame in the target video, the target feature sequence of the target video is determined; The target feature sequence is divided according to the time dimension to obtain multiple sub-feature sequences, wherein each sub-feature sequence includes features of at least one video frame; The first fusion process is performed on each sub-feature sequence of the plurality of sub-feature sequences along the temporal direction to obtain intermediate features corresponding to each sub-feature sequence. This includes: performing the first fusion process on each sub-feature sequence of the plurality of sub-feature sequences along the first temporal direction to obtain multiple fused features corresponding to each of the plurality of sub-feature sequences; and performing the first fusion process on each fused feature of the plurality of fused features along the second temporal direction to obtain multiple intermediate features corresponding to each of the plurality of fused features. The first temporal direction is forward, and the second temporal direction is reverse. A second fusion process is performed on multiple intermediate features to obtain the video features of the target video.

2. The method according to claim 1, characterized in that, The first fusion process is performed on each sub-feature sequence of the plurality of sub-feature sequences along the first temporal direction to obtain a plurality of fused features corresponding to the plurality of sub-feature sequences, including: Each of the sub-feature sequences is sequentially used as the first target sub-feature according to the first temporal direction; Based on the first target sub-feature and the first reference sub-feature, a first feature to be processed is obtained, wherein the first reference sub-feature is the previous sub-feature sequence of the first target sub-feature in the first temporal direction; The first feature to be processed is convolved using the first convolution kernel to obtain the fused feature corresponding to the first target sub-feature. The first fusion process is performed on each of the plurality of fusion features along the second time sequence direction to obtain a plurality of intermediate features corresponding to the plurality of fusion features, including: According to the second temporal direction, the fused features corresponding to each of the sub-feature sequences are sequentially used as the second target sub-features; Based on the second target sub-feature and the second reference sub-feature, a second feature to be processed is obtained, wherein the second reference sub-feature is the previous fused feature of the second target sub-feature in the second temporal direction; The second convolution kernel is used to perform convolution processing on the second feature to be processed to obtain the intermediate feature corresponding to the second target sub-feature.

3. The method according to claim 2, characterized in that, The step of obtaining the first feature to be processed based on the first target sub-feature and the first reference sub-feature includes: If the first target sub-feature is the first sub-feature sequence in the first temporal direction, then the first target sub-feature is taken as the first feature to be processed. If the first target sub-feature is not the first sub-feature sequence in the first temporal direction, then the first target sub-feature and the fused feature corresponding to the first reference sub-feature are superimposed to obtain the first feature to be processed. The step of obtaining the second feature to be processed based on the second target sub-feature and the second reference sub-feature includes: If the second target sub-feature is the first fusion feature in the second temporal direction, then the second target sub-feature is used as the second feature to be processed; If the second target sub-feature is not the first fusion feature in the second temporal direction, then the second target sub-feature and the intermediate feature corresponding to the second reference sub-feature are superimposed to obtain the second feature to be processed.

4. The method according to claim 2, characterized in that, The first convolution kernel is different from the second convolution kernel, and the first convolution kernel and the second convolution kernel may have the same or different sizes.

5. The method according to claim 1, characterized in that, The second fusion process, which involves fusing the multiple intermediate features to obtain the video features of the target video, includes: The multiple intermediate features are subjected to a second fusion process according to the time sequence dimension to obtain the video features of the target video.

6. The method according to claim 1, characterized in that, Determining the target feature sequence of the target video based on the features of each video frame in the target video includes: Using a sliding window, select several groups of video frames from the target video, where each group of video frames includes one or more video frames. Extract the color features or optical flow features of each group of video frames to obtain the target feature sequence.

7. A motion localization method, characterized in that, The method includes: The features of the action to be detected in the target video are determined as the target feature sequence; The target feature sequence is processed using the method described in any one of claims 1 to 6 to obtain the video features of the target video; Based on the video features of the target video, determine the video frame location information of the action to be detected in the target video; wherein, the video frame location information is used to indicate that the target video contains the video frame of the action to be detected.

8. The method according to claim 7, characterized in that, The step of determining the video frame location information of the action to be detected in the target video based on the target features includes: Based on the video features of the target video, a first estimate of each candidate video frame in the target video as the starting video frame and a second estimate of each candidate video frame as the ending video frame are predicted. Based on the first estimate and the second estimate of each of the candidate video frames, a portion of the candidate video frames are selected; Based on the selected candidate video frames, the location information of the video frames is obtained.

9. The method according to claim 8, characterized in that, The step of predicting a first estimate of each candidate video frame as a start video frame and a second estimate of each candidate video frame as an end video frame based on the video features of the target video includes: The video features of the target video are subjected to a first convolution process and a second convolution process respectively to obtain the first target feature and the second target feature. The first target feature and the second target feature are respectively determined as new target feature sequences, and the new target feature sequences are processed using the method described in any one of claims 1 to 6 to obtain new first target features and second target features; The new first target feature and second target feature are predicted to obtain the first estimate and the second estimate for each of the candidate video frames.

10. The method according to claim 8, characterized in that, The step of selecting a subset of the candidate video frames based on the first and second estimates of each candidate video frame includes: The candidate video frames whose first estimated value satisfies the first condition are determined as candidate start video frames; and the candidate video frames whose second estimated value satisfies the second condition are determined as candidate end video frames.

11. The method according to claim 8, characterized in that, The step of obtaining the video frame location information based on a selected subset of the candidate video frames includes: From the selected candidate start video frames and candidate end video frames, at least one candidate start video frame and candidate end video frame corresponding to a target video segment are selected; wherein, each target video segment corresponds to one candidate start video frame and one candidate end video frame, and the candidate start video frame corresponding to the target video segment is earlier in the time sequence of the target video than the candidate end video frame corresponding to the target video segment. Based on the candidate start video frame and candidate end video frame corresponding to the at least one target video segment, the video frame positioning information is determined.

12. The method according to claim 11, characterized in that, The step of determining the video frame positioning information based on the candidate start video frame and candidate end video frame corresponding to the at least one target video segment includes: Based on the video features of the target video, obtain the region features of each of the target video segments; Based on the corresponding regional features, each target video segment is identified to obtain the action evaluation value of each target video segment.

13. The method as described in claim 12, characterized in that, After identifying each target video segment based on corresponding regional features and obtaining the action evaluation value of each target video segment, the method further includes: If the action evaluation value includes a first evaluation value indicating whether the action to be detected exists, then the target video segment whose first evaluation value is greater than a first threshold is determined as the final target video segment; or, If the action evaluation value includes a second evaluation value representing the regressivity of the action to be detected, then the target video segment whose second evaluation value is greater than a second threshold is determined as the final target video segment; or, If the action evaluation value includes the first evaluation value and the second evaluation value, then the target video segment whose first evaluation value is greater than the first threshold and whose second evaluation value is greater than the second threshold is determined as the final target video segment.

14. The method as described in claim 7, characterized in that, After obtaining the video frame location information of the action to be detected in the target video, the method further includes: Based on the video frame positioning information, determine the video frame in the target video corresponding to the video frame positioning information; The action to be detected is identified based on the video frame corresponding to the video frame positioning information.

15. The method according to claim 7, characterized in that, The steps of processing the target feature sequence to obtain the video features of the target video and locating the video features of the target video to obtain the video frame location information of the action to be detected in the target video are implemented by an action localization network.

16. A computer device, characterized in that, The method includes a memory and a processor coupled to each other, the memory storing program data and the processor executing the program data to implement the steps of the method according to any one of claims 1 to 15.

17. A storage device, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Group abnormal behavior identification method and system based on multi-scale time information fusion

    CN112016500A