Video scene retrieval method, device, electronic device and readable storage medium

By performing dense deep learning feature fusion and aggregation on multi-frame images of video sequences, global feature descriptors are generated, which solves the problems of environmental changes and dynamic object interference in scene re-identification, and improves accuracy and user experience.

CN114743139BActive Publication Date: 2025-07-22CHANGCHUN YIHANG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210339794.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-07-22
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

The prior art has low accuracy due to environmental changes and dynamic object interference in scene re-identification, which affects the user experience.

Method used

By obtaining the multi-frame image of the current video sequence, dense deep learning feature maps are extracted, and time-domain feature fusion and spatiotemporal feature aggregation are performed to generate a global feature descriptor, which is used to search in the global database.

Benefits of technology

It improves the accuracy of scene re-identification, reduces the impact of environmental changes and local occlusion, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743139B_ABST
    Figure CN114743139B_ABST
Patent Text Reader

Abstract

The present application relates to a video scene retrieval method, apparatus, electronic device, and readable storage medium, and relates to the field of computer technology. The method includes: obtaining a current video sequence, where the current video sequence includes multiple frames of images, then respectively extracting dense deep learning feature maps corresponding to each frame of the multiple frames of images, then respectively performing time-domain feature fusion based on the dense deep learning feature maps corresponding to each frame of the multiple frames of images to obtain respective fused features, then performing spatio-temporal feature aggregation processing based on the fused features corresponding to each frame of the multiple frames of images to obtain a global feature descriptor corresponding to the current video sequence, and then performing retrieval from a global database based on the global feature descriptor corresponding to the current video sequence to obtain a first preset number of video sequences. A video scene retrieval method, apparatus, electronic device, and readable storage medium provided by the present application can improve the accuracy of the retrieved video sequences, and thus can enhance the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a video scene retrieval method, apparatus, electronic device, and readable storage medium. Background Art

[0002] In recent years, with the emergence of application scenarios such as autonomous memory parking, intelligent logistics carts, restaurant intelligent robot food delivery, and unmanned aerial vehicle autonomous cruising, it is very important to identify the scenes that have been reached. In these application scenarios, usually when performing the task for the first time (such as parking a car in its own parking space), a correct movement path is pre-planned manually and a scene map is established. When performing the task autonomously later, the intelligent robot or the autonomous driving vehicle perceives which position in the scene map it is in according to the currently observed scene, and then moves along the pre-planned path autonomously, or navigates autonomously to avoid obstacles according to the scene map. Therefore, the accuracy of scene re-identification is crucial for the operation of subsequent positioning and path tracking navigation algorithm modules.

[0003] In the above application scenarios, during the execution of the autonomous navigation task and the establishment of the scene map, there may be a long time span in the middle, resulting in a large change in the surrounding environment of the scene. For example, it is in the morning when building the map and at night when performing autonomous navigation; it is sunny when building the map, and it is rainy, foggy or snowy when performing autonomous navigation, and there may even be a situation across seasons, resulting in a large change in the appearance of the scenes observed in the two cases. In addition, the scenes of these applications are often very complex. For example, during map building and autonomous navigation, there are interferences from dynamic objects such as pedestrians and vehicles, which further increases the difference in the appearance of the scenes observed twice, and even these dynamic objects may cause partial occlusion of the scene; at the same time, the repeated appearance of some empty scenes or objects with the same texture is also a major challenge, such as empty parking lots, similar design styles of different garages, and almost identical lamp posts and fences repeatedly appearing on the road.

[0004] The inventors found during the research process that the above situation may lead to a low accuracy of scene re-identification, and thus a poor user experience. Summary of the Invention

[0005] The purpose of the present application is to provide a video scene retrieval method, apparatus, electronic device, and readable storage medium, which are used to solve at least one of the above technical problems.

[0006] The above invention purpose of the present application is achieved through the following technical solutions:

[0007] In a first aspect, a video scene retrieval method is provided, including:

[0008] Obtain a current video sequence, where the current video sequence includes multiple frames of images;

[0009] Extract dense deep learning feature maps corresponding to each frame image from the multi-frame images respectively;

[0010] Perform temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image respectively to obtain the fused features of each;

[0011] Perform spatio-temporal feature aggregation processing based on the fused features corresponding to each frame image respectively to obtain the global feature descriptor corresponding to the current video sequence;

[0012] Retrieve from the global database based on the global feature descriptor corresponding to the current video sequence to obtain the first preset number of video sequences.

[0013] In a possible implementation manner, the performing temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image respectively to obtain the fused features of each includes: performing temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image respectively through a self-attention mechanism to obtain the fused features of each.

[0014] In another possible implementation manner, the performing spatio-temporal feature aggregation processing based on the fused features corresponding to each frame image respectively to obtain the global feature descriptor corresponding to the current video sequence includes:

[0015] Perform splicing processing on the temporal feature maps corresponding to each frame image respectively to obtain the spliced feature map;

[0016] Perform pointwise convolution processing on the spliced feature map to obtain the convolution processing result;

[0017] Perform normalization processing on the convolution processing result to obtain the result after normalization processing;

[0018] Determine the global feature descriptor corresponding to the current video sequence based on the result after normalization processing and the spliced feature map.

[0019] In another possible implementation manner, the spliced feature map includes multiple feature points;

[0020] The determining the global feature descriptor corresponding to the current video sequence based on the result after normalization processing and the spliced feature map includes:

[0021] Perform clustering processing on the multiple feature points to obtain at least one clustering center;

[0022] Determine the distances between each feature point and each clustering center respectively, determine the distance information corresponding to each clustering center, and the distance information corresponding to any clustering center is the distances between each feature point and the any clustering center;

[0023] Determine the global representation corresponding to each clustering cluster based on the distance information corresponding to each clustering center and the result after the normalization process;

[0024] Perform regularization processing on the global representations corresponding to each clustering cluster;

[0025] Concatenate the global representations after the regularization processing;

[0026] Perform regularization processing on the concatenated global representations to obtain the global feature descriptor corresponding to the current video sequence.

[0027] In another possible implementation, after separately extracting the dense deep learning feature maps corresponding to each frame image from the multiple frame images, the following steps are further included:

[0028] Perform regional feature extraction on the dense deep learning feature maps corresponding to each frame image respectively to obtain the multi-scale regional features corresponding to each of them;

[0029] Perform regional matching based on the multi-scale regional features corresponding to each of them to obtain the spatio-temporal feature descriptor corresponding to the current video sequence;

[0030] Perform regional matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence to obtain the second preset number of video sequences.

[0031] In another possible implementation, performing regional feature extraction on the dense deep learning feature map corresponding to any one frame image to obtain the multi-scale regional feature corresponding to the any one frame image includes:

[0032] Determine a weighted residual feature map based on the dense deep learning feature map corresponding to the any one frame image;

[0033] Divide the weighted residual feature map into multiple regional blocks;

[0034] Determine the regional feature representations corresponding to each regional block respectively to obtain the multi-scale regional feature corresponding to the any one frame image.

[0035] In another possible implementation, the determining the weighted residual feature map based on the dense deep learning feature map corresponding to the any one frame image includes:

[0036] Perform pointwise convolution processing on the dense deep learning feature map corresponding to the any one frame image to obtain a convolution result;

[0037] Perform normalization processing on the convolution result to obtain a normalization result;

[0038] Determine the weighted residual feature map based on the normalization result and the distance information corresponding to each clustering center.

[0039] In another possible implementation, the multi-scale region features are characterized by region descriptors;

[0040] Performing region matching based on the respective corresponding multi-scale region features to obtain a spatio-temporal feature descriptor corresponding to the current video sequence, including:

[0041] Perform region feature matching between the region descriptors corresponding to each frame image in the current video sequence and the region descriptors corresponding to each other frame image in the current video sequence to obtain a region matching result corresponding to the current video sequence;

[0042] Select region descriptors that meet preset conditions from the region matching result corresponding to the current video sequence as the region descriptors corresponding to the current video sequence;

[0043] Perform region feature matching between the region descriptors corresponding to the current video sequence and each video sequence stored in the global database to obtain region matching results corresponding to the current video sequence and each of the video sequences.

[0044] In another possible implementation, perform region feature matching between the region descriptor corresponding to any frame image in the current video sequence and the region description corresponding to any other frame image in the current video sequence to obtain a corresponding matching result, including:

[0045] Determine the distances between the region descriptor corresponding to the any frame image and each region descriptor in each region of any other frame image in the current video sequence;

[0046] Perform region feature matching between the region descriptor corresponding to any frame image in the current video sequence and the region description corresponding to any other frame image in the current video sequence through the following formula to obtain a corresponding matching result:

[0047] ;

[0048] where the element D in the matrix ij represents the distance between the i-th region descriptor in the Tm-th frame and the j-th region descriptor in the Tn-th frame, and the matrix D is used to represent the distances between all region descriptors in the Tm-th frame and all region descriptors in the Tn-th frame of the video sequence. The Tm-th frame represents the any frame image, and the Tn is used to represent any other frame image in the current video sequence; D ij k represents the element with the minimum distance value in the j-th column of the matrix D, D i kj The element with the smallest distance value in the i-th row of the characterization matrix D, where t is used to represent the threshold parameter, and the matching items (i, j) that meet the conditions constitute the matching set P between the Tm frame and the Tn frame mn 。

[0049] In another possible implementation, selecting the region descriptor that meets the first preset condition from the region matching results corresponding to any of the regions as the region descriptor corresponding to any of the regions includes:

[0050] Determining the average value of the distances that meet the preset conditions;

[0051] Based on the average value of the distances that meet the first preset condition, determining the region descriptor corresponding to any of the regions.

[0052] In another possible implementation, the determining the region descriptor corresponding to any of the regions based on the average value of the distances that meet the first preset condition includes:

[0053] Based on the average value of the distances that meet the first preset condition, and determining the region descriptor corresponding to any of the regions through the following formula:

[0054] ;

[0055] where x is the region belonging to S i in, P x is the set of matching items corresponding to region x in the frame matching set where region x is located, D x is all D extracted from the set of matching items P x set, ij set, is the average value of the elements in the D ij set, and x' is used to represent the region descriptors of all regions determined from the set S i in.

[0056] In another possible implementation, performing region feature matching between the region descriptor corresponding to the current video sequence and any video sequence includes:

[0057] Performing region feature matching between the region descriptors corresponding to each frame image in the current video sequence and each frame image in any video sequence respectively.

[0058] In another possible implementation, performing region feature matching between the region descriptor corresponding to each frame image and the region descriptor corresponding to any frame image includes:

[0059] Based on the region descriptors corresponding to each region in each frame of the image and the region descriptors corresponding to each region in any one frame of the image, determine the distance vector corresponding to each frame of the image. The distance vector corresponding to each frame of the image contains multiple elements, and any one element is the distance between the region descriptor corresponding to any region in each frame of the image and the region descriptor corresponding to any region in any one frame of the image.

[0060] In another possible implementation manner, the region matching of the first preset number of video sequences based on the spatio-temporal feature descriptors corresponding to the current video sequence to obtain the second preset number of video sequences includes:

[0061] Based on the region matching results corresponding to the current video sequence and each of the video sequences respectively, determine the spatial consistency scores corresponding to the current video sequence and each of the video sequences respectively. Each of the video sequences belongs to the first preset number of video sequences;

[0062] Based on the spatial consistency scores corresponding to the current video sequence and each of the video sequences respectively, reorder the first preset number of video sequences;

[0063] Extract the second preset number of video sequences from the first preset number of video sequences after sorting.

[0064] In another possible implementation manner, determining the spatial consistency score corresponding to the current video sequence and any one video sequence includes:

[0065] Determine the spatial consistency scores between each frame of the image in the current video sequence and each frame of the image in any one of the video sequences;

[0066] Determine the weight information of each frame of the image in the current video sequence;

[0067] Based on the weight information of each frame of the image in the current video sequence and the spatial consistency scores between each frame of the image in the current video sequence and each frame of the image in any one of the video sequences respectively, determine the spatial consistency score corresponding to the current video sequence and any one video sequence.

[0068] In another possible implementation manner, the determination of the spatial consistency score between each frame of the image in the current video sequence and any one frame of the image includes:

[0069] Determine the region matching spatial consistency scores for regions of each size;

[0070] Determine the weight information corresponding to regions of each size;

[0071] Determine the spatial consistency score between each frame image in the current video sequence and any frame image based on the regional matching spatial consistency scores of each size and the weight information corresponding to the regions of each size.

[0072] In another possible implementation, determining the regional matching spatial consistency score for any size includes:

[0073] Determine the regional matching spatial consistency score for any size through the following formula:

[0074] ;

[0075] where SS p represents the regional matching spatial consistency score for regions of scale size p, n p represents the number of regional blocks of scale size p extracted from this frame image, P p is the regional matching set of regional features of scale size p, (r p , c p ) is the matching offset stored in P p ; and respectively represent the average column offset and average row offset in the set P p ; i, j represent the numbers during the traversal of the set P p , and the dist(·) function is the distance function, and max(·) is the maximum value function;

[0076] Among them, based on the regional matching spatial consistency scores of each size and the weight information corresponding to the regions of each size, and determining the spatial consistency score between each frame image in the current video sequence and any frame image through the following formula, includes:

[0077] ;

[0078] where SS represents the spatial consistency score between each frame image in the current video sequence and any frame image, i is the traversal of the scale set, n s is the number of scales, w i is the weight information corresponding to size i, and w i ∈[0,1].

[0079] In another possible implementation, based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any video sequence, determining the spatial consistency score corresponding to the current video sequence and any video sequence, includes:

[0080] Based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any one of the video sequences, the spatial consistency score corresponding to the current video sequence and any one of the video sequences is determined by the following formula:

[0081] ;

[0082] where VSS represents the spatial consistency score corresponding to the current video sequence and any one of the video sequences, V ref belongs to a video sequence with the first preset number, m is used to represent the frames in the current video sequence, and k is used to represent the frames in V ref ; is the weight information used to represent m.

[0083] In a second aspect, a video scene retrieval device is provided, including:

[0084] An acquisition module, configured to acquire a current video sequence, where the current video sequence includes multiple frame images;

[0085] A feature map extraction module, configured to respectively extract dense deep learning feature maps corresponding to each frame image from the multiple frame images;

[0086] A temporal feature fusion module, configured to respectively perform temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image to obtain respective fused features;

[0087] A spatio-temporal feature aggregation processing module, configured to perform spatio-temporal feature aggregation processing based on the fused features corresponding to each frame image to obtain a global feature descriptor corresponding to the current video sequence;

[0088] A first retrieval module, configured to retrieve from a global database based on the global feature descriptor corresponding to the current video sequence to obtain a video sequence with the first preset number.

[0089] In a possible implementation manner, when the temporal feature fusion module respectively performs temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image to obtain respective fused features, it is specifically configured to:

[0090] Perform temporal feature fusion based on the dense deep learning feature maps corresponding to each frame image through a self-attention mechanism to obtain respective fused features.

[0091] In another possible implementation manner, when the spatio-temporal feature aggregation processing module performs spatio-temporal feature aggregation processing based on the fused features corresponding to each frame image to obtain a global feature descriptor corresponding to the current video sequence, it is specifically configured to:

[0092] Perform splicing processing on the time-domain feature maps corresponding to each frame image to obtain a spliced feature map;

[0093] Perform point-by-point convolution processing on the spliced feature map to obtain a convolution processing result;

[0094] Perform normalization processing on the convolution processing result to obtain a normalized result;

[0095] Based on the normalized result and the spliced feature map, determine the global feature descriptor corresponding to the current video sequence.

[0096] In another possible implementation, the spliced feature map includes multiple feature points;

[0097] When the spatio-temporal feature aggregation processing module determines the global feature descriptor corresponding to the current video sequence based on the normalized result and the spliced feature map, it is specifically used for:

[0098] Perform clustering processing on the multiple feature points to obtain at least one clustering center;

[0099] Determine the distances between each feature point and each clustering center respectively, and determine the distance information corresponding to each clustering center. The distance information corresponding to any clustering center is the distances between each feature point and the any clustering center;

[0100] Based on the distance information corresponding to each clustering center and the normalized result, determine the global representations corresponding to each clustering cluster respectively;

[0101] Perform regularization processing on the global representations corresponding to each clustering cluster respectively;

[0102] Perform splicing processing on the regularized global representations;

[0103] Perform regularization processing on the spliced global representations to obtain the global feature descriptor corresponding to the current video sequence.

[0104] In another possible implementation, the device further includes: a multi-scale region feature extraction module, a spatio-temporal region feature matching module, and a second retrieval module, where,

[0105] The multi-scale region extraction module is used to perform region feature extraction on the dense deep learning feature maps corresponding to each frame image respectively to obtain the corresponding multi-scale region features;

[0106] The spatio-temporal region feature matching module is used to perform region matching based on the respective corresponding multi-scale region features to obtain a spatio-temporal feature descriptor corresponding to the current video sequence;

[0107] The second retrieval module is used to perform region matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence to obtain a second preset number of video sequences.

[0108] In another possible implementation manner, when the multi-scale region feature extraction module performs region feature extraction based on the dense deep learning feature map corresponding to any frame image to obtain the multi-scale region feature corresponding to the any frame image, it specifically is used for:

[0109] Determine a weighted residual feature map based on the dense deep learning feature map corresponding to the any frame image;

[0110] Divide the weighted residual feature map into multiple region blocks;

[0111] Determine the region feature representations respectively corresponding to each region block to obtain the multi-scale region feature corresponding to the any frame image.

[0112] In another possible implementation manner, when the multi-scale region feature extraction module determines the weighted residual feature map based on the dense deep learning feature map corresponding to the any frame image, it specifically is used for:

[0113] Perform pointwise convolution processing on the dense deep learning feature map corresponding to the any frame image to obtain a convolution result;

[0114] Perform normalization processing on the convolution result to obtain a normalization result;

[0115] Determine the weighted residual feature map based on the normalization result and the distance information corresponding to each clustering center.

[0116] In another possible implementation manner, the multi-scale region feature is represented by a region descriptor;

[0117] When the spatio-temporal region feature matching module performs region matching based on the respective corresponding multi-scale region features to obtain a spatio-temporal feature descriptor corresponding to the current video sequence, it specifically is used for:

[0118] Perform region feature matching on the region descriptors corresponding to each frame image in the current video sequence and the region descriptors corresponding to other respective frame images in the current video sequence to obtain a region matching result corresponding to the current video sequence;

[0119] Select the region descriptors that meet the preset conditions from the region matching results corresponding to the current video sequence as the region descriptors corresponding to the current video sequence;

[0120] Perform region feature matching between the region descriptors corresponding to the current video sequence and each video sequence stored in the global database respectively to obtain the region matching results corresponding to the current video sequence and each of the video sequences.

[0121] In another possible implementation, when the spatio-temporal region feature matching module performs region feature matching between the region descriptor corresponding to any frame image in the current video sequence and the region descriptor corresponding to any other frame image in the current video sequence to obtain the corresponding matching result, it specifically is used for:

[0122] Determine the distances between the region descriptor corresponding to the any frame image and each region descriptor in each region of any other frame image in the current video sequence;

[0123] Perform region feature matching between the region descriptor corresponding to any frame image in the current video sequence and the region descriptor corresponding to any other frame image in the current video sequence through the following formula to obtain the corresponding matching result:

[0124] ;

[0125] Wherein, the element D in the matrix ij represents the distance between the i-th region descriptor in the Tm frame and the j-th region descriptor in the Tn frame, and the matrix D is used to represent the distances between all region descriptors in the Tm frame and all region descriptors in the Tn frame of the video sequence. The Tm frame represents the any frame image, and the Tn is used to represent any other frame image in the current video sequence; D ij k represents the element with the smallest distance value in the j-th column of the matrix D, D i k j represents the element with the smallest distance value in the i-th row of the matrix D, and t is used to represent the threshold parameter. The matching items (i, j) that meet the conditions constitute the matching set P between the Tm frame and the Tn frame mn .

[0126] In another possible implementation, when the spatio-temporal region feature matching module selects the region descriptors that meet the preset conditions from the region matching results corresponding to the any region as the region descriptors corresponding to the any region, it specifically is used for:

[0127] Determine the average value of the distances that meet the preset conditions;

[0128] Determine the region descriptor corresponding to any region based on the average value of the distances that meet the preset conditions.

[0129] In another possible implementation manner, when the spatio-temporal region feature matching module determines the region descriptor corresponding to any region based on the average value of the distances that meet the first preset condition, it is specifically configured to:

[0130] Based on the average value of the distances that meet the first preset condition, determine the region descriptor corresponding to any region through the following formula:

[0131] ;

[0132] where x is the region belonging to S i in, P x is the set of matching items corresponding to region x in the frame matching set where region x is located, and D x is all D x extracted from the set of matching items P ij set, is the average value of the elements in the D ij set, and x' is used to represent the region descriptors of all regions determined from the set S i in.

[0133] In another possible implementation manner, when the spatio-temporal region feature matching module performs region feature matching between the region descriptor corresponding to the current video sequence and any video sequence, it is specifically configured to:

[0134] Perform region feature matching between the region descriptors corresponding to each frame image in the current video sequence and each frame image in the any video sequence respectively.

[0135] In another possible implementation manner, when the spatio-temporal region feature matching module performs region feature matching between the region descriptor corresponding to each frame image and the region descriptor corresponding to any frame image, it is specifically configured to:

[0136] Based on the region descriptors corresponding to each region in each frame image and the region descriptors corresponding to each region in the any frame image, determine the distance vector corresponding to each frame image. The distance vector corresponding to each frame image contains multiple elements, and any element is the distance between the region descriptor corresponding to any region in each frame image and the region descriptor corresponding to any region in the any frame image.

[0137] In another possible implementation manner, when the second retrieval module performs region matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence and obtains the second preset number of video sequences, it is specifically configured to:

[0138] Based on the region matching results corresponding to the current video sequence and each of the video sequences, determine the spatial consistency scores corresponding to the current video sequence and each of the video sequences, where each of the video sequences belongs to the first preset number of video sequences;

[0139] Based on the spatial consistency scores corresponding to the current video sequence and each of the video sequences, reorder the first preset number of video sequences;

[0140] Extract the second preset number of video sequences from the reordered first preset number of video sequences.

[0141] In another possible implementation manner, when the second retrieval module determines the spatial consistency score corresponding to the current video sequence and any one of the video sequences, it specifically is used for:

[0142] Determine the spatial consistency scores between each frame image in the current video sequence and each frame image in any one of the video sequences;

[0143] Determine the weight information of each frame image in the current video sequence;

[0144] Based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any one of the video sequences, determine the spatial consistency score corresponding to the current video sequence and any one of the video sequences.

[0145] In another possible implementation manner, when the second retrieval module determines the spatial consistency score between each frame image in the current video sequence and any one frame image, it specifically is used for:

[0146] Determine the region matching spatial consistency scores of each size;

[0147] Determine the weight information corresponding to each region of each size;

[0148] Based on the region matching spatial consistency scores of each size and the weight information corresponding to each region of each size, determine the spatial consistency score between each frame image in the current video sequence and any one frame image.

[0149] In another possible implementation manner, when the second retrieval module determines the region matching spatial consistency score of any one size, it specifically is used for:

[0150] Determine the region matching spatial consistency score of any one size through the following formula:

[0151] ;

[0152] where SSp Characterize the regional matching spatial consistency score of the region with scale size p, n p Characterize the number of region blocks with scale size p extracted from this frame of image, P p Is the regional matching set of the regional features with scale size p, (r p , c p ) is the matching offset stored in P p ; and Respectively characterize the average column offset and average row offset in the P p Set; i, j characterize the number during the traversal of the set P p , dist(·) function is the distance function, and max(·) is the maximum value function;

[0153] Among them, when the second retrieval module determines the spatial consistency score between each frame image in the current video sequence and any frame image based on the regional matching spatial consistency scores of the regions of each size and the weight information corresponding to the regions of each size, it specifically is used for:

[0154] ;

[0155] Among them, SS characterizes the spatial consistency score between each frame image in the current video sequence and any frame image, i is the traversal of the scale set, n s Is the number of scales, w i Is the weight information corresponding to size i, and w i ∈[0, 1].

[0156] In another possible implementation manner, when the second retrieval module determines the spatial consistency score corresponding to the current video sequence and any video sequence based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in the any video sequence, it specifically is used for:

[0157] Based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in the any video sequence, and determine the spatial consistency score corresponding to the current video sequence and any video sequence through the following formula:

[0158] ;

[0159] Among them, VSS characterizes the spatial consistency score corresponding to the current video sequence and any video sequence, V ref Belongs to the video sequences with the first preset number, m is used to characterize the frames in the current video sequence, and k is used to characterize V refthe frames in for characterizing the weight information of m.

[0160] In a third aspect, an electronic device is provided, and the electronic device includes:

[0161] one or more processors;

[0162] a memory;

[0163] one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute operations corresponding to the video scene retrieval method shown in any possible implementation manner of the first aspect.

[0164] In a fourth aspect, a computer-readable storage medium is provided, and the storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement the video scene retrieval method shown in any possible implementation manner of the first aspect.

[0165] In summary, the present application includes at least one of the following beneficial technical effects:

[0166] The present application provides a video scene retrieval method, apparatus, electronic device and readable storage medium. Compared with the related art, in the present application, temporal feature fusion is performed based on the dense deep learning feature maps respectively corresponding to the frames in the current video sequence, and then spatio-temporal feature aggregation processing is performed according to the fused features to obtain a global feature descriptor corresponding to the current video sequence, that is, the spatio-temporal features of the current video sequence can be reflected in the global feature descriptor corresponding to the current video sequence. Therefore, retrieval can be performed from the global database based on the global feature descriptor corresponding to the current video sequence, which can reduce the influence of changes in the surrounding environment of the scene and local occlusion, etc. on scene re-identification, thereby improving the accuracy of the retrieved video sequence and further enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0167] Figure 1 is a schematic flowchart of a method for video scene retrieval in an embodiment of the present application;

[0168] Figure 2 is a schematic structural diagram of a temporal feature fusion network based on a self-attention mechanism in an embodiment of the present application;

[0169] Figure 3 is a schematic diagram of the network model architecture of TemproalVLAD in an embodiment of the present application;

[0170] Figure 4 is an example diagram of a video scene retrieval in an embodiment of the present application;

[0171] Figure 5 It is a schematic structural diagram of a device for video scene retrieval in an embodiment of the present application;

[0172] Figure 6 It is a schematic structural diagram of a device of an electronic device in an embodiment of the present application. Detailed implementation manners

[0173] The following further describes the present application in detail with reference to the accompanying drawings.

[0174] This specific embodiment is only an explanation of the present application, and it does not limit the present application. After reading this specification, those skilled in the art can make modifications without creative contributions to this embodiment as needed, but as long as it is within the scope of the claims of the present application, it is protected by the patent law.

[0175] The embodiment of the present application provides a video scene retrieval method. The main purpose of vision-based scene retrieval is to find the observation information (image or video sequence) at the same geographical location when establishing a scene map based on the current observation information;

[0176] There are mainly three differences between vision-based scene retrieval and general image retrieval / video retrieval:

[0177] 1. The main benchmarks for measuring similarity in general image retrieval / video retrieval are "whether it is the same object category", "whether it has a similar appearance", etc., while the main benchmark for measuring similarity in vision-based scene retrieval is "whether it is the same geographical location". Even if there are large appearance differences due to external factors such as weather and season, as long as the positions of the two are close enough, the similarity should also be high;

[0178] 2. General image retrieval / video retrieval mainly targets foreground objects in the image, while vision-based scene retrieval mainly targets background regions in the image;

[0179] 3. General image retrieval / video retrieval can be performed offline, while vision-based scene retrieval is often applied in fields with strong real-time requirements, such as the relocalization and loop detection links in SLAM. Therefore, in addition to requiring a low complexity of the scene retrieval algorithm, it is also necessary to perform an efficient global representation of the observation information (image or video sequence) to make it easier to calculate and store, such as converting the observation information into a vector or a matrix.

[0180] In related technologies, most visual scene re-identification technologies calculate similarity based on single-frame images, such as methods like Bag of Words (BoW), Fisher Vectors (FV), and Vector of Locally Aggregated Descriptors (VLAD). However, for methods of scene retrieval based on single-frame images, when there is a change in perspective during two observations, the accuracy drops significantly. In addition, the video-based scene retrieval methods in related technologies mainly perform information aggregation based on the global representation of single-frame images, ignoring the spatio-temporal information of video sequences, and thus are still limited by the retrieval accuracy of single-frame images.

[0181] Aiming at the problem of low recall rate in image scene retrieval when there is a change in perspective, the embodiments of this application provide a method for extracting spatio-temporal hierarchical features from short video sequences. According to the currently observed video sequence, a video sequence located at the same geographical location is retrieved from another large video sequence database (which can be established during the first map construction). In the embodiments of this application, spatio-temporal hierarchical features are used to perform video scene retrieval from coarse to fine, and the specific details are as follows in the following embodiments:

[0182] (1) First, perform coarse-grained fast video scene retrieval: For each frame image in the video sequence, dense deep learning feature points are extracted by a neural network. Then, a self-attention mechanism is used to aggregate the feature points co-visible between different frames and optimize the descriptors. Finally, the TemproalVLAD layer proposed by us is used to cluster the feature points on multiple frames in the time domain, retain the unique observation information of each frame image, and remove the redundant observation information between frames, thereby generating a high-dimensional vector as the global representation of the video sequence, and using the distance between the global representations of video sequences for scene retrieval. This coarse-grained fast video scene retrieval branch has advantages such as fast retrieval speed and efficient storage.

[0183] (2) Then, perform fine-grained optimization and sorting. Use the regional block features on the feature map for regional matching, and construct an image pyramid to extract multi-scale regional features; for the regional matching results, use the relative offsets between all matching pairs to define the image similarity, thereby optimizing the retrieval sorting results of the coarse-grained branch. Among them, this fine-grained optimization and sorting branch has advantages such as high retrieval recall rate and robustness to perspective changes / local occlusions.

[0184] (3) The two branches of coarse-grained fast retrieval and fine-grained optimization and sorting share the CNN (Convolutional Neural Network) backbone network, thereby reducing the computational amount, lowering the algorithm complexity, and having advantages such as high real-time performance and small latency.

[0185] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.

[0186] In addition, the term "and / or" in this text is merely an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0187] The following further describes the embodiments of this application in detail with reference to the accompanying drawings of the specification.

[0188] The embodiments of this application provide a video scene retrieval method. This video scene retrieval method can be executed by an electronic device, which can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this.

[0189] It should be noted that the electronic device for executing the video scene retrieval method may further include: a driverless vehicle and a smart robot.

[0190] Furthermore, as Figure 1 shown, this method may include:

[0191] Step 101: Obtain the current video sequence.

[0192] For the embodiments of this application, the current video sequence is the video sequence to be retrieved for the scene. In the embodiments of this application, the current video sequence can be collected by a camera set in a driverless vehicle or a smart robot, or can be obtained from other devices, and this application does not limit this.

[0193] Specifically, in the embodiments of this application, the current video sequence contains multiple frames of images.

[0194] Step S102: Extract the dense deep learning feature maps corresponding to each frame of the images from the multiple frames of images respectively.

[0195] Specifically, in the embodiments of the present application, extracting the dense deep learning feature maps corresponding to each frame image from multiple frame images can be specifically implemented through a feature extraction network, or can also be extracted from each frame image through other means.

[0196] Step S103: Perform temporal feature fusion on the dense deep learning feature maps corresponding to each frame image respectively to obtain the fused features of each.

[0197] Step S104: Perform spatio-temporal feature aggregation processing on the fused features corresponding to each frame image respectively to obtain the global feature descriptor corresponding to the current video sequence.

[0198] Step S105: Retrieve from the global database based on the global feature descriptor corresponding to the current video sequence to obtain the first preset number of video sequences.

[0199] For the embodiments of the present application, the global database stores global feature descriptors and regional feature descriptors. In the embodiments of the present application, the global feature descriptors and regional feature descriptors stored in the global database are the global feature descriptors and regional feature descriptors corresponding to different video sequences respectively.

[0200] It should be noted that: the different video sequences here may include: video sequences corresponding to different geographical locations. Among them, the same geographical location may correspond to one video sequence, may correspond to multiple video sequences, and of course may also correspond to one video sequence corresponding to at least two geographical locations, which is not limited in the embodiments of the present application. In addition, the video sequence construction methods used when building the scene map and when performing scene retrieval may be the same or different. Since the process of constructing the video sequence when building the scene can be performed offline, key frames that are roughly in the same location can also be clustered according to geographical location relationships (such as translation distance, rotation angle, gps coordinate distance) to construct the video sequence.

[0201] The embodiments of the present application provide a video scene retrieval method. Compared with the related art, in the embodiments of the present application, temporal feature fusion is performed on the dense deep learning feature maps corresponding to each frame image in the current video sequence, and then spatio-temporal feature aggregation processing is performed according to the fused features to obtain the global feature descriptor corresponding to the current video sequence, that is, the spatio-temporal features of the current video sequence can be reflected in the global feature descriptor corresponding to the current video sequence. Therefore, retrieving from the global database based on the global feature descriptor corresponding to the current video sequence can reduce the influence of changes in the surrounding environment of the scene and local occlusion, etc. on scene re-identification, thereby improving the accuracy of the retrieved video sequence and further enhancing the user experience.

[0202] Further, after obtaining the current video sequence, in order to avoid waste of computing resources caused by excessive overlapping regions observed between video frames in the current video sequence, before step S102, it may further include: extracting key frames from the current video sequence. In the embodiments of the present application, the extraction criteria for key frames may be based on the overlapping percentage of the observation region, the number of inliers in feature point matching, the geographical location relationship between two frames (such as translation distance, rotation angle, gps coordinate distance), etc., or the key frame extraction strategy in the existing Simultaneous Localization And Mapping (SLAM) framework may also be used, such as the feature point-based extraction strategy of ORB-SLAM (English full name: Oriented FAST and Rotated BRIEF-SLAM) and the optical flow-based extraction strategy in DSO-SLAM (English full name: Direct Sparse Odometry-SLAM), etc.

[0203] Further, in order to further reduce waste of computing resources, after extracting key frames from the current video sequence, M frames may also be selected from the extracted key frames to form a new video sequence. In the embodiments of the present application, the M frames selected from the extracted key frames may be M consecutive key frames, or key frames may be selected at equal intervals, or the co-visibility relationship between key frames may be judged by feature point matching and optical flow method, so as to dynamically select key frames with a longer time interval. Further, if multiple new video sequences need to be constructed, the multiple constructed new video sequences may be of equal length or of unequal length, and M may be between [3, 15].

[0204] Further, after extracting M frames from the key frames, in step S102, extracting the dense deep learning feature maps corresponding to each frame image from the multiple frame images may specifically include: extracting the dense deep learning feature maps corresponding to each frame image from the M frame images.

[0205] Specifically, in the embodiments of the present application, the dense deep learning feature maps corresponding to each frame image are extracted from the M frame images through a feature extraction network.

[0206] Specifically, in order to improve the recall rate of scene retrieval when external factors such as perspective, illumination, and scene appearance change, and at the same time in order to better fuse local information on a single frame image, a neural network is used as the feature extraction network in the embodiments of the present application. Among them, the feature extraction network includes but is not limited to common deep learning backbone networks such as VGG, Unet, ResNet, RegNet, AlexNet, GoogLeNet, and MobileNet.

[0207] Further, before inputting each frame image into the feature extraction network for feature extraction, it is also necessary to adjust each frame image to the same size, and then input each frame image of the same size into the feature extraction network for feature extraction. In addition, in order to use the feature extraction network as a common network for the subsequent two branches of coarse-grained fast retrieval and fine-grained optimization sorting, it is necessary to move multiple frame images in a video sequence to the batch dimension before inputting them into the feature extraction network, so as not to randomly shuffle the inside of the same video sequence, but take a video sequence as a whole and only randomly shuffle the arrangement of different video sequences. In addition, when using a multi-Graphic Processing Unit (GPU), it is also necessary to ensure that the images in the same video sequence are assigned to the same GPU for processing. In the embodiment of the present application, extracting dense deep learning feature maps of each frame image of the same size through the feature extraction network may specifically include: extracting features of each frame image of the same size through the feature extraction network, and / or adjusting each frame image of the same size to different sizes through image pyramid operations, and then performing feature extraction. Among them, the image sizes of the corresponding layers of the image pyramid images corresponding to different frame images are the same.

[0208] Further, in order to make the model more robust to changes in the viewing angle and reduce the impact of local occlusion, and at the same time to make the model pay more attention to the stable features observed multiple times in the video sequence and reduce the interference of dynamic objects, the embodiment of the present application performs temporal fusion on the extracted feature maps through a self-attention mechanism to update the features repeatedly observed in the video sequence.

[0209] Specifically, performing temporal feature fusion based on the dense deep learning feature maps respectively corresponding to each frame image in step S103 to obtain the fused features respectively may specifically include: performing temporal feature fusion based on the dense deep learning feature maps respectively corresponding to each frame image through a self-attention mechanism to obtain the fused features respectively. In the embodiment of the present application, the temporal feature fusion network based on the self-attention mechanism is implemented by using a 3D Non-local (3D non-local network) network, and at the same time fuses the features in space and time, as shown in formula (1):

[0210] Formula (1);

[0211] Among them, x is the input feature map, i and j are different coordinates on the feature map, then x i 、x jValues representing different points on the feature map, the f() function measures the similarity between two points, such as using Gaussian similarity or Dot product similarity, etc., the g() function is used to calculate the feature value of the feature map at position j, C(x) is the normalization parameter, and y i is the value of the output feature map at coordinate i.

[0212] Specifically, the detailed structure of the time-domain feature fusion network based on the self-attention mechanism used in the embodiments of the present application is as follows Figure 2 shown. First, a linear mapping is performed on the input feature map X (T×H×W×1024), and a 1×1×1 convolution is used to compress the number of channels and streamline the original information, thereby obtaining (T×H×W×512), (T×H×W×512), (T×H×W×512) features. Then, all dimensions of the above three features except the number of channels are merged. Then, in order to calculate the self-correlation between the features, for and a matrix dot product operation is performed, so as to obtain the relationship between each pixel in each frame and all pixels in all other frames. Then, the calculated self-correlation result is Softmax-normalized to obtain a result with a value range of [0,1], which is used as the weight of self-attention. Then, the weight of self-attention is multiplied corresponding to the feature matrix and upsampled. Finally, a residual operation is performed with the original input feature map X, thereby obtaining the output Z (T×H×W×1024) of the time-domain feature fusion network.

[0213] It should be noted that in addition to the 3D Non-local network, other network variants based on the self-attention mechanism can also be used for time-domain feature fusion, all of which are within the protection scope of the embodiments of the present application, including but not limited to the Transformer network, the Temproal Non-local network, the graph convolutional neural network (Graph Neural Networks, GNN) based on the self-attention mechanism, etc.

[0214] Further, in order to aggregate all the features in the same video sequence, retain the unique observation information of each frame image, and remove the redundant observation information between frames, so as to generate a high-dimensional vector as the global representation of a video sequence. Furthermore, using a vector to represent a video sequence not only facilitates fast scene retrieval but also efficient storage. In the embodiments of the present application, through spatio-temporal feature aggregation processing on the result of spatio-temporal feature fusion, a global feature descriptor corresponding to the current video sequence is obtained. For details, see the following embodiments.

[0215] Specifically, in step S104, spatio-temporal feature aggregation processing is performed based on the fused features corresponding to each frame image to obtain a global feature descriptor corresponding to the current video sequence, which may specifically include: step S1041 (not shown in the figure), step S1042 (not shown in the figure), step S1043 (not shown in the figure), and step S1044 (not shown in the figure), where

[0216] Step S1041: Concatenate the temporal feature maps corresponding to each frame image to obtain a concatenated feature map.

[0217] Specifically, in the embodiments of the present application, the temporal feature maps corresponding to each frame image in the same video sequence may be concatenated along the long side to obtain a concatenated feature map, or may be concatenated along the short side to obtain a concatenated feature map. In the embodiments of the present application, taking the concatenation of the temporal feature maps corresponding to each frame image along the long side to obtain a concatenated feature map as an example for illustration.

[0218] Step S1042: Perform pointwise convolution processing on the concatenated feature map to obtain a convolution processing result.

[0219] Step S1043: Perform normalization processing on the convolution processing result to obtain a normalized result.

[0220] Specifically, in the embodiments of the present application, performing normalization processing on the convolution processing result may specifically include: performing exponential normalization processing on the convolution processing result through a normalization exponential function. In the embodiments of the present application, the normalization exponential function, or Softmax function, is a generalization of the logistic function. It can "compress" a K-dimensional vector z containing arbitrary real numbers into another K-dimensional real vector σ(z), such that the range of each element is between (0,1), and the sum of all elements is 1.

[0221] Step S1044: Based on the normalized result and the concatenated feature map, determine the global feature descriptor corresponding to the current video sequence.

[0222] Specifically, in the embodiments of the present application, the feature map after splicing processing includes multiple feature points; based on the result after normalization processing and the spliced feature map in step S1044, determining the global feature descriptor corresponding to the current video sequence may specifically include: performing clustering processing on the multiple feature points to obtain at least one clustering center; determining the distances between each feature point and each clustering center respectively, and determining the distance information corresponding to each clustering center; based on the distance information corresponding to each clustering center and the result after normalization processing, determining the global representations corresponding to each clustering cluster respectively; performing regularization processing on the global representations corresponding to each clustering cluster respectively; performing splicing processing on the globally regularized global representations; and performing regularization processing on the globally represented spliced processing results to obtain the global feature descriptor corresponding to the current video sequence.

[0223] Among them, the distance information corresponding to any one clustering center is the distances between each feature point and any one clustering center respectively.

[0224] Specifically, in the embodiments of the present application, taking the example of using TemproalVLAD as the spatio-temporal feature aggregation network to perform spatio-temporal feature aggregation, the network model architecture of TemproalVLAD is as Figure 3 shown. The temporally fused feature map is subjected to temporal splicing processing, and then the temporal splicing result is sequentially passed through pointwise convolution and exponential normalization processing to obtain the normalization processing result. Then, the temporal feature splicing result and the normalization processing result are processed through a residual calculation module, and intra-cluster regularization and global regularization processing are sequentially performed to obtain the global descriptor of the video sequence. Specifically, the feature maps of the same video sequence in the feature map output by the above-mentioned temporal feature fusion network based on the self-attention mechanism are spliced along the long side to obtain the spliced feature map F. Then, a 1×1 convolutional kernel is used to perform pointwise convolution on the feature map F, and Softmax is used to perform exponential normalization on the result to obtain the result a. Regarding the spliced feature map F (regarded as dense feature points), all the feature points in all the feature maps are taken out, and the Kmeans++ clustering algorithm is used for unsupervised clustering to obtain K clustering centers. Then, in the residual calculation module, the distances between each point in the feature map F and the K clusters are calculated respectively, and the result a output by the exponential normalization unit is used as the weight for weighted summation to obtain K vectors, which respectively correspond to the global representations of the K clustering clusters. Then, regularization is performed on the vectors of each cluster, and then the vectors of the K clusters are spliced together and globally regularized. Finally, a high-dimensional vector is obtained and used as the global descriptor of the entire video sequence.

[0225] Among them, K is between [16, 128], and the regularization operations include but are not limited to L1 regularization, L2 regularization, etc.

[0226] Further, after obtaining the global descriptor of the entire video sequence through spatio-temporal aggregation processing based on the above embodiments, the global descriptor of the current observed video sequence is used for coarse-grained and fast retrieval in the global database, reducing the number of video sequences that need to be precisely retrieved, so as to reduce the calculation time-consuming of the fine-grained optimization sorting branch.

[0227] Specifically, in the embodiments of the present application, the distances between the global descriptor of the current video sequence and the global descriptors of each video sequence stored in the global database are calculated in sequence to determine the similarity between the current video sequence and each video sequence stored in the global database. Then, the TopK1 video sequences that are most similar to the current observed scene are retrieved from the global database through coarse-grained and fast retrieval. Among them, the distance between the global descriptor of the current video sequence and the global descriptor of any video sequence stored in the global database can be characterized by Manhattan distance, Euclidean distance, Minkowski distance, etc., and the smaller the distance between the global descriptors of the video sequences, the higher the similarity.

[0228] It should be noted that in the embodiments of the present application, the value of TopK1 can be input by the user or preset, which is not limited in the embodiments of the present application. For example, TopK1 is between [20, 100].

[0229] It should be noted that, as can be seen from the above embodiments: the global database stores the global descriptors corresponding to multiple video sequences respectively. Among them, the global descriptors corresponding to the multiple video sequences stored in the global database are the same as the method for determining the global feature descriptor corresponding to the current video sequence based on the current video sequence in the above embodiments, and will not be elaborated here.

[0230] Further, in order to further improve the recall rate of the retrieval and improve the robustness to view changes and local occlusions, fine-grained optimization sorting is performed in the embodiments of the present application to optimize the retrieval sorting results of the coarse-grained branch. Further, after respectively extracting the dense deep learning feature maps corresponding to each frame image from multiple frame images in step S105, the following steps may further include: step S106 (not shown in the figure), step S107 (not shown in the figure), and step S108 (not shown in the figure). Among them, step S106 and step S107 can be executed before step S103-step S105, can also be executed after step S103-step S105, and can also be executed simultaneously with at least one step in step S103-step S105. Any possible execution order is within the protection scope of the embodiments of the present application and is not limited in the embodiments of the present application. The details of step S106-step S108 are shown in the following embodiments:

[0231] Step S106: Region feature extraction is performed on the dense deep learning feature maps corresponding to each frame of images respectively to obtain the multi-scale region features corresponding to each of them.

[0232] Specifically, in the embodiments of the present application, the region feature extraction of the dense deep learning feature maps corresponding to each frame of images can be performed through a multi-scale region feature extraction model, or can be performed without a region feature extraction model. In the embodiments of the present application, the case where the region feature extraction of the dense deep learning feature maps corresponding to each frame of images is performed through a region feature extraction model is taken as an example for illustration.

[0233] In the embodiments of the present application, when performing region feature extraction on the dense deep learning feature maps corresponding to each frame of images through a region feature extraction model, similar to the design of the spatio-temporal feature aggregation network, the multi-scale region feature extraction model first performs pointwise convolution on the feature map output by the feature extraction network and performs exponential normalization. The obtained result is used to weight the residual result between the original feature map and K clustering centers, and finally the weighted residual feature map R is obtained. The difference is that the multi-scale region feature extraction model does not perform time-domain feature stitching, and no summation and subsequent regularization operations are performed on the weighted residual feature map R either.

[0234] Specifically, based on the dense deep learning feature map corresponding to any frame of image, region feature extraction is performed to obtain the multi-scale region feature corresponding to any frame of image, which may specifically include: determining a weighted residual feature map based on the dense deep learning feature map corresponding to any frame of image; dividing the weighted residual feature map into multiple region blocks; determining the region feature representations corresponding to each region block respectively to obtain the multi-scale region feature corresponding to any frame of image.

[0235] Specifically, determining a weighted residual feature map based on the dense deep learning feature map corresponding to any frame of image may specifically include: performing pointwise convolution processing on the dense deep learning feature map corresponding to any frame of image to obtain a convolution result; performing normalization processing on the convolution result to obtain a normalization result; determining a weighted residual feature map based on the normalization result and the dense deep learning feature map corresponding to any frame of image.

[0236] For the embodiments of the present application, the specific methods of performing pointwise convolution and normalization processing on the dense deep learning feature map corresponding to any frame of image are specifically described in the above embodiments of spatio-temporal feature aggregation, and will not be elaborated in the embodiments of the present application.

[0237] Further, based on the normalization result and the dense deep learning feature map corresponding to any frame of image, a weighted residual feature map is determined, which specifically may include: obtaining K clustering centers for all feature points in the dense deep learning feature map corresponding to any frame of image extracted by the feature extraction model through a clustering algorithm; determining the distances between each feature point and each clustering center respectively, and determining the distance information corresponding to each clustering center; based on the distance information corresponding to each clustering center and the result after the normalization process, determining the weighted residual feature map R.

[0238] Among them, the distance information corresponding to any clustering center is the distances between each feature point and any clustering center respectively.

[0239] Further, after obtaining the weighted residual feature map R, a sliding window with a size of p×p is used to divide the region blocks on the weighted residual feature map R. The mean value of the residuals in each region block is used as the descriptor of this region after regularization. The regularization operations include but are not limited to L1 regularization and L2 regularization, etc. To enhance the robustness of the region descriptor to view changes, the size p of the sliding window can be changed, or region blocks can be divided using sliding windows of different sizes, and each region block generates a corresponding region descriptor. Since the region descriptor is generated by calculating the mean value, the dimensions of the region descriptors in different regions are the same.

[0240] It should be noted that in the above manner, multi-scale region features are extracted for each frame of image in the video sequence, and the region descriptor is used as the representation and storage method of the region features.

[0241] Further, in the embodiments of the present application, the multi-scale region feature extraction module extracts region features from each frame of image in the video sequence, which can also be considered as the fusion of the dense feature points extracted by the previous feature extraction network in the spatial domain, so as to optimize and sort the results of the fast retrieval branch using the region features subsequently.

[0242] Further, after extracting the multi-scale region features corresponding to each frame of image respectively, in order to further enhance the robustness of the region features to view changes, region matching is continued for the multi-scale region features corresponding to each frame of image respectively. For details, see the following embodiments.

[0243] Step S107: Perform region matching based on the respective corresponding multi-scale region features to obtain the spatio-temporal feature descriptor corresponding to the current video sequence.

[0244] Among them, the multi-scale region features are characterized by region descriptors.

[0245] Specifically, in the embodiments of the present application, region matching based on respective multi-scale region features includes: region matching between frames of the same video sequence, aiming to update the original region descriptor using region descriptors that are robust under different perspectives (see the following steps S1071 - S1072 for details), and then perform region matching in the time domain of different video sequences (see the following step S1073 for details), that is, the region matching based on respective multi-scale region features in step S107 to obtain the spatio-temporal feature descriptor corresponding to the current video sequence, which may specifically include: step S1071 (not shown in the figure), step S1072 (not shown in the figure), and step S1073 (not shown in the figure), where,

[0246] Step S1071: Perform region feature matching between the region descriptors corresponding to each frame image in the current video sequence and the region descriptors corresponding to each of the other frame images in the current video sequence to obtain the region matching result corresponding to the current video sequence.

[0247] For the embodiments of the present application, performing region feature matching between the region descriptors corresponding to each frame image in the current video sequence and the region descriptors corresponding to each of the other frame images in the current video sequence may specifically be based on the method of two-way matching and ratio test, or may also be through other region matching methods, including but not limited to K-nearest neighbor matching, Greedy-Nearest Neighbor (Greedy-NN) matching, k-dimensional tree (k-d tree) matching, etc. In the embodiments of the present application, the method based on two-way matching and ratio test is taken as an example for introduction.

[0248] Specifically, performing region feature matching between the region descriptor corresponding to each frame image and the region descriptor corresponding to any one frame image may specifically include: determining the distance vector corresponding to each frame image based on the region descriptors corresponding to each region in each frame image and the region descriptors corresponding to each region in the any one frame image. Wherein, the distance vector corresponding to each frame image contains multiple elements, and any one element is the distance between the region descriptor corresponding to any region in each frame image and the region descriptor corresponding to any region in the any one frame image.

[0249] Specifically, in the embodiments of the present application, taking frame Tm and frame Tn as an example to introduce the specific method of region feature matching, where (m, n ∈ [1, M] and m ≠ n), in the embodiments of the present application, calculate the distances between all region descriptors of frame Tm and all region descriptors of frame Tn in the video sequence to form a distance matrix D, and the element D in the matrix ijThat is, it represents the distance between the $i$-th region descriptor in the $T_m$ frame and the $j$-th region descriptor in the $T_n$ frame. This distance includes, but is not limited to, Manhattan distance, Euclidean distance, and Minkowski distance, etc.

[0250] Further, through the above method, the region descriptors corresponding to each frame image in the current video sequence can be calculated and region feature matching can be performed with the region descriptors corresponding to each other frame image in the current video sequence to obtain the region matching result corresponding to the current video sequence.

[0251] Step S1072: Select the region descriptors that meet the preset conditions from the region matching result corresponding to the current video sequence as the region descriptors corresponding to the current video sequence.

[0252] Specifically, selecting the region descriptors that meet the preset conditions from the distance vectors corresponding to each frame image may specifically include: determining the region descriptors that meet the preset conditions through the following formula (2), where

[0253] , formula (2);

[0254] where $X'$ represents the region descriptor that meets the preset conditions, $D$ ij k represents the element with the smallest distance value in the $j$-th column of the distance matrix $D$, $D$ i k j represents the element with the smallest distance value in the $i$-th row of the distance matrix $D$, $t$ is a threshold parameter, and the matching items $(i, j)$ that meet the conditions constitute the matching set $P$ between the $T_m$ frame and the $T_n$ frame mn , and at the same time, the distance value $D$ is also stored in the matching set ij . Among them, the threshold $t$ is between $[0.5, 0.9]$.

[0255] After exhausting all $(m, n)$ combinations for the same video sequence, the region $i$ of each frame is matched with the regions $j$ in several frames, denoted as the set $S$ i $=\{i, j1, j2, \ldots, j$ L $\}$, and the region $j$ is in turn matched with $j'$ in several frames in its own frame, and $j'$ is recursively added to the set $S_i$; select the region descriptors that meet the following formula (3) to replace the original descriptors of all regions in the set $S$ i so that it is more robust to view changes.

[0256] formula (3);

[0257] where $x$ is the region belonging to $S$ i , $P$ xThe set of matching items corresponding to region x in the matching set for the frame where region x is located, D x For all D x extracted from the set of matching items P ij set, For the above D ij is the average value of the elements in the set. Select the descriptor corresponding to region x' in S i in as the new region descriptors for all regions in the set S i Since the average distance of region x' to other matching regions in the video sequence is the smallest, it is considered that the features of region x' are more robust under different observation perspectives.

[0258] Step S1073: Perform region feature matching between the region descriptors corresponding to the current video sequence and each video sequence stored in the global database, and obtain the region matching results corresponding to the current video sequence and each video sequence respectively.

[0259] Specifically, in the embodiment of the present application, performing region feature matching between the region descriptors corresponding to the current video sequence and any video sequence may specifically include: performing region feature matching between the region descriptors corresponding to each frame image in the current video sequence and each frame image in any video sequence. In the embodiment of the present application, the method of performing region feature matching between the region descriptors corresponding to each frame image in the current video sequence and each frame image in any video sequence is based on the method of two-way matching and ratio test, and may also be performed through other region matching methods, including but not limited to K-nearest neighbor matching, Greedy-Nearest Neighbor (Greedy-NN) matching, k-dimensional tree (k-d tree) matching, etc. In the embodiment of the present application, the method based on two-way matching and ratio test is taken as an example for introduction.

[0260] Specifically, in the embodiment of the present application, performing region feature matching between the region descriptor corresponding to each frame image and the region descriptor corresponding to any frame image may specifically include: determining the distance vector corresponding to each frame image based on the region descriptors corresponding to each region in each frame image and the region descriptors corresponding to each region in any frame image. Among them, the distance vector corresponding to each frame image contains multiple elements, and any element is the distance between the region descriptor corresponding to any region in each frame image and the region descriptor corresponding to any region in any frame image.

[0261] Specifically, during video scene retrieval, for the current video sequence Vqry and a video sequence Vref stored in the database (global database) when the scene map is established, according to the above-mentioned region matching method based on bidirectional matching and ratio test, the region matching set P between each frame image in Vqry and each frame image in Vref is calculated in sequence, and at the same time, the coordinates (r, c) of each pair of matching regions in the region matching set P in the residual feature map R described above are recorded.

[0262] Furthermore, the region feature matching between the current video sequence Vqry and each video sequence stored in the database (global database) when the scene map is established is performed in the above-mentioned manner, and the details are not described herein again. It should be noted that the region descriptors corresponding to each video sequence Vref are also stored in the global database. In the embodiments of the present application, the determination method of the region descriptors corresponding to each video sequence Vref is the same as that of the region descriptors corresponding to the current video sequence Vqry, and the details are not described herein again.

[0263] Step S108: Perform region matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence to obtain the second preset number of video sequences.

[0264] For the embodiments of the present application, the spatio-temporal region descriptor extracted from the current observed video sequence optimizes the arrangement order of the above-mentioned TopK1 video sequences, and the optimized TopK2 video sequences are selected as the final result of scene retrieval. In the embodiments of the present application, the value of TopK2 can be input by the user or can be preset, and it is not limited in the embodiments of the present application. For example, TopK2 is between [1, 10].

[0265] Specifically, in the embodiments of the present application, in step S108, region matching is performed on the first preset number of video sequences based on the spatio-temporal feature descriptors corresponding to the current video sequence to obtain the second preset number of video sequences, which may specifically include: determining the spatial consistency scores between the current video sequence and each of the video sequences respectively based on the region matching results corresponding to the current video sequence and each video sequence; reordering the first preset number of video sequences based on the spatial consistency scores between the current video sequence and each of the video sequences respectively; and extracting the second preset number of video sequences from the reordered first preset number of video sequences. Among them, each video sequence belongs to the first preset number of video sequences. That is to say, in the embodiments of the present application, the spatial consistency scores between two frames are calculated successively for the results of region matching, and the TopK1 video sequences are reordered according to the magnitudes of the overall spatial consistency scores of the video sequences. Specifically, determining the spatial consistency score between the current video sequence and any one video sequence includes: determining the spatial consistency scores between each frame image in the current video sequence and each frame image in the any one video sequence; determining the weight information of each frame image in the current video sequence; and determining the spatial consistency score between the current video sequence and any one video sequence based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in the any one video sequence. In the embodiments of the present application, for the spatial consistency score SS between every two frames in two video sequences, the calculation formula (4) for the spatial consistency score of the two video sequences is as follows:

[0266] Formula (4);

[0267] Wherein, VSS represents the spatial consistency score of the two video sequences; m and k are the frames in the current observed video sequence V qry and the retrieved video sequence V ref respectively, where V ref ∈{TopK1 video sequences obtained from the fast retrieval branch}; is the weight of the observed frame, and the is between (0, 1], and its selection strategy is: select a certain frame (such as the first frame, the middle frame or the last frame) in the observed video sequence as the reference frame, the weight of this frame is 1, and the weights of other frames decay exponentially from the reference frame.

[0268] Further, determine the spatial consistency score between each frame image in the current video sequence and any frame image, which specifically may include: determining the region matching spatial consistency scores of regions of various sizes; determining the weight information corresponding to the regions of various sizes; and determining the spatial consistency score between each frame image in the current video sequence and any frame image based on the region matching spatial consistency scores of the regions of various sizes and the weight information corresponding to the regions of various sizes.

[0269] Specifically, the calculation formula (5) for the overall spatial consistency score of region matching between two frames is as follows:

[0270] Formula (5);

[0271] Among them, SS represents the overall spatial consistency score of region matching between two frames; i is the traversal of the scale set; n s is the number of scales; w i is the scale weight, with one weight for each scale, and w i ∈[0,1]. Specifically, determining the spatial consistency score between each frame image in the current video sequence and any frame image includes: determining the region matching spatial consistency scores of regions of various sizes; determining the weight information corresponding to the regions of various sizes; and determining the spatial consistency score between each frame image in the current video sequence and any frame image based on the region matching spatial consistency scores of the regions of various sizes and the weight information corresponding to the regions of various sizes.

[0272] Specifically, the formula (6) for the region matching spatial consistency score of scale p is as follows:

[0273] Formula (6);

[0274] Among them, SS p represents the region matching spatial consistency score with a scale size of p; n p is the number of region blocks with a scale size of p extracted from a frame image in the multi-scale region feature extraction module; P p is the region matching set of region features with a scale size of p; (r p , c p ) is the matching offset stored in P p , which is the spatial position offset of region matching calculated by the spatio-temporal region feature matching module; and respectively represent the average column offset and average row offset in the P p set; i, j represent the traversal of the set P pThe numbers during traversal; the dist(·) function is a distance function, including but not limited to Manhattan distance, Euclidean distance, Minkowski distance, etc., and max(·) is the maximum value function.

[0275] Furthermore, the following introduces a method for video scene retrieval through specific examples, as Figure 4 shown, obtain the current video stream sequence, and then based on the key frame extraction and feature map extraction processes described above, obtain the dense deep learning feature map corresponding to the current video stream sequence, and then execute the coarse-grained branch and the fine-grained branch,

[0276] Among them, the specific execution process of the coarse-grained branch: determine the global feature descriptor corresponding to the current video sequence based on the dense deep learning feature map corresponding to the current video stream sequence, and then retrieve from the database constructed during map building based on the global feature descriptor corresponding to the current video sequence to obtain the TopK1 retrieval results;

[0277] Among them, the specific execution process of the fine-grained branch: obtain the corresponding region descriptor based on the dense deep learning feature map corresponding to the current video stream sequence, then update its own region descriptor, and then perform region matching based on the updated region descriptor and the region descriptors of each video in the database constructed during map building, calculate the spatial consistency scores between the current video sequence and each video sequence in the Top1 retrieval results based on the matching results, and optimize and sort each video sequence in the TopK1 retrieval results based on the spatial consistency scores of each video sequence to obtain the final TopK2 retrieval results.

[0278] The above embodiments introduce a video scene retrieval method from the perspective of the method process. The following embodiments introduce a video scene retrieval device from the perspective of virtual modules. For details, see the following embodiments.

[0279] An embodiment of the present application provides a video scene retrieval device, as Figure 5 shown, the video scene retrieval device 50 may include: an acquisition module 51, a feature map extraction module 52, a time-domain feature fusion module 53, a spatio-temporal feature aggregation processing module 54, and a first retrieval module 55, where,

[0280] The acquisition module 51 is used to acquire the current video sequence, and the current video sequence includes multiple frames of images;

[0281] The feature map extraction module 52 is used to extract the dense deep learning feature maps corresponding to each frame of image from multiple frames of images;

[0282] The time-domain feature fusion module 53 is used to perform time-domain feature fusion based on the dense deep learning feature maps corresponding to each frame of image respectively to obtain the respective fused features;

[0283] A spatio-temporal feature aggregation processing module 54, configured to perform spatio-temporal feature aggregation processing based on the fused features respectively corresponding to each frame of image, so as to obtain a global feature descriptor corresponding to the current video sequence;

[0284] A first retrieval module 55, configured to retrieve from a global database based on the global feature descriptor corresponding to the current video sequence, so as to obtain a first preset number of video sequences.

[0285] In a possible implementation manner of the embodiment of the present application, when the time-domain feature fusion module 53 performs time-domain feature fusion based on the dense deep learning feature maps respectively corresponding to each frame of image to obtain the fused features respectively, it is specifically configured to: perform time-domain feature fusion based on the dense deep learning feature maps respectively corresponding to each frame of image through a self-attention mechanism, so as to obtain the fused features respectively.

[0286] In another possible implementation manner of the embodiment of the present application, when the spatio-temporal feature aggregation processing module 54 performs spatio-temporal feature aggregation processing based on the fused features respectively corresponding to each frame of image to obtain a global feature descriptor corresponding to the current video sequence, it is specifically configured to: splice the time-domain feature maps respectively corresponding to each frame of image to obtain a spliced feature map; perform point-by-point convolution processing on the spliced feature map to obtain a convolution processing result; perform normalization processing on the convolution processing result to obtain a normalized result; determine a global feature descriptor corresponding to the current video sequence based on the normalized result and the spliced feature map.

[0287] In another possible implementation manner of the embodiment of the present application, the spliced feature map includes multiple feature points; when the spatio-temporal feature aggregation processing module 54 determines a global feature descriptor corresponding to the current video sequence based on the normalized result and the spliced feature map, it is specifically configured to: perform clustering processing on the multiple feature points to obtain at least one clustering center; determine the distances between each feature point and each clustering center respectively, and determine the distance information corresponding to each clustering center, where the distance information corresponding to any clustering center is the distances between each feature point and any clustering center respectively; determine the global representations respectively corresponding to each clustering cluster based on the distance information corresponding to each clustering center and the normalized result; perform regularization processing on the global representations respectively corresponding to each clustering cluster; splice the globally represented regularized processing; perform regularization processing on the globally represented spliced processing to obtain a global feature descriptor corresponding to the current video sequence.

[0288] In another possible implementation manner of the embodiment of the present application, the apparatus 50 further includes: a multi-scale regional feature extraction module, a spatio-temporal regional feature matching module, and a second retrieval module, where,

[0289] The multi-scale region extraction module is used to perform region feature extraction on the dense deep learning feature maps corresponding to each frame of the image respectively, so as to obtain the multi-scale region features corresponding to each of them;

[0290] The spatio-temporal region feature matching module is used to perform region matching based on the multi-scale region features corresponding to each of them, so as to obtain the spatio-temporal feature descriptor corresponding to the current video sequence;

[0291] The second retrieval module is used to perform region matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence, so as to obtain the second preset number of video sequences.

[0292] For the embodiments of the present application, the first retrieval module 55 and the second retrieval module may be the same retrieval module or different retrieval modules, which is not limited in the embodiments of the present application.

[0293] Another possible implementation manner of the embodiments of the present application is that when the multi-scale region extraction module performs region feature extraction on the dense deep learning feature map corresponding to any frame of the image to obtain the multi-scale region feature corresponding to any frame of the image, it is specifically used for: determining a weighted residual feature map based on the dense deep learning feature map corresponding to any frame of the image; dividing the weighted residual feature map into multiple region blocks; determining the region feature representations corresponding to each region block respectively, so as to obtain the multi-scale region feature corresponding to any frame of the image.

[0294] Another possible implementation manner of the embodiments of the present application is that when the multi-scale region extraction module determines the weighted residual feature map based on the dense deep learning feature map corresponding to any frame of the image, it is specifically used for: performing point-by-point convolution processing on the dense deep learning feature map corresponding to any frame of the image to obtain a convolution result; performing normalization processing on the convolution result to obtain a normalization result; determining the weighted residual feature map based on the normalization result and the distance information corresponding to each clustering center.

[0295] Another possible implementation manner of the embodiments of the present application is that the multi-scale region features are represented by region descriptors; when the spatio-temporal region feature matching module performs region matching based on the multi-scale region features corresponding to each of them to obtain the spatio-temporal feature descriptor corresponding to the current video sequence, it is specifically used for: performing region feature matching on the region descriptors corresponding to each frame of the image in the current video sequence and the region descriptors corresponding to the other frames of the image in the current video sequence respectively, so as to obtain the region matching result corresponding to the current video sequence; selecting the region descriptors that meet the preset conditions from the region matching result corresponding to the current video sequence as the region descriptors corresponding to the current video sequence; performing region feature matching on the region descriptors corresponding to the current video sequence and each video sequence stored in the global database respectively, so as to obtain the region matching results corresponding to the current video sequence and each video sequence respectively.

[0296] Another possible implementation of the embodiment of the present application. When the spatio-temporal region feature matching module performs region feature matching on the region descriptor corresponding to any frame image in the current video sequence and the region descriptors corresponding to any other frame image in the current video sequence to obtain the corresponding matching result, it is specifically used for:

[0297] Determine the distance between the region descriptor corresponding to any frame image and each region descriptor in each region of any other frame image in the current video sequence;

[0298] Through the following formula, perform region feature matching on the region descriptor corresponding to any frame image in the current video sequence and the region descriptors corresponding to any other frame image in the current video sequence to obtain the corresponding matching result:

[0299] ;

[0300] Wherein, the element D in the matrix ij represents the distance between the i-th region descriptor in the Tm frame and the j-th region descriptor in the Tn frame. The matrix D is used to represent the distances between all region descriptors in the Tm frame and all region descriptors in the Tn frame of the video sequence. The Tm frame represents any frame image, and Tn is used to represent any other frame image in the current video sequence; D ij k represents the element with the smallest distance value in the j-th column of the matrix D, D i k j represents the element with the smallest distance value in the i-th row of the matrix D. t represents the threshold parameter. The matching items (i, j) that meet the conditions form the matching set P between the Tm frame and the Tn frame mn .

[0301] Another possible implementation of the embodiment of the present application. When the spatio-temporal region feature matching module selects the region descriptors meeting the preset conditions from the region matching results corresponding to any region as the region descriptors corresponding to any region, it is specifically used for:

[0302] Determine the average value of the distances that meet the preset conditions;

[0303] Based on the average value of the distances that meet the preset conditions, determine the region descriptors corresponding to any region.

[0304] Another possible implementation of the embodiment of the present application. When the spatio-temporal region feature matching module determines the region descriptors corresponding to any region based on the average value of the distances that meet the first preset condition, it is specifically used for:

[0305] Based on the average value of the distances that meet the first preset condition, and determine the region descriptors corresponding to any region through the following formula:

[0306] ;

[0307] where x is a region belonging to S i in, P x is the set of matching items corresponding to region x in the frame matching set where region x is located, and D x is all D x extracted from the set of matching items P ij set, is the average value of the elements in the D ij set, and x' is used to represent the region descriptors of all regions determined in the set S i in.

[0308] Another possible implementation manner of the embodiment of the present application is that when the spatio-temporal region feature matching module performs region feature matching between the region descriptor corresponding to the current video sequence and any video sequence, it is specifically configured to: perform region feature matching between the region descriptors corresponding to each frame image in the current video sequence and each frame image in any video sequence respectively.

[0309] Another possible implementation manner of the embodiment of the present application is that when the spatio-temporal region feature matching module performs region feature matching between the region descriptor corresponding to each frame image and the region descriptor corresponding to any frame image, it is specifically configured to: determine the distance vector corresponding to each frame image based on the region descriptors corresponding to each region in each frame image and the region descriptors corresponding to each region in any frame image. Each distance vector corresponding to each frame image contains multiple elements, and any element is the distance between the region descriptor corresponding to any region in each frame image and the region descriptor corresponding to any region in any frame image.

[0310] Another possible implementation manner of the embodiment of the present application is that when the second retrieval module performs region matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence and obtains the second preset number of video sequences, it is specifically configured to: determine the spatial consistency scores corresponding to the current video sequence and each video sequence respectively based on the region matching results corresponding to the current video sequence and each video sequence respectively. Each video sequence belongs to the first preset number of video sequences; reorder the first preset number of video sequences based on the spatial consistency scores corresponding to the current video sequence and each video sequence respectively; extract the second preset number of video sequences from the reordered first preset number of video sequences.

[0311] In another possible implementation manner of the embodiment of the present application, when the second retrieval module determines the spatial consistency score corresponding to the current video sequence and any video sequence, it specifically is used for: determining the spatial consistency scores between each frame image in the current video sequence and each frame image in any video sequence; determining the weight information of each frame image in the current video sequence; and determining the spatial consistency score corresponding to the current video sequence and any video sequence based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any video sequence.

[0312] In another possible implementation manner of the embodiment of the present application, when the second retrieval module determines the spatial consistency score between each frame image in the current video sequence and any frame image, it specifically is used for: determining the region matching spatial consistency scores of each size; determining the weight information corresponding to each region of each size; and determining the spatial consistency score between each frame image in the current video sequence and any frame image based on the region matching spatial consistency scores of each size and the weight information corresponding to each region of each size.

[0313] In another possible implementation manner of the embodiment of the present application, when the second retrieval module determines the region matching spatial consistency score of any size, it specifically is used for:

[0314] Determining the region matching spatial consistency score of any size through the following formula:

[0315] ;

[0316] wherein, SS p represents the region matching spatial consistency score of the region with the scale size of p, n p represents the number of region blocks with the scale size of p extracted from this frame image, P p is the region matching set of the region features with the scale size of p, (r p , c p ) is the matching offset stored in P p ; and respectively represent the average column offset and the average row offset in the set P p ; i and j represent the numbers during traversing the set P p , and the dist(·) function is a distance function, and max(·) is a maximum value function;

[0317] wherein, when the second retrieval module determines the spatial consistency score between each frame image in the current video sequence and any frame image based on the region matching spatial consistency scores of each size and the weight information corresponding to each region of each size and through the following formula, it specifically is used for:

[0318] ;

[0319] Among them, SS represents the spatial consistency score between each frame image in the current video sequence and any frame image. i is the traversal of the scale set, and n s is the number of scales, and w i is the weight information corresponding to the size i, and w i ∈[0,1].

[0320] In another possible implementation manner of the embodiment of the present application, when the second retrieval module determines the spatial consistency score corresponding to the current video sequence and any video sequence based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any video sequence, it is specifically used for:

[0321] Based on the weight information of each frame image in the current video sequence and the spatial consistency scores between each frame image in the current video sequence and each frame image in any video sequence, and determine the spatial consistency score corresponding to the current video sequence and any video sequence through the following formula:

[0322] ;

[0323] Among them, VSS represents the spatial consistency score corresponding to the current video sequence and any video sequence. V ref belongs to the video sequences of the first preset number. m is used to represent the frame in the current video sequence, and k is used to represent the frame in V ref in, is the weight information used to represent m.

[0324] The embodiment of the present application provides a video scene retrieval device. Compared with the related technology, in the embodiment of the present application, time-domain feature fusion is performed based on the dense deep learning feature maps respectively corresponding to each frame image in the current video sequence, and then spatio-temporal feature aggregation processing is performed according to the fused features to obtain the global feature descriptor corresponding to the current video sequence, that is, the spatio-temporal features of the current video sequence can be reflected in the global feature descriptor corresponding to the current video sequence. Therefore, retrieval can be performed from the global database based on the global feature descriptor corresponding to the current video sequence, which can reduce the influence of changes in the surrounding environment of the scene and local occlusion on scene re-identification, thereby improving the accuracy of the retrieved video sequence, and further improving the user experience.

[0325] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0326] An electronic device is provided in an embodiment of the present application, such as Figure 6 shown. Figure 6 The electronic device 600 shown in Figure 6 includes: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiment of the present application.

[0327] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 601 may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0328] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus.

[0329] The memory 603 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0330] The memory 603 is used to store the application program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0331] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 6 The shown electronic device is only an example and should not bring any restrictions to the functions and usage scopes of the embodiments of this application.

[0332] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments. Compared with the related art, in the embodiments of this application, time-domain feature fusion is performed based on the dense deep learning feature maps respectively corresponding to each frame image in the current video sequence, and then spatio-temporal feature aggregation processing is performed according to the fused features to obtain the global feature descriptor corresponding to the current video sequence, that is, the spatio-temporal features of the current video sequence can be reflected in the global feature descriptor corresponding to the current video sequence. Therefore, retrieving from the global database based on the global feature descriptor corresponding to the current video sequence can reduce the influence of changes in the surrounding environment of the scene and local occlusion, etc. on scene re-identification, thereby improving the accuracy of retrieving the video sequence and further enhancing the user experience.

[0333] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0334] In the embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0335] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0336] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0337] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.

[0338] As described above, the above embodiments are only used to introduce the technical solutions of the present application in detail. However, the description of the above embodiments is only used to help understand the method and its core idea of the present application, and should not be construed as a limitation of the present application. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application.

Claims

1. A video scene retrieval method, characterized in that, Including: Obtain a current video sequence, where the current video sequence contains multiple frames of images; Extract dense deep learning feature maps corresponding to each frame of images from the multiple frames of images respectively; Perform temporal feature fusion based on the dense deep learning feature maps corresponding to each frame of images respectively to obtain respective fused features; Perform spatio-temporal feature aggregation processing based on the fused features corresponding to each frame of images to obtain a global feature descriptor corresponding to the current video sequence, specifically including: The performing spatio-temporal feature aggregation processing based on the fused features corresponding to each frame of images to obtain a global feature descriptor corresponding to the current video sequence includes: Perform splicing processing on the temporal feature maps corresponding to each frame of images to obtain a spliced feature map; Perform pointwise convolution processing on the spliced feature map to obtain a convolution processing result; Perform normalization processing on the convolution processing result to obtain a result after normalization processing; Based on the result after normalization processing and the spliced feature map, determine the global feature descriptor corresponding to the current video sequence; The spliced feature map includes multiple feature points; The determining the global feature descriptor corresponding to the current video sequence based on the result after normalization processing and the spliced feature map includes: Perform clustering processing on the multiple feature points to obtain at least one clustering center; Determine the distances between each feature point and each clustering center respectively, and determine the distance information corresponding to each clustering center. The distance information corresponding to any one clustering center is the distances between each feature point and the any one clustering center; Based on the distance information corresponding to each clustering center and the result after normalization processing, determine the global representations corresponding to each clustering cluster respectively; Perform regularization processing on the global representations corresponding to each clustering cluster respectively; Perform splicing processing on the globally regularized representations; Perform regularization processing on the globally represented spliced result to obtain the global feature descriptor corresponding to the current video sequence; Retrieve from a global database based on the global feature descriptor corresponding to the current video sequence to obtain a first preset number of video sequences.

2. The method according to claim 1, wherein The performing temporal feature fusion based on the dense deep learning feature maps corresponding to each frame of images respectively to obtain respective fused features includes: Perform temporal feature fusion based on the dense deep learning feature maps corresponding to each frame of images respectively through a self-attention mechanism to obtain respective fused features.

3. The method according to claim 1, characterized in that, After the extracting dense deep learning feature maps corresponding to each frame of images from the multiple frames of images respectively, it further includes: Perform regional feature extraction on the dense deep learning feature maps corresponding to each frame of images respectively to obtain respective multi-scale regional features; Perform regional matching based on the respective multi-scale regional features to obtain a spatio-temporal feature descriptor corresponding to the current video sequence; Perform regional matching on the first preset number of video sequences based on the spatio-temporal feature descriptor corresponding to the current video sequence to obtain a second preset number of video sequences.

4. The method according to claim 3, characterized in that Region feature extraction is performed based on the dense deep learning feature map corresponding to any frame of image to obtain the multi-scale region features corresponding to the any frame of image, including: Determine a weighted residual feature map based on the dense deep learning feature map corresponding to the any frame of image; Divide the weighted residual feature map into multiple region blocks; Determine the region feature representations corresponding to the respective region blocks to obtain the multi-scale region features corresponding to the any frame of image.

5. The method according to claim 4, wherein The determining a weighted residual feature map based on the dense deep learning feature map corresponding to the any frame of image includes: Perform pointwise convolution processing on the dense deep learning feature map corresponding to the any frame of image to obtain a convolution result; Perform normalization processing on the convolution result to obtain a normalization result; Determine the weighted residual feature map based on the normalization result and the distance information corresponding to each clustering center.

6. The method according to any one of claims 3-5, characterized in that, The multi-scale region features are represented by region descriptors; The performing region matching based on the respective corresponding multi-scale region features to obtain the spatio-temporal feature descriptor corresponding to the current video sequence includes: Perform region feature matching between the region descriptors corresponding to each frame of image in the current video sequence and the region descriptors corresponding to each other frame of image in the current video sequence to obtain the region matching result corresponding to the current video sequence; Select the region descriptors that meet the preset conditions from the region matching result corresponding to the current video sequence as the region descriptors corresponding to the current video sequence; Perform region feature matching between the region descriptors corresponding to the current video sequence and each video sequence stored in the global database to obtain the region matching results corresponding to the current video sequence and the respective video sequences.

7. The method according to claim 6, characterized in that The performing region feature matching between the region descriptor corresponding to any frame of image in the current video sequence and the region description corresponding to any other frame of image in the current video sequence to obtain the corresponding matching result includes: Determine the distances between the region descriptor corresponding to the any frame of image and the region descriptors of each region in any other frame of image in the current video sequence; Perform region feature matching between the region descriptor corresponding to any frame of image in the current video sequence and the region description corresponding to any other frame of image in the current video sequence through the following formula to obtain the corresponding matching result: Among them, the element D in the matrix ij represents the distance between the i-th region descriptor in the Tm frame and the j-th region descriptor in the Tn frame. The matrix D is used to represent the distances between all region descriptors in the Tm frame and all region descriptors in the Tn frame of the video sequence. The Tm frame represents any one of the frame images, and the Tn is used to represent any other frame image in the current video sequence; D ij k represents the element with the minimum distance value in the j-th column of the matrix D, D i k j represents the element with the minimum distance value in the i-th row of the matrix D. t is used to represent the threshold parameter. The matching items (i, j) that meet the conditions form the matching set P between the Tm frame and the Tn frame mn .

8. The method according to claim 7, wherein The selecting the region descriptors that meet the preset conditions from the region matching result corresponding to any region as the region descriptors corresponding to the any region includes: Determine the average value of the distances that meet the preset conditions; Determine the region descriptors corresponding to the any region based on the average value of the distances that meet the preset conditions.

9. The method according to claim 8, wherein The determining the region descriptors corresponding to the any region based on the average value of the distances that meet the preset conditions includes: Determine the region descriptors corresponding to the any region based on the average value of the distances that meet the preset conditions and through the following formula: where x belongs to S i in the region, P x is the set of matching items corresponding to region x in the frame matching set where region x is located, D x is all D x extracted from the set of matching items P ij set, is the average value of the elements in the D ij set, and x' is used to represent the region descriptors of all regions determined in the set S i set.

10. The method according to claim 9, wherein The performing region feature matching between the region descriptors corresponding to the current video sequence and any video sequence includes: Perform region feature matching between the region descriptors corresponding to each frame of image in the current video sequence and each frame of image in the any video sequence respectively.

11. The method according to any one of claims 7 to 10, characterized in that, Performing region feature matching between the region descriptors corresponding to each frame of image and the region descriptors corresponding to any one frame of image, including: Based on the region descriptors corresponding to each region in each frame of image and the region descriptors corresponding to each region in the any one frame of image, determining a distance vector corresponding to each frame of image, where the distance vector corresponding to each frame of image contains multiple elements, and any one element is the distance between the region descriptor corresponding to any region in each frame of image and the region descriptor corresponding to any region in the any one frame of image.

12. The method according to claim 11, wherein The performing region matching on the first preset number of video sequences based on the spatio-temporal feature descriptors corresponding to the current video sequence to obtain a second preset number of video sequences includes: Based on the region matching results corresponding to the current video sequence and each of the video sequences, determining the spatial consistency scores corresponding to the current video sequence and each of the video sequences, where each of the video sequences belongs to the first preset number of video sequences; Based on the spatial consistency scores corresponding to the current video sequence and each of the video sequences, reordering the first preset number of video sequences; Extracting a second preset number of video sequences from the reordered first preset number of video sequences.

13. The method according to claim 12, wherein Based on the region matching result corresponding to the current video sequence and any one video sequence, determining the spatial consistency score corresponding to the current video sequence and the any one video sequence includes: Based on the region matching results corresponding to the current video sequence and any one video sequence, determining the spatial consistency scores between each frame of image in the current video sequence and each frame of image in the any one video sequence; Determining the weight information of each frame of image in the current video sequence; Based on the weight information of each frame of image in the current video sequence and the spatial consistency scores between each frame of image in the current video sequence and each frame of image in the any one video sequence, determining the spatial consistency score corresponding to the current video sequence and any one video sequence.

14. The method according to claim 13, wherein The determining the spatial consistency score between each frame of image in the current video sequence and any one frame of image includes: Determining the spatial consistency scores of region matching for each size; Determining the weight information corresponding to each region of each size; Based on the spatial consistency scores of region matching for each size and the weight information corresponding to each region of each size, determining the spatial consistency score between each frame of image in the current video sequence and any one frame of image.

15. The method according to claim 14, characterized in that The based on the weight information of each frame of image in the current video sequence and the spatial consistency scores between each frame of image in the current video sequence and each frame of image in the any one video sequence, determining the spatial consistency score corresponding to the current video sequence and any one video sequence includes: Based on the weight information of each frame of image in the current video sequence and the spatial consistency scores between each frame of image in the current video sequence and each frame of image in the any one video sequence, and determining the spatial consistency score corresponding to the current video sequence and any one video sequence through the following formula: Among them, VSS represents the spatial consistency score corresponding to the current video sequence and any video sequence, V ref belongs to the video sequences of the first preset number, m is used to represent the frames in the current video sequence, and k is used to represent the frames in V ref in, λ m is used to represent the weight information of m, and SS represents the spatial consistency score between each frame image in the current video sequence and any frame image.

16. A video scene retrieval device for performing the video scene retrieval method according to any one of claims 1 to 15, characterized in that, Including: An obtaining module, configured to obtain a current video sequence, where the current video sequence contains multiple frames of images; A feature map extraction module, configured to respectively extract dense deep learning feature maps corresponding to each frame image from the multi-frame images; A time-domain feature fusion module, configured to respectively perform time-domain feature fusion based on the dense deep learning feature maps corresponding to each frame image to obtain respective fused features; Spatio-temporal A feature aggregation processing module, configured to perform spatio-temporal feature aggregation processing based on the fused features corresponding to each frame image to obtain a global feature descriptor corresponding to the current video sequence; A first retrieval module, configured to retrieve from a global database based on the global feature descriptor corresponding to the current video sequence to obtain a first preset number of video sequences.

17. An electronic device, characterized in that, It includes: One or more processors; A memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to: execute a video scene retrieval method according to any one of claims 1 to 15.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements a video scene retrieval method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Video matching, retrieving, classifying and recommending method and device and electronic device

    CN108416013A

  • Video behavior identification method based on local feature aggregation descriptor and sequential relationship network

    CN110097000A