Fast Optical Flow Estimation Method for Small Moving Targets Based on Local Search and Storage Medium

By using local search in the optical flow estimation network instead of global search, the problem of large and time-consuming calculations of the optical flow estimation network is solved, and more efficient optical flow estimation is achieved and real-time performance is improved.

CN115170826BActive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210797411.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2025-06-20
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

The existing deep learning-based optical flow estimation network adopts global search in the feature search stage, resulting in large calculations and long time, affecting the real-time nature of optical flow estimation.

Method used

The fast optical flow estimation method of moving small targets based on local search is adopted to match the corresponding local search area through each feature vector in the feature map, reducing the computational amount and time-consuming search match.

Benefits of technology

While ensuring the estimation accuracy of the small-target optical flow, the speed of the optical flow estimation is significantly improved, the calculation amount is reduced, and the real-time performance of the optical flow estimation network is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170826B_ABST
    Figure CN115170826B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of deep learning, and particularly to a fast optical flow estimation method and storage medium for moving small targets based on local search. First, two adjacent images are obtained, and corresponding feature maps and context information are respectively obtained through feature extraction, and the context information is obtained by encoding the first frame image alone; then, for each feature vector in the feature map, a corresponding local search area is matched in the feature map, and corresponding similar information is sequentially matched from the corresponding local search area according to each feature vector in the feature map, and all the similar information is aggregated into matching information; finally, the context information and the matching information are used to perform iterative optical flow estimation through a preset recurrent network. By changing the feature search and matching from global search to search within a proper and reasonable local range, the search time and computational amount are reduced, and to a certain extent, the problem of increased computational amount caused by downsampling is avoided. At the same time, while ensuring the accuracy of the optical flow estimation of moving small targets, the speed of the optical flow estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning, and in particular, to a method and storage medium for fast optical flow estimation of moving small targets based on local search. Background Art

[0002] Optical flow estimation is an important direction in computer vision research. The so-called optical flow provides a motion vector for each pixel in the image, that is, the motion speed in the x-axis and the motion speed in the y-axis. This motion vector matrix characterizes the motion field of the entire image and contains potential dynamic information. Analyzing the information of the motion vector matrix can accurately obtain the position of the moving small target, which helps in the detection and recognition of small targets in the field of computer vision.

[0003] Currently, optical flow estimation networks based on deep learning have exceeded traditional optical flow algorithms in terms of speed and accuracy, which has greatly promoted the practical application of optical flow estimation in various engineering fields. However, existing optical flow estimation networks based on deep learning usually adopt global search in the feature search stage. Generally speaking, in order to accelerate the inference speed of the network, it is necessary to perform at least 3 downsamplings on the original image. In this way, the information of some moving small targets will be lost. In order to obtain accurate optical flow estimation of moving small targets, it is necessary to reduce the number of downsamplings of the network. However, whenever the number of downsamplings is reduced by one, the computational amount of the feature extraction and iterative optical flow estimation module will be expanded by 4 times, and the computational amount of the feature search module will be expanded by 16 times. This will lead to a significant increase in the optical flow estimation time. Therefore, existing optical flow estimation networks based on deep learning usually need to increase the search time and computational amount as a cost to ensure the accuracy of small target optical flow estimation, which reduces the real-time performance of optical flow estimation. Summary of the Invention

[0004] The purpose of the present application is to provide a method and storage medium for fast optical flow estimation of moving small targets based on local search, which can improve the speed of optical flow estimation while ensuring the accuracy of optical flow estimation of moving small targets.

[0005] In the first aspect, the present application provides a method for fast optical flow estimation of moving small targets based on local search, adopting the following technical solutions:

[0006] A method for fast optical flow estimation of moving small targets based on local search includes the following steps:

[0007] Obtain two adjacent frames of images, extract features from the two frames of images, and obtain feature maps corresponding to the two frames of images respectively and , and separately encode the first frame of image to obtain context information;

[0008] For each feature vector in the feature map in the feature map Match the corresponding local search area;

[0009] According to the feature map Each feature vector in the matching local search area matches the corresponding similar information, and all similar information is aggregated into matching information;

[0010] Using the context information and the matching information, performing iterative optical flow estimation through a preset recurrent network, and outputting an optical flow estimation result;

[0011] Through the above technical solution, the feature search and matching is changed from a global search to a search within an appropriate and reasonable local range. Compared with the global search, the local search greatly reduces the time and computational complexity of the search and matching stage, and to a certain extent avoids the problem of increased computational complexity due to downsampling. At the same time, while ensuring the accuracy of optical flow estimation of small moving targets, the speed of optical flow estimation is improved, further improving the real-time performance of the optical flow estimation network.

[0012] Optionally, the first frame image is separately encoded to obtain context information, including:

[0013] Multi-scale feature extraction is performed on the first frame image according to the preset context network to obtain local features and global features respectively.

[0014] The local features and global features are fused to obtain context information.

[0015] Optional, feature map Each feature vector in the feature map The corresponding local search area is matched in, including:

[0016] For feature maps Divide to form multiple contiguous regions ,in is a positive integer;

[0017] For any area In the feature map Find the same location in the mapping area ;

[0018] For the mapping area Expand to obtain the expanded area ,area That is the feature map middle All feature vectors in the region are in the feature map The local search area in .

[0019] Optionally, for the mapped region Expand to obtain the expanded area , including:

[0020] Obtain the size of the small moving target and the displacement amount between two frames of images;

[0021] Calculate the extended side length based on the size of the small moving target and the displacement amount between two frames of images;

[0022] Obtain the mapping area in the feature map position information;

[0023] Based on the position information of the mapping area in the feature map and the extended side length, obtain the extended area .

[0024] Optionally, based on the position information of the mapping area in the feature map and the extended side length, obtain the extended area , including:

[0025] Based on the position information of the mapping area in the feature map and the extended side length, obtain the position information of the extended area in the feature map ;

[0026] Based on the position information of the extended area in the feature map , determine whether the extended area exceeds the range of the feature map . If there is an exceeding part, readjust the position information of the extended area .

[0027] Optionally, match the corresponding similar information from the corresponding local search area for each feature vector in the feature map , including:

[0028] Based on the feature vector in the feature map , obtain all feature vectors in its corresponding local search area and form a feature vector set ;

[0029] Match the k feature vectors with the highest similarity to the feature vector from the feature vector set , and obtain the position information and similarity of the k matched feature vectors;

[0030] Integrate the position information and similarity of the k matched feature vectors to form a feature map The eigenvector of similarity information.

[0031] Optionally, match from the eigenvector set the k eigenvectors with the highest similarity to the eigenvector , and obtain the position information and similarity of the k matched eigenvectors, including:

[0032] Construct an index for the eigenvector set ;

[0033] Perform a k-nearest neighbor search on the eigenvector and the eigenvector set to obtain the region index values and similarities of the k eigenvectors with the highest similarity to the eigenvector ;

[0034] Convert the index value of the region to the index value of the entire image;

[0035] According to the index value of the entire image, obtain the position information of the k eigenvectors matched with the eigenvector from the feature map .

[0036] Optionally, utilize the context information and matching information to perform iterative optical flow estimation through a preset recurrent network, including:

[0037] Obtain the initial value of the optical flow estimation;

[0038] According to the initial value of the optical flow estimation, matching information, and context information, obtain the input information;

[0039] According to the input information, perform iteration through the recurrent network to obtain the optical flow estimation.

[0040] Optionally, perform iteration through the recurrent network to obtain the optical flow estimation, including:

[0041] Obtain the optical flow estimation of the previous iteration;

[0042] According to the optical flow estimation of the previous iteration, matching information, and context information, obtain the input of the current iteration and the historical hidden layer state ;

[0043] According to the input of the current iteration and the historical hidden layer state , obtain the updated hidden layer state through the recurrent network;

[0044] For the updated hidden layer state The residual optical flow is obtained through several convolutions, and the optical flow estimate of the previous iteration is updated according to the residual optical flow to obtain the optical flow estimate of the current iteration.

[0045] In a second aspect, the present application provides a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform the above-mentioned fast optical flow estimation method for small moving targets based on local search.

[0046] In summary, the present application changes the feature search and matching from global search to search within a proper and reasonable local range, reducing the search time and computational amount, effectively solving the problem of increased computational amount caused by downsampling, and the matching information obtained through local search can constrain the iteration of optical flow estimation. Combined with the context optical association of context information, to a certain extent, it avoids the influence brought by the possible loss of some effective information due to the reduction of the search range. While ensuring the accuracy of the optical flow estimation of small moving targets, the speed of optical flow estimation is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flowchart of an embodiment of the present application;

[0048] Figure 2 is a schematic diagram of local search;

[0049] Figure 3 is a schematic diagram of the effective range limitation of the area in local search;

[0050] Figure 4 Schematic diagram of obtaining optical flow estimation through iteration by a recurrent network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will further describe the present application in detail Figure 1 - attached Figure 4 drawings,

[0052] The present application provides a fast optical flow estimation method for small moving targets based on local search. Refer to Figure 1 , including the following steps:

[0053] S100. Obtain two adjacent frames of images, extract features from the two frames of images to obtain feature maps and corresponding to the two frames of images respectively, and separately encode the first frame of image to obtain context information.

[0054] Among them, the two adjacent frames of images are obtained from an image sequence, and the image sequence is a set of continuously arranged images, usually multiple consecutive images obtained by converting a video or a dynamic image according to a set number of frames. The feature maps and The deep features are extracted from the two input frames of images by the same preset multi-layer convolutional neural network. The context information is the environmental information where the target is located, including the position information of the target in the image and the correlation information between the target and other surrounding objects.

[0055] In the embodiment of the present application, the image sequence is a dynamic background image sequence collected by a drone. The resolution of the original image is 1080P, and after appropriate cropping and reduction, its image size is 700*980. The context information is obtained by encoding the first frame image alone, and specifically includes the following steps:

[0056] S110. Perform multi-scale feature extraction on the first frame image according to the preset context network to obtain local features and global features respectively.

[0057] In the embodiment of the present application, dilated convolution is used on the first frame image to capture feature information of different scales. The feature information obtained through a large receptive field is global information, which focuses on the description of the position information of the pixels included in the target in the image. The small receptive field is local information, which focuses on the description of the correlation information between the pixels included in the target and other surrounding pixels. By combining the global information and the local information, the context information of the target is constructed, providing information support for the subsequent iterative optical flow estimation.

[0058] S120. Fuse the local features and the global features to obtain context information.

[0059] S200. Match a corresponding local search area in the feature map for each feature vector in the feature map

[0060] The local search area is the search range in the feature search and matching process, and the purpose is to narrow the search range, reduce the search time consumption and the amount of calculation. Because in the actual scenario, especially in the video scenario, the size of the moving small target is relatively small, and the relative movement between two frames of images is also very small. Then the search range of the feature vector can actually be limited to a small area, rather than performing a full-image search. This can improve the inference speed of the network and enhance the real-time performance of the optical flow estimation.

[0061] In the embodiment of the present application, match a corresponding local search area in the feature map for each feature vector in the feature map , see Figure 2 , and specifically includes the following steps:

[0062] S210. Divide the feature map to form a plurality of continuous regions , where ​is a positive integer.

[0063] Among them, the divided regions are required to be continuous because it is necessary to ensure that for any feature vector in the feature map it is within the divided regions , that is, the divided regions can completely piece together the feature map .

[0064] S220. For any region find the mapped region at the same position in the feature map .

[0065] Among them, , is the number of regions obtained by dividing the feature map , and the value of can be 4, 16, 64, etc. Because the feature maps and are features extracted by the same convolutional network, their resolutions are the same. Therefore, according to the relative position of the region in the feature map , the mapped region at the same position can be found in the feature map .

[0066] S230. Expand the mapped region to obtain the expanded region , and the region is the local search region of all feature vectors in the region in the feature map .

[0067] Among them, expanding the mapped region is to ensure that there is enough search scope for the edge part of the region in the feature map , because there is a displacement change of the target between two frames. Therefore, it is necessary to consider that after the target moves, the positions corresponding to the feature vectors included in the target will also change. Therefore, expanding the mapped region can try to ensure that after the displacement change of the target between two frames, the corresponding feature vectors are still within the matching range.

[0068] Expand the mapped region to obtain the expanded region , which specifically includes the following steps:

[0069] ​​S231. Obtain the size of the small moving target and the displacement amount between two frames of images.

[0070] In one embodiment, the size of the small moving target is less than 30*30, and the displacement amount of the small moving target between two frames rarely exceeds its size. Therefore, the size of the small moving target is estimated to be 30, and the displacement amount of the small moving target is estimated to be 10.

[0071] S232. Calculate the extended side length according to the size of the small moving target and the displacement amount between two frames of images.

[0072] In the embodiment of the present application, the calculation method of the extended side length is as follows:

[0073]

[0074] where, is the number of downsamplings in the multi-layer feature extraction network.

[0075] S233. Obtain the position information of the mapping area in the feature map .

[0076] In the embodiment of the present application, the position information is the upper left corner coordinate and the width. Taking the upper left corner of the feature map as the coordinate origin, according to the position information of the mapping area in the feature map , the upper left corner coordinate ( ) of the mapping area can be obtained. According to the width of the feature map and the number of divided regions , the width of the mapping area can be calculated.

[0077] S234. Obtain the extended area according to the position information of the mapping area in the feature map and the extended side length. .

[0078] In the embodiment of the present application, according to the position information of the mapping area in the feature map and the extended side length, the extended area is obtained, specifically including:

[0079] S2341. Obtain the extended area according to the position information of the mapping area in the feature map and the extended side length, and obtain the extended area in the feature map Position information

[0080] According to the mapping area The upper left corner coordinates of ( ), width And the extended side length , the extended area can be obtained The upper left corner coordinates are ( ), and the width is , thus determining the extended area In the feature map Position in

[0081] S2342. According to the position information of the extended area In the feature map , judge whether the extended area Exceeds the range of the feature map . If there is an exceeding part, readjust the position information of the extended area Position information

[0082] In the embodiment of the present application, see Figure 3 , because when dividing the area of the feature map , there will be some areas at the boundary position of the feature map Or overlapping with the boundary position of the feature map . When expanding these areas, there will be a situation where the extended area exceeds the range of the feature map . Therefore, it is necessary to readjust the position information of the extended area . While ensuring that the width of the extended area Remains unchanged, adjust the upper left corner coordinates of the extended area Position information

[0083] Such as Figure 3 In the first case, the originally calculated upper left corner coordinates of the extended area Are ( ), after adjustment, the upper left corner coordinates of the extended area Are ( ). In the second case, after adjustment, the upper left corner coordinates of the extended area Are ( ).

[0084] The specific judgment calculation method is as follows. First, judge whether the coordinate value is negative according to the upper left corner coordinates. If only the x-axis is negative, set the x-axis coordinate to 0, and then judge whether the y-axis coordinate plus the width of the extended area Width Is greater than the width of the feature map Width , if not, the y-axis coordinate remains unchanged; if so, the y-axis coordinate is subtracted by a width.

[0085] If only the y-axis is negative, set the y-axis coordinate to 0, and then determine whether the x-axis coordinate plus the expanded area width is greater than the width of the feature map . If not, the x-axis coordinate remains unchanged; if so, the x-axis coordinate is subtracted by a width.

[0086] If both the x-axis coordinate and the y-axis coordinate are negative, set both the x-axis coordinate and the y-axis coordinate to 0.

[0087] If the x-axis coordinate and the y-axis coordinate respectively plus the expanded area width are both greater than the width of the feature map , then subtract a width from both the x-axis coordinate and the y-axis coordinate on the original basis.

[0088] Performing the above operations on the region in sequence, a local search region corresponding one-to-one with the region can be obtained. .

[0089] S300. Match the corresponding similar information from the corresponding local search region for each feature vector in the feature map , and aggregate all the similar information into matching information.

[0090] Among them, the similar information includes similarity and position information. The similarity is the degree of similarity between two feature vectors, and the position information is the position index of the feature vector in the feature map. The similar information reflects the change state of pixel points between two frames, and it helps to estimate the optical flow more accurately by associating with the pixel motion vector matrix.

[0091] In the embodiments of the present application, the Euclidean distance is used to represent the similarity degree between two feature vectors. The similarity calculates the similarity degree between the feature vector to be matched in the feature map and the feature vectors in the corresponding local search region. The position information refers to the position index of the feature vector with high similarity to the feature vector to be matched in the feature map in the local search region in the feature map .

[0092] Specifically, according to the feature map ​​Each eigenvector matches corresponding similar information from the corresponding local search area, including the following steps:

[0093] S310. According to the feature map the eigenvectors in , obtain all the eigenvectors within its corresponding local search area, and form an eigenvector set .

[0094] S320. Perform k-nearest neighbor search on the eigenvector and the eigenvector set to obtain the position information and similarity of the k eigenvectors with the highest similarity to .

[0095] Among them, the eigenvector set is equivalent to a matching information library, and similar information to it is searched and matched according to the eigenvector .

[0096] In one embodiment, perform k-nearest neighbor search on the eigenvector and the eigenvector set to obtain the position information and similarity of the k eigenvectors with the highest similarity to , specifically including the following steps:

[0097] S321. Build an index for the eigenvector set .

[0098] Among them, the index used is IndexFlatL2 in the Faiss library. Building an index for the eigenvector set specifically includes: building an index according to the capacity of the eigenvector set , and adding the eigenvector set into the index.

[0099] S322. Perform k-nearest neighbor search on the eigenvector and the eigenvector set to obtain the region index values and similarity of the k eigenvectors with the highest similarity to .

[0100] Among them, performing k-nearest neighbor search on the eigenvector and the eigenvector set is performed through the established index, and the returned information is the region index values of the k eigenvectors with the highest similarity to the eigenvector and the similarity between the two.

[0101] S323. Convert the index value of the region to the index value of the full map.

[0102] In the case of the embodiments of the present application, since the k-nearest neighbor search is performed in the local search area of the feature map , the returned index value is the area index value. Therefore, it is still necessary to convert the index value into the full-map index value. The specific conversion process includes: according to the upper-left coordinate ( , , ) of the local search area in the entire feature map , the width of the local search area, the width of the entire feature map , and the area index value returned by the k-nearest neighbor search, the full-map index value

[0103]

[0104] is obtained. The specific calculation method is: S324. According to the full-map index value, obtain the position information of the corresponding feature vector from the feature map

[0105] S330. Integrate the position information and the similarity to form the similarity information of the feature vector in the feature map .

[0106] In the embodiments of the present application, the above operations are performed on all the feature vectors in the feature map , and the similarity information corresponding to all the feature vectors can be obtained. The set of all the similarity information forms the matching information, which provides support for the subsequent iterative estimation of the optical flow.

[0107] S400. Utilize the context information and the matching information, and perform iterative optical flow estimation through a preset recurrent network, and output the optical flow estimation result.

[0108] In the present application embodiment, the context information includes the deep feature information obtained by feature extraction for each pixel point and the environmental information where the pixel point is located. The environmental information includes the position information of the pixel point in the image and the association information between pixel points; the matching information is a set of similarity information obtained through local search matching, which can help the optical flow estimation to bias towards the position with the highest matching degree, that is, the position closer to the change state of the pixel motion vector matrix; the preset recurrent network uses GRU (Gated Recurrent Unit). GRU can better help the network to perform iteration by selectively retaining the historical node information. The iterative optical flow estimation through the recurrent network specifically includes the following steps:

[0109] S410. Obtain the initial value of the optical flow estimation.

[0110] In the embodiment of the present application, an initial value of optical flow estimation is obtained, that is, an initial value assignment is performed for optical flow estimation. Since the optical flow estimation of each iteration will be used as the input of the next iteration, it is necessary to perform an initial value assignment for optical flow estimation. In this way, in the first iteration, the initial value of optical flow estimation can be used as input information. At the same time, the residual optical flow obtained through the recurrent network combined with the initial value of optical flow estimation is the optical flow estimation of the first iteration. The process of continuous iteration through the recurrent network is actually a process of approaching the initial value of optical flow estimation to the true optical flow.

[0111] S420. Obtain the input information for the first iteration according to the initial value of optical flow estimation, matching information, and context information.

[0112] In the embodiment of the present application, the initial value of optical flow estimation is assigned as 0. The initial value of optical flow estimation represents a pixel motion vector matrix, and the initial assignment is 0, that is, each element in the matrix is assigned 0.

[0113] S430. Obtain the optical flow estimation through iteration by the recurrent network according to the input information.

[0114] In the embodiment of the present application, the optical flow estimation is obtained through iteration by the recurrent network. Refer to Figure 4 , and it specifically includes the following steps:

[0115] S431. Obtain the optical flow estimation of the previous iteration.

[0116] If it is the first iteration, the optical flow estimation of the previous iteration is the initial value of optical flow estimation.

[0117] S432. Obtain the input and the historical hidden layer state of the current iteration according to the optical flow estimation of the previous iteration, matching information, and context information;

[0118] Among them, is the result of the previous optical flow estimation , matching information and context information fusion, and the specific manifestation is:

[0119]

[0120]

[0121] is the accumulated information after several previous iterations. If it is the first iteration, then the initial value is the context information .

[0122] S433. Obtain the updated hidden layer state through a recurrent network according to the input of the current iteration and the historical hidden layer state . .

[0123] In the embodiment of the present application, the updated hidden layer state is obtained through a recurrent network , which specifically includes:

[0124] According to and , obtain the update gate state and the reset gate state respectively, which is specifically manifested as:

[0125]

[0126]

[0127] According to , and the reset gate state obtain the candidate hidden layer state . The candidate hidden layer state includes the current input information and the reserved information of the hidden layer state of the previous node targetedly. The state of the reset gate determines the amount of reserved information, which is specifically manifested as:

[0128]

[0129] According to the candidate hidden layer state , the hidden layer state of the previous node and the update gate state , obtain the updated hidden layer state . The updated hidden layer state includes the selective reservation of the hidden layer state of the previous node and the candidate hidden layer state of the current node , which is specifically manifested as:

[0130]

[0131] S434. After several convolutions on the updated hidden layer state , obtain the residual optical flow, and update the optical flow estimate of the previous iteration according to the residual optical flow to obtain the optical flow estimate of the current iteration

[0132] Among them, the residual optical flow , which can be understood as the update direction, for the optical flow estimate value of the previous iteration Perform an update to obtain the optical flow estimation value for the current iteration, which is specifically manifested as follows:

[0133]

[0134] To better illustrate the technical effects of the present invention, the inventor also conducted the following experiments:

[0135] The datasets used in the experiments include: publicly available large-scale datasets such as FlyingChairs, Sintel, MPI-Sintel, etc.

[0136] The evaluation metric used in the experiment is EPE (Endpoint error), which represents the average of the Euclidean distances between the estimated optical flow and the true optical flow for all pixels.

[0137] Experiment 1: Under the condition that the GRU network loops 2 times and top_k is set to 2, the local search region sizes are successively defined as 1 / 4, 1 / 16, and 1 / 64 of the original image, and then experiments are conducted separately. The time consumption of each experimental group is as follows:

[0138] Table 1:

[0139]

[0140] Experiment 2: Under the condition that the GRU network loops 2 times and top_k is set to 2 and 8 respectively, the local search region sizes are successively defined as 1 / 4, 1 / 16, and 1 / 64 of the original image, and then experiments are conducted separately. The optical flow estimation accuracy of each experimental group is as follows:

[0141] Table 2:

[0142]

[0143] It is not difficult to see from the above two groups of experiments that without changing the network structure and network weights, and only changing the search method, the search time consumption of local search is greatly reduced compared with global search. Among them, when the search range is 1 / 16 of the original image, its search time consumption is only 16.1% of that of global search, and the overall time consumption of its optical flow estimation is 78.6% of that of global search. And in the case of local search, the accuracy of optical flow estimation does not decrease significantly, but instead improves in some cases. Among them, when top_k = 2, the optical flow estimation accuracy when the search range is Figure 1 1 / 16 is significantly better than that of full-image search, which indicates that the local search strategy can improve the real-time performance of optical flow estimation, and at the same time, the accuracy of optical flow estimation does not decrease significantly.

[0144] The embodiment of the present application further provides a computer-readable storage medium, storing a computer program that can be loaded and executed by a processor to perform any one of the above-mentioned fast optical flow estimation methods for moving small targets based on local search.

[0145] The embodiments of the specific implementation manners are all preferred embodiments of the present application, and do not limit the protection scope of the present application. Therefore, all equivalent changes made according to the principle of the present application shall be covered within the protection scope of the present application.

Claims

1. A fast optical flow estimation method for small moving targets based on local search, characterized in that, Including: Obtain two adjacent frames of images, extract features from the two frames of images, obtain feature maps f1 and f2 corresponding to the two frames of images respectively, and separately encode the first frame of image to obtain context information; Match a corresponding local search area in the feature map f2 for each feature vector in the feature map f1; Sequentially match corresponding similarity information from the corresponding local search area according to each feature vector in the feature map f1, and aggregate all the similarity information into matching information; Utilize the context information and the matching information to perform iterative optical flow estimation through a preset recurrent network, and output an optical flow estimation result; Among them, matching a corresponding local search area in the feature map f2 for each feature vector in the feature map f1 includes: Divide the feature map f1 to form multiple consecutive regions a1 to a n , where n is a positive integer; For any region a i Find the mapped region a at the same position in the feature map f2 i ′ ; Expand the mapped region a i ′ to obtain the expanded region b i , where region b i is the local search region in feature map f2 for all feature vectors within the a i region in feature map f1; For the mapped area a i ′ Expand it to obtain the expanded area b i , including: Obtain the size of the moving small target and the displacement amount between the two frames of images; Calculate the extended side length according to the size of the moving small target and the displacement amount between the two frames of images; Obtain the mapping area a i ′ Position information in the feature map f2; According to the mapping region a i ′ Based on the position information in the feature map f2 and the extended side length, the extended region b is obtained i ; Among them, the extended side length w add is calculated as follows: Among them, target w is the size of the small motion target, and target x is the displacement of the small motion target. down_sampling is the number of downsamplings in the multi-layer feature extraction network.

2. The fast optical flow estimation method for small moving targets based on local search according to claim 1, characterized in that, Separately encoding the first frame of image to obtain context information includes: Perform multi-scale feature extraction on the first frame of image according to a preset context network to separately obtain local features and global features; Fuse the local features and the global features to obtain context information.

3. The fast optical flow estimation method for small moving targets based on local search according to claim 1, characterized in that, According to the mapping region a i ′ Based on the position information in the feature map f2 and the extended side length, the extended region b is obtained i , including: According to the mapping region a i ′ Obtain the expanded region b based on the position information of the feature map f2 and the expanded side length i The position information of the feature map f2; According to the expanded region b i Based on the position information of the feature map f2, determine whether the expanded region b i exceeds the range of the feature map f2. If there is an exceeding part, readjust the position information of the expanded region b i .

4. The fast optical flow estimation method for small moving targets based on local search according to claim 1, characterized in that, Matching corresponding similarity information from the corresponding local search area according to each feature vector in the feature map f1 includes: According to the feature vector α in the feature map f1 ij , obtain all the feature vectors within its corresponding local search area and form a feature vector set β; Match k feature vectors with the highest similarity to the feature vector α from the feature vector set β ij and obtain the position information and similarity of the matched k feature vectors; Integrate the position information and similarity of the k matched feature vectors to form the similarity information of the feature vector α in the feature map f1. ij ​ 5. The fast optical flow estimation method for small moving targets based on local search according to claim 4, characterized in that, Match the k feature vectors with the highest similarity to the feature vector α from the feature vector set β ij and obtain the position information and similarity of the k matched feature vectors, including: Construct an index for the set of feature vectors β; Perform a k-nearest neighbor search on the feature vector α ij and the set of feature vectors β to obtain the region index values and similarities of the k feature vectors with the highest similarity to the feature vector α ij ; Convert the index value of the region to the index value of the full image; According to the full-image index value, obtain the position information of k feature vectors matching the feature vector α from the feature map f2 ij ​ 6. A fast optical flow estimation method for small moving targets based on local search, characterized in that Utilize the context information and the matching information to perform iterative optical flow estimation through a preset recurrent network, including: Obtain an initial value of optical flow estimation; According to the initial value of optical flow estimation, the matching information and the context information, obtain input information; According to the input information, perform iteration through a recurrent network to obtain optical flow estimation.

7. A fast optical flow estimation method for small moving targets based on local search according to claim 6, characterized in that Performing iteration through a recurrent network to obtain optical flow estimation includes: Obtain the optical flow estimation of the previous iteration; Obtain the input x of the current iteration based on the optical flow estimation, matching information, and context information of the previous iteration t and the historical hidden layer state h t-1 ; Based on the input x of the current iteration t and the historical hidden layer state h t-1 , obtain the updated hidden layer state h through a recurrent network t ; For the updated hidden layer state h t The residual optical flow is obtained through several convolutions. Based on the residual optical flow, the optical flow estimate of the previous iteration is updated to obtain the optical flow estimate of the current iteration.

8. A readable storage medium, characterized in that , storing a computer program that can be loaded and executed by a processor to perform the fast optical flow estimation method for moving small targets based on local search according to any one of claims 1 to 7.