Video image feature point detection method and system, equipment and storage medium
By dividing the video image into chunks and extracting the initial feature points in combination with the feature points in the previous frame, the problem of uneven distribution of feature points in the ORB algorithm in image matching is solved, and more accurate feature point detection and higher image matching success rate are achieved.
Patent Information
- Application Number
- CN202311444736.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the ORB algorithm has the phenomenon of accidentally detecting edge points and pseudo-corner points in image matching, resulting in uneven distribution of feature points and affecting the success rate of image matching.
By dividing the video image into multiple blocks, the initial feature points of each block are generated, and the initial feature points are extracted in combination with the feature points acquisition situation in the previous video image to obtain the feature points of the block.
This achieves more accurate detection of feature points of the current frame video image, avoids the problem of uneven distribution of feature points, and improves the success rate of image matching.
Smart Images

Figure CN119942393A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a method and system, device and storage medium for detecting feature points of a video image. Background Art
[0002] Image matching algorithm is an essential key step in computer vision and is widely used in areas such as drone visual navigation, target detection and tracking. Image matching algorithms can generally be divided into two categories: region-based image matching algorithms and feature-based matching algorithms. Feature-based matching algorithms have become a hot topic of research due to their low computational complexity and good robustness. However, feature extraction is a crucial step in feature matching algorithms, which directly affects the success rate of image matching.
[0003] Common feature extraction algorithms include SIFT, SURF, and ORB algorithms. Among them, SIFT has good distinguishability, but the algorithm is too complex and requires a lot of calculations; the SURF algorithm has the advantages of accurate parameter estimation and low calculations, but the number of matching point pairs obtained is small; the ORB image feature extraction algorithm uses FAST to calculate key points, so it has the fastest computing speed and consumes the least storage space. Its computing time is about one hundredth of SIFT and one tenth of SURF. However, its robustness is not as good as SIFT, and it does not have scale invariance, which can easily cause mismatches in images.
[0004] Most existing ORB algorithms use FAST9-16 to extract feature points. Although the calculation speed is fast, it will cause false detection of some edge points, resulting in the existence of some pseudo corner points, which directly interferes with the matching effect and causes false matching. Moreover, since features are extracted for the entire image, there are often phenomena such as too concentrated feature points in texture-rich areas and no feature points can be extracted in texture-deficient areas, which leads to uneven distribution of feature points and indirectly affects the success rate of image matching. Summary of the invention
[0005] The problem solved by the embodiments of the present invention is to provide a method and system, a device and a storage medium for detecting feature points of a video image, which are conducive to realizing more accurate detection of feature points of a video image of a current frame.
[0006] To solve the above problems, an embodiment of the present invention provides a method for detecting feature points of a video image, comprising: acquiring a video image of a current frame; dividing the video image of the current frame into a plurality of blocks; acquiring feature points for each block in turn; wherein the feature point acquisition comprises: generating initial feature points of the blocks; extracting the initial feature points in combination with the feature point acquisition in the video image of the previous frame to obtain feature points of the blocks.
[0007] Correspondingly, an embodiment of the present invention also provides a video image feature point detection system, including: a video image acquisition module, used to acquire a video image of a current frame; a division module, used to divide the video image of the current frame into multiple blocks; a feature point acquisition module, used to acquire feature points for each block in turn; wherein the feature point acquisition module includes: an initial feature point generation unit, used to generate initial feature points for blocks; a feature point extraction unit, used to extract the initial feature points in combination with the feature point acquisition situation in the video image of the previous frame, to obtain feature points for the blocks.
[0008] Correspondingly, an embodiment of the present invention also provides a device, comprising at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for detecting feature points of video images provided in an embodiment of the present invention.
[0009] Correspondingly, an embodiment of the present invention further provides a storage medium, which stores one or more computer instructions, and the one or more computer instructions are used to implement the method for detecting feature points of a video image provided by an embodiment of the present invention.
[0010] Compared with the prior art, the technical solution of the embodiment of the present invention has the following advantages:
[0011] In the method for detecting feature points of video images provided by an embodiment of the present invention, feature point acquisition includes: generating initial feature points of blocks, extracting the initial feature points in combination with the feature point acquisition situation in the video image of the previous frame, and obtaining feature points of the blocks; in the embodiment of the present invention, in a video image sequence, two adjacent frames of video images have a large correlation, then the initial feature points are extracted in combination with the feature point acquisition situation in the video image of the previous frame to obtain the feature points of the blocks, which is conducive to real-time processing of the extraction of the initial feature points according to the actual situation of the video image, obtaining feature points that are more closely matched with the actual image feature situation of the video image, and facilitating more accurate detection of the feature points of the video image of the current frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flow chart of an embodiment of a method for detecting feature points of a video image according to the present invention;
[0013] Figures 2 to 7 It is a schematic diagram corresponding to each step in an embodiment of a method for detecting feature points of a video image of the present invention;
[0014] Figures 8 to 9 is a functional block diagram of an embodiment of a detection system for feature points of a video image according to the present invention;
[0015] Fig.10It is a hardware structure diagram of an embodiment of the device provided by the present invention. DETAILED DESCRIPTION
[0016] As can be seen from the background technology, most images are stored in the format of pixel value matrix (RGB) in computers. It is difficult to accurately find the same object in multiple video images only by the pixel value of the image. The main reason is that the pixel value of the image itself is very easily affected by the change of illumination during shooting, and when the shooting angle changes (such as rotation, translation, etc.), the pixel value of the same object will also change accordingly. Therefore, it is necessary to find a feature that can remain stable when the camera rotates, translates and the illumination changes to identify these objects, and finally use the robust feature to find the same object in different video images.
[0017] To this end, computer vision researchers have designed many highly robust feature points that do not change with the shooting distance, camera rotation, and lighting changes. For example, common feature extraction algorithms include SIFT, SURF, and ORB algorithms.
[0018] Although the traditional ORB algorithm has fast processing speed and good real-time performance, it has the following problem: the feature points finally extracted tend to be concentrated in areas with rich textures or strong features, while for weak texture areas with few features, the number of feature points will be very small or even non-existent, which will cause some feature points in the feature-rich area to be useless. Originally, one feature point can clearly express a small area, and most of the remaining feature points are redundant. The low-texture area lacks a feature to describe this information. From a global perspective, the distribution of overall feature points is also uneven.
[0019] In order to solve the technical problem, an embodiment of the present invention provides a method for detecting feature points of a video image. Figure 1 , shows a flow chart of an embodiment of a method for detecting feature points of a video image of the present invention.
[0020] In this embodiment, the method for detecting feature points of a video image includes the following basic steps:
[0021] Step S1: Obtain the video image of the current frame;
[0022] Step S2: Divide the video image of the current frame into multiple blocks;
[0023] Step S3: acquiring feature points for each block in turn;
[0024] The step S3 of acquiring feature points includes:
[0025] Step S31: generating initial feature points of the blocks;
[0026] Step S32: extracting initial feature points based on the feature point acquisition in the video image of the previous frame to obtain feature points of the blocks.
[0027] In an embodiment of the present invention, in a video image sequence, two adjacent frames of video images have a large correlation. In this way, the initial feature points are extracted in combination with the feature point acquisition in the previous frame of the video image to obtain the feature points of the blocks. This is beneficial for real-time processing of the extraction of the initial feature points according to the actual situation of the video image, and for obtaining feature points that are more closely matched with the actual image feature situation of the video image, which is beneficial for achieving more accurate detection of the feature points of the video image of the current frame.
[0028] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present invention more obvious and understandable, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0029] Figures 2 to 7 It is a schematic diagram corresponding to each step in an embodiment of a method for detecting feature points of a video image of the present invention.
[0030] refer to Figure 2 , execute step S1: obtain the video image 100 of the current frame.
[0031] In computer vision applications, the primary problem is how to efficiently and accurately identify the same object in multiple video images.
[0032] Specifically, in video images, selecting representative areas (areas with strong features) can better match the same objects between multiple video images, such as feature points, strong edges, line features, etc. in video images. One of the representative features in video images is the corner point, and its recognizable image algorithm is well understood and intuitive. Therefore, in most computer vision processing, corner points are extracted as features, also known as "feature points", and feature points are used to match video images, such as motion image stabilization, panoramic image stitching, visual SLAM, etc.
[0033] In this embodiment, the video image 100 of the current frame is obtained to obtain feature points of the video image 100 of the current frame to perform image capture.
[0034] In this embodiment, when acquiring the video image 100 of the current frame, it is assumed that the video image 100 has a preset number of feature points.
[0035] The video image 100 is assumed to have a preset number of feature points, which are used as an evaluation benchmark for determining a mode to be adopted for subsequent intra-frame measurement of the feature points.
[0036] It should be noted that, in this embodiment, the preset number of feature points of the current frame of video image 100 is set according to the total number of feature points finally obtained in the previous frame of video image.
[0037] refer to Figure 3 , execute step S2: divide the video image 100 of the current frame into multiple blocks 100a.
[0038] The multiple blocks 100a are used for subsequent feature point extraction of each block 100a. Compared with the subsequent direct feature point extraction of the entire current frame video image, in this embodiment, the current frame video image 100 is first divided into multiple blocks 100a, and then feature points are extracted from each block 100a respectively. This is conducive to extracting feature points from the current frame video image 100 globally as much as possible, and is conducive to adjusting the feature point extraction of each block 100a. For example, by adjusting, some feature points can be extracted in the weak texture area of the video image 100 to describe the area.
[0039] In this embodiment, the video image 100 of the current frame is divided into a plurality of blocks 100 a , and the size of each divided block 100 a is set according to the size of the video image 100 of the current frame.
[0040] It should be noted that the size of each divided block 100a should not be too large or too small. If the size of the divided blocks 100a is set too small, the complexity of detecting each block 100a increases, and it is also easy to cause the video image 100 of the current frame to be divided too densely, increasing the amount of calculation, reducing the detection efficiency, and increasing the detection cost; if the size of the divided blocks 100a is set too large, it is easy to cause large unevenness in the subsequent feature point extraction of the blocks 100a, affecting the accuracy of the feature point extraction of the video image 100 of the current frame. For this reason, in this embodiment, the size of each divided block 100a is set according to the size of the video image 100 of the current frame.
[0041] As an example, for images larger than 1080P, the size of block 100a of 30×30 is too small, which will produce an aperture effect, and the size of block 100a needs to be set larger than 30×30. For 720P images, the size of block 100a of 60×60 is too large, which may easily cause unevenness within the block and lead to uneven distribution of overall feature points. The size of block 100a needs to be set smaller than 60×60. Therefore, the size of block 100a needs to be adjusted according to the size of the video image.
[0042] In this embodiment, the video image 100 of the current frame is divided into a plurality of blocks 100a. The video image 100 is evenly divided into the plurality of blocks 100a, which is conducive to making the subsequent extraction of feature points of the blocks 100a more uniform.
[0043] It should be noted that, in this embodiment, the multiple blocks 100a into which the video image 100 of the current frame is divided do not overlap.
[0044] Combined with reference Figures 4 to 7 , execute step S3, and obtain feature points 100b for each block 100a in turn.
[0045] It should be noted that acquiring feature points 100b for each block 100a in turn refers to the step of acquiring feature points 100a for the next block 100a after acquiring feature points 100b for one block 100a along a specified direction, wherein the specified direction may be scanning horizontally, and after acquiring feature points 100b for a row of blocks 100a, scanning the next row of blocks 100a is performed until feature points 100b are acquired for all blocks 100a.
[0046] Feature points 100b are acquired for each block 100a. The feature points 100b are used as representative points in the video image and are used for subsequent matching of the same object in multiple video images.
[0047] Specifically, the feature point algorithms usually include: SIFT, SURF and ORB algorithms, etc. The feature point algorithm of video images mainly consists of two parts: the location of feature points (Keypoint Location) and the generation of descriptors (Descriptor Generation). The location of feature points refers to obtaining the position (i.e., coordinates) of the feature point in the video image through the definition of the feature point algorithm, and some key information with direction and scale will be obtained; the generation of descriptors will generate a descriptor for each located feature point. The result of the descriptor is usually a multi-dimensional vector, designed according to an artificial algorithm. The descriptor is used to describe the information of the pixels around the feature point, which is equivalent to giving the feature point an independent ID card to facilitate its identification or distinguish different feature points. When matching feature points, the feature point descriptors are generated according to the obtained feature point positions, and the differences between the descriptors are compared for matching.
[0048] In this embodiment, the sampling ORB algorithm is used to obtain feature points as an example for explanation.
[0049] The step S3 of acquiring the feature point 100b includes: executing step S31 to generate the initial feature point of the block 100a.
[0050] The initial feature points of the block 100a are generated for subsequent extraction of the final feature points of the block 100a.
[0051] In this embodiment, the FAST algorithm is used to generate the initial feature points of the block 100a, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold.
[0052] Among them, the FAST (Features from accelerated segment test) algorithm is an algorithm for corner detection. The principle of the algorithm is to take a detection point in the image and determine whether the detection point is a corner point based on the pixels in the surrounding neighborhood with the point as the center. That is to say, if a certain number of pixels around a pixel have different pixel values from the point, it is considered to be a corner point, that is, the initial feature point.
[0053] It should be noted that the detection threshold in the FAST algorithm includes a high threshold and a low threshold, which means that the detection threshold in the FAST algorithm is a dual threshold, and in the dual threshold, the higher value is the high threshold, and the lower value is the low threshold.
[0054] Specifically, in this embodiment, the FAST algorithm is used to detect block 100a, and the pixel point with a grayscale value difference between the surrounding pixels and the pixel point is greater than or equal to the detection threshold is generated as the initial feature point. Correspondingly, the detection threshold is the grayscale difference, the high threshold is a larger grayscale difference, and the low threshold is a smaller grayscale difference.
[0055] The FAST algorithm is used to generate the initial feature points of the block 100a, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold. This enables the initial feature points to be generated by the low threshold when the initial feature points cannot be generated by the high threshold in the weak texture area of the video image 100. The generation of the initial feature points correspondingly affects the number and positioning of the final feature points 100b. In the extraction of the feature points 100b, if the extracted feature points 100b are only concentrated in the area with rich texture or strong features, then for a block 100a, one feature point 100b can clearly express an area, and the remaining feature points 100b are redundant. This may easily lead to unnecessary waste and increased costs. In addition, the low-texture area lacks feature points 100b to describe the information of the block 100a. From a global perspective, the distribution of the overall feature points 100b is also uneven. Therefore, the detection threshold in the FAST algorithm includes a high threshold and a low threshold, which is beneficial to avoid as much as possible that the final extracted feature points 100b are concentrated only in areas with rich textures or strong features, while the number of feature points 100b in the weak-texture area lacking features will be very small or even non-existent. This is beneficial to better improve the extraction of feature points 100b in weak-texture areas, so that the distribution of feature points 100b in the current frame video image 100 is more uniform.
[0056] Specifically, in this embodiment, the FAST algorithm is used to generate the initial feature points of the block 100a, including: using a high threshold and a low threshold to simultaneously detect the block 100a.
[0057] It should be noted that the use of a high threshold and a low threshold to simultaneously detect the block 100a means that high threshold detection and low threshold detection are performed on the block 100a at the same time.
[0058] By using a high threshold and a low threshold to detect the block 100a simultaneously, there is no need to perform the step of detecting all the blocks 100a twice. Only one detection is required for all the blocks 100a, and the high threshold detection and the low threshold detection for each block 100a can be completed, which is beneficial to improving the efficiency of generating initial feature points.
[0059] In this embodiment, when the initial feature points can be detected by using the high threshold, the low threshold detection is stopped, and the initial feature points detected by the high threshold are obtained.
[0060] When the initial feature point can be detected by using a high threshold, stopping the low threshold detection is beneficial to avoid the step of repeatedly reading the content of the same block 100a, reducing unnecessary calculations in real time, and improving detection efficiency and reducing detection costs.
[0061] In this embodiment, when the initial feature points cannot be detected by using a high threshold value, but can be detected by using a low threshold value, the initial feature points detected by using the low threshold value are obtained.
[0062] When the initial feature points cannot be detected by using a high threshold value, but can be detected by using a low threshold value, that is, when the texture of block 100a is weak, low threshold detection can still be performed in real time to obtain the initial feature points.
[0063] In this embodiment, when the initial feature point cannot be detected by using the high threshold and the low threshold, the step of acquiring the feature point 100b for the block 100a is skipped.
[0064] When the initial feature point cannot be detected by using both the high threshold and the low threshold, that is, the feature of block 100b is too unclear, and there is no need to extract feature point 100b, thereby directly skipping the step of acquiring feature point 100b for block 100a, saving computing power.
[0065] The step S3 of acquiring the feature point 100b further includes: executing step S32, extracting the initial feature point in combination with the acquisition of the feature point 101b in the video image 101 of the previous frame, and obtaining the feature point 100b of the block 100a.
[0066] Combined with the acquisition of the feature point 101b in the video image 101 of the previous frame, the initial feature point is extracted to obtain the feature point 100b of the block 100a, which is the inter-frame feature point allocation algorithm.
[0067] In the present embodiment, in the video image sequence, two adjacent frames of video images have a large correlation, and the initial feature points are extracted in combination with the acquisition of the feature points 101b in the video image 101 of the previous frame to obtain the feature points 100b of the block 100a. This is beneficial for real-time processing of the extraction of the initial feature points according to the actual situation of the video image, and obtaining the feature points 100a that are more closely matched with the actual image features of the video image, which is beneficial for achieving more accurate detection of the feature points 100a of the video image 100 of the current frame.
[0068] In this embodiment, the initial feature points are extracted in combination with the acquisition of the feature points 101b in the previous frame of the video image 101 to obtain the feature points 100b of the block 100a, including: obtaining information based on the feature points 101b of the corresponding block 101a in the previous frame of the video image 101 of the current block 100a to obtain the estimated number of feature points 100b of the current block 100a.
[0069] The estimated number of feature points 100b of the current block 100a is obtained, which is used as a reference for obtaining the actual feature points 100b of the current block 100a. Moreover, information is obtained based on the feature points 101b of the corresponding block 101a in the previous frame of video image 101 of the current block 100a, which is beneficial to obtain the estimated number of feature points 100a that are more closely matched with the actual image features of the video image based on the actual situation of the video image.
[0070] Specifically, refer to Figure 4 , obtaining information based on the feature point 101b of the current block 100a corresponding to the block 101a in the previous frame of video image 101, and obtaining the estimated number of feature points 100b of the current block 100a, including: obtaining the block 101a corresponding to the current block in the previous frame of video image 101 as a reference block.
[0071] The reference block is used as a basis for obtaining an estimated number of feature points 100b of the current block 100a.
[0072] In this embodiment, blocks in the neighborhood of the reference block in the previous frame of video image 101 are obtained as surrounding blocks.
[0073] The surrounding blocks are used together with the reference blocks as a basis for obtaining the estimated number of feature points 100b of the current block 100a.
[0074] In this embodiment, the block 101a in the neighborhood of the reference block in the previous frame of video image 101 is obtained as the surrounding block, and the neighborhood range of the reference block is set according to the global motion vector of the previous frame of video image 101, wherein the size of the global motion vector is proportional to the size of the neighborhood range.
[0075] It should be noted that the size of the global motion vector is proportional to the size of the neighborhood range, which means that the larger the global motion vector is, the larger the neighborhood range is set, and the smaller the global motion vector is, the smaller the neighborhood range is set.
[0076] The global motion vector (GMV) represents the motion speed of the previous frame of video image 101. The larger the global motion vector is, the greater the motion speed of the previous frame of video image 101 is, and the greater the image change is, so a larger range needs to be referenced. The smaller the global motion vector is, the smaller the motion speed of the previous frame of video image 101 is, and the smaller the image change is, so a smaller range can be referenced. Therefore, the size of the global motion vector is proportional to the size of the neighborhood range.
[0077] As an example, when the global motion vector is relatively large, the reference block neighborhood is selected as the surrounding 24 blocks 101a as the reference block, and when the global motion vector is relatively small, the reference block neighborhood is selected as the surrounding 8 blocks 101a as the reference block.
[0078] In this embodiment, the feature point acquisition information of the reference block and the surrounding blocks is combined to obtain the estimated number of feature points 100b of the current block 100a.
[0079] If there is a certain spatial correlation in the video image, the feature point acquisition information of the surrounding blocks will also affect the estimated number of feature points 100b of the current block 100a. Therefore, the estimated number of feature points 100b of the current block 100a is obtained by combining the feature point acquisition information of the reference block and the surrounding blocks.
[0080] In this embodiment, information is obtained by combining the feature points 101b of the reference block and the surrounding blocks to obtain an estimated number of feature points 100b of the current block 100a, including: information is obtained by combining the feature points 101b of the reference block and the surrounding blocks, and the blocks 100a of the video image 100 of the current frame are graded to obtain the grade corresponding to each block 100a, wherein each grade has a corresponding estimated number.
[0081] By classifying the blocks 100a of the video image 100 of the current frame, all the blocks 100a in the video image 100 of the current frame can be classified, and the blocks 100a with similar features can be classified into the same category. Then, the same estimated number can be assigned to the blocks 100a of the same category, which is beneficial for obtaining the estimated numbers of all the blocks 100a at the same time, and there is no need to assign an estimated number to each block 100a separately, thereby reducing the number of times the corresponding estimated number is assigned to the blocks 100a, which is beneficial for improving computing efficiency and reducing computing costs.
[0082] Accordingly, in this embodiment, the estimated number of feature points 100b of each block 100a is obtained according to the level corresponding to each block 100a.
[0083] In this embodiment, information is obtained by combining the feature points 101b of the reference block and the surrounding blocks to obtain the estimated number of feature points 100b of the current block 100a, including: obtaining whether the reference block has the feature point 101b as the first information.
[0084] Whether the reference block has the feature point 101 b represents the texture of the reference block, and can correspondingly provide feedback on the texture of the current block 100 a in the video image 100 of the current frame.
[0085] In this embodiment, when the initial feature points of the reference block generated by the FAST algorithm are obtained, the situation in which the detection threshold is a high threshold or a low threshold is used as the second information.
[0086] When the FAST algorithm is used to generate the initial feature points of the reference block, the detection threshold is a high threshold or a low threshold, which further finely characterizes the texture of the reference block, and correspondingly can further finely feedback the texture of the current block 100a in the video image 100 of the current frame.
[0087] In this embodiment, whether the number of feature points 101b in the surrounding blocks is greater than or equal to a set threshold is obtained as the third information.
[0088] The case where the number of feature points 101b in the surrounding blocks is greater than or equal to the set threshold means that some blocks 101a in the surrounding blocks may not have feature points 101b, and some blocks 101a have feature points 101b. The case where the number of feature points 101b in the surrounding blocks is greater than or equal to the set threshold is obtained.
[0089] Whether the number of feature points 101b in the surrounding blocks is greater than or equal to a set threshold represents the texture distribution of the surrounding blocks, and can correspondingly feedback the texture distribution around the current block 100a in the video image 100 of the current frame.
[0090] In this embodiment, when obtaining whether the number of the characteristic points 101b in the surrounding blocks is greater than or equal to the set threshold, the set threshold is half of the total number of the surrounding blocks.
[0091] The threshold is set to half of the total number of surrounding blocks, that is, whether the number of surrounding blocks with feature points 101b exceeds half. If it exceeds half, it indicates that the texture of the surrounding blocks is strong. If it does not exceed half, it indicates that the texture of the surrounding blocks is weak.
[0092] In this embodiment, the estimated number of feature points 100b of the current block 100a is obtained by combining the first information, the second information and the third information.
[0093] Combining the first information, the second information and the third information to obtain the estimated number of feature points 100b of the current block 100a is beneficial to more comprehensively consider all factors affecting the distribution of feature points 100b of the current block 100a and obtain a more accurate estimated number of feature points 100b.
[0094] In this embodiment, according to the estimated number of feature points 100b of the current block 100a, a corresponding number of initial feature points are extracted as the feature points 100b of the block 100a.
[0095] Extracting a corresponding number of initial feature points based on the estimated number of feature points 100b of the current block 100a is conducive to obtaining feature points 100b that are relatively matched with the actual texture distribution of the video image 100 of the current frame.
[0096] refer to Figure 5 , according to the estimated number of feature points 100b of the current block 100a, extracting a corresponding number of initial feature points as feature points 100b of the block 100a, including: obtaining an adjustment value for the number of feature points 100b of the current block 100a according to the number of feature points 100b of all blocks 100a before the current block 100a and the estimated number of feature points 100b.
[0097] According to the number of feature points 100b of all blocks 100a before the current block 100a and the estimated number of feature points 100b, an adjustment value of the number of feature points 100b of the current block 100a is obtained, which is an intra-frame feature point allocation algorithm.
[0098] In the process of extracting initial feature points to obtain feature points 100a, there may be a mismatch between the number of initial feature points actually produced in each block 100a and the estimated number of feature points 100b. In particular, when the actual number of initial feature points produced is less than the estimated number of feature points 100b, it will make it difficult for the block 100a to extract a sufficient number of initial feature points as feature points 100b, resulting in a shortage of feature points 100b in the video image 100 of the current frame. Therefore, according to the number of feature points 100b of all blocks 100a before the current block 100a and the estimated number of feature points 100b, an adjustment value for the number of feature points 100b of the current block 100a is obtained, which is used to adjust the number of feature points 100b in the current block 100a and the subsequently processed blocks 100a, thereby reducing the probability of a shortage of feature points 100b in the video image 100 of the current frame.
[0099] refer to Figure 6 , according to the number of feature points 100b of all blocks 100a before the current block 100a and the estimated number of feature points 100b, obtaining an adjustment value for the number of feature points 100b of the current block 100a, including: according to the distribution of feature points 101b of the previous frame of video image 101, setting a selection mode for obtaining the adjustment value for the number of feature points 100b of the current block 100a, wherein the selection mode includes a uniform averaging mode or a greedy acquisition mode.
[0100] The distribution of the feature points 101b of the previous frame of video image 101 provides some feedback on the distribution of the feature points 101b of the current frame of video image 100 to a certain extent. Therefore, a selection mode for obtaining an adjustment value for the number of feature points 100b of the current block 100a is set according to the distribution of the feature points 101b of the previous frame of video image 101, so that a selection mode that is more in line with the actual texture of the current frame of video image 100 can be obtained.
[0101] Among them, the uniform averaging mode means that the adjustment value is uniformly assigned to the current block 100a and other blocks 100a to be processed subsequently, and the greedy acquisition mode means that for the current block 100a and other blocks 100a to be processed subsequently, the blocks 100a with stronger texture, i.e., stronger characteristics, are assigned larger adjustment values, and the blocks 100a with weaker texture, i.e., weaker characteristics, are assigned smaller adjustment values.
[0102] Correspondingly, in this embodiment, according to the selection mode, the adjustment value of the number of feature points 100b of the current block 100a is obtained.
[0103] In this embodiment, according to the distribution of the feature points 101b of the previous frame of video image 101, a selection mode for obtaining the feature point 100b quantity adjustment value of the current block 100a is set, including: determining whether the feature points 101b of the previous frame of video image 101 are evenly distributed.
[0104] Whether the feature points 101b of the previous frame video image 101 are evenly distributed indicates whether the feature points 100b of the current frame video image 100 are evenly distributed. Therefore, the selection mode for obtaining the adjustment value of the number of feature points 100b of the current block 100a is set according to whether the feature points 101b of the previous frame video image 101 are evenly distributed.
[0105] In this embodiment, if the feature points 101 b of the previous frame of video image 101 are evenly distributed, the selection mode for obtaining the adjustment value of the number of feature points 100 b of the current block 100 a is set to the even distribution mode.
[0106] If the feature points 101b of the previous frame video image 101 are evenly distributed, then the feature points 100b representing the current frame video image 100 are evenly distributed, so that the uniform averaging mode can be used to more evenly adjust the number of extracted initial feature points.
[0107] In this embodiment, otherwise (ie, if the feature points 101b of the previous frame of video image 101 are unevenly distributed), the selection mode for obtaining the adjustment value of the number of feature points 100b of the current block 100a is set to the greedy acquisition mode.
[0108] Otherwise, the feature points 100b representing the current frame video image 100 are unevenly distributed, so a greedy acquisition mode is adopted to fill in some areas where the feature points 100b are unevenly distributed, so that the number of extracted initial feature points can be adjusted more evenly.
[0109] In this embodiment, determining whether the feature points 101b of the previous frame of video image 101 are evenly distributed includes: performing one or more division processes on the previous frame of video image 101, and obtaining two sub-areas 10z with equal areas each time the division process.
[0110] The previous frame of video image 101 is divided once or multiple times, and each division process obtains two sub-regions 10z with equal areas. Then the area of each sub-region 10z is equal. In this embodiment, uniform distribution means that the number of feature points 100b in each sub-region 10z is as similar as possible.
[0111] As an example, Figure 6 As shown, the previous frame of video image 101 is divided into five parts, including: Figure 6 (a) Equally divided into upper and lower parts, such as Figure 6(b) Divide equally on the left and right sides, such as Figure 6 (c) The upper left and lower right are equally divided, such as Figure 6 (d) The upper right and lower left parts are equally divided, and Figure 6 (e) The inner and outer circles are divided equally to obtain 10 sub-regions 10z of equal area.
[0112] In this embodiment, the variance of the number of feature points 101b in each sub-region 10z among all the sub-regions 10z obtained by the division process is obtained.
[0113] As an example, the variance of the number of feature points 101b in 10 sub-regions 10z is obtained, which represents the degree of dispersion between the feature points 101b in the 10 sub-regions 10z and the average value.
[0114] In this embodiment, the total number of feature points 101b of the plurality of sub-regions 10z obtained by all division processes is obtained.
[0115] As an example, the total number of feature points 101b in 10 sub-areas 10z is obtained.
[0116] In this embodiment, the ratio of the variance to the sum is obtained.
[0117] As an example, the ratio of the variance of the number of feature points 101b in 10 sub-regions 10z to the total number is obtained, which represents the uniformity of the feature points 101b in the 10 sub-regions 10z.
[0118] It can be seen that the smaller the variance is, the larger the sum of the number of feature points 101b in multiple sub-areas 10z is, the smaller the ratio of the variance to the sum is, and the more uniform the distribution of the feature points 101b representing the previous frame of video image 101 is.
[0119] Correspondingly, in this embodiment, when the ratio of the variance to the sum is less than or equal to the uniformity threshold, it is judged that the feature points 101b of the previous frame of video image 101 are evenly distributed; when the ratio of the variance to the sum is greater than the uniformity threshold, it is judged that the feature points 101b of the previous frame of video image 101 are unevenly distributed.
[0120] In this embodiment, when the selection mode for obtaining the adjustment value of the number of feature points 100b of the current block 100a is set to the uniform averaging mode, the adjustment value of the number of feature points of the current block is obtained according to the selection mode, including: obtaining the difference between the estimated total number of feature points 100b of all blocks 100a before the current block 100a and the total number of feature points 100b as a deviation value.
[0121] In the process of extracting initial feature points to obtain feature points 100a, if the actual number of initial feature points generated is less than the estimated number of feature points 100b, it will cause the block 100a to be difficult to extract a sufficient number of initial feature points as feature points 100b, so that some blocks 100a before the current block 100a are missing feature points 100a. Therefore, the difference between the total estimated number of feature points 100b of all blocks 100a before the current block 100a and the total number of feature points 100b is obtained, that is, the total number of feature points 100b missing in all blocks 100a before the current block 100a is obtained, which is used to evenly distribute them to the current block 100a and all blocks 100a after the current block 100a. The pre-allocated plan is dynamically optimized again within the frame according to the actual situation, so that the number of extracted initial feature points further meets the demand.
[0122] In this embodiment, the number of the current block 100a and all the blocks 100a after the current block 100a is obtained.
[0123] The number of blocks 100a of the current block 100a and all blocks 100a after the current block 100a is obtained as the total number of feature points 100b for amortizing the deficit of all blocks 100a before the current block 100a.
[0124] In this embodiment, the ratio of the deviation value to the number of blocks 100a is used as the adjustment value of the number of feature points 100b of the current block 100a.
[0125] The ratio of the deviation value to the number of blocks 100a is used as the adjustment value of the number of feature points 100b of the current block 100a, so that all blocks 100a after the current block 100a are evenly adjusted.
[0126] In this embodiment, when the selection mode for obtaining the adjustment value of the number of feature points 100b of the current block 100a is set to the greedy acquisition mode, the adjustment value of the number of feature points 100b of the current block 100a is obtained according to the selection mode, including: obtaining the difference between the estimated total number of feature points 100b of all blocks 100a before the current block 100a and the total number of feature points 100b as a deviation value.
[0127] The explanation of the deviation value is the same as that when the uniform amortization mode is adopted, and will not be repeated here.
[0128] In this embodiment, the ratio of a preset number to the number of blocks 100a of the video image 100 of the current frame is used as a quantity threshold. When the estimated number of feature points 100b of the current block 100a is less than the quantity threshold, and the estimated number of feature points 100b of the current block 100a is less than the number of initial feature points, the difference between the quantity threshold and the estimated number of feature points 100b of the current block 100a is obtained as an adjustment value for the number of feature points 100b of the current block 100a.
[0129] The ratio of the preset number to the number of blocks 100a of the video image 100 of the current frame is the preset average number of feature points 100b of each block 100a, that is, the number threshold. The estimated number of feature points 100b of the current block 100a is less than the number threshold, and the estimated number of feature points 100b of the current block 100a is less than the number of initial feature points, which indicates that the features of the current block 100a are more obvious, and there is a margin of initial feature points that can be selected. Therefore, the current block 100a is allocated a larger feature point 100b adjustment value as much as possible. Specifically, the difference between the number threshold and the estimated number of feature points 100b of the current block 100a is obtained as the feature point 100b number adjustment value of the current block 100a. The operation is simple and the distribution tends to be uniform.
[0130] In this embodiment, when the estimated number of feature points 100b of the current block 100a is greater than or equal to the number threshold, or the estimated number of feature points 100b of the current block 100a is less than the number threshold, and the estimated number of feature points 100b of the current block 100a is greater than or equal to the number of initial feature points, the adjustment value of the number of feature points 100b of the current block 100a is 0.
[0131] The estimated number of feature points 100b of the current block 100a is greater than or equal to the number threshold, indicating that the features of the current block 100a are not obvious enough to allocate more feature points 100b. Therefore, the number adjustment value of feature points 100b of the current block 100a is 0. Alternatively, the estimated number of feature points 100b of the current block 100a is less than the number threshold, and the estimated number of feature points 100b of the current block 100a is greater than or equal to the number of initial feature points, indicating that the current block 100a does not have a surplus of initial feature points that can be selected. Therefore, the number adjustment value of feature points 100b of the current block 100a is 0.
[0132] In this embodiment, the target number of feature points 100b of the current block 100a is obtained by combining the adjustment value of the number of feature points 100b of the current block 100a and the estimated number of feature points 100b.
[0133] Combined with the adjustment value of the number of feature points 100b of the current block 100a and the estimated number of feature points 100b, the number of initial feature points to be extracted in the current block 100a is adjusted to obtain the target number of feature points 100b of the current block 100a, so that the feature points 100b of the current frame video image 100 are more uniform.
[0134] Specifically, in this embodiment, the target number of feature points 100b of the current block 100a is obtained by combining the adjustment value of the number of feature points 100b of the current block 100a and the estimated number of feature points 100b, including: obtaining the sum of the adjustment value of the number of feature points 100b of the current block 100a and the estimated number of feature points 100b as the target number of feature points 100b of the current block 100a.
[0135] refer to Figure 7 , according to the target quantity, the initial feature points are extracted as the feature points 100b of the block 100a.
[0136] According to the number of targets, initial feature points are extracted as feature points 100b of block 100a, and finally feature points 100b with relatively uniform distribution in the current frame video image 100 can be obtained.
[0137] In this embodiment, initial feature points are extracted as feature points 100b of block 100a according to the target number, including: when the target number of feature points 100b of the current block 100a is less than or equal to the number of initial feature points, an equal number of initial feature points is extracted as feature points 100b of block 100a.
[0138] When the target number of feature points 100b of the current block 100a is less than or equal to the number of initial feature points, initial feature points equal to the target number are extracted to meet the detection requirements.
[0139] In this embodiment, when the target number of feature points 100b of the current block 100a is greater than the number of initial feature points, all initial feature points are extracted as feature points 100b of the block 100a.
[0140] When the target number of feature points 100b of the current block 100a is greater than the number of initial feature points, the current block 100a does not have a surplus of initial feature points that can be selected, and therefore all initial feature points are extracted as feature points 100b of the block 100a.
[0141] Correspondingly, the present invention also provides a system for detecting feature points of video images. Figure 8 and Fig. 9 It is a functional block diagram of an embodiment of a detection system for video image feature points of the present invention.
[0142] In this embodiment, the detection system 50 of video image feature points includes: a video image acquisition module 501, used to acquire the video image of the current frame; a division module 502, used to divide the video image of the current frame into multiple blocks; a feature point acquisition module 503, used to acquire feature points for each block in turn; wherein the feature point acquisition module 503 includes: an initial feature point generation unit 5031, used to generate initial feature points for the blocks; a feature point extraction unit 5032, used to extract the initial feature points in combination with the feature point acquisition situation in the video image of the previous frame, so as to obtain the feature points of the blocks.
[0143] The video image acquisition module 501 is used to acquire the video image of the current frame.
[0144] In computer vision applications, the primary problem is how to efficiently and accurately identify the same object in multiple video images.
[0145] Specifically, in video images, selecting representative areas (areas with strong features) can better match the same objects between multiple video images, such as feature points, strong edges, line features, etc. in video images. One of the representative features in video images is the corner point, and its recognizable image algorithm is well understood and intuitive. Therefore, in most computer vision processing, corner points are extracted as features, also known as "feature points", and feature points are used to match video images, such as motion image stabilization, panoramic image stitching, visual SLAM, etc.
[0146] In this embodiment, a video image of the current frame is obtained to obtain feature points of the video image of the current frame to perform image capture.
[0147] In this embodiment, when acquiring a video image of a current frame, it is set that the video image has a preset number of feature points.
[0148] The video image is set to have a preset number of feature points, which are used as an evaluation benchmark for subsequent determination of a mode to be adopted for intra-frame measurement of the feature points.
[0149] It should be noted that, in this embodiment, the preset number of feature points of the current frame of video image is set according to the total number of feature points finally obtained in the previous frame of video image.
[0150] The division module 502 is used to divide the video image of the current frame into a plurality of blocks.
[0151] Multiple blocks are used for subsequent feature point extraction for each block. Compared with the subsequent direct feature point extraction for the entire current frame video image, in this embodiment, the current frame video image is first divided into multiple blocks, and then feature points are extracted for each block. This is conducive to extracting feature points from the current frame video image as globally as possible, and is conducive to adjusting the feature point extraction of each block. For example, by adjusting, some feature points can be extracted in the weak texture area of the video image to describe the area.
[0152] In this embodiment, the video image of the current frame is divided into a plurality of blocks, and the size of each divided block is set according to the size of the video image of the current frame.
[0153] It should be noted that the size of each divided block should not be too large or too small. If the size of the divided blocks is set too small, the complexity of detecting each block increases, and it is also easy to cause the video image of the current frame to be divided too densely, increasing the amount of calculation, reducing the detection efficiency, and increasing the detection cost; if the size of the divided blocks is set too large, it is easy to cause large unevenness in the subsequent feature point extraction of the blocks, affecting the accuracy of the feature point extraction of the video image of the current frame. For this reason, in this embodiment, the size of each divided block is set according to the size of the video image of the current frame.
[0154] As an example, for images larger than 1080P, a block size of 30×30 is too small and will produce an aperture effect. Therefore, the block size needs to be set to be larger than 30×30. For 720P images, a block size of 60×60 is too large and can easily cause unevenness within the block, leading to uneven distribution of overall feature points. Therefore, the block size needs to be set to be smaller than 60×60. Therefore, the block size needs to be adjusted according to the size of the video image.
[0155] In this embodiment, the video image of the current frame is divided into a plurality of blocks. The video image is evenly divided into a plurality of blocks, which is conducive to making the subsequent extraction of feature points of the blocks more uniform.
[0156] It should be noted that, in this embodiment, the multiple blocks into which the video image of the current frame is divided do not overlap.
[0157] The feature point acquisition module 503 is used to acquire feature points for each block in turn.
[0158] It should be noted that acquiring feature points for each block in turn refers to the step of acquiring feature points for one block along a specified direction, and then acquiring feature points for the next block. The specified direction may be scanning horizontally, and after acquiring feature points for one row of blocks, scanning the next row of blocks is performed until feature points are acquired for all blocks.
[0159] Feature points are acquired for each block. Feature points are used as representative points in the video image and are used to subsequently match the same object in multiple video images.
[0160] Specifically, the feature point algorithms usually include: SIFT, SURF and ORB algorithms, etc. The feature point algorithm of video images mainly consists of two parts: the location of feature points (Keypoint Location) and the generation of descriptors (Descriptor Generation). The location of feature points refers to obtaining the position (i.e., coordinates) of the feature point in the video image through the definition of the feature point algorithm, and some key information with direction and scale will be obtained; the generation of descriptors will generate a descriptor for each located feature point. The result of the descriptor is usually a multi-dimensional vector, designed according to an artificial algorithm. The descriptor is used to describe the information of the pixels around the feature point, which is equivalent to giving the feature point an independent ID card to facilitate its identification or distinguish different feature points. When matching feature points, the feature point descriptors are generated according to the obtained feature point positions, and the differences between the descriptors are compared for matching.
[0161] In this embodiment, the sampling ORB algorithm is used to obtain feature points as an example for explanation.
[0162] The feature point acquisition module 503 includes: an initial feature point generation unit 5031 for generating initial feature points of the blocks.
[0163] Generate the initial feature points of the block, which are then used to extract the final feature points of the block.
[0164] In this embodiment, the FAST algorithm is used to generate initial feature points of the blocks, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold.
[0165] Among them, the FAST (Features from accelerated segment test) algorithm is an algorithm for corner detection. The principle of the algorithm is to take a detection point in the image and determine whether the detection point is a corner point based on the pixels in the surrounding neighborhood with the point as the center. That is to say, if a certain number of pixels around a pixel have different pixel values from the point, it is considered to be a corner point, that is, the initial feature point.
[0166] It should be noted that the detection threshold in the FAST algorithm includes a high threshold and a low threshold, which means that the detection threshold in the FAST algorithm is a dual threshold, and in the dual threshold, the higher value is the high threshold, and the lower value is the low threshold.
[0167] Specifically, in this embodiment, the FAST algorithm is used to detect the blocks, and the pixel point with a grayscale value difference between the surrounding pixels and the pixel point is greater than or equal to the detection threshold is generated as the initial feature point. Correspondingly, the detection threshold is the grayscale difference, the high threshold is a larger grayscale difference, and the low threshold is a smaller grayscale difference.
[0168] The FAST algorithm is used to generate the initial feature points of the block, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold, so that when the initial feature points cannot be generated by the high threshold in the weak texture area of the video image, the initial feature points can still be generated by the low threshold. The generation of the initial feature points accordingly affects the number and positioning of the final feature points. In the extraction of feature points, if the extracted feature points are only concentrated in the area with rich texture or strong features, then for a block, one feature point can clearly express an area, and the remaining feature points are redundant, which is easy to cause unnecessary waste and cost increase. For the low texture area, there is a lack of feature points to describe the information of the block. From a global perspective, the distribution of the overall feature points is also uneven. Therefore, the detection threshold in the FAST algorithm includes a high threshold and a low threshold, which is conducive to avoiding the situation where the final extracted feature points are only concentrated in the area with rich texture or strong features, and the number of feature points in the weak texture area lacking features will be small or even no, which is conducive to better improving the extraction of feature points in the weak texture area, so that the distribution of feature points in the current frame video image is more uniform.
[0169] Specifically, in this embodiment, the FAST algorithm is used to generate initial feature points of the blocks, including: using a high threshold and a low threshold to detect the blocks simultaneously.
[0170] It should be noted that using a high threshold and a low threshold to detect blocks simultaneously means that high threshold detection and low threshold detection are performed on the blocks simultaneously.
[0171] The high threshold and the low threshold are used to detect the blocks simultaneously. There is no need to perform the steps of detecting all the blocks twice. Only one detection is required for all the blocks, and the high threshold detection and the low threshold detection for each block can be completed, which is beneficial to improve the efficiency of generating initial feature points.
[0172] In this embodiment, when the initial feature points can be detected by using the high threshold, the low threshold detection is stopped, and the initial feature points detected by the high threshold are obtained.
[0173] When the initial feature points can be detected using a high threshold, the low threshold detection is stopped, which helps to avoid the step of repeatedly reading the content of the same block, reduces unnecessary calculations in real time, and helps to improve detection efficiency and reduce detection costs.
[0174] In this embodiment, when the initial feature points cannot be detected by using a high threshold value, but can be detected by using a low threshold value, the initial feature points detected by using the low threshold value are obtained.
[0175] When the initial feature points cannot be detected by using a high threshold, but can be detected by using a low threshold, that is, when the texture of the block is weak, low threshold detection can still be performed in real time to obtain the initial feature points.
[0176] In this embodiment, when the initial feature points cannot be detected by using the high threshold and the low threshold, the step of acquiring feature points for the blocks is skipped.
[0177] When the initial feature points cannot be detected using both the high threshold and the low threshold, that is, the features of the block are too unclear, and there is no need to extract feature points, thus directly skipping the step of acquiring feature points for the block and saving computing power.
[0178] The feature point acquisition module 503 further includes: a feature point extraction unit 5032 for extracting initial feature points in combination with the feature point acquisition situation in the video image of the previous frame to obtain feature points of the blocks.
[0179] Combined with the feature point acquisition in the previous frame of the video image, the initial feature points are extracted to obtain the block feature points, which is the inter-frame feature point allocation algorithm.
[0180] In the present embodiment, in the video image sequence, two adjacent frames of video images have a large correlation, and the initial feature points are extracted in combination with the feature point acquisition in the previous frame of the video image to obtain the block feature points. This is beneficial for real-time processing of the extraction of the initial feature points according to the actual situation of the video image, and obtaining feature points that are more closely matched with the actual image features of the video image, which is beneficial for achieving more accurate detection of the feature points of the video image of the current frame.
[0181] In this embodiment, the initial feature points are extracted in combination with the feature point acquisition in the previous frame of the video image to obtain the feature points of the blocks, including: obtaining information based on the feature points of the current block corresponding to the block in the previous frame of the video image to obtain the estimated number of feature points of the current block.
[0182] The estimated number of feature points of the current block is obtained as a reference for obtaining the actual feature points of the current block. Moreover, obtaining information based on the feature points of the corresponding block of the current block in the previous frame of video image is beneficial to obtaining the estimated number of feature points that is more closely matched with the actual image features of the video image based on the actual situation of the video image.
[0183] Specifically, obtaining information based on feature points of a corresponding block of the current block in a previous frame of video image to obtain an estimated number of feature points of the current block includes: obtaining a corresponding block of the current block in the previous frame of video image as a reference block.
[0184] The reference block is used as a basis for obtaining the estimated number of feature points in the current block.
[0185] In this embodiment, blocks in the neighborhood of the reference block in the previous frame of video image are obtained as surrounding blocks.
[0186] The surrounding blocks are used together with the reference blocks as a basis for obtaining the estimated number of feature points of the current block.
[0187] In this embodiment, the blocks in the neighborhood of the reference block in the previous frame of video image are obtained as surrounding blocks, and the neighborhood range of the reference block is set according to the global motion vector of the previous frame of video image, wherein the size of the global motion vector is proportional to the size of the neighborhood range.
[0188] It should be noted that the size of the global motion vector is proportional to the size of the neighborhood range, which means that the larger the global motion vector is, the larger the neighborhood range is set, and the smaller the global motion vector is, the smaller the neighborhood range is set.
[0189] The global motion vector (GMV) represents the motion speed of the previous frame of video image. The larger the global motion vector is, the greater the motion speed of the previous frame of video image is, and the greater the image change is, so a larger range needs to be referenced. The smaller the global motion vector is, the smaller the motion speed of the previous frame of video image is, and the smaller the image change is, so a smaller range can be referenced. Therefore, the size of the global motion vector is proportional to the size of the neighborhood range.
[0190] As an example, when the global motion vector is relatively large, the reference block neighborhood is selected to be the surrounding 24 blocks as the reference block, and when the global motion vector is relatively small, the reference block neighborhood is selected to be the surrounding 8 blocks as the reference block.
[0191] In this embodiment, the feature point acquisition information of the reference block and the surrounding blocks is combined to obtain the estimated number of feature points of the current block.
[0192] If there is a certain spatial correlation in the video image, the feature point acquisition information of the surrounding blocks will also affect the estimated number of feature points of the current block. Therefore, the feature point acquisition information of the reference block and the surrounding blocks is combined to obtain the estimated number of feature points of the current block.
[0193] In this embodiment, feature point acquisition information of the reference block and surrounding blocks is combined to obtain an estimated number of feature points of the current block, including: obtaining feature point information of the reference block and surrounding blocks, dividing the blocks of the video image of the current frame into levels, and obtaining the level corresponding to each block, wherein each level has a corresponding estimated number.
[0194] By hierarchically dividing the blocks of the video image of the current frame, all blocks in the video image of the current frame can be classified, and blocks with similar features can be divided into the same category. Then, the same estimated number can be assigned to blocks of the same category, which is conducive to obtaining the estimated number of all blocks at the same time, and there is no need to assign an estimated number to each block separately, reducing the number of times the estimated number corresponding to the block is assigned, which is conducive to improving computing efficiency and reducing computing costs.
[0195] Accordingly, in this embodiment, the estimated number of feature points of each block is obtained according to the level corresponding to each block.
[0196] In this embodiment, the feature point acquisition information of the reference block and the surrounding blocks is combined to obtain the estimated number of feature points of the current block, including: obtaining whether the reference block has feature points as the first information.
[0197] Whether the reference block has feature points represents the texture of the reference block, and can correspondingly provide feedback on the texture of the current block in the video image of the current frame.
[0198] In this embodiment, when the initial feature points of the reference block generated by the FAST algorithm are obtained, the situation in which the detection threshold is a high threshold or a low threshold is used as the second information.
[0199] When the FAST algorithm is used to generate the initial feature points of the reference block, the detection threshold is a high threshold or a low threshold, which further finely characterizes the texture of the reference block, and accordingly can further finely feedback the texture of the current block in the video image of the current frame.
[0200] In this embodiment, whether the number of feature points in the surrounding blocks is greater than or equal to a set threshold is obtained as the third information.
[0201] The case where the number of feature points in the surrounding blocks is greater than or equal to the set threshold means that some blocks in the surrounding blocks may not have feature points, while some blocks have feature points. The case where the number of feature points in the surrounding blocks is greater than or equal to the set threshold is obtained.
[0202] Whether the number of feature points in the surrounding blocks is greater than or equal to a set threshold represents the texture distribution of the surrounding blocks, and can correspondingly feedback the texture distribution around the current block in the video image of the current frame.
[0203] In this embodiment, when obtaining whether the number of feature points in the surrounding blocks is greater than or equal to a set threshold, the set threshold is half of the total number of the surrounding blocks.
[0204] The threshold is set to half of the total number of surrounding blocks, that is, whether the number of feature points in the surrounding blocks exceeds half. If it exceeds half, it indicates that the texture of the surrounding blocks is strong. If it does not exceed half, it indicates that the texture of the surrounding blocks is weak.
[0205] In this embodiment, the estimated number of feature points of the current block is obtained by combining the first information, the second information and the third information.
[0206] Combining the first information, the second information and the third information to obtain the estimated number of feature points of the current block is conducive to more comprehensively taking into account the factors affecting the distribution of feature points of the current block and obtaining a more accurate estimated number of feature points.
[0207] In this embodiment, according to the estimated number of feature points of the current block, a corresponding number of initial feature points are extracted as feature points of the block.
[0208] Extracting a corresponding number of initial feature points based on the estimated number of feature points of the current block is conducive to obtaining feature points that are more closely matched with the actual texture distribution of the video image of the current frame.
[0209] In this embodiment, according to the estimated number of feature points of the current block, a corresponding number of initial feature points are extracted as the feature points of the block, including: obtaining an adjustment value for the number of feature points of the current block according to the number of feature points of all blocks before the current block and the estimated number of feature points.
[0210] According to the number of feature points of all blocks before the current block and the estimated number of feature points, the adjustment value of the number of feature points of the current block is obtained, which is the intra-frame feature point allocation algorithm.
[0211] In the process of extracting initial feature points to obtain feature points, there may be a mismatch between the actual number of initial feature points produced in each block and the estimated number of feature points. In particular, when the actual number of initial feature points produced is less than the estimated number of feature points, it will be difficult for the block to extract a sufficient number of initial feature points as feature points, resulting in a lack of feature points in the video image of the final current frame. Therefore, based on the number of feature points of all blocks before the current block and the estimated number of feature points, an adjustment value for the number of feature points of the current block is obtained, which is used to adjust the number of feature points in the current block and subsequent processed blocks, thereby reducing the probability of a lack of feature points in the video image of the final current frame.
[0212] In this embodiment, based on the number of feature points of all blocks before the current block and the estimated number of feature points, an adjustment value of the number of feature points of the current block is obtained, including: according to the distribution of feature points of the previous frame of video image, a selection mode for obtaining the adjustment value of the number of feature points of the current block is set, wherein the selection mode includes a uniform averaging mode or a greedy acquisition mode.
[0213] The distribution of feature points of the previous frame of video image feeds back the distribution of feature points of the current frame of video image to a certain extent. Therefore, according to the distribution of feature points of the previous frame of video image, a selection mode for obtaining the adjustment value of the number of feature points of the current block is set, so that a selection mode that is more in line with the actual texture of the current frame of video image can be obtained.
[0214] Among them, the uniform averaging mode means that the adjustment value is uniformly assigned to the current block and other blocks to be processed subsequently, and the greedy acquisition mode means that for the current block and other blocks to be processed subsequently, blocks with stronger textures, that is, blocks with stronger characteristics, are assigned larger adjustment values, and blocks with weaker textures, that is, blocks with weaker characteristics, are assigned smaller adjustment values.
[0215] Correspondingly, in this embodiment, the feature point quantity adjustment value of the current block is obtained according to the selection mode.
[0216] In this embodiment, according to the distribution of feature points of the previous frame of video image, setting a selection mode for obtaining the feature point quantity adjustment value of the current block includes: determining whether the feature points of the previous frame of video image are evenly distributed.
[0217] Whether the feature points of the previous frame of video image are evenly distributed indicates whether the feature points of the current frame of video image are evenly distributed. Therefore, the selection mode for obtaining the feature point quantity adjustment value of the current block is set according to whether the feature points of the previous frame of video image are evenly distributed.
[0218] In this embodiment, if the feature points of the previous frame of video image are evenly distributed, the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the even distribution mode.
[0219] If the feature points of the previous frame of video image are evenly distributed, then the feature points representing the current frame of video image are evenly distributed, so that the uniform averaging mode is adopted, and the number of extracted initial feature points can be adjusted more evenly.
[0220] In this embodiment, otherwise, the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the greedy acquisition mode.
[0221] Otherwise, the feature points representing the current frame video image are unevenly distributed, so a greedy acquisition mode is used to fill in some areas where the feature points are unevenly distributed, so that the number of extracted initial feature points can be adjusted more evenly.
[0222] In this embodiment, determining whether the feature points of the previous frame of video image are evenly distributed includes: performing one or more division processes on the previous frame of video image, and obtaining two sub-areas of equal area each time the division process is performed.
[0223] The previous frame of video image is divided once or multiple times, and each division process obtains two sub-regions of equal area, then the area of each sub-region is equal. In this embodiment, uniform distribution means that the number of feature points in each sub-region is as similar as possible.
[0224] As an example, Figure 6 As shown, the previous frame of video image is divided into five parts, including Figure 6 (a) Equally divided into upper and lower parts, such as Figure 6 (b) Divide equally on the left and right sides, such as Figure 6 (c) The upper left and lower right are equally divided, such as Figure 6 (d) The upper right and lower left parts are equally divided, and Figure 6 (e) The inner and outer circles are divided equally to obtain 10 sub-regions of equal area.
[0225] In this embodiment, the variance of the number of feature points in each sub-region among all the sub-regions obtained by the division process is obtained.
[0226] As an example, the variance of the number of feature points in 10 sub-regions is obtained, which represents the degree of dispersion between the feature points in the 10 sub-regions and the average value.
[0227] In this embodiment, the total number of feature points of the multiple sub-regions obtained by all the division processes is obtained.
[0228] As an example, the total number of feature points in 10 sub-areas is obtained.
[0229] In this embodiment, the ratio of the variance to the sum is obtained.
[0230] As an example, the ratio of the variance of the number of feature points in 10 sub-regions to the total is obtained, which represents the uniformity of the feature points in the 10 sub-regions.
[0231] It can be seen that the smaller the variance is, the larger the sum of the number of feature points in multiple sub-regions is, the smaller the ratio of the variance to the sum is, and the more uniform the distribution of the feature points representing the previous frame of the video image is.
[0232] Correspondingly, in this embodiment, when the ratio of the variance to the sum is less than or equal to the uniformity threshold, it is judged that the feature points of the previous frame of video image are evenly distributed; when the ratio of the variance to the sum is greater than the uniformity threshold, it is judged that the feature points of the previous frame of video image are unevenly distributed.
[0233] In this embodiment, when the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the uniform averaging mode, the feature point quantity adjustment value of the current block is obtained according to the selection mode, including: obtaining the difference between the total estimated number of feature points of all blocks before the current block and the total number of feature points as the deviation value.
[0234] In the process of extracting initial feature points to obtain feature points, if the actual number of initial feature points generated is less than the estimated number of feature points, it will cause difficulty in extracting a sufficient number of initial feature points as feature points for the block, resulting in a situation where some blocks before the current block are short of feature points. Therefore, the difference between the total estimated number of feature points of all blocks before the current block and the total number of feature points is obtained, that is, the total number of feature points that are short of all blocks before the current block is obtained, which is used to evenly distribute them to the current block and all blocks after the current block. The pre-allocated scheme is dynamically optimized again within the frame according to the actual situation, so that the number of extracted initial feature points can further meet the demand.
[0235] In this embodiment, the block numbers of the current block and all blocks after the current block are obtained.
[0236] Get the number of blocks of the current block and all blocks after the current block as the total number of feature points for amortizing the deficit of all blocks before the current block.
[0237] In this embodiment, the ratio of the deviation value to the number of blocks is used as the adjustment value of the number of feature points of the current block.
[0238] The ratio of the deviation value to the number of blocks is used as the adjustment value of the number of feature points of the current block, so that all blocks after the current block are evenly adjusted.
[0239] In this embodiment, when the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the greedy acquisition mode, the feature point quantity adjustment value of the current block is obtained according to the selection mode, including: obtaining the difference between the estimated total number of feature points of all blocks before the current block and the total number of feature points as the deviation value.
[0240] The explanation of the deviation value is the same as that when the uniform amortization mode is adopted, and will not be repeated here.
[0241] In this embodiment, the ratio of a preset number to the number of blocks of the video image of the current frame is used as a quantity threshold. When the estimated number of feature points of the current block is less than the quantity threshold, and the estimated number of feature points of the current block is less than the number of initial feature points, the difference between the quantity threshold and the estimated number of feature points of the current block is obtained as an adjustment value for the number of feature points of the current block.
[0242] The ratio of the preset number to the number of blocks of the video image of the current frame is the preset average number of feature points of each block, that is, the number threshold. The estimated number of feature points of the current block is less than the number threshold, and the estimated number of feature points of the current block is less than the number of initial feature points. The characteristics of the current block are more obvious, and there is a margin of initial feature points that can be selected. Therefore, the current block is allocated a larger feature point adjustment value as much as possible. Specifically, the difference between the number threshold and the estimated number of feature points of the current block is obtained as the feature point number adjustment value of the current block. The operation is simple and the distribution tends to be uniform.
[0243] In this embodiment, when the estimated number of feature points of the current block is greater than or equal to the number threshold, or the estimated number of feature points of the current block is less than the number threshold, and the estimated number of feature points of the current block is greater than or equal to the number of initial feature points, the feature point number adjustment value of the current block is 0.
[0244] The estimated number of feature points of the current block is greater than or equal to the number threshold, indicating that the features of the current block are not obvious enough to allocate more feature points. Therefore, the adjustment value of the number of feature points of the current block is 0. Alternatively, the estimated number of feature points of the current block is less than the number threshold, and the estimated number of feature points of the current block is greater than or equal to the number of initial feature points, indicating that the current block does not have a surplus of initial feature points that can be selected. Therefore, the adjustment value of the number of feature points of the current block is 0.
[0245] In this embodiment, the target number of feature points in the current block is obtained by combining the feature point quantity adjustment value of the current block with the estimated number of feature points.
[0246] Combined with the feature point quantity adjustment value of the current block and the estimated number of feature points, the number of initial feature points to be extracted in the current block is adjusted to obtain the target number of feature points in the current block, so that the feature points of the current frame video image are more uniform.
[0247] Specifically, in this embodiment, the feature point quantity adjustment value of the current block and the estimated number of feature points are combined to obtain the target number of feature points of the current block, including: obtaining the sum of the feature point quantity adjustment value of the current block and the estimated number of feature points as the target number of feature points of the current block.
[0248] In this embodiment, initial feature points are extracted as feature points of the blocks according to the number of targets.
[0249] According to the number of targets, initial feature points are extracted as feature points of the blocks, and finally feature points with relatively even distribution in the current frame video image can be obtained.
[0250] In this embodiment, initial feature points are extracted as feature points of the block according to the target number, including: when the target number of feature points of the current block is less than or equal to the number of initial feature points, an equal number of initial feature points is extracted as feature points of the block.
[0251] When the number of feature point targets in the current block is less than or equal to the number of initial feature points, the same number of initial feature points as the number of targets is extracted to meet the detection requirements.
[0252] In this embodiment, when the number of feature point targets of the current block is greater than the number of initial feature points, all initial feature points are extracted as feature points of the block.
[0253] When the target number of feature points of the current block is greater than the number of initial feature points, the current block does not have a surplus of initial feature points that can be selected, and therefore all initial feature points are extracted as feature points of the block.
[0254] The embodiment of the present invention further provides a device, which can implement the method for detecting feature points of video images provided by the embodiment of the present invention by loading the above-mentioned method for detecting feature points of video images in the form of a program. An optional hardware structure of the terminal device provided by the embodiment of the present invention can be as follows Fig.10 As shown, it includes: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.
[0255] In this embodiment, the number of processor 01, communication interface 02, memory 03, and communication bus 04 is at least one, and the processor 01, communication interface 02, and memory 03 communicate with each other through the communication bus 04. The communication interface 02 can be an interface of a communication module for network communication, such as an interface of a GSM module. The processor 01 may be a central processing unit CPU, or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present invention. The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory (NVM), such as at least one disk storage. Among them, the memory 03 stores one or more computer instructions, and the one or more computer instructions are executed by the processor 01 to implement the method for detecting feature points of a video image provided in an embodiment of the present invention.
[0256] It should be noted that the above-mentioned terminal device may also include other devices (not shown) that may not be necessary for understanding the contents disclosed in the embodiments of the present invention; given that these other devices may not be necessary for understanding the contents disclosed in the embodiments of the present invention, the embodiments of the present invention will not introduce them one by one.
[0257] An embodiment of the present invention further provides a storage medium, wherein the storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the method for detecting feature points of a video image provided by an embodiment of the present invention.
[0258] In an embodiment of the present invention, in a video image sequence, two adjacent frames of video images have a large correlation. In this way, the initial feature points are extracted in combination with the feature point acquisition in the previous frame of the video image to obtain the feature points of the blocks. This is beneficial for real-time processing of the extraction of the initial feature points according to the actual situation of the video image, and for obtaining feature points that are more closely matched with the actual image feature situation of the video image, which is beneficial for achieving more accurate detection of the feature points of the video image of the current frame.
[0259] The above-mentioned embodiments of the present invention are the combination of elements and features of the present invention. Unless otherwise mentioned, elements or features may be considered as optional. Each element or feature may be put into practice without being combined with other elements or features. In addition, embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some constructions of any one embodiment may be included in another embodiment, and may be replaced by the corresponding construction of another embodiment. It is obvious to those skilled in the art that claims that do not have a clear reference relationship to each other in the attached claims may be combined into embodiments of the present invention, or may be included as new claims in the amendment after submitting this application.
[0260] Embodiments of the present invention may be implemented by various means such as hardware, firmware, software or a combination thereof. In a hardware configuration, the method according to an exemplary embodiment of the present invention may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc. In a firmware or software configuration, embodiments of the present invention may be implemented in the form of modules, processes, functions, etc. The software code may be stored in a memory unit and executed by a processor. The memory unit is located inside or outside the processor and may send data to the processor and receive data from the processor via various known means.
[0261] The above description of the disclosed embodiments enables one skilled in the art to make or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0262] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A method for detecting feature points of a video image, characterized in that: include: Acquire the video image of the current frame; Dividing the video image of the current frame into a plurality of blocks; Acquiring feature points for each of the blocks in turn; The feature point acquisition includes: generating initial feature points of the block; extracting the initial feature points in combination with the feature point acquisition situation in the video image of the previous frame to obtain the feature points of the block.
2. The method for detecting feature points of a video image according to claim 1, wherein: The video image of the current frame is divided into a plurality of blocks, and the size of each of the divided blocks is set according to the size of the video image of the current frame.
3. The method for detecting feature points of a video image according to claim 1, wherein: The FAST algorithm is used to generate the initial feature points of the blocks, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold.
4. The method for detecting feature points of a video image according to claim 3, wherein: The FAST algorithm is used to generate initial feature points of the blocks, including: using a high threshold and a low threshold to simultaneously detect the blocks; When the initial feature point can be detected by using the high threshold, the low threshold detection is stopped, and the initial feature point detected by the high threshold is obtained; When the initial feature point cannot be detected by using the high threshold value, and the initial feature point can be detected by using the low threshold value, acquiring the initial feature point detected by the low threshold value; When the initial feature point cannot be detected by using the high threshold and the low threshold, the step of acquiring feature points for the block is skipped.
5. The method for detecting feature points of a video image according to claim 1, wherein: Extracting the initial feature points in combination with the feature point acquisition in the video image of the previous frame to obtain the feature points of the block, including: obtaining the estimated number of feature points of the current block according to the feature point acquisition information of the corresponding block in the video image of the previous frame; According to the estimated number of feature points of the current block, a corresponding number of the initial feature points are extracted as the feature points of the block.
6. The method for detecting feature points of a video image according to claim 5, wherein: Acquiring information based on feature points of a corresponding block of the current block in the video image of the previous frame to obtain an estimated number of feature points of the current block, including: acquiring a corresponding block of the current block in the video image of the previous frame as a reference block; Obtaining blocks in the neighborhood of the reference block in the previous frame of the video image as surrounding blocks; The feature point acquisition information of the reference block and surrounding blocks is combined to obtain the estimated number of feature points of the current block.
7. The method for detecting feature points of a video image according to claim 6, wherein: Obtain the blocks in the neighborhood of the reference block in the video image of the previous frame as the surrounding blocks, and set the neighborhood range of the reference block according to the global motion vector of the video image of the previous frame, wherein the size of the global motion vector is proportional to the size of the neighborhood range.
8. The method for detecting feature points of a video image according to claim 6, wherein: Acquiring feature point information of the reference block and surrounding blocks to obtain an estimated number of feature points of the current block, including: acquiring feature point information of the reference block and surrounding blocks to grade the blocks of the video image of the current frame to obtain a grade corresponding to each of the blocks, wherein each grade has a corresponding estimated number; According to the level corresponding to each of the blocks, the estimated number of feature points of each block is obtained.
9. The method for detecting feature points of a video image according to claim 6, wherein: The FAST algorithm is used to generate the initial feature points of the blocks, wherein the detection threshold in the FAST algorithm includes a high threshold and a low threshold; Acquiring information on feature points of the reference block and surrounding blocks to obtain an estimated number of feature points of the current block, including: obtaining whether the reference block has feature points as first information; When the initial feature points of the reference block are generated by using the FAST algorithm, the detection threshold is a high threshold or a low threshold, as the second information; Obtaining whether the number of feature points in the surrounding blocks is greater than or equal to a set threshold as third information; The first information, the second information and the third information are combined to obtain an estimated number of feature points of the current block.
10. The method for detecting feature points of a video image according to claim 9, wherein: In the case of obtaining whether the number of characteristic points in the surrounding blocks is greater than or equal to a set threshold, the set threshold is half of the total number of the surrounding blocks.
11. The method for detecting feature points of a video image according to claim 5, wherein: According to the estimated number of feature points of the current block, extracting a corresponding number of the initial feature points as the feature points of the block, including: obtaining an adjustment value of the number of feature points of the current block according to the number of feature points of all blocks before the current block and the estimated number of feature points; Combining the feature point quantity adjustment value of the current block with the estimated number of feature points, obtaining the target number of feature points of the current block; According to the target quantity, the initial feature points are extracted as feature points of the blocks.
12. The method for detecting feature points of a video image according to claim 11, wherein: According to the number of feature points of all blocks before the current block and the estimated number of feature points, obtaining the feature point number adjustment value of the current block, including: according to the distribution of feature points of the video image of the previous frame, setting a selection mode for obtaining the feature point number adjustment value of the current block, wherein the selection mode includes a uniform averaging mode or a greedy acquisition mode; According to the selection mode, an adjustment value of the number of feature points of the current block is obtained.
13. The method for detecting feature points of a video image according to claim 12, wherein: According to the distribution of feature points of the video image of the previous frame, setting a selection mode for obtaining the feature point quantity adjustment value of the current block includes: determining whether the feature points of the video image of the previous frame are evenly distributed; If the feature points of the video image of the previous frame are evenly distributed, setting the selection mode for obtaining the feature point quantity adjustment value of the current block to the evenly distributed mode; Otherwise, the selection mode for obtaining the adjustment value of the number of feature points of the current block is set to the greedy acquisition mode.
14. The method for detecting feature points of a video image according to claim 13, wherein: Determining whether the feature points of the video image of the previous frame are evenly distributed, including: dividing the video image of the previous frame once or multiple times, and obtaining two sub-areas with equal areas each time; Obtain the variance of the number of feature points in each sub-region among the multiple sub-regions obtained by all the division processes; Obtain the total number of feature points of multiple sub-regions obtained through all division processes; Obtaining a ratio of the variance to the sum; When the ratio of the variance to the sum is less than or equal to the uniformity threshold, it is determined that the feature points of the video image of the previous frame are evenly distributed; When the ratio of the variance to the sum is greater than the uniformity threshold, it is determined that the feature points of the video image of the previous frame are unevenly distributed.
15. The method for detecting feature points of a video image according to claim 13, wherein: When the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the uniform amortization mode, obtaining the feature point quantity adjustment value of the current block according to the selection mode includes: obtaining the difference between the sum of the estimated numbers of feature points of all blocks before the current block and the sum of the feature points as a deviation value; Get the number of blocks of the current block and all blocks after the current block; The ratio of the deviation value to the number of blocks is used as an adjustment value for the number of feature points of the current block.
16. The method for detecting feature points of a video image according to claim 13, wherein: Acquire the video image of the current frame, and set the video image to have a preset number of feature points; When the selection mode for obtaining the feature point quantity adjustment value of the current block is set to the greedy acquisition mode, obtaining the feature point quantity adjustment value of the current block according to the selection mode includes: obtaining the difference between the sum of the estimated numbers of feature points of all blocks before the current block and the sum of the feature points as a deviation value; The ratio of the preset number to the number of blocks of the video image of the current frame is used as a number threshold, and when the estimated number of feature points of the current block is less than the number threshold, and the estimated number of feature points of the current block is less than the number of initial feature points, the difference between the number threshold and the estimated number of feature points of the current block is obtained as an adjustment value for the number of feature points of the current block; When the estimated number of feature points of the current block is greater than or equal to the number threshold, or the estimated number of feature points of the current block is less than the number threshold, and the estimated number of feature points of the current block is greater than or equal to the number of initial feature points, the feature point number adjustment value of the current block is 0.
17. The method for detecting feature points of a video image according to claim 11, wherein: Combining the feature point quantity adjustment value and the estimated number of feature points of the current block to obtain the target number of feature points of the current block includes: obtaining the sum of the feature point quantity adjustment value and the estimated number of feature points of the current block as the target number of feature points of the current block.
18. The method for detecting feature points of a video image according to claim 11, wherein: Extracting the initial feature points as the feature points of the block according to the target number includes: when the target number of feature points of the current block is less than or equal to the number of the initial feature points, extracting the initial feature points equal to the target number as the feature points of the block; When the target number of feature points of the current block is greater than the number of initial feature points, all of the initial feature points are extracted as feature points of the block.
19. A video image feature point detection system, characterized in that: include: A video image acquisition module, used to acquire the video image of the current frame; A division module, used for dividing the video image of the current frame into a plurality of blocks; A feature point acquisition module, used to acquire feature points for each of the blocks in turn; Among them, the feature point acquisition module includes: an initial feature point generation unit, which is used to generate the initial feature points of the block; a feature point extraction unit, which is used to extract the initial feature points in combination with the feature point acquisition situation in the video image of the previous frame to obtain the feature points of the block.
20. A device, characterized in that It includes at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for detecting feature points of a video image as described in any one of claims 1-18.
21. A storage medium, characterized in that: The storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the method for detecting feature points of a video image as described in any one of claims 1-18.
Citation Information
Cited By
Key point identification method and related equipment
CN121392307A
Slope deformation visual detection method fusing improved YOLO and sub-pixel extraction
CN121962160A