Lens displacement detection method and device, electronic equipment and storage medium
By extracting key feature points in video frames and calculating their descriptors, the lens displacement is detected based on the feature distance between the descriptors, and the problem of low robustness in the prior art is solved, and more accurate lens displacement detection is achieved.
Patent Information
- Application Number
- CN202510083264.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is not very robust when detecting lens displacement, and it is easy to misjudgment the lens displacement due to changes in light.
By obtaining the video frame to be detected and the video frame to be compared, the key feature points in the video frame are extracted, the descriptor of the key feature points is calculated, and whether the lens has displacement is determined based on the feature distance between the descriptors.
It improves the robustness of lens displacement detection, reduces misjudgment caused by light changes, and enhances the accuracy of detection results.
Smart Images

Figure CN119922299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of displacement detection technology, and in particular to a lens displacement detection method, device, electronic equipment and storage medium. Background Art
[0002] The lens is an optical component installed on the image acquisition device. By adjusting the position of the lens, the field of view of the image acquisition device can be adjusted, such as by setting the position of the lens so that the field of view of the image acquisition device covers a specified area. When the lens is displaced, the field of view of the image acquisition device changes accordingly, and the field of view may not cover the specified area. In this case, the lens needs to be adjusted.
[0003] In the prior art, a technician can set a fixed marker in a designated area, and the marker will be displayed in the image captured by the lens, such as the marker being displayed in the lower left corner of the image. If the lens is displaced, the marker may no longer be displayed in the lower left corner of the captured image. Then, the pixel value of each pixel point in the lower left corner area of the currently captured image will have a large change compared to the pixel value of each pixel point in the lower left corner area of the first image captured by the lens after the marker is set in the designated area (hereinafter referred to as the initial pixel value). Therefore, when the electronic device detects that the pixel value of the pixel point in the lower left corner area of the currently captured image has a large change compared to the initial pixel value, a detection result indicating that the lens has been displaced can be obtained.
[0004] However, as the light changes, the pixel value of each pixel in the image captured by the image acquisition device will change, resulting in a detection result indicating that the lens has shifted even if the lens has not shifted when the light changes too much. It can be seen that the robustness of lens shift detection based on the prior art is not high. Summary of the invention
[0005] The purpose of the embodiments of the present invention is to provide a lens displacement detection method, device, electronic device and storage medium to improve the robustness of lens displacement detection. The specific technical solution is as follows:
[0006] In a first aspect of an embodiment of the present invention, a lens displacement detection method is first provided, the method comprising: obtaining a video frame to be detected and a video frame to be compared; wherein the video frame to be detected and the video frame to be compared are obtained based on video frames in a video captured by a lens to be detected; for each acquired video frame, extracting each key feature point in the video frame; for each key feature point to be detected in the video frame to be detected, determining, from each key feature point to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected; obtaining a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance; and determining whether the lens to be detected is displaced based on each reference distance obtained.
[0007] Optionally, for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the matching key feature point corresponding to the key feature point to be detected, including: for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected; wherein the hash bucket to which a key feature point belongs is: determined by mapping the descriptor of the key feature point using a local sensitive hashing algorithm; calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point; based on the minimum value of each feature distance, determine the candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected from the candidate key feature points corresponding to the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected.
[0008] Optionally, based on the minimum value among the feature distances, determining, from the candidate key feature points corresponding to the key feature point to be detected, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected, includes: when there is only one video frame to be compared, calculating the ratio of the minimum value to the second minimum value among the feature distances corresponding to the key feature point to be detected as a first ratio; if the first ratio is less than a first ratio threshold, using the candidate key feature point with the minimum feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected; based on the obtained reference distances, determining whether the lens to be detected is displaced includes: calculating the average level of the obtained reference distances; if the average level is less than the first distance threshold, determining that the lens to be detected is not displaced; otherwise, determining that the lens to be detected is displaced.
[0009] Optionally, before obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, the method also includes: if the first ratio is not less than the first ratio threshold, determining that there is no matching key feature point corresponding to the key feature point to be detected; obtaining the number of key feature points to be detected for which there are corresponding matching key feature points as the first matching number; when the first matching number is less than the first matching point number threshold, outputting a prompt message; obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance includes: when the first matching number is not less than the first matching point number threshold, obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance.
[0010] Optionally, based on the minimum value among each feature distance, determining, from the candidate key feature points corresponding to the key feature point to be detected, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected, including: when there are multiple video frames to be compared, for each video frame to be compared, determining the candidate key feature point corresponding to the key feature point to be detected in the video frame to be compared as the key feature point to be screened; calculating the ratio of the minimum value and the second minimum value in the feature distances corresponding to the determined key feature points to be screened as the second ratio corresponding to the video frame to be compared; if the second ratio corresponding to the video frame to be compared is less than a second ratio threshold, determining the key feature point to be screened with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected is the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; the determining whether the lens to be detected is displaced based on the acquired reference distances comprises: for each video frame to be compared, calculating the average level of the reference distances corresponding to the matching key feature points in the video frame to be compared, as the average level corresponding to the video frame to be compared; if the average level corresponding to the video frame to be compared is less than a second distance threshold, determining that the video frame to be compared indicates that the lens to be detected is not displaced; otherwise, determining that the video frame to be compared indicates that the lens to be detected is displaced; if the number of video frames to be compared indicating that the lens to be detected is not displaced is less than the number of video frames to be compared indicating that the lens to be detected is displaced, determining that the lens to be detected is not displaced; otherwise, determining that the lens to be detected is displaced.
[0011] Optionally, before obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, the method also includes: for each video frame to be compared, if the second ratio corresponding to the video frame to be compared is not less than the second ratio threshold, determining that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; obtaining the number of key feature points to be detected for which there are corresponding matching key feature points in the video frame to be compared, as the second matching number corresponding to the video frame to be compared; when the second matching number corresponding to any video frame to be compared is less than the second matching point number threshold, outputting prompt information; obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance includes: when the second matching number corresponding to each video frame to be compared is not less than the second matching point number threshold, for each video frame to be compared, obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point in the video frame to be compared as the reference distance corresponding to the key feature point to be detected in the video frame to be compared.
[0012] Optionally, the video frame to be detected is obtained based on the current video frame captured by the lens to be detected; when the field of view of the lens to be detected is fixed, the multiple video frames to be compared are obtained based on the following method: sampling multiple video frames from all video frames currently captured by the lens to be detected; and obtaining multiple video frames to be compared based on the sampled video frames; or, starting from the first video frame captured by the lens to be detected, sampling multiple consecutive video frames; and obtaining multiple video frames to be compared based on the sampled video frames;
[0013] And / or, when the field of view of the lens to be detected is variable, the multiple video frames to be compared are obtained based on the following method: determining a sampling interval according to the frame rate of the video captured by the lens to be detected; determining a starting video frame that is separated from the current video frame by the sampling interval of video frames from the video captured by the lens to be detected; sampling multiple video frames from all video frames between the starting video frame and the current video frame; and obtaining multiple video frames to be compared based on the sampled video frames.
[0014] Optionally, the video frame to be detected is obtained based on the following method: acquiring the current video frame captured by the lens to be detected; performing image enhancement on the current video frame to obtain the video frame to be detected; and obtaining multiple video frames to be compared based on the sampled video frame, including: performing image enhancement on each sampled video frame to obtain multiple video frames to be compared.
[0015] Optionally, for any video frame, image enhancement is performed on the video frame by at least one of the following: image compression of the video frame; conversion of the video frame into a grayscale image; division of the video frame, and based on a histogram equalization algorithm, pixel value equalization of each divided image area.
[0016] Optionally, the histogram-based equalization algorithm performs pixel value equalization on each divided image area, including: for each divided image area, dividing the pixel points with the same pixel value among the pixel points included in the image area into a group; for each group of pixel points, determining the number of pixel points whose pixel values are not greater than the pixel values of the pixel points in the group among the pixel points included in the image area, as a first number; calculating the ratio of the first number to the total number of pixel points included in the image area to obtain the proportion corresponding to the group of pixel points; calculating the product of the proportion corresponding to the group of pixel points and a specified value to obtain the to-be-restricted equalization value corresponding to the group of pixel points; wherein the specified value is: The method comprises the following steps: the method comprises: determining the length of the maximum distribution range of the pixel values of the video frame; for each group of pixels, calculating the product of the pixel value of the group of pixels and the contrast limiting coefficient to obtain the pixel threshold of the group of pixels; when the to-be-limited equalization value of the group of pixels is greater than the pixel threshold of the group of pixels, calculating the difference between the to-be-limited equalization value of the group of pixels and the pixel threshold of the group of pixels; calculating the quotient of the difference and the total number to obtain the correction value of the group of pixels; for each pixel in the image area, determining the minimum value between the to-be-limited equalization value of the pixel and the pixel threshold of the pixel; calculating the sum of the minimum value and the correction value of each group of pixels; and obtaining the pixel value equalization processing result corresponding to the pixel based on the sum.
[0017] Optionally, based on the sum value, a pixel value equalization processing result corresponding to the pixel point is obtained, including: when the pixel point is not adjacent to other image areas, using the sum value corresponding to the pixel point as the pixel value equalization processing result corresponding to the pixel point; when the pixel point is adjacent to other image areas, determining the pixels adjacent to the pixel point in other image areas; and interpolating the sum value corresponding to the pixel point with the sum value corresponding to the determined adjacent pixel points to obtain the pixel value equalization processing result corresponding to the pixel point.
[0018] Optionally, before, for each key feature point to be detected in the video frame to be detected, determining, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected as a matching key feature point corresponding to the key feature point to be detected, the method further comprises: obtaining the number of key feature points to be detected extracted from the video frame to be detected as a reference number; if the reference number is less than a detection quantity threshold, outputting a prompt message; for each key feature point to be detected in the video frame to be detected, determining, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected as a matching key feature point corresponding to the key feature point to be detected The method comprises the following steps: determining, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected, including: when the reference number is not less than the detection quantity threshold, for each key feature point to be detected in the video frame to be detected, determining, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected.
[0019] In a second aspect of an embodiment of the present invention, a lens displacement detection device is provided, the device comprising:
[0020] The video frame acquisition module is used to acquire the video frame to be detected and the video frame to be compared; wherein the video frame to be detected and the video frame to be compared are obtained based on the video frame in the video captured by the lens to be detected;
[0021] An extraction module is used to extract key feature points in each acquired video frame;
[0022] A determination module is used to determine, for each key feature point to be detected in the video frame to be detected, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected;
[0023] A distance acquisition module is used to acquire a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance;
[0024] The detection module is used to determine whether the lens to be detected is displaced based on the acquired reference distances.
[0025] Optionally, the determination module is specifically used to: for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected; wherein the hash bucket to which a key feature point belongs is: determined by mapping the descriptor of the key feature point using a local sensitive hashing algorithm; calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point; based on the minimum value of each feature distance, determine, from the candidate key feature points corresponding to the key feature point to be detected, the candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected.
[0026] Optionally, the determination module is specifically used to: when there is only one video frame to be compared, calculate the ratio of the minimum value and the second minimum value in the feature distance corresponding to the key feature point to be detected as a first ratio; if the first ratio is less than a first ratio threshold, use the candidate key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected;
[0027] The detection module is specifically used to: calculate the average level of each reference distance obtained; if the average level is less than a first distance threshold, determine that the lens to be detected has no displacement; otherwise, determine that the lens to be detected has displacement.
[0028] Optionally, the device further includes: a first prompt module, which is used to execute, before the distance acquisition module executes to acquire the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, if the first ratio is not less than the first ratio threshold, determining that there is no matching key feature point corresponding to the key feature point to be detected; acquiring the number of key feature points to be detected that have corresponding matching key feature points as the first matching number; and outputting a prompt message when the first matching number is less than the first matching point number threshold;
[0029] The distance acquisition module is specifically used to: when the first matching number is not less than the first matching point quantity threshold, obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance.
[0030] Optionally, the determination module is specifically used to: when there are multiple video frames to be compared, for each video frame to be compared, determine the candidate key feature point corresponding to the key feature point to be detected in the video frame to be compared as the key feature point to be screened; calculate the ratio of the minimum value and the second minimum value in the feature distance corresponding to the determined key feature point to be screened as the second ratio corresponding to the video frame to be compared; if the second ratio corresponding to the video frame to be compared is less than a second ratio threshold, use the key feature point to be screened with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared;
[0031] The detection module is specifically used to: for each video frame to be compared, calculate the average level of the reference distance corresponding to the matching key feature points in the video frame to be compared as the average level corresponding to the video frame to be compared; if the average level corresponding to the video frame to be compared is less than a second distance threshold, determine that the video frame to be compared indicates that the lens to be detected has not been displaced; otherwise, determine that the video frame to be compared indicates that the lens to be detected has been displaced; if the number of video frames to be compared indicating that the lens to be detected has not been displaced is less than the number of video frames to be compared indicating that the lens to be detected has been displaced, determine that the lens to be detected has not been displaced; otherwise, determine that the lens to be detected has been displaced.
[0032] Optionally, the device further includes: a second prompt module, which is used to execute, for each video frame to be compared, before the distance acquisition module executes to obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance, if the second ratio corresponding to the video frame to be compared is not less than the second ratio threshold, determine that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; obtain the number of key feature points to be detected that have corresponding matching key feature points in the video frame to be compared as the second matching number corresponding to the video frame to be compared; and output a prompt message when the second matching number corresponding to any video frame to be compared is less than the second matching point number threshold;
[0033] The distance acquisition module is specifically used to: when the second matching number corresponding to each video frame to be compared is not less than the second matching point number threshold, for each video frame to be compared, obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point in the video frame to be compared, as the reference distance corresponding to the key feature point to be detected in the video frame to be compared.
[0034] Optionally, the video frame to be detected is: obtained based on a current video frame captured by the lens to be detected;
[0035] The device further includes: a sampling module, configured to obtain a plurality of video frames to be compared based on the following method:
[0036] In the case where the field of view of the lens to be detected is fixed, multiple video frames are sampled from all video frames currently captured by the lens to be detected; and multiple video frames to be compared are obtained based on the sampled video frames; or, starting from the first video frame captured by the lens to be detected, multiple consecutive video frames are sampled; and multiple video frames to be compared are obtained based on the sampled video frames; and / or, in the case where the field of view of the lens to be detected is variable, the multiple video frames to be compared are obtained based on the following method: determining a sampling interval according to the frame rate of the video captured by the lens to be detected; determining a starting video frame with an interval of the sampling interval of video frames between the starting video frame and the current video frame from the video captured by the lens to be detected; sampling multiple video frames from all video frames between the starting video frame and the current video frame; and multiple video frames to be compared are obtained based on the sampled video frames.
[0037] Optionally, the video frame to be detected is obtained based on the following method: acquiring the current video frame captured by the lens to be detected; performing image enhancement on the current video frame to obtain the video frame to be detected; the sampling module is specifically used to: perform image enhancement on each sampled video frame to obtain multiple video frames to be compared.
[0038] Optionally, the device also includes: an image enhancement module, used to perform image enhancement on any video frame by at least one of the following: compressing the video frame; converting the video frame into a grayscale image; dividing the video frame and, based on a histogram equalization algorithm, performing pixel value equalization on each divided image area.
[0039] Optionally, the image enhancement module is specifically used to: for each image area obtained by division, divide the pixel points with the same pixel value among the pixel points included in the image area into a group; for each group of pixel points, determine the number of pixel points whose pixel values are not greater than the pixel value of the pixel points in the group among the pixel points included in the image area, as a first number; calculate the ratio of the first number to the total number of pixel points included in the image area to obtain the proportion corresponding to the group of pixel points; calculate the product of the proportion corresponding to the group of pixel points and a specified value to obtain the corresponding equalization value to be restricted for the group of pixel points; wherein the specified value is: the maximum distribution of pixel values representing the video frame in the shot to be detected The length of the range; for each group of pixels, calculate the product of the pixel value of the group of pixels and the contrast limit coefficient to obtain the pixel threshold of the group of pixels; when the to-be-limited equalization value of the group of pixels is greater than the pixel threshold of the group of pixels, calculate the difference between the to-be-limited equalization value of the group of pixels and the pixel threshold of the group of pixels; calculate the quotient of the difference and the total number to obtain the correction value of the group of pixels; for each pixel in the image area, determine the minimum value of the to-be-limited equalization value of the pixel and the pixel threshold of the pixel; calculate the sum of the minimum value and the correction value of each group of pixels; based on the sum, obtain the pixel value equalization processing result corresponding to the pixel.
[0040] Optionally, based on the image enhancement module, it is specifically used to: when the pixel point is not adjacent to other image areas, use the sum value corresponding to the pixel point as the pixel value equalization processing result corresponding to the pixel point; when the pixel point is adjacent to other image areas, determine the pixels in other image areas adjacent to the pixel point; interpolate the sum value corresponding to the pixel point with the sum value corresponding to the determined adjacent pixel points to obtain the pixel value equalization processing result corresponding to the pixel point.
[0041] Optionally, the device also includes: a third prompt module, which is used to obtain the number of key feature points to be detected extracted from the video frame to be detected as a reference number before the determination module executes the determination of each key feature point to be detected in the video frame to be detected, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected; and output a prompt message when the reference number is less than the detection quantity threshold; the determination module is specifically used to: when the reference number is not less than the detection quantity threshold, for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the matching key feature point corresponding to the key feature point to be detected.
[0042] In a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement any lens displacement detection method described in the first aspect when executing the program stored in the memory.
[0043] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the lens displacement detection method described in any one of the first aspects is implemented.
[0044] An embodiment of the present invention further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the above-mentioned lens displacement detection methods.
[0045] The embodiment of the present invention provides a lens displacement detection method, which obtains a video frame to be detected and a video frame to be compared; the video frame to be detected and the video frame to be compared are obtained based on video frames in a video captured by a lens to be detected; for each obtained video frame, each key feature point in the video frame is extracted; for each key feature point to be detected in the video frame to be detected, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected from each key feature point to be compared included in the video frame to be compared is determined as a matching key feature point corresponding to the key feature point to be detected; a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point is obtained as a reference distance; based on each obtained reference distance, it is determined whether the lens to be detected is displaced.
[0046] Based on the above processing, after the electronic device obtains the video frame to be detected and the video frame to be compared, for each obtained video frame, it can extract each key feature point in the video frame and calculate the descriptor of each extracted key feature point. Further, for each key feature point to be detected in the video frame to be detected, the electronic device can determine the matching key feature point corresponding to the key feature point to be detected from the key feature points to be compared included in the video frame to be compared according to the characteristic distance between the descriptors of the key feature points, that is, determine the matching key feature point indicating the same position in the three-dimensional space as the key feature point to be detected. Then the reference distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point can indicate: in different video frames collected by the lens to be detected, the pixel position change indicating the same position in the three-dimensional space. Obviously, when the lens to be detected is displaced, the probability of a larger reference distance is greater; when the lens to be detected is not displaced, the probability of a smaller reference distance is greater. Therefore, based on the obtained reference distances, it can be determined whether the lens to be detected is displaced.
[0047] In addition, since the key feature points are determined based on multiple pixels, the descriptors of the key feature points are also calculated based on multiple pixels, that is, the key feature points and descriptors are obtained by integrating multiple pixels. When some pixels change greatly, the key feature points and descriptors change less, that is, the robustness of the key feature points and descriptors is higher than the robustness of the pixel values of the pixels. Therefore, lens shift detection based on more robust key feature points and descriptors is more robust than lens shift detection directly based on changes in pixel values.
[0048] In addition, by obtaining the detection results based on multiple reference distances, the interference on the detection results caused by excessive changes in a few key feature points can be reduced, and the robustness of lens shift detection can be further improved.
[0049] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0051] Figure 1 A first flow chart of a lens displacement detection method provided by an embodiment of the present invention;
[0052] Figure 2 A second flow chart of the lens displacement detection method provided by an embodiment of the present invention;
[0053] Figure 3 A third flow chart of the lens displacement detection method provided by an embodiment of the present invention;
[0054] Figure 4 A fourth flow chart of the lens displacement detection method provided by an embodiment of the present invention;
[0055] Figure 5 A structural diagram of a lens displacement detection device provided by an embodiment of the present invention;
[0056] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field based on the present invention belong to the scope of protection of the present invention.
[0058] The lens is an optical component set on the image acquisition device. By adjusting the position of the lens, the field of view of the image acquisition device can be adjusted. When the lens is displaced, the field of view of the image acquisition device changes accordingly. In actual scenarios, the lens can be displaced within a certain range. For example, if the field of view of the image acquisition device is large, when the lens has a small displacement, the field of view of the image acquisition device can still cover the specified area. Obviously, there is no need to adjust the lens at this time, and the displacement of the lens in this case can be called normal displacement. When the lens is displaced, causing the field of view of the image acquisition device to be unable to cover the specified area, it can be considered that the lens has an abnormal displacement, and at this time, the lens needs to be adjusted. In order to adjust the lens in time, that is, to detect the abnormal displacement of the lens in time, lens displacement detection can be performed, that is, abnormal lens displacement detection. With the development of image processing technology and computer vision technology, abnormal lens displacement detection has important application value in monitoring systems, photographic equipment, industrial detection and other fields.
[0059] In the prior art, there are many ways to detect abnormal lens displacement. For example, detection can be performed based on inertial sensors (such as gyroscopes, accelerometers) or mechanical calibration equipment. However, this method requires additional hardware support and is costly; and the actual scene is usually more complicated. For example, there may be interference from factors such as people walking and ground vibration, which makes this method susceptible to noise and has low detection accuracy.
[0060] For example, detection can be performed in an image analysis manner based on specific scene constraints (such as fixed markers, specific background textures). A technician can set a fixed marker in a specified area, and the marker will be displayed in the image collected by the lens, such as the marker displayed in the lower left corner of the image. If the lens is displaced, the lower left corner of the collected image may no longer display the marker, and the pixel value of each pixel in the lower left corner area of the currently collected image has a large change compared to the initial pixel value. Therefore, when the electronic device detects that the pixel value of the pixel in the lower left corner area of the currently collected image has a large change compared to the initial pixel value, a detection result indicating that the lens has been displaced can be obtained. However, although this method does not rely on additional hardware support, it relies on specific scene constraints and has poor versatility. In addition, as the light changes, the pixel value of each pixel in the image collected by the image acquisition device will change, resulting in a detection result indicating that the lens has been displaced when the light changes too much, even if the lens has not been displaced, a detection result indicating that the lens has been displaced will be obtained, that is, the robustness of this method is not high. Alternatively, the user may select the target area, that is, select an area in the video frame captured by the lens, and then detect whether the lens is displaced based on the area selected by the user. That is, the user pre-sets the area, which requires the user to perform additional operations, increases the user's usage cost and usage threshold, and lacks universality.
[0061] In order to solve the above problems, the present invention provides a lens displacement detection method, which can be applied to electronic devices. For example, the video captured by the lens can be uploaded to a server, and the electronic device can be a server or other device that communicates with the server, which is not limited by the present invention. After the electronic device obtains the video frame to be detected and the video frame to be compared, for each acquired video frame, each key feature point in the video frame can be extracted, and the descriptor of each extracted key feature point can be calculated. Furthermore, for each key feature point to be detected in the video frame to be detected, the electronic device can determine the matching key feature point corresponding to the key feature point to be detected from the key feature points to be compared included in the video frame to be compared according to the feature distance between the descriptors of the key feature points, that is, determine the matching key feature point indicating the same position in the three-dimensional space as the key feature point to be detected. Then the reference distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point can indicate: in different video frames captured by the lens to be detected, indicating the change of the pixel position of the same position in the three-dimensional space. Obviously, when the lens to be detected is displaced, the probability of a larger reference distance is higher; when the lens to be detected is not displaced, the probability of a smaller reference distance is higher. Therefore, based on the obtained reference distances, it is possible to determine whether the lens to be detected is displaced. Moreover, lens displacement detection based on key feature points and descriptors with higher robustness is more robust than lens displacement detection directly based on changes in pixel values. In addition, obtaining detection results based on multiple reference distances can reduce the interference caused by excessive changes in a few key feature points on the detection results, further improving the robustness of lens displacement detection.
[0062] See also Figure 1 , Figure 1 A first flow chart of a lens displacement detection method provided by an embodiment of the present invention, the method may include the following steps:
[0063] S101: Obtain a video frame to be detected and a video frame to be compared.
[0064] The video frames to be detected and the video frames to be compared are obtained based on the video frames in the video captured by the lens to be detected.
[0065] S102: For each acquired video frame, extract key feature points in the video frame.
[0066] S103: For each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the matching key feature point corresponding to the key feature point to be detected.
[0067] S104: Acquire a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance.
[0068] S105: Determine whether the lens to be detected is displaced based on the acquired reference distances.
[0069] Based on the lens displacement detection method provided by the present invention, after the electronic device obtains the video frame to be detected and the video frame to be compared, for each obtained video frame, each key feature point in the video frame can be extracted, and the descriptor of each extracted key feature point can be calculated. Further, for each key feature point to be detected in the video frame to be detected, the electronic device can determine the matching key feature point corresponding to the key feature point to be detected from the key feature points to be compared included in the video frame to be compared according to the feature distance between the descriptors of the key feature points, that is, determine the matching key feature point indicating the same position in the three-dimensional space as the key feature point to be detected. Then the reference distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point can indicate: in different video frames collected by the lens to be detected, the pixel position change indicating the same position in the three-dimensional space. Obviously, when the lens to be detected is displaced, the probability of a larger reference distance is greater; when the lens to be detected is not displaced, the probability of a smaller reference distance is greater. Therefore, based on the obtained reference distances, it can be realized to determine whether the lens to be detected is displaced.
[0070] In addition, since the key feature points are determined based on multiple pixels, the descriptors of the key feature points are also calculated based on multiple pixels, that is, the key feature points and descriptors are obtained by integrating multiple pixels. When some pixels change greatly, the key feature points and descriptors change less, that is, the robustness of the key feature points and descriptors is higher than the robustness of the pixel values of the pixels. Therefore, lens shift detection based on more robust key feature points and descriptors is more robust than lens shift detection directly based on changes in pixel values.
[0071] In addition, by obtaining the detection results based on multiple reference distances, the interference on the detection results caused by excessive changes in a few key feature points can be reduced, and the robustness of lens shift detection can be further improved.
[0072] For step S101, the electronic device can obtain a video frame to be detected and a video frame to be compared based on a video frame (which can be called a candidate video frame) in the video captured by the lens to be detected. There is usually one video frame to be detected, and there can be one or more video frames to be compared. The video frame to be detected and the video frame to be compared are not the same video frame.
[0073] In actual application scenarios, the lenses to be tested usually include two types: lenses with a fixed field of view (hereinafter referred to as fixed lenses) and lenses with a variable field of view (hereinafter referred to as variable lenses). For example, fixed lenses are usually fixed; variable lenses can rotate at a specified speed, or technicians can walk with variable lenses in their hands. Obviously, fixed lenses always capture videos of the same area, while variable lenses can capture videos of different areas.
[0074] The electronic device can obtain the video frames to be detected and the video frames to be compared according to the categories of the shots to be detected.
[0075] When the lens to be detected is a fixed lens, in one implementation method, the electronic device can randomly sample a video frame from the alternative video frames, and obtain the video frame to be detected based on the sampled video frame; and then sample video frames from other alternative video frames except the video frame to be detected, and obtain the video frame to be compared based on the video frame sampled this time.
[0076] In another implementation, the electronic device can obtain the video frame to be detected based on the first video frame captured by the lens to be detected; obtain the video frame to be compared based on the video frame currently captured by the lens to be detected, or based on multiple video frames most recently captured by the lens to be detected.
[0077] In another implementation, the electronic device can obtain the video frame to be detected based on the video frame currently captured by the lens to be detected; obtain the video frame to be compared based on the first video frame captured by the lens to be detected, or based on multiple video frames captured earliest by the lens to be detected.
[0078] But obviously, the method of obtaining the video frame to be detected and the video frame to be compared in the actual scene is not limited to this, and the present invention is not limited to this.
[0079] When the lens to be detected is a variable lens, the electronic device can randomly sample a video frame from the candidate video frames, and obtain the video frame to be detected based on the sampled video frame; then, within the preset video frame range of the sampled video frame, another video frame is sampled, and the video frame to be compared is obtained based on the video frame sampled this time. In this case, the way in which the electronic device obtains the video frame to be detected and the video frame to be compared can be referred to the detailed description of the subsequent embodiments.
[0080] After sampling video frames from the video captured by the lens to be detected, the electronic device can obtain the video frames to be detected and the video frames to be compared based on the sampled video frames.
[0081] In one implementation, the electronic device may directly use multiple video frames captured by the to-be-detected lens as the to-be-detected video frames and the to-be-compared video frames.
[0082] In another implementation, the electronic device may preprocess the multiple video frames captured by the detection lens, and use the preprocessed video frames as the detection video frames and the comparison video frames. The manner in which the electronic device preprocesses the multiple video frames captured by the detection lens is described in detail in the subsequent embodiments.
[0083] For step S102, for each acquired video frame, the electronic device can extract key feature points in the video frame based on the feature point extraction operator. The electronic device can extract key feature points from multiple video frames in parallel, or can extract key feature points from each video frame in sequence according to the acquisition order of the video frames, which is not limited in the present invention.
[0084] The feature point extraction operator may be a SIFT (Scale-Invariant Feature Transform) operator, an ORB (Oriented FAST and Rotated BRIEF) operator, etc., and the present invention is not limited to this. It can be understood that the ORB operator is a chip-friendly operator, that is, an operator designed based on a basic calculation method rather than a complex function. The requirements for hardware performance to run the ORB operator are low, so that the ORB operator can be run on most chips. Therefore, when extracting key feature points in a video frame based on the ORB operator, the hardware requirements can be reduced on the basis of improving the efficiency and robustness of the extracted key feature points, and the deployment is easy, thereby improving the versatility of the lens detection method.
[0085] For step S103, the key feature points extracted from the video frame to be detected can be called key feature points to be detected; the key feature points extracted from the video frame to be compared can be called key feature points to be compared. Obviously, there can be multiple key feature points in any video frame.
[0086] For each acquired video frame, after extracting the key feature points in the video frame, the electronic device can also calculate the descriptor of each extracted key feature point based on the descriptor calculation operator. The descriptor of a key feature point can describe the features of the image area within the preset neighborhood of the key feature point. For example, the descriptor calculation operator can be a SURF (Speeded Up Robust Features) operator, an ORB operator, etc., which is not limited in the present invention.
[0087] The descriptor of a key feature point to be detected matches the descriptor of a key feature point to be compared, that is, the feature distance between the descriptor of the key feature point to be detected and the descriptor of the key feature point to be compared is small, that is, the image area within the preset neighborhood range of the key feature point to be detected is similar to the image area within the preset neighborhood range of the key feature point to be compared, and the probability that the key feature point to be detected and the key feature point to be compared indicate the same position in the three-dimensional space is high. The feature distance can be Hamming distance, Euclidean distance, etc., and the present invention is not limited to this.
[0088] For each key feature point to be detected in the video frame to be detected, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected (i.e., a matching key feature point) is determined from the key feature points to be compared included in the video frame to be compared. The matching key feature point and the key feature point to be detected indicate the same position in the three-dimensional space.
[0089] For step S104 and step S105, after determining the matching key feature point corresponding to the key feature point to be detected, the electronic device can obtain the characteristic distance (i.e., reference distance) between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point. The reference distance can indicate the degree of similarity between the key feature point to be detected and the corresponding matching key feature point, and can also indicate the pixel distance between the key feature point to be detected and the corresponding matching key feature point.
[0090] After obtaining each reference distance, that is, obtaining the reference distance corresponding to each of the multiple key feature points to be detected, the electronic device can comprehensively obtain the multiple reference distances to jointly determine whether the lens to be detected is displaced. If the minimum value in each reference distance is also greater than the maximum distance threshold, that is, for the key feature point to be detected that is most similar to the corresponding matching key feature point, the pixel position of the key feature point to be detected is still far away from the pixel position of the matching key feature point indicating the same position in the three-dimensional space, the electronic device can determine that the lens to be detected is displaced. The maximum distance threshold can be determined according to the actual scene. For example, when the lens to be detected is a fixed lens, the maximum distance threshold can be set to a smaller value, such as 3; when the lens to be detected is a variable lens, the maximum distance threshold can be set to a larger value, such as 10.
[0091] In some embodiments, the video frame to be detected is obtained based on the current video frame captured by the lens to be detected; when the field of view of the lens to be detected is fixed, multiple video frames to be compared are obtained based on the following methods: sampling multiple video frames from all video frames currently captured by the lens to be detected; obtaining multiple video frames to be compared based on the sampled video frames; or, starting from the first video frame captured by the lens to be detected, sampling multiple consecutive video frames; obtaining multiple video frames to be compared based on the sampled video frames; and / or, when the field of view of the lens to be detected is variable, multiple video frames to be compared are obtained based on the following methods: determining the sampling interval according to the frame rate of the video captured by the lens to be detected; determining the starting video frame with a sampling interval of video frames between the current video frame and the starting video frame from the video captured by the lens to be detected; sampling multiple video frames from all video frames between the starting video frame and the current video frame; obtaining multiple video frames to be compared based on the sampled video frames.
[0092] It is understandable that when the lens to be detected is displaced, it is necessary to adjust the lens to be detected. In order to improve the timeliness of the adjustment, the electronic device can obtain the video frame to be detected based on the video frame currently collected by the lens to be detected. Compared with sampling video frames from the candidate video frames collected by the lens to be detected, the method of obtaining the video frame to be detected based on the sampled video frame (hereinafter referred to as the sampling method), since the video frame to be detected determined based on the sampling method is a historical video frame, if the lens to be detected is determined to be displaced based on the historical video frame, that is, after the lens to be detected is displaced for a period of time, it can be determined that the lens to be detected is displaced, and then the lens to be detected can be adjusted. Obviously, the real-time performance of determining the video frame to be detected based on the sampling method is not good, resulting in the lens to be detected that is displaced cannot be adjusted in time. When the electronic device obtains the video frame to be detected based on the video frame currently collected by the lens to be detected, the electronic device can detect in real time whether the lens to be detected is displaced, so that when the lens to be detected is displaced, the lens to be detected can be adjusted in time.
[0093] In one implementation, each time a video frame is captured by the lens to be detected, the electronic device can detect whether the lens to be detected is displaced based on the lens displacement detection method provided by the present invention and the video frame currently captured by the lens to be detected. That is, whether the lens to be detected is displaced is detected in real time, so that when the lens to be detected is displaced, the lens to be detected can be adjusted in time.
[0094] In another implementation, in order to reduce the amount of calculation and reduce the calculation pressure of the electronic device, the electronic device can periodically detect whether the lens to be detected is displaced based on the lens displacement detection method provided by the present invention. Each time the preset period is reached, the electronic device can obtain the video frame to be detected based on the current video frame captured by the lens to be detected. The preset period can be determined according to the business needs of the actual scene and the hardware performance of the electronic device. If the actual scene requires timely adjustment when the lens to be detected is displaced, and the hardware performance of the electronic device is good, the preset period can be set to a smaller value, such as 5 seconds; if the hardware performance of the electronic device is poor, the preset period can be set to a larger value, such as 1 minute, and the present invention is not limited to this.
[0095] In the case where there are multiple video frames to be compared, the electronic device can determine the method of obtaining the multiple video frames to be compared according to the category of the lens to be detected. The number of video frames to be compared (hereinafter referred to as the specified number) can be set based on the business needs of the actual scenario and the performance of the electronic device. If the business needs of the actual scenario require a lower false alarm rate, the specified number can be set to a larger value, such as 6; when it is required to reduce the computing pressure of the electronic device, the specified number can be set to a smaller value, such as 2. Normally, in order to reduce the false alarm rate and avoid excessive computing pressure on the electronic device, the specified number can be set to 4. The specified number can be recorded as frame_set (frame-setting).
[0096] When the field of view of the lens to be detected is fixed, that is, when the lens to be detected is a fixed lens, since the video captured by the fixed lens is always of the same area, as long as the lens to be detected does not move, each video frame in the video captured by the lens to be detected displays the same area.
[0097] Therefore, in this case, the electronic device can start sampling multiple consecutive video frames from the first video frame captured by the lens to be detected, that is, sampling multiple consecutive video frames first captured by the lens to be detected. This method can also be called sampling based on a sliding window, including multiple consecutive video frames of the first video frame, which can also be called video frames in the first window. Correspondingly, the specified number can be called the window size of the sliding window, and the window size can be recorded as WINDOW_SIZE (window-size). Alternatively, the electronic device can also sample multiple video frames from all video frames currently captured by the lens to be detected. For example, the electronic device can randomly sample multiple video frames, or it can sample multiple video frames according to a specified sampling interval. The present invention is not limited to this. The video frames sampled by the electronic device can be called frames added to the sliding window; sampling multiple video frames can be called filling the sliding window.
[0098] In the case where the field of view of the lens to be detected is variable, that is, in the case where the lens to be detected is a variable lens, the electronic device can determine the sampling interval according to the frame rate of the video captured by the lens to be detected. Since it is necessary to capture videos of different areas through a variable lens, that is, even if the variable lens does not have abnormal displacement, the video frames captured by the variable lens in different time periods may also show different areas. Therefore, in this case, in order to reduce the false alarm rate, the video frames to be detected and the video frames to be compared are obtained based on the video frames collected in the same time period. That is, in the case where the video frame to be detected is obtained based on the current video frame, the video frame interval between the video frame to be compared and the video frame to be detected needs to be no greater than the sampling interval.
[0099] For example, the sampling interval can be proportional to the frame rate of the video captured by the lens to be detected. For example, when the frame rate of the video captured by the lens to be detected is 30 frames per second, the sampling interval can be 60; when the frame rate of the video captured by the lens to be detected is 20 frames per second, the sampling interval can be 40. Alternatively, the electronic device can input the frame rate of the video captured by the lens to be detected and the frame number of the video frame currently captured by the lens to be detected in the video into a preset interval acquisition function, and use the calculation result of the function as the sampling interval. For example, the interval acquisition function can be MATCH_RESOLUTION (matching-frame rate). That is, the frequency of frame comparison is controlled based on the frame sampling logic.
[0100] Furthermore, the electronic device may determine a starting video frame that is separated from the current video frame by a sampling interval of video frames. Then, multiple video frames are sampled from all video frames between the starting video frame and the current video frame (i.e., the preset video frame range in the aforementioned embodiment). For example, the electronic device may randomly sample multiple video frames; sample multiple video frames according to a specified sampling interval; sample multiple consecutive video frames starting from the starting video frame, etc. The present invention is not limited to this.
[0101] In some embodiments, the electronic device can detect whether the lens to be detected is displaced based on the lens displacement detection method provided by the present invention every time the lens to be detected collects sampling interval video frames, that is, perform lens displacement detection every MATCH_RESOLUTION video frames.
[0102] Based on the above processing, lens displacement detection is performed by sampling and determining the sampling interval based on the frame rate, which can reduce the amount of calculation and reduce the calculation pressure of the electronic device, that is, further optimize the calculation efficiency, so that the lens detection method provided by the present invention can be implemented in a resource-constrained environment. Furthermore, the lens detection method provided by the present invention is suitable for multiple application fields such as dynamic scene monitoring, moving target detection, video anomaly analysis, etc., and provides an efficient solution for intelligent video analysis.
[0103] For example, see Figure 2 , Figure 2 A second flow chart of the lens displacement detection method provided by an embodiment of the present invention. Figure 2 The vertical axis represents the time axis, and the horizontal axis represents the play axis. Figure 2 The squares in represent video frames, and the squares filled with black represent: at a moment, the current video frame currently captured by the lens to be detected.
[0104] At time t1, the lens to be detected has currently captured 8 video frames. The last video frame is the current video frame currently captured by the lens to be detected at time t1. Based on the first 4 video frames captured by the lens to be detected, the video frame to be compared can be obtained. At time t2, the lens to be detected has currently captured multiple video frames. The last video frame is the current video frame currently captured by the lens to be detected at time t2; based on the first 4 video frames captured after the lens to be detected captures the sampling interval video frames, the video frame to be compared can be obtained.
[0105] In some embodiments, the video frame to be detected is obtained based on the following method: obtaining the current video frame captured by the lens to be detected; performing image enhancement on the current video frame to obtain the video frame to be detected. Accordingly, the aforementioned obtaining multiple video frames to be compared based on the sampled video frame includes: performing image enhancement on each sampled video frame to obtain multiple video frames to be compared.
[0106] After determining the current video frame and each sampled video frame (hereinafter collectively referred to as the original video frame), the electronic device can perform image enhancement on the original video frame to obtain the video frame to be detected and the video frame to be compared. Compared with directly using the original video frame as the video frame to be detected and the video frame to be compared, the image features in the video frame to be detected and the video frame to be compared are prominent, and the key feature points extracted subsequently are more accurate. Therefore, the accuracy of the subsequent detection result of whether the lens to be detected is displaced based on the more accurate key feature points is higher, which can further reduce the false alarm rate.
[0107] In some embodiments, for any video frame, image enhancement is performed on the video frame by at least one of the following: image compression is performed on the video frame; the video frame is converted into a grayscale image; the video frame is divided, and based on a histogram equalization algorithm, pixel value equalization is performed on each divided image area.
[0108] When the video frame is compressed, the compressed video frame has a smaller data volume than the original video frame. Therefore, lens displacement detection based on the compressed video frame can reduce subsequent calculation costs and improve the efficiency of lens displacement detection. For example, the video frame can be compressed to 30% of the original video frame, that is, the reduction ratio is 30% of the original size.
[0109] When the video frame is converted into a grayscale image, since the data volume of the grayscale image is smaller than that of the RGB (red, green, and blue) image, lens displacement detection based on the grayscale image can reduce subsequent calculation costs and improve the efficiency of lens displacement detection.
[0110] When the video frame is divided and the pixel value of each divided image area is equalized based on the histogram equalization algorithm, the electronic device first divides the video frame into multiple image areas and then equalizes the pixel value of each image area. The divided image area can be recorded as Tile Grid (region grid); the size of each divided image area (hereinafter referred to as grid size) can be set according to the business requirements of the actual scene and the performance of the electronic device. If the lens shift detection is required to be more robust in the actual scene, the image details in the enhanced video frame are required to be clearer, then the grid size can be set to a smaller value, such as 4×4, and accordingly, the number of divided image areas is larger. When the performance of the electronic device is poor, in order to reduce the amount of calculation when performing image enhancement on the video frame, the grid size can be set to a larger value, such as 8×8, and the grid size can be recorded as Tile Grid Size=(8,8). Correspondingly, the number of divided image areas is smaller.
[0111] By equalizing the pixel values of each divided image area, the electronic device can perform equalization processing on areas with uneven brightness distribution in the video frame, such as shadow areas and bright light areas, to reduce the probability of image detail loss in shadow areas or bright light areas. Compared with the original video frame, the video frame after pixel value equalization is more robust to different lighting conditions and has higher local contrast. Lens shift detection based on the video frame after pixel value equalization can reduce the impact of different lighting conditions on the extraction and matching of key feature points, and has higher robustness under different lighting conditions, further reducing the false alarm rate.
[0112] In some embodiments, the above-mentioned equalization of pixel values of each divided image region based on the histogram equalization algorithm includes the following steps:
[0113] Step 1: For each image region obtained by division, the pixel points with the same pixel value among the pixel points included in the image region are divided into a group.
[0114] Step 2: For each group of pixels, determine the number of pixels in the image area whose pixel values are not greater than the pixel values of the group of pixels, as the first number.
[0115] Step 3: Calculate the ratio of the first number to the total number of pixels included in the image area to obtain the proportion corresponding to the group of pixels.
[0116] Step 4: Calculate the product of the proportion corresponding to the group of pixels and the specified value to obtain the equalization value to be restricted corresponding to the group of pixels.
[0117] The specified value is: the length of the maximum distribution range of pixel values representing the video frame in the shot to be detected.
[0118] Step 5: For each group of pixels, calculate the product of the pixel value of the group of pixels and the contrast limit coefficient to obtain the pixel threshold of the group of pixels.
[0119] Step 6: When the to-be-limited equalization value of the group of pixels is greater than the pixel threshold of the group of pixels, the difference between the to-be-limited equalization value of the group of pixels and the pixel threshold of the group of pixels is calculated.
[0120] Step 7: Calculate the quotient of the difference and the total number to obtain the correction value of the group of pixels.
[0121] Step 8: For each pixel in the image area, determine the minimum value of the pixel to be limited equalization and the pixel threshold of the pixel.
[0122] Step 9: Calculate the sum of the minimum value and the correction value of each group of pixels.
[0123] Step 10: Based on the sum value, obtain the pixel value equalization processing result corresponding to the pixel point.
[0124] For each image area obtained by division, the electronic device can obtain the pixel value of each pixel point included in the image area, and then divide the pixel points with the same pixel value into a group. Subsequently, the electronic device can jointly calculate the equalization value to be limited and the pixel threshold corresponding to each pixel point in the group of pixel points.
[0125] The specified value is: the length of the maximum distribution range of the pixel values of the video frame in the shot to be detected. If the video frame is a grayscale image, the maximum distribution range of the pixel values of the video frame is: [0, 255], then the specified value is 256; if the video frame is a binary image, the pixel values of the video frame only include 0 or 255, then the specified value is 2.
[0126] According to the business needs of the actual scenario, you can set Clip Limit (contrast limit). When the contrast limit coefficient is large, the image details in the enhanced video frame are more prominent; when the contrast limit coefficient is small, it can avoid excessive contrast enhancement, limit the cumulative frequency of high-frequency pixel values, and reduce the probability of overexposure or artifact problems. Exemplarily, under normal circumstances, the contrast limit coefficient can be set to 2.0. The above-mentioned histogram equalization algorithm can also be CLAHE (Contrast Limited Adaptive Histogram Equalization).
[0127] For example, the maximum distribution range of the pixel value of a video frame is [0, 255], that is, the specified value is 256; the grid size is 8×8, that is, an image area includes 64 pixels; the contrast limit coefficient is 2. Among them, there are 8 pixels with a pixel value of 40 (hereinafter referred to as the first pixel), 16 pixels with a pixel value of 50 (hereinafter referred to as the second pixel), and 40 pixels with a pixel value of 60 (hereinafter referred to as the third pixel).
[0128] For the first pixel, it is obvious that the first number corresponding to the first pixel is 8, and the proportion corresponding to the first pixel is 0.125. The equalization value to be limited corresponding to the first pixel is: 0.125×256, that is, 32; the pixel threshold of the first pixel is: 40×2, that is, 80. Since 32 is less than 80, the first pixel does not correspond to the correction value.
[0129] For the second pixel, it is obvious that the first number corresponding to the second pixel is 24, and the proportion corresponding to the second pixel is 0.375. The equalization value to be limited corresponding to the second pixel is: 0.375×256, that is, 96; the pixel threshold of the first pixel is: 50×2, that is, 100. Since 96 is less than 100, the second pixel does not correspond to the correction value.
[0130] For the third pixel, it is obvious that the first number corresponding to the third pixel is 64, so the proportion corresponding to the third pixel is 1. The equalization value to be limited corresponding to the third pixel is: 1×256, that is, 256; the pixel threshold of the third pixel is: 60×2, that is, 120. Since 256 is greater than 120, the electronic device calculates the difference between 256 and 120, that is, 136; then calculates the quotient of the difference and the total number of pixels included in the image area, that is, 136 / 64, and obtains 2.125. That is, the correction value of the third pixel is 2.125.
[0131] Then, for each pixel in the image area, for example, for the first pixel, the minimum value of the first pixel's to-be-restricted equalization value (i.e., 32) and the first pixel's pixel threshold (i.e., 80) is 32; then the sum of 32 and 2.125 is calculated, i.e., 34.125; based on 34.125, the pixel value equalization processing result corresponding to the first pixel can be obtained.
[0132] Similarly, for the second pixel, the minimum value of the second pixel's to-be-restricted equalization value (i.e., 96) and the second pixel's pixel threshold (i.e., 100) is 96; then the sum of 96 and 2.125 is calculated, i.e., 98.125; based on 98.125, the pixel value equalization processing result corresponding to the second pixel can be obtained.
[0133] For the third pixel, the minimum value of the third pixel's to-be-restricted equalization value (i.e., 256) and the third pixel's pixel threshold (i.e., 120) is 120; then the sum of 120 and 2.125 is calculated, i.e., 122.125; based on 122.125, the pixel value equalization processing result corresponding to the third pixel can be obtained.
[0134] In some embodiments, after obtaining the sum value corresponding to each pixel point, the electronic device can directly use the calculated sum value as the pixel value equalization processing result corresponding to the pixel point.
[0135] In some embodiments, the aforementioned step 10 may include the following steps: when the pixel point is not adjacent to other image areas, using the sum value corresponding to the pixel point as the pixel value equalization processing result corresponding to the pixel point; when the pixel point is adjacent to other image areas, determining the pixel points adjacent to the pixel point in other image areas; interpolating the sum value corresponding to the pixel point with the sum value corresponding to the determined adjacent pixel points to obtain the pixel value equalization processing result corresponding to the pixel point.
[0136] Since the electronic device performs pixel value equalization on each image area separately, when the image areas are adjacent, for the pixels at the boundary of the adjacent image areas, the difference in the sum values corresponding to these pixel points may be large. If the sum value corresponding to the pixel point is directly used as the result of the pixel value equalization processing corresponding to the pixel point, the enhanced video frame obtained may have obvious block division, resulting in low accuracy of the key feature points extracted based on the video frame with obvious block division. Therefore, in order to improve the accuracy of the key feature points extracted subsequently, the electronic device can perform fusion boundary processing to make the adjacent image areas excessively smooth and reduce the probability of obvious block division problems.
[0137] Specifically, for each pixel point in the video frame, when the pixel point is not adjacent to other image areas, the electronic device can directly use the sum value corresponding to the pixel point as the pixel value equalization processing result corresponding to the pixel point.
[0138] When the pixel point is adjacent to other image areas, that is, the pixel point is located at the boundary between image area 1 and an adjacent image area (hereinafter referred to as image area 2) in the image area to which it belongs (hereinafter referred to as image area 1). The electronic device can determine the pixel points adjacent to the pixel point in image area 2. Obviously, there may be multiple determined pixel points. Then, the electronic device can interpolate the sum value corresponding to the pixel point and the sum value corresponding to all the determined adjacent pixel points according to the interpolation algorithm to obtain the pixel value equalization processing result corresponding to the pixel point. For example, the interpolation algorithm can be the nearest neighbor method, the bilinear interpolation method, etc., and the present invention is not limited to this.
[0139] In some embodiments, Figure 1 Based on Figure 3 Before step S103, the method may further include the following steps:
[0140] S106: Obtain the number of key feature points to be detected extracted from the video frame to be detected as a reference number.
[0141] S107: When the reference number is less than the detection quantity threshold, output a prompt message.
[0142] Accordingly, the above step S103 may include the following steps:
[0143] S1031: When the reference number is not less than the detection quantity threshold, for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the matching key feature point corresponding to the key feature point to be detected.
[0144] The actual application scenario is relatively complex, and the features contained in the video frame to be detected may be very few, such as the video of a solid-color wall captured by the lens to be detected. In this case, the number of key feature points to be detected (i.e., the reference number) extracted will also be small, that is, the reference number is less than the detection quantity threshold. Subsequently, the accuracy of the detection result of whether the lens is displaced based on fewer key feature points to be detected is not high. The key feature points extracted at this time can be called invalid data. In this case, the electronic device can directly output a prompt message, such as a prompt message that can be a text "The current video contains too little information and is difficult to detect." Subsequently, the technician can perform corresponding processing based on the text, such as directly determining that the lens to be detected has not been displaced.
[0145] The detection quantity threshold can be determined according to the business needs of the actual scenario. For example, if the business needs indicate that the false alarm rate needs to be reduced, the detection quantity threshold can be set to a larger value, such as 500; if the business needs indicate that the displacement of the lens to be detected needs to be adjusted in time, the detection quantity threshold can be set to a smaller value, such as 100.
[0146] When the reference number is not less than the detection quantity threshold, a relatively accurate detection result can be obtained based on the extracted key feature points to be detected. At this time, the electronic device determines whether the lens to be detected is displaced based on the extracted key feature points to be detected according to the lens displacement detection method provided by the present invention.
[0147] Based on the above processing, the electronic device can perform pre-processing according to the reference number of key feature points to be detected that are extracted, and directly output prompt information when the reference number is less than the detection quantity threshold; subsequent processing is performed only when the reference number is not less than the detection quantity threshold, which can reduce the probability of subsequent processing of invalid data, reduce the amount of calculation of the electronic device, and reduce the calculation pressure of the electronic device.
[0148] In some embodiments, for each key feature point to be detected in the video frame to be detected, the electronic device can calculate the feature distance between the descriptor of the key feature point to be detected and the descriptors of each key feature point to be compared, and then determine the matching key feature point corresponding to the key feature point to be detected based on the minimum value of each feature distance.
[0149] In some embodiments, Figure 1 Based on Figure 4 The above step S103 may include the following steps:
[0150] S1032: For each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected.
[0151] The hash bucket to which a key feature point belongs is determined by mapping the descriptor of the key feature point using a local sensitive hashing algorithm.
[0152] S1033: Calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point.
[0153] S1034: Based on the minimum value among the feature distances, determine, from the candidate key feature points corresponding to the key feature point to be detected, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected.
[0154] The electronic device may determine matching key feature points corresponding to the key feature points to be detected based on a FLANN (Fast Library for Approximate Nearest Neighbors) matcher.
[0155] Specifically, after extracting the key feature points in each video frame, the electronic device can first map the descriptor of the key feature point based on the LSH (Locality-Sensitive Hashing) algorithm for each extracted key feature point, and determine the hash bucket to which the key feature point belongs. The probability that the key feature points belonging to the same hash bucket have similar descriptors is high. Furthermore, for each key feature point to be detected in the video frame to be detected, the electronic device can determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected. In this way, the electronic device also selects the candidate key feature points from the key feature points to be compared based on the LSH algorithm. Then, the electronic device calculates the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point, and determines the matching key feature point corresponding to the key feature point to be detected based on the minimum value of each feature distance.
[0156] Based on the above processing, the electronic device first screens each key feature point to be compared based on the LSH algorithm, and then determines the matching key feature point corresponding to the key feature point to be detected based on the screened candidate key feature points. There is no need to calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each key feature point to be compared, which reduces the amount of calculation and improves the efficiency and quality of search and matching.
[0157] In some embodiments, when the minimum value of the feature distance corresponding to the key feature point to be detected is not greater than the feature distance threshold, the electronic device can directly use the alternative key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected.
[0158] In some embodiments, in order to filter out false matches and further improve the robustness of the lens displacement detection method, the electronic device may determine the matching key feature points corresponding to the key feature points to be detected based on the kNN (k-Nearest Neighbor) algorithm. Specifically, the above step S1034 may include the following steps: when there is only one video frame to be compared, calculate the ratio of the minimum value and the second minimum value in the feature distance corresponding to the key feature point to be detected as the first ratio; if the first ratio is less than the first ratio threshold, use the candidate key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected. Correspondingly, the above step S105 may include the following steps: calculate the average level of each reference distance obtained; if the average level is less than the first distance threshold, determine that the lens to be detected has not been displaced; otherwise, determine that the lens to be detected has been displaced.
[0159] In the case where there is only one video frame to be compared, that is, each feature distance corresponding to the key feature point to be detected is: the feature distance between the descriptor of the key feature point to be detected and the descriptor of the candidate key feature point in the single video frame to be compared. The electronic device can determine the minimum feature distance (i.e., the minimum value) and the second smallest feature distance (i.e., the second smallest value) from the feature distances corresponding to the key feature point to be detected. Then, the ratio of the minimum value to the second smallest value is calculated as the first ratio. In the case where the first ratio is less than the first ratio threshold, that is, the key feature point to be detected is more matched with the best matching candidate key feature point, and the difference with the second matching candidate key feature point is large. At this time, the electronic device can use the candidate key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected, that is, the best matching candidate key feature point, as the matching key feature point corresponding to the key feature point to be detected. The first ratio threshold can be set by the technician according to the actual scene, such as 0.7.
[0160] In some embodiments, before the above step S104, the method may further include the following steps: if the first ratio is not less than the first ratio threshold, determine that there is no matching key feature point corresponding to the key feature point to be detected; obtain the number of key feature points to be detected for which there are corresponding matching key feature points as the first matching number; if the first matching number is less than the first matching point number threshold, output prompt information. Accordingly, the above step S104 may include the following steps: if the first matching number is not less than the first matching point number threshold, obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance.
[0161] In the case where the first ratio is not less than the first ratio threshold, that is, the key feature point to be detected is relatively close to multiple candidate key feature points, that is, it is difficult to determine the matching key feature point that accurately corresponds to the key feature point to be detected in the video frame to be compared. Therefore, in order to filter out wrong matches, reduce the probability of false detection, and improve the robustness of the lens displacement detection method, in this case, the electronic device can directly determine that there is no matching key feature point corresponding to the key feature point to be detected. Furthermore, the electronic device can count the number of key feature points to be detected that have corresponding matching key feature points as the first matching number.
[0162] In the case where the number of first matches is less than the first matching point number threshold, that is, the number of key feature points to be detected (hereinafter referred to as first valid key feature points) with corresponding matching key feature points is small, then when determining whether the lens is displaced based on the small number of first valid key feature points, the detection results may be disturbed due to the large changes in a few first valid key feature points, resulting in reduced robustness of lens displacement detection. Therefore, in order to reduce the probability of the above problem, in this case, the electronic device can also directly output prompt information, such as a text message "The number of valid points is too small and difficult to detect". Subsequently, the technician can perform corresponding processing based on the text, such as directly determining that the lens to be detected has not been displaced.
[0163] The first matching point number threshold can be recorded as: MIN_MATCH_COUNT, which can be determined according to the business needs of the actual scene. For example, if the business needs indicate that the false alarm rate should be reduced, the first matching point number threshold can be set to a larger value, such as 15; if the business needs indicate that the displacement of the to-be-detected lens should be adjusted in time, the first matching point number threshold can be set to a smaller value, such as 5. Experimental results show that when the first matching point number threshold is set to 10, most actual scenes can be covered.
[0164] When the first matching number is not less than the first matching point number threshold, that is, when the number of valid key feature points is large, the electronic device can filter out erroneous matches based on subsequent processing, reduce the probability of false detection, and improve the robustness of the lens displacement detection method.
[0165] Furthermore, after obtaining each reference distance, the electronic device can calculate the average level of the obtained reference distances, such as determining the median of each reference distance obtained, or determining the average of each reference distance obtained. When the average level is less than the first distance threshold, it means that the change of the video frame to be detected is small compared with the video frame to be compared. At this time, it can be determined that the lens to be detected has not been displaced. When the average level is not less than the first distance threshold, it means that the change of the video frame to be detected is large compared with the video frame to be compared. At this time, it can be determined that the lens to be detected has been displaced. The first distance threshold can be recorded as GOOD_DISTANCE (good matching distance threshold). The first distance threshold can be determined based on the business needs of the actual scene. If the business needs indicate that the false alarm rate should be reduced, the first distance threshold can be set to a larger value, such as 10; if the business needs indicate that the lens to be detected that has been displaced should be adjusted in time, the first distance threshold can be set to a smaller value, such as 6. Experimental results show that when the first matching point number threshold is set to 8, most actual scenes can be covered.
[0166] In some embodiments, the above-mentioned step S1034 may include the following steps: in the case that there are multiple video frames to be compared, for each video frame to be compared, determining the alternative key feature point corresponding to the key feature point to be detected in the video frame to be compared as the key feature point to be screened; calculating the ratio of the minimum value and the second minimum value in the feature distance corresponding to the determined key feature point to be screened as the second ratio corresponding to the video frame to be compared; if the second ratio corresponding to the video frame to be compared is less than the second ratio threshold, the key feature point to be screened with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected is used as the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared. Accordingly, the above-mentioned step S105 may include the following steps: for each video frame to be compared, calculating the average level of the reference distance corresponding to the matching key feature points in the video frame to be compared, as the average level corresponding to the video frame to be compared; if the average level corresponding to the video frame to be compared is less than the second distance threshold, determining that the video frame to be compared indicates that the lens to be detected has not been displaced; otherwise, determining that the video frame to be compared indicates that the lens to be detected has been displaced; if the number of video frames to be compared indicating that the lens to be detected has not been displaced is less than the number of video frames to be compared indicating that the lens to be detected has been displaced, determining that the lens to be detected has not been displaced; otherwise, determining that the lens to be detected has been displaced.
[0167] In the case where there are multiple video frames to be compared, that is, the candidate key feature points corresponding to the key feature points to be detected may be candidate key feature points in different video frames to be compared. At this time, for each video frame to be compared, the electronic device can determine the candidate key feature points corresponding to the key feature points to be detected in the video frame to be compared as the key feature points to be screened. At this time, the feature distances corresponding to the key feature points to be screened, that is, the feature distances between the descriptors of the key feature points to be detected and the descriptors of the candidate key feature points in the video frame to be compared.
[0168] Furthermore, the electronic device can determine the smallest characteristic distance (i.e., the minimum value) and the second smallest characteristic distance (i.e., the second smallest value) from the characteristic distances corresponding to the determined key characteristic points to be screened. Then, the ratio of the smallest value and the second smallest value of the characteristic distances corresponding to the determined key characteristic points to be screened is calculated as the second ratio.
[0169] When the second ratio is less than the second ratio threshold, that is, the key feature point to be detected is relatively matched with the most matched key feature point to be screened, and the difference between it and the second most matched key feature point to be screened is relatively large. At this time, the electronic device can use the key feature point to be screened with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected, that is, the most matched key feature point to be screened, as the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared. The second ratio threshold can be set by a technician according to the actual scenario, such as 0.7. The first ratio threshold and the second ratio threshold can be the same or different, and the present invention is not limited to this.
[0170] In some embodiments, before the above step S104, the method may further include the following steps: for each video frame to be compared, if the second ratio corresponding to the video frame to be compared is not less than the second ratio threshold, determine that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; obtain the number of key feature points to be detected that have corresponding matching key feature points in the video frame to be compared, as the second matching number corresponding to the video frame to be compared; when the second matching number corresponding to any video frame to be compared is less than the second matching point number threshold, output prompt information. Accordingly, the above step S104 may include the following steps: when the second matching number corresponding to each video frame to be compared is not less than the second matching point number threshold, for each video frame to be compared, obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point in the video frame to be compared, as the reference distance corresponding to the key feature point to be detected in the video frame to be compared.
[0171] In the case where the second ratio is not less than the second ratio threshold, that is, the key feature point to be detected is relatively matched with multiple key feature points to be screened, that is, it is difficult to determine the matching key feature point that accurately corresponds to the key feature point to be detected in the video frame to be compared. Therefore, in order to filter out wrong matches, reduce the probability of false detection, and improve the robustness of the lens displacement detection method, in this case, the electronic device can directly determine that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared. Furthermore, the electronic device can count the number of key feature points to be detected that have corresponding matching key feature points in the video frame to be compared as the second matching number.
[0172] In the case where the second matching number is less than the second matching point number threshold, that is, the number of key feature points to be detected (hereinafter referred to as second valid key feature points) with corresponding matching key feature points is small, then when determining whether the lens is displaced based on the small number of second valid key feature points, the detection results may be disturbed due to the large changes in a few second valid key feature points, resulting in reduced robustness of lens displacement detection. Therefore, in order to reduce the probability of the above problem, in this case, the electronic device can also directly output prompt information, such as a text message "The number of valid points is too small and difficult to detect". Subsequently, the technician can perform corresponding processing based on the text, such as directly determining that the lens to be detected has not been displaced.
[0173] The second matching point number threshold can be determined according to the business needs of the actual scene. For example, if the business needs indicate that the false alarm rate should be reduced, the second matching point number threshold can be set to a larger value, such as 15; if the business needs indicate that the displacement of the to-be-detected lens should be adjusted in time, the second matching point number threshold can be set to a smaller value, such as 5. Experimental results show that when the first matching point number threshold is set to 10, most actual scenes can be covered. The first matching point number threshold and the second matching point number threshold can be the same or different, and the present invention is not limited to this.
[0174] When the second matching number is not less than the second matching point number threshold, that is, when the number of second valid key feature points is large, the electronic device can filter out erroneous matches based on subsequent processing, reduce the probability of false detection, and improve the robustness of the lens displacement detection method.
[0175] Furthermore, for each video frame to be compared, after determining the matching key feature points in the video frame to be compared, the electronic device can obtain the average level of the reference distance (hereinafter referred to as the matching reference distance) corresponding to the matching key feature points as the average level corresponding to the video frame to be compared. For example, the median of the matching reference distance is determined, or the average of the matching reference distance is determined. When the average level is less than the second distance threshold, it means that the video frame to be compared indicates that the change of the video frame to be detected is small compared with the video frame to be compared, that is, it indicates that the lens to be detected has not been displaced. When the average level is not less than the second distance threshold, it means that the video frame to be compared indicates that the change of the video frame to be detected is large compared with the video frame to be compared, that is, it indicates that the lens to be detected has been displaced.
[0176] The second distance threshold can be recorded as GOOD_DISTANCE and can be determined based on the business needs of the actual scene. For example, if the business needs indicate that the false alarm rate should be reduced, the second distance threshold can be set to a larger value, such as 10; if the business needs indicate that the displacement of the to-be-detected lens should be adjusted in time, the second distance threshold can be set to a smaller value, such as 6. Experimental results show that when the first matching point number threshold is set to 8, most actual scenes can be covered. The first distance threshold and the second distance threshold can be the same or different, and the present invention is not limited to this.
[0177] Then, the electronic device can compare the number of video frames to be compared that indicate that the lens to be detected has not been displaced (hereinafter referred to as the first number) and the number of video frames to be compared that indicate that the lens to be detected has been displaced (hereinafter referred to as the second number). When the first number is less than the second number, that is, more video frames to be compared indicate that the lens to be detected has not been displaced, at this time, it can be determined that the lens to be detected has not been displaced. When the first number is not less than the second number, that is, more video frames to be compared indicate that the lens to be detected has been displaced, at this time, it can be determined that the lens to be detected has been displaced. That is, the electronic device can combine the inter-frame feature matching and the voting mechanism to determine whether there is motion or abnormality in the current frame.
[0178] For example, when a video frame to be compared indicates that the lens to be detected has not been displaced, a normal matching frame pair returns a result of "0"; when the video frame to be compared indicates that the lens to be detected has been displaced, an abnormal matching frame pair returns a result of "1". The electronic device can summarize all the returned results into results (result set) for voting and statistics. The number of results "0" can be recorded as: count-0; the number of results "1" can be recorded as: count-1. When count-1 is not less than count-0, it can be determined that the current frame (i.e., the current video frame) is an abnormal frame, that is, the lens to be detected has been displaced; when count-1 is less than count-0, it can be determined that the current frame is a normal frame, that is, the lens to be detected has not been displaced.
[0179] Based on the above processing, the electronic device can determine the matching key features corresponding to the key feature points to be detected based on the kNN algorithm, and then determine whether the lens is displaced based on the reference distance corresponding to the matching key feature points. False matches can be filtered to further improve the robustness of the lens displacement detection method. In addition, when there are multiple video frames to be compared, the multiple video frames are combined to jointly detect whether the lens is displaced, which can reduce the probability of single-frame misjudgment and improve the accuracy of the obtained lens displacement detection results.
[0180] Furthermore, based on the technical solution provided by the present invention, the lens displacement detection results shown in the following table (1) can be obtained:
[0181] Table (1)
[0182]
[0183] Among them, ".mp4" indicates the video format of the lens to be detected, "night" indicates that the video is a video captured at night; "day" indicates that the video is a video captured during the day; "displacement" indicates that the lens to be detected is displaced; "no displacement" indicates that the lens to be detected is not displaced; "no illumination change" indicates that the lens to be detected is capturing video under constant illumination conditions; "illumination change" indicates that the lens to be detected is capturing video under variable illumination conditions; "with pedestrians" indicates that there are pedestrians in the video captured by the lens to be detected; "(2)" indicates the second video under the same conditions; "(3)" indicates the third video under the same conditions. "True value" indicates whether the lens to be detected is actually displaced; "benchmark" indicates: the detection result of whether the lens to be detected is displaced based on the lens displacement detection method provided by the present invention when no image enhancement is performed and there is only one video frame to be compared; "further optimization" indicates: the detection result of whether the lens to be detected is displaced based on the lens displacement detection method provided by the present invention when image enhancement is performed and there are multiple video frames to be compared.
[0184] As shown in Table (1), in the above 14 videos collected in actual scenes, the lens to be detected actually moved when collecting 3 of the videos, and the “benchmark” method can detect all 3 situations, that is, the recall rate of displacement is 100%. However, the “benchmark” method has 3 false alarms, that is, when the lens to be detected actually did not move, it was still determined that the lens to be detected was displaced. However, based on the “further optimization” method, compared with the “benchmark” method, that is, compared with only detecting whether the lens is displaced based on one video frame to be compared, that is, judging abnormal displacement based on single-frame feature matching, the false alarm rate can be further reduced, and the robustness of the lens detection method can be improved. And through the video visualization method, it can be found that the slowly changing environmental scene changes in the scene, such as sunlight (changes between day and night), have a low impact on the algorithm. However, for fast environmental scene changes, such as nearby flashlights, changes in light and darkness on buildings, etc., the robustness of the benchmark algorithm is reduced and false alarms occur. However, after image enhancement and lens displacement detection based on multiple video frames to be compared, accurate warnings and no false alarms can be given in the presence of pedestrians and changes in light intensity. That is, based on the "further optimization" method, it is robust to changes in light intensity and to flashlights at near and far distances.
[0185] Compared with the method of detecting whether the lens is displaced based on a neural network, since the neural network needs to detect the position of a fixed object in the video captured by the lens, and then determine that the lens is displaced when the position of the detected fixed object changes, a large amount of calculation is required, which is time-consuming, and the detection efficiency is low, making it difficult to meet the needs of large-scale deployment. The lens displacement detection method provided by the present invention does not need to use a neural network algorithm, is more lightweight, reduces memory loss, and has a lower implementation cost.
[0186] Based on the same inventive concept as the above lens displacement detection method, the present invention also provides a lens displacement detection device, see Figure 5 , Figure 5 A structural diagram of a lens displacement detection device provided by an embodiment of the present invention, the device includes:
[0187] The video frame acquisition module 501 is used to acquire the video frame to be detected and the video frame to be compared; wherein the video frame to be detected and the video frame to be compared are obtained based on the video frame in the video captured by the lens to be detected;
[0188] The extraction module 502 is used to extract key feature points in each acquired video frame;
[0189] A determination module 503 is used to determine, for each key feature point to be detected in the video frame to be detected, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected;
[0190] A distance acquisition module 504 is used to acquire a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance;
[0191] The detection module 505 is used to determine whether the lens to be detected is displaced based on the acquired reference distances.
[0192] Optionally, the determination module 503 is specifically used to: for each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected; wherein the hash bucket to which a key feature point belongs is: determined by mapping the descriptor of the key feature point using the local sensitive hashing algorithm; calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point; based on the minimum value of each feature distance, determine the candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected from the candidate key feature points corresponding to the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected.
[0193] Optionally, the determination module 503 is specifically used to: when there is only one video frame to be compared, calculate the ratio of the minimum value to the second minimum value in the feature distance corresponding to the key feature point to be detected as a first ratio; if the first ratio is less than a first ratio threshold, use the alternative key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected; the detection module 505 is specifically used to: calculate the average level of each reference distance obtained; if the average level is less than the first distance threshold, determine that the lens to be detected has not been displaced; otherwise, determine that the lens to be detected has been displaced.
[0194] Optionally, the device further includes: a first prompt module, which is used to execute, before the distance acquisition module 504 executes to obtain the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, if the first ratio is not less than the first ratio threshold, determine that there is no matching key feature point corresponding to the key feature point to be detected; obtain the number of key feature points to be detected that have corresponding matching key feature points as the first matching number; and output a prompt message when the first matching number is less than the first matching point number threshold;
[0195] The distance acquisition module 504 is specifically used to: when the first matching number is not less than the first matching point quantity threshold, obtain the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance.
[0196] Optionally, the determination module 503 is specifically used to: in the case where there are multiple video frames to be compared, for each video frame to be compared, determine the candidate key feature point corresponding to the key feature point to be detected in the video frame to be compared as the key feature point to be screened; calculate the ratio of the minimum value and the second minimum value in the feature distance corresponding to the determined key feature point to be screened as the second ratio corresponding to the video frame to be compared; if the second ratio corresponding to the video frame to be compared is less than a second ratio threshold, use the key feature point to be screened with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared;
[0197] The detection module 505 is specifically used to: for each video frame to be compared, calculate the average level of the reference distance corresponding to the matching key feature points in the video frame to be compared as the average level corresponding to the video frame to be compared; if the average level corresponding to the video frame to be compared is less than the second distance threshold, determine that the video frame to be compared indicates that the lens to be detected has not been displaced; otherwise, determine that the video frame to be compared indicates that the lens to be detected has been displaced; if the number of video frames to be compared indicating that the lens to be detected has not been displaced is less than the number of video frames to be compared indicating that the lens to be detected has been displaced, determine that the lens to be detected has not been displaced; otherwise, determine that the lens to be detected has been displaced.
[0198] Optionally, the device also includes: a second prompt module, which is used to execute, before the distance acquisition module 504 executes to obtain the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, for each video frame to be compared, if the second ratio corresponding to the video frame to be compared is not less than the second ratio threshold, determine that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; obtain the number of key feature points to be detected that have corresponding matching key feature points in the video frame to be compared, as the second matching number corresponding to the video frame to be compared; when the second matching number corresponding to any video frame to be compared is less than the second matching point number threshold, output prompt information; the distance acquisition module 504 is specifically used to: when the second matching number corresponding to each video frame to be compared is not less than the second matching point number threshold, for each video frame to be compared, obtain the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point in the video frame to be compared, as the reference distance corresponding to the key feature point to be detected in the video frame to be compared.
[0199] Optionally, the video frame to be detected is obtained based on the current video frame captured by the lens to be detected; the device also includes a sampling module, which is used to obtain multiple video frames to be compared based on the following method: when the field of view of the lens to be detected is fixed, multiple video frames are sampled from all video frames currently captured by the lens to be detected; based on the sampled video frames, multiple video frames to be compared are obtained; or, starting from the first video frame captured by the lens to be detected, multiple consecutive video frames are sampled; based on the sampled video frames, multiple video frames to be compared are obtained; and / or, when the field of view of the lens to be detected is variable, the multiple video frames to be compared are obtained based on the following method: determining the sampling interval according to the frame rate of the video captured by the lens to be detected; determining the starting video frame with the sampling interval of video frames between the current video frame and the video captured by the lens to be detected; sampling multiple video frames from all video frames between the starting video frame and the current video frame; and obtaining multiple video frames to be compared based on the sampled video frames.
[0200] Optionally, the video frame to be detected is obtained based on the following method: acquiring the current video frame captured by the lens to be detected; performing image enhancement on the current video frame to obtain the video frame to be detected; the sampling module is specifically used to: perform image enhancement on each sampled video frame to obtain multiple video frames to be compared.
[0201] Optionally, the device also includes: an image enhancement module, used to perform image enhancement on any video frame by at least one of the following: compressing the video frame; converting the video frame into a grayscale image; dividing the video frame and, based on a histogram equalization algorithm, performing pixel value equalization on each divided image area.
[0202] Optionally, the image enhancement module is specifically used to: for each image area obtained by division, divide the pixel points with the same pixel value among the pixel points included in the image area into a group; for each group of pixel points, determine the number of pixel points whose pixel values are not greater than the pixel value of the pixel points in the group among the pixel points included in the image area, as a first number; calculate the ratio of the first number to the total number of pixel points included in the image area to obtain the proportion corresponding to the group of pixel points; calculate the product of the proportion corresponding to the group of pixel points and a specified value to obtain the corresponding equalization value to be restricted for the group of pixel points; wherein the specified value is: the maximum distribution of pixel values representing the video frame in the shot to be detected The length of the range; for each group of pixels, calculate the product of the pixel value of the group of pixels and the contrast limit coefficient to obtain the pixel threshold of the group of pixels; when the to-be-limited equalization value of the group of pixels is greater than the pixel threshold of the group of pixels, calculate the difference between the to-be-limited equalization value of the group of pixels and the pixel threshold of the group of pixels; calculate the quotient of the difference and the total number to obtain the correction value of the group of pixels; for each pixel in the image area, determine the minimum value of the to-be-limited equalization value of the pixel and the pixel threshold of the pixel; calculate the sum of the minimum value and the correction value of each group of pixels; based on the sum, obtain the pixel value equalization processing result corresponding to the pixel.
[0203] Optionally, based on the image enhancement module, it is specifically used to: when the pixel point is not adjacent to other image areas, use the sum value corresponding to the pixel point as the pixel value equalization processing result corresponding to the pixel point; when the pixel point is adjacent to other image areas, determine the pixels in other image areas adjacent to the pixel point; interpolate the sum value corresponding to the pixel point with the sum value corresponding to the determined adjacent pixel points to obtain the pixel value equalization processing result corresponding to the pixel point.
[0204] Optionally, the device also includes: a third prompt module, which is used to obtain the number of key feature points to be detected extracted from the video frame to be detected as a reference number before the determination module 503 executes, for each key feature point to be detected in the video frame to be detected, to determine, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected; and output a prompt message when the reference number is less than the detection quantity threshold; the determination module 503 is specifically used to: when the reference number is not less than the detection quantity threshold, for each key feature point to be detected in the video frame to be detected, to determine, from among the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as the matching key feature point corresponding to the key feature point to be detected.
[0205] Based on the lens displacement detection device provided by the embodiment of the present invention, after acquiring the video frame to be detected and the video frame to be compared, the electronic device can extract each key feature point in the video frame for each acquired video frame, and calculate the descriptor of each extracted key feature point. Furthermore, for each key feature point to be detected in the video frame to be detected, the electronic device can determine the matching key feature point corresponding to the key feature point to be detected from the key feature points to be compared included in the video frame to be compared according to the feature distance between the descriptors of the key feature points, that is, determine the matching key feature point indicating the same position in the three-dimensional space as the key feature point to be detected. Then the reference distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point can indicate: the change of the pixel position indicating the same position in the three-dimensional space in different video frames collected by the lens to be detected. Obviously, when the lens to be detected is displaced, the probability of a larger reference distance is greater; when the lens to be detected is not displaced, the probability of a smaller reference distance is greater. Therefore, based on the acquired reference distances, it can be realized to determine whether the lens to be detected is displaced.
[0206] In addition, since the key feature points are determined based on multiple pixels, the descriptors of the key feature points are also calculated based on multiple pixels, that is, the key feature points and descriptors are obtained by integrating multiple pixels. When some pixels change greatly, the key feature points and descriptors change less, that is, the robustness of the key feature points and descriptors is higher than the robustness of the pixel values of the pixels. Therefore, lens shift detection based on more robust key feature points and descriptors is more robust than lens shift detection directly based on changes in pixel values.
[0207] In addition, by obtaining the detection results based on multiple reference distances, the interference on the detection results caused by excessive changes in a few key feature points can be reduced, and the robustness of lens shift detection can be further improved.
[0208] The embodiment of the present invention further provides an electronic device, such as Figure 6 As shown, the system includes a processor 601, a communication interface 602, a memory 603 and a communication bus 604, wherein the processor 601, the communication interface 602 and the memory 603 communicate with each other via the communication bus 604. The memory 603 is used to store computer programs; the processor 601 is used to implement the steps of the lens displacement detection method described in any of the above embodiments when executing the program stored in the memory 603.
[0209] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0210] The communication interface is used for communication between the above electronic device and other devices.
[0211] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0212] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0213] In another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned lens displacement detection methods are implemented.
[0214] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes any lens displacement detection method in the above embodiments.
[0215] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0216] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0217] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device, computer-readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0218] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A lens displacement detection method, characterized in that: The method comprises: Obtaining a video frame to be detected and a video frame to be compared; wherein the video frame to be detected and the video frame to be compared are obtained based on video frames in a video captured by a lens to be detected; For each acquired video frame, extract each key feature point in the video frame; For each key feature point to be detected in the video frame to be detected, determine, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected; Obtaining a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance; Based on the acquired reference distances, it is determined whether the lens to be detected is displaced.
2. The method according to claim 1, characterized in that For each key feature point to be detected in the video frame to be detected, determining, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected as a matching key feature point corresponding to the key feature point to be detected, including: For each key feature point to be detected in the video frame to be detected, determine the key feature point to be compared that belongs to the same hash bucket as the key feature point to be detected from the key feature points to be compared included in the video frame to be compared, as the candidate key feature point corresponding to the key feature point to be detected; wherein the hash bucket to which a key feature point belongs is determined by mapping the descriptor of the key feature point using a local sensitive hashing algorithm; Calculate the feature distance between the descriptor of the key feature point to be detected and the descriptor of each corresponding candidate key feature point; Based on the minimum value of each feature distance, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected is determined from the candidate key feature points corresponding to the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected.
3. The method according to claim 2, characterized in that The step of determining, based on the minimum value among the feature distances, from the candidate key feature points corresponding to the key feature point to be detected, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected, includes: In the case where there is only one video frame to be compared, the ratio of the minimum value to the second minimum value in the feature distance corresponding to the key feature point to be detected is calculated as a first ratio; if the first ratio is less than a first ratio threshold, the candidate key feature point with the smallest feature distance between the descriptor and the descriptor of the key feature point to be detected is used as the matching key feature point corresponding to the key feature point to be detected; The step of determining whether the lens to be detected is displaced based on the acquired reference distances includes: Calculate the average level of each reference distance obtained; If the average level is less than the first distance threshold, it is determined that the lens to be detected has not been displaced; otherwise, it is determined that the lens to be detected has been displaced.
4. The method according to claim 3, characterized in that Before obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, the method further includes: If the first ratio is not less than the first ratio threshold, it is determined that there is no matching key feature point corresponding to the key feature point to be detected; Obtaining the number of key feature points to be detected that have corresponding matching key feature points as a first matching number; When the first matching number is less than a first matching point quantity threshold, outputting prompt information; The step of obtaining a characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance includes: When the first matching number is not less than the first matching point quantity threshold, a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point is obtained as a reference distance.
5. The method according to claim 2, characterized in that: The step of determining, based on the minimum value among the feature distances, from the candidate key feature points corresponding to the key feature point to be detected, a candidate key feature point whose descriptor matches the descriptor of the key feature point to be detected as the matching key feature point corresponding to the key feature point to be detected, includes: In the case where there are multiple video frames to be compared, for each video frame to be compared, determining a candidate key feature point corresponding to the key feature point to be detected in the video frame to be compared as the key feature point to be screened; Calculate the ratio of the minimum value and the second minimum value in the feature distances corresponding to the key feature points to be screened, as the second ratio corresponding to the video frame to be compared; If the second ratio corresponding to the video frame to be compared is less than the second ratio threshold, the key feature point to be screened whose descriptor has the smallest feature distance with the descriptor of the key feature point to be detected is used as the matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; The step of determining whether the lens to be detected is displaced based on the acquired reference distances includes: For each video frame to be compared, the average level of the reference distances corresponding to the matching key feature points in the video frame to be compared is calculated as the average level corresponding to the video frame to be compared; If the average level corresponding to the video frame to be compared is less than the second distance threshold, it is determined that the video frame to be compared indicates that the lens to be detected has not been displaced; otherwise, it is determined that the video frame to be compared indicates that the lens to be detected has been displaced; If the number of compared video frames indicating that the lens to be detected has not been displaced is less than the number of compared video frames indicating that the lens to be detected has been displaced, it is determined that the lens to be detected has not been displaced; otherwise, it is determined that the lens to be detected has been displaced.
6. The method according to claim 5, characterized in that Before obtaining the characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as the reference distance, the method further includes: For each video frame to be compared, if the second ratio corresponding to the video frame to be compared is not less than the second ratio threshold, it is determined that there is no matching key feature point corresponding to the key feature point to be detected in the video frame to be compared; Obtaining the number of key feature points to be detected having corresponding matching key feature points in the video frame to be compared as the second matching number corresponding to the video frame to be compared; When the second matching number corresponding to any video frame to be compared is less than the second matching point number threshold, outputting prompt information; The step of obtaining a characteristic distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance includes: When the number of second matches corresponding to each video frame to be compared is not less than the second matching point number threshold, for each video frame to be compared, the feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point in the video frame to be compared is obtained as the reference distance of the key feature point to be detected in the video frame to be compared.
7. The method according to claim 5 or 6, characterized in that: The video frame to be detected is obtained based on the current video frame collected by the lens to be detected; When the field of view of the lens to be detected is fixed, the multiple video frames to be compared are obtained based on the following method: Sampling multiple video frames from all video frames currently captured by the lens to be detected; obtaining multiple video frames to be compared based on the sampled video frames; or, starting from the first video frame captured by the lens to be detected, sampling multiple consecutive video frames; obtaining multiple video frames to be compared based on the sampled video frames; and / or, In the case where the field of view of the lens to be detected is variable, the multiple video frames to be compared are obtained based on the following method: Determine a sampling interval according to the frame rate of the video captured by the to-be-detected shot; Determine, from the video captured by the to-be-detected shot, a starting video frame that is separated from the current video frame by the sampling interval of video frames; Sampling a plurality of video frames from all video frames between the starting video frame and the current video frame; Based on the sampled video frames, multiple video frames to be compared are obtained.
8. The method according to claim 7, characterized in that The video frame to be detected is obtained based on the following method: Obtaining a current video frame captured by the shot to be detected; Performing image enhancement on the current video frame to obtain a video frame to be detected; The method of obtaining a plurality of video frames to be compared based on the sampled video frames includes: Image enhancement is performed on each sampled video frame to obtain multiple video frames to be compared.
9. The method according to claim 8, characterized in that For any video frame, perform image enhancement on the video frame by at least one of the following: Performing image compression on the video frame; Convert the video frame into a grayscale image; The video frame is divided, and based on the histogram equalization algorithm, the pixel values of each divided image area are equalized.
10. The method according to claim 9, characterized in that The histogram equalization algorithm is based on which pixel values of each divided image area are equalized, including: For each image region obtained by division, pixel points with the same pixel value among the pixel points included in the image region are divided into a group; For each group of pixels, determine the number of pixels whose pixel values are not greater than the pixel values of the group of pixels in the image area as the first number; Calculate the ratio of the first number to the total number of pixels in the image area to obtain a proportion corresponding to the group of pixels; Calculate the product of the proportion corresponding to the group of pixels and the specified value to obtain the equalization value to be restricted corresponding to the group of pixels; wherein the specified value is: the length of the maximum distribution range of the pixel values representing the video frame in the shot to be detected; For each group of pixels, the product of the pixel value of the group of pixels and the contrast limit coefficient is calculated to obtain the pixel threshold of the group of pixels; When the to-be-limited equalization value of the group of pixels is greater than the pixel threshold of the group of pixels, calculating the difference between the to-be-limited equalization value of the group of pixels and the pixel threshold of the group of pixels; Calculate the quotient of the difference and the total number to obtain the correction value of the group of pixels; For each pixel in the image area, determining a minimum value between the to-be-limited equalization value of the pixel and the pixel threshold of the pixel; Calculate the sum of the minimum value and the correction value of each group of pixels; Based on the sum value, a pixel value equalization processing result corresponding to the pixel point is obtained.
11. The method according to claim 10, characterized in that Based on the sum value, a pixel value equalization processing result corresponding to the pixel point is obtained, including: When the pixel point is not adjacent to other image regions, the sum value corresponding to the pixel point is used as the pixel value equalization processing result corresponding to the pixel point; When the pixel point is adjacent to other image areas, the pixel points in other image areas adjacent to the pixel point are determined; the sum value corresponding to the pixel point and the sum value corresponding to the determined adjacent pixel points are interpolated to obtain a pixel value equalization processing result corresponding to the pixel point.
12. The method according to claim 1, characterized in that Before determining, for each key feature point to be detected in the video frame to be detected, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected as a matching key feature point corresponding to the key feature point to be detected, the method further comprises: Obtaining the number of key feature points to be detected extracted from the video frame to be detected as a reference number; When the reference number is less than the detection quantity threshold, outputting a prompt message; For each key feature point to be detected in the video frame to be detected, determining, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected as a matching key feature point corresponding to the key feature point to be detected, including: When the reference number is not less than the detection quantity threshold, for each key feature point to be detected in the video frame to be detected, the key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected is determined from the key feature points to be compared included in the video frame to be compared, as the matching key feature point corresponding to the key feature point to be detected.
13. A lens displacement detection device, characterized in that: The device comprises: The video frame acquisition module is used to acquire the video frame to be detected and the video frame to be compared; wherein the video frame to be detected and the video frame to be compared are obtained based on the video frame in the video captured by the lens to be detected; An extraction module is used to extract key feature points in each acquired video frame; A determination module is used to determine, for each key feature point to be detected in the video frame to be detected, from the key feature points to be compared included in the video frame to be compared, a key feature point to be compared whose descriptor matches the descriptor of the key feature point to be detected, as a matching key feature point corresponding to the key feature point to be detected; A distance acquisition module is used to acquire a feature distance between the descriptor of the key feature point to be detected and the descriptor of the corresponding matching key feature point as a reference distance; The detection module is used to determine whether the lens to be detected is displaced based on the acquired reference distances.
14. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-12 when executing a program stored in a memory.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
16. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method according to any one of claims 1 to 12.