Pole tower distance measuring method and sensor structure based on visual identification and automatic focusing
By combining a visual recognition and autofocus-based approach with an image dataset-trained model and a stepper motor drive mechanism, the problem of accurately determining the optimal image distance point in existing machine vision ranging technologies has been solved. This has enabled high-precision and fast object recognition and ranging, improving the system's intelligence and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-04-10
AI Technical Summary
Existing machine vision ranging technology cannot meet the high-precision and intelligent ranging requirements in the power and industrial fields. In particular, it is difficult to achieve fast and accurate frame drawing and image ranging of specific objects in complex scenes, and existing cameras cannot accurately determine the optimal image distance point.
By constructing a method based on visual recognition and autofocus, a weight model is trained using an image dataset to recognize objects and label recognition boxes. The camera image distance is adjusted by combining a stepper motor-driven screw walking mechanism. A sharpness evaluation operator is used to evaluate the optimal image distance point, and the object distance is calculated by combining the image distance measurement with a grating ruler.
It enables accurate identification and ranging of specific objects, improves image quality and measurement accuracy, reduces manual intervention, increases work efficiency and system reliability, and reduces measurement errors.
Smart Images

Figure CN121829464A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a tower ranging method and sensor structure based on visual recognition and automatic focusing, belonging to the field of machine vision ranging technology. Background Technology
[0002] With the development of technology, natural language processing has given rise to artificial intelligence, which has a wide range of applications. Machine vision, as an important branch, has already achieved object recognition. Among them, machine vision ranging is a key direction with huge application potential. It can mimic the human eye and use cameras to measure object distances, achieving the effect of "what you see is what you measure," which is of great significance to the development of multiple fields.
[0003] In the power sector, current drone ranging primarily relies on electronic maps to mark hovering positions, achieving distance measurement at a constant speed end-to-end. While this method solves the problem of spatial distance measurement, it faces numerous technical limitations in practical applications. For example, its application scenarios are relatively limited, and the process is not convenient enough, failing to meet the increasingly diverse and precise ranging needs of the power industry. In the industrial sector, although some ranging and robot operation technologies exist, significant shortcomings remain. For instance, existing robot operations often require additional guidance facilities. For example, AVG robot handling operations typically require auxiliary means such as tracks, marking lines, or pre-set positions. This increases equipment costs and operational complexity, limiting the robot's flexibility and autonomy, and failing to meet the demands of efficient, flexible handling and intelligent operation in industrial production.
[0004] Currently, ordinary cameras have functional limitations in distance measurement. The image distance of ordinary cameras is usually fixed or cannot be precisely measured; their structure is relatively simple and their level of intelligence is low. While autofocus cameras can achieve high overall image clarity, they lack a dedicated evaluation system for a specific object being measured within the overall image. In other words, although the overall image clarity is high, the clarity of a specific object within the image may not be exactly the highest. This means that during focusing, the image distance of the object being measured may not be at the optimal image distance point, making it impossible to accurately obtain the distance information between the object and the camera, thus failing to meet the requirements of high-precision distance measurement. Therefore, it is necessary to use technical means to extract the specific object being measured from the overall image, or to use technical means to block out objects other than the specific object being measured.
[0005] In machine vision applications, images captured by cameras contain objects at various distances. Without a fixed reference point to determine the object being measured, existing evaluation system algorithms struggle to accurately determine the optimal image distance. Even technologies capable of image recognition fall short in combining object recognition with precise ranging, failing to achieve rapid and accurate bounding of specific objects and accurate ranging based on the image within the bounding frame. This limits the effectiveness of machine vision ranging technology in complex scenarios. In summary, existing ranging technologies and related equipment have numerous limitations in applications in the power and industrial sectors, failing to meet the growing demand for high-precision, intelligent ranging. Therefore, developing a novel machine vision ranging evaluation algorithm and sensor is of significant practical importance. Summary of the Invention
[0006] The purpose of this invention is to provide a tower ranging method and sensor structure based on visual recognition and autofocus. By combining the marking of the recognition box with the real-time image distance change of the camera, and the image sharpness evaluation algorithm, the image distance point with the best imaging effect can be accurately found, which significantly improves the image quality and measurement accuracy, and effectively meets the requirements of high precision and fast response ranging.
[0007] To achieve the above objectives, the present invention employs the following technical solution: A tower ranging method based on visual recognition and automatic focusing includes the following steps: The weight model file of the object under test is obtained by training the image dataset, and the file is then transmitted to the sensor main control board. The weight model file is used to perform real-time object recognition on images captured by the camera, and recognition boxes are marked in the images; The stepper motor drives the screw walking mechanism to move on the fixed screw, thereby adjusting the image distance of the camera. A sharpness evaluation operator is used to evaluate the sharpness of the image area within the recognition frame to determine the optimal image distance point. Based on the optimal image distance point, the image distance is measured using a grating ruler, and the tower object distance is calculated according to the lens imaging principle.
[0008] Preferably, the sharpness evaluation operator includes the Tenengrad operator, the Brenner operator, or the SMD operator.
[0009] Preferably, the weight model file is a single object recognition model trained based on the YOLO model, which can recognize multiple similar objects within the field of view; When the YOLO model identifies a single target object and marks the bounding box, the sharpness evaluation operator only evaluates the sharpness of the image within that specific bounding box.
[0010] Images outside the bounding box are not evaluated; in this invention, besides labeling objects, another significant use of the bounding box is to filter out sharpness evaluation of images outside the bounding box; when the YOLO model identifies multiple objects and labels them with bounding boxes, the sharpness evaluation operator will evaluate the sharpness of each image within the bounding box according to the identified image sequence, and images outside the bounding box are also not evaluated; in this invention, besides labeling objects, another significant use of the bounding box is to filter out sharpness evaluation of images outside the bounding box.
[0011] Preferably, the method for calculating the distance between the tower and the object is as follows: , in, The focal length of the lens. For object distance, The distance is the image distance.
[0012] Preferably, it also includes a step of sensing the relative positions of multiple objects being measured: Identify at least two objects and calculate the object distance between the camera and each object. Based on the lines connecting the camera and each object in the image, the angles between the lines are calculated using the built-in angular coordinate system. The relative distances between the measured objects are calculated using the triangle relationship formula.
[0013] A sensor structure for visual ranging includes: Sensor housing; An image acquisition module includes a lens and an image sensor disposed within the sensor housing; A focusing drive module includes a fixed screw, a fixed smooth slide bar arranged parallel to the fixed screw, and a walking mechanism assembly slidably mounted on the fixed screw and the fixed smooth slide bar; The walking mechanism assembly includes a fixed screw walking mechanism and a fixed smooth slide bar walking mechanism. The image acquisition module is mounted on the walking mechanism assembly and moves with it to achieve image distance adjustment. The fixed smooth slide rod walking mechanism is slidably sleeved on the fixed smooth slide rod, providing auxiliary support and sliding guidance for the walking mechanism assembly; The fixed screw traveling mechanism is equipped with a stepper motor, which drives the transmission gear through the motor output shaft. The transmission gear is connected to a screw transmission threaded ring, which is sleeved on and meshes with the fixed screw. The measurement module includes a grating ruler detector mounted on the walking mechanism assembly and a grating ruler mounted on the sensor housing, used to detect the displacement of the image acquisition module in real time to determine the image distance.
[0014] Preferably, the fixed screw traveling mechanism further includes a limiting rod and a sliding ball. The screw drive threaded ring is clamped and fixed in position by the limiting rod, and the relative rotation and axial sliding between the screw and the fixed screw are realized by the sliding ball.
[0015] Preferably, it also includes a sensor main control board, which is communicatively connected to the image sensor, the stepper motor, and the grating ruler detector; the sensor main control board has a built-in or loaded trained object recognition weight model for performing specific object recognition and annotation on the images acquired by the image sensor.
[0016] Preferably, the stepper motor is connected to the sensor's built-in battery via a micro stepper motor power cable.
[0017] Preferably, the sensor main control board is also connected to a sensor switch, which is fixed to the sensor housing.
[0018] The advantages of this invention are: By constructing a data weighting model and using it for real-time object recognition and bounding box annotation, the accuracy and real-time performance of visual ranging are ensured. A fixed screw walking mechanism and a stepper motor drive system are introduced to achieve automatic selection of the optimal image distance point for the object. Through the coordination of bounding box annotation and real-time image distance changes from the camera, combined with an image sharpness evaluation algorithm, the optimal image distance point for imaging effect can be accurately found, significantly improving image quality and measurement accuracy.
[0019] It can not only accurately identify and measure the distance of a single object, but also handle the problem of relative position sensing between multiple objects. By constructing a relative position sensing model between multiple objects and combining it with the precise image distance data measured by the grating ruler, the relative distance between the objects can be accurately calculated.
[0020] The automated and intelligent design of this solution greatly reduces manual intervention during operation. Users can complete complex ranging and identification tasks with just simple settings. This not only improves work efficiency but also reduces measurement errors caused by human factors, enhancing the overall reliability and stability of the system. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0022] Figure 1 This is a schematic diagram of the method flow of the present invention.
[0023] Figure 2 This is a schematic diagram of the device structure of the present invention.
[0024] Figure 3 This is a detailed schematic diagram of the fixed screw traveling mechanism of the present invention.
[0025] In the diagram, 1 is the fixed screw traveling mechanism, 2 is the sensor casing, 3 is the image sensor, 4 is the fixed smooth slide rod traveling mechanism, 5 is the fixed smooth slide rod, 6 is the lens, 7 is the fixed screw, 8 is the micro stepper motor power supply, 9 is the stepper motor, 10 is the motor output shaft, 11 is the transmission gear, 12 is the screw transmission threaded ring, 13 is the sliding ball, 14 is the limit rod, 15 is the grating ruler, 16 is the grating ruler detector, 17 is the sensor switch, 18 is the sensor built-in battery, and 19 is the sensor main control board. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] This invention achieves its purpose by encapsulating a sensor through the cooperation of hardware and a ranging method. While the sensor's appearance is not significantly different from a camera module, the visual ranging sensor designed in this solution differs from ordinary cameras in that the image distance of the camera in this solution is known and variable. It has a more complex and intelligent structure than ordinary cameras. The visual evaluation system encapsulated within the sensor can identify and evaluate the edge sharpness of the image on the photosensitive area of the moving image sensor, and can identify the image with the highest edge sharpness. This image is considered the optimal image distance point for the object, and there will not be a second such point of highest image sharpness on the lens imaging side. The main difference from an autofocus camera is that the evaluation system of this invention can lock onto the object to be measured. While an autofocus camera achieves the highest overall sharpness, the image distance of the object being measured in the image may not be the optimal image distance point because the autofocus camera does not have a dedicated evaluation system for locking onto the object. The grating ruler encapsulated in the sensor can also accurately output the position of this optimal image distance point from the focal length via a standard USB interface, i.e., the optimal image distance (hereinafter referred to as image distance). With the measured image distance and the known focal length, the object distance (the distance from the object to the focal length) can be calculated using the image distance imaging principle, thus realizing machine vision ranging.
[0028] Example 1 like Figure 1 As shown, a tower ranging method based on visual recognition and automatic focusing includes the following steps: S1: Train the weight model file of the tower under test using the image dataset, and then transfer the file to the sensor main control board; S2: Use the weight model file to perform real-time object recognition on the images captured by the camera, and mark the recognition boxes in the images; S3: The stepper motor drives the screw walking mechanism to move on the fixed screw, thereby adjusting the image distance of the camera. The sharpness evaluation operator is used to evaluate the sharpness of the image area within the recognition box to determine the optimal image distance point. S4: Based on the optimal image distance point, measure the image distance using a grating ruler, and calculate the tower object distance according to the lens imaging principle.
[0029] As a refinement of the above embodiments, the image dataset of the tower under test is used to learn and extract the object features of the measured object to form a data weight model .PT file. This .PT file is then transmitted to the sensor main control board via an API interface, an RJ45 network cable, or a storage device.
[0030] As a refinement of the above embodiments, the sharpness evaluation operator includes the Tenengrad operator, the Brenner operator, or the SMD operator.
[0031] As a refinement of the above embodiments, the weight model file is a single object recognition model trained based on the YOLO model, which can recognize multiple similar objects within the field of view; When the YOLO model identifies a single target object and marks the bounding box, the sharpness evaluation operator only evaluates the sharpness of the image within that specific bounding box.
[0032] Images outside the bounding box are not evaluated; in this invention, besides labeling objects, another significant use of the bounding box is to filter out sharpness evaluation of images outside the bounding box; when the YOLO model identifies multiple objects and labels them with bounding boxes, the sharpness evaluation operator will evaluate the sharpness of each image within the bounding box according to the identified image sequence, and images outside the bounding box are also not evaluated; in this invention, besides labeling objects, another significant use of the bounding box is to filter out sharpness evaluation of images outside the bounding box.
[0033] As a refinement of the above embodiments, the method for calculating the distance between poles and objects is as follows: , in, The focal length of the lens. For object distance, The distance is the image distance.
[0034] The difference between this invention and ordinary cameras lies in the fact that the number of lenses in a camera can be fixed or a variable high focal length can be achieved by stacking lenses. Furthermore, the distance between the photosensitive area of the image sensor and the lens focal point is variable, and the image distance can be precisely measured. By controlling a single lens or multiple stacked lenses to move forward or backward on a screw via a micro stepper motor, the image distance of the image sensor behind the lens is changed.
[0035] When multiple lenses are moved, that is, when multiple lenses are added together, the focal length is calculated first: Suppose there are two thin lenses with focal lengths of respectively and The distance between them is The equivalent focal length of the combined lens is... It is given by the following formula: , Calculate the focal length Then, the object distance can be calculated using the lens imaging principle. Thus, distance measurement using machine vision has been achieved. The object distance measured by this method has high accuracy and is simple and easy to implement.
[0036] As a refinement of the above embodiments, a step of sensing the relative positions of multiple objects under test is also included: Identify at least two objects and add sequence labels to the identified objects. Calculate the object distance between the camera and the objects with different sequence labels. Based on the lines connecting the camera and each object in the image, the angles between the lines are calculated using the built-in angular coordinate system. The relative distances between the measured objects are calculated using the triangle relationship formula.
[0037] Example 2 like Figures 2-3 As shown, a sensor structure based on visual recognition and autofocus includes: Sensor housing 2; The image acquisition module includes a lens 6 and an image sensor 3 disposed inside the sensor housing 2; The focusing drive module includes a fixed screw 7, a fixed smooth slide bar 5 arranged parallel to the fixed screw 7, and a walking mechanism assembly slidably mounted on the fixed screw 7 and the fixed smooth slide bar 5; The walking mechanism assembly includes a fixed screw walking mechanism 1 and a fixed smooth slide bar walking mechanism 4. The image acquisition module is mounted on the walking mechanism assembly and moves with it to achieve image distance adjustment. The fixed smooth slide rod walking mechanism 4 is slidably sleeved on the fixed smooth slide rod 5, providing auxiliary support and sliding guidance for the walking mechanism assembly; The fixed screw traveling mechanism 1 is equipped with a stepper motor 9. The stepper motor 9 drives the transmission gear 11 through the motor output shaft 10. The transmission gear 11 is connected to a screw transmission threaded ring 12. The screw transmission threaded ring 12 is sleeved and meshed on the fixed screw 7. The measurement module includes a grating ruler detector 16 disposed on the walking mechanism assembly and a grating ruler 15 disposed on the sensor housing 2, for real-time detection of the displacement of the image acquisition module to determine the image distance.
[0038] As a refinement of the above embodiment, the fixed screw traveling mechanism 1 further includes a limiting rod 14 and a sliding ball 13. The screw drive threaded ring 12 is fixed in position by the limiting rod 14, and the relative rotation and axial sliding between it and the fixed screw 7 are realized by the sliding ball 13.
[0039] As a refinement of the above embodiments, a sensor main control board 19 is also included. The sensor main control board 19 is communicatively connected to the image sensor 3, the stepper motor 9, and the grating ruler detector 16. The sensor main control board 19 has a built-in or loaded trained object recognition weight model for performing specific object recognition and annotation on the images acquired by the image sensor 3. As a refinement of the above embodiment, the stepper motor 9 is connected to the sensor's built-in battery 18 via a micro stepper motor power line 8.
[0040] As a refinement of the above embodiment, the sensor main control board 19 is also connected to a sensor switch 17, which is fixed to the sensor housing 2.
[0041] Specifically, the implementation method is as follows: Point the sensor at the tower being measured, turn on the sensor switch 17, and the sensor main control board 19 uses the built-in software environment and .PT weight files to perform real-time recognition of learned objects. It then draws a bounding box around the recognized objects in the image. The recognition frame rate can be ensured to be above 30fps, ensuring the accuracy and real-time performance of the sensor's visual ranging. When this invention is executed, the .PT weight file only learns the weights of a single object, such as a tower. Only the tower's .PT file needs to be transferred to the sensor main control board 19. Although it learns the weights of a single object, it can recognize all towers within the camera's field of view.
[0042] The fixed screw traveling mechanism 1 can move back and forth on the fixed screw 7 inside the sensor. Specifically, the stepper motor 9 inside the fixed screw traveling mechanism 1 rotates at a suitable step distance, driving the screw drive threaded ring 12 to move on the fixed screw 7 via the transmission gear 11. More specifically, the screw drive threaded ring 12 is clamped and fixed in position by the limiting rod 14, and its fixed position is slidably rotated by the sliding ball 13. Since the fixed screw 7 cannot rotate, the torque of the transmission gear 11 is converted into the sliding rotation force of the fixed screw traveling mechanism 1, allowing it to move on the fixed screw 7. The result of the fixed screw traveling mechanism 1 moving back and forth is the selection of the optimal image distance point for the object with the recognition box drawn above.
[0043] The sensor main control board 19 can use the onboard YOLO environment or other models and the transmitted .PT file to mark the object with a bounding box. All subsequent operations will expand the object within the bounding box. The combination of the bounding box marking and the real-time image distance change of the camera is crucial for selecting the optimal image distance point. Without a bounding box, a zoom camera alone cannot find the optimal image distance point because the camera's image includes all objects and scenes, making it impossible to distinguish which object or scene has the optimal image distance point. However, with a bounding box, all objects and scenes outside the bounding box are effectively masked. The evaluation system within the sensor main control board 19 uses the Tenengrad operator (which calculates the sum of squares of gradient magnitudes, with larger values indicating clearer images), or the Brenner operator (which calculates the sum of squares of grayscale differences between adjacent pixels, with larger values indicating clearer images), or SMD (Sum of Modified Differences) (the sum of absolute values of grayscale differences between adjacent pixels, with larger values indicating clearer images). The image distance point with the best imaging effect found is the optimal image distance point of this invention.
[0044] Calculate the object distance by using the grating ruler 15 to measure the precise image distance, and then directly determine the object distance based on the lens imaging principle.
[0045] Sensing the relative position of multiple objects is relatively complex. For example, if two or more towers are identified, the straight-line distance between the camera and the identified object can be calculated using the method described above. The relative position sensing model between multiple objects in the sensor main control board 19 connects the objects and the camera with lines, then compares the angles between the lines with the built-in angular coordinate system. Finally, the relative distance between the objects is calculated using the polygon relationship formula.
[0046] For example, when two towers A and B are identified, the distance from the camera to A and the distance from the camera to B are calculated using the method described above. The camera and A are connected in the displayed image, and the camera and B are connected in the displayed image. Then, the angle between the two connecting lines is calculated using the built-in angular coordinate system. In the triangle, the distance of the third side is calculated using the two adjacent sides and the included angle. This distance is the relative distance between towers A and B.
[0047] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A tower ranging method based on visual recognition and automatic focusing, characterized in that, Includes the following steps: The weight model file of the tower under test is obtained by training the image dataset, and the file is transmitted to the sensor main control board. The weight model file is used to perform real-time object recognition on images captured by the camera, and recognition boxes are marked in the images; The stepper motor drives the screw walking mechanism to move on the fixed screw, thereby adjusting the image distance of the camera. A sharpness evaluation operator is used to evaluate the sharpness of the image area within the recognition frame to determine the optimal image distance point. Based on the optimal image distance point, the image distance is measured using a grating ruler, and the tower object distance is calculated according to the lens imaging principle.
2. The tower ranging method based on visual recognition and automatic focusing according to claim 1, characterized in that, The sharpness evaluation operators include the Tenengrad operator, the Brenner operator, or the SMD operator.
3. The tower ranging method based on visual recognition and automatic focusing according to claim 1, characterized in that, The weight model file is a single object recognition model trained based on the YOLO model, which can recognize multiple similar objects within the field of view, and the object is a pole or tower. When the YOLO model identifies a single target object and marks the bounding box, the sharpness evaluation operator only evaluates the sharpness of the image within that specific bounding box.
4. The tower ranging method based on visual recognition and automatic focusing according to claim 1, characterized in that, The method for calculating the distance between the tower and the object is as follows: , in, The focal length of the lens. For object distance, The distance is the image distance.
5. The tower ranging method based on visual recognition and automatic focusing according to claim 1, characterized in that, It also includes a step of sensing the relative positions of multiple objects being measured: Identify at least two objects and calculate the object distance between the camera and each object. Based on the lines connecting the camera and each object in the image, the angles between the lines are calculated using the built-in angular coordinate system. The relative distances between the measured objects are calculated using the triangle relationship formula.
6. A sensor structure based on visual recognition and autofocus, characterized in that, For implementing the method as described in any one of claims 1-5, comprising: Sensor housing (2); The image acquisition module includes a lens (6) and an image sensor (3) disposed inside the sensor housing (2); The focusing drive module includes a fixed screw (7), a fixed smooth slide bar (5) arranged parallel to the fixed screw (7), and a walking mechanism assembly slidably mounted on the fixed screw (7) and the fixed smooth slide bar (5); The walking mechanism assembly includes a fixed screw walking mechanism (1) and a fixed smooth slide rod walking mechanism (4). The image acquisition module is mounted on the walking mechanism assembly and moves with it to achieve image distance adjustment. The fixed smooth slide rod walking mechanism (4) is slidably sleeved on the fixed smooth slide rod (5) to provide auxiliary support and sliding guidance for the walking mechanism assembly; The fixed screw traveling mechanism (1) is equipped with a stepper motor (9). The stepper motor (9) drives a transmission gear (11) through a motor output shaft (10). The transmission gear (11) is connected to a screw transmission threaded ring (12). The screw transmission threaded ring (12) is sleeved and meshed on the fixed screw (7). The measurement module includes a grating ruler detector (16) disposed on the walking mechanism assembly and a grating ruler (15) disposed on the sensor housing (2) for real-time detection of the displacement of the image acquisition module to determine the image distance.
7. The sensor structure based on visual recognition and autofocus according to claim 5, characterized in that, The fixed screw traveling mechanism (1) also includes a limiting rod (14) and a sliding ball (13). The screw drive threaded ring (12) is fixed in position by the limiting rod (14) and achieves relative rotation and axial sliding with the fixed screw (7) by the sliding ball (13).
8. The sensor structure based on visual recognition and autofocus according to claim 5, characterized in that, It also includes a sensor main control board (19), which is communicatively connected to the image sensor (3), the stepper motor (9) and the grating ruler detector (16); the sensor main control board (19) has a built-in or loaded trained object recognition weight model, which is used to perform specific object recognition and annotation on the images collected by the image sensor (3).
9. The sensor structure based on visual recognition and autofocus according to claim 5, characterized in that, The stepper motor (9) is connected to the sensor’s built-in battery (18) via a micro stepper motor power cable (8).
10. The sensor structure based on visual recognition and autofocus according to claim 6, characterized in that, The sensor main control board (19) is also connected to a sensor switch (17), which is fixed to the sensor housing (2).