Robot snapshot method and system based on template matching and holder fine tuning

Through template matching and gimbal fine-tuning, the problem of low recognition accuracy caused by gimbal error during robot capture is solved, and image quality and recognition accuracy are improved.

CN120583206APending Publication Date: 2025-09-02JIANGSU GAOXINXING ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510731572.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The robot has a gimbal error during the capture process, resulting in low recognition accuracy, especially in scenarios such as meter recognition and smoke sensor detection.

Method used

Through template matching and gimbal fine-tuning, the robot matches the capture image with the template image when capturing, calculates the offset distance and adjusts the gimbal position parameters to ensure that the image quality meets the requirements.

Benefits of technology

It improves the reliability of the snapshot image and the accuracy of subsequent image recognition, reduces errors, and ensures the accuracy of algorithm recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583206A_ABST
    Figure CN120583206A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robots, in particular to a robot snapshot method and system based on template matching and holder fine adjustment. The robot snapshot method comprises the following steps: in response to a received inspection task, carrying out snapshot when the robot travels to a snapshot position according to the inspection task to obtain a snapshot image; the inspection task comprises a template drawing, a snapshot position and a holder preset position; template matching is conducted on the snapshot image and a template image in the inspection task, in response to successful matching, the offset distance between the snapshot image and the center point of a target sub-image is calculated, and the target sub-image represents the sub-image, most similar to the template image, in all the sub-images of the snapshot image; and according to the offset distance and the current position parameter of the holder, adjusting the holder to perform snapshot. According to the robot snapshot method, the holder error during snapshot can be reduced, and the snapshot accuracy and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology. More specifically, the present invention relates to a robot capture method and system based on template matching and pan / tilt fine-tuning. Background Art

[0002] Daily inspection tasks require different work methods due to the needs of different scenarios, and are generally divided into manual inspection and robot-assisted inspection. Manual inspection involves a person carrying a camera to photograph and inspect the target object, and manually compiling the data. Robot-assisted inspection involves a robot arriving at a capture point based on its positioning, using the pan / tilt camera presets to capture data at a fixed point. Assisted positioning technology is used to label the target objects and equipment in each scene. Upon arrival, the robot uses the positioning device to directly aim the pan / tilt camera at the relevant location to capture data. This approach utilizes a single device, single camera solution, with each camera corresponding to multiple points on a single device, capturing data at fixed points and at regular intervals.

[0003] Currently, in various scene inspection systems, robots are mostly composed of a visible light camera and an infrared thermal imaging camera. During equipment inspection, a visible light camera is needed to capture the equipment. After obtaining the captured image, the equipment-related conditions are judged based on the captured image, including but not limited to equipment readings, equipment dirt or equipment damage, etc. Therefore, the captured image quality requirements are relatively high, especially for meter recognition and smoke sensor detection.

[0004] However, when the robot moves to the target position through navigation, there will be certain deviations in coordinates and angles. This capture error has a great impact on algorithm recognition with high precision requirements. For example, in the meter recognition algorithm, it is necessary to accurately read the readings in the meter, so it is necessary to ensure the orientation of the dial and the clarity of the pointer in the captured image. Once the robot has coordinate and angle deviations at the target position, the captured image may not guarantee that the algorithm can accurately recognize the results.

[0005] In this regard, how to reduce the pan-tilt error of the robot when taking pictures is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] In order to solve the above-mentioned technical problem of low recognition accuracy caused by pan / tilt error during the capture process, the present invention provides solutions in the following aspects.

[0007] In a first aspect, the present invention provides a robot snapshot method based on template matching and gimbal fine-tuning, comprising: in response to receiving an inspection task, according to the inspection task, driving to a snapshot position, calling a gimbal preset position to snapshot, and obtaining a snapshot image; the inspection task includes a template image, a snapshot position, and a gimbal preset position; performing template matching on the snapshot image and the template image in the inspection task, and in response to a successful match, calculating an offset distance between the snapshot image and a center point of a target sub-image, the target sub-image representing a sub-image of all sub-images of the snapshot image that is most similar to the template image; adjusting the gimbal according to the offset distance and the current position parameters of the gimbal to snapshot.

[0008] Furthermore, the method for obtaining the template image includes: in response to receiving a capture instruction, capturing the target object according to the capture instruction to obtain the template image; saving the template image, the capture position when capturing the template image, and the pan / tilt position parameters to the inspection task.

[0009] Furthermore, template matching is performed on the snapshot image and the template image in the inspection task, including: superimposing the template image on the snapshot image and translating the template image. During the translation process, the area in the snapshot image covered by the template image is defined as a sub-image of the snapshot image; for any sub-image, the similarity between the sub-image and the template image is calculated; in response to the maximum similarity among all similarities being greater than a preset threshold, it is determined that the match is successful; in response to the maximum similarity among all similarities being less than or equal to the preset threshold, it is determined that the match fails.

[0010] Furthermore, the calculation expression of the similarity is:

[0011]

[0012] Where R(i,j) is the similarity, M is the number of pixels in the vertical direction of the template image, N is the number of pixels in the horizontal direction of the template image, and S ij (m,n) is the sub-image of the captured image, and T(m,n) is the template image.

[0013] Furthermore, the template matching process further includes: in response to a matching failure, returning to the capture position and re-capturing.

[0014] Furthermore, the offset distance includes a first difference and a second difference, and calculating the offset distance between the center point of the captured image and the target sub-image includes: subtracting the horizontal coordinate of the center point of the template image from the horizontal coordinate of the center point of the target sub-image to obtain the first difference; subtracting the vertical coordinate of the center point of the template image from the vertical coordinate of the center point of the target sub-image to obtain the second difference.

[0015] Furthermore, the gimbal is adjusted to capture images according to the offset distance and the current position parameters of the gimbal, including: determining the horizontal field of view angle and the vertical field of view angle according to the current focal length of the camera; wherein the horizontal field of view angle is negatively correlated with the focal length, and the vertical field of view angle is negatively correlated with the focal length; determining the horizontal angle according to the horizontal field of view angle and the first difference; the horizontal angle is positively correlated with the horizontal field of view angle and negatively correlated with the first difference; determining the vertical angle according to the vertical field of view angle and the second difference; the vertical angle is positively correlated with the vertical field of view angle and negatively correlated with the second difference; based on the horizontal angle and the vertical angle, the gimbal is rotated according to the current position parameters to capture images.

[0016] Furthermore, the calculation expression of the horizontal field angle is:

[0017]

[0018] The calculation expression of the vertical field angle is:

[0019]

[0020] Where A ver is the vertical field of view, H t is the camera target surface size height, F is the current focal length value, A level is the horizontal field of view, W t is the width of the camera target surface.

[0021] Furthermore, before matching, the method further includes: performing grayscale processing on the template image and the captured image.

[0022] In a second aspect, the present invention provides a robot capture system comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a robot capture method based on template matching and gimbal fine-tuning according to any one of the first aspects is implemented.

[0023] The beneficial effects of the present invention are as follows: the present invention can avoid identifying invalid images by re-capturing images that fail to match, thereby improving the reliability of obtaining captured images and the accuracy of subsequent image recognition; by adjusting the gimbal according to the offset distance between the template image and the target sub-image and the position parameters of the gimbal, and shooting according to the adjusted gimbal, the error between the captured captured image and the template image can be reduced, thereby improving the accuracy of subsequent image recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0025] Figure 1 is a flow chart schematically illustrating a robot capture method based on template matching and pan / tilt fine-tuning according to an embodiment of the present invention;

[0026] Figure 2 is a flow chart schematically illustrating a robot capture method based on template matching and pan / tilt fine-tuning according to another embodiment of the present invention;

[0027] Figure 3 Schematically illustrates a structural block diagram of a robot capture system based on template matching and pan / tilt fine-tuning according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0029] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 The flowchart schematically illustrates a robot capture method based on template matching and gimbal fine-tuning according to an embodiment of the present invention.

[0031] Because there are certain deviations in coordinates and angles when the robot moves to the target position through navigation to take pictures, this deviation reduces the accuracy of the algorithm recognition with high precision requirements, thereby affecting the inspection effect.

[0032] In this regard, in the first aspect, the present invention provides a robot capture method based on template matching and pan-tilt fine-tuning, which is mainly used for a robot carrying a controllable pan-tilt and a visible light camera, and is applied to scenarios where the robot needs to obtain accurate snapshots at a fixed point when inspecting in a relatively complex environment. It should be noted that the pan-tilt is an installation platform composed of two AC motors or DC motors, which can move horizontally and vertically. Specifically, as Figure 1 As shown, the method of the present invention includes:

[0033] S101 , in response to receiving an inspection task, driving to a capture position according to the inspection task, calling a pan / tilt preset position to capture, and obtaining a captured image.

[0034] Specifically, during the inspection process, the robot conducts inspections according to the inspection route. Every time it arrives at an inspection point, it obtains the inspection task corresponding to the inspection point. According to the snapshot position in the inspection task, the robot drives to the corresponding snapshot position and controls the pan-tilt head to snapshot according to the pan-tilt head preset position in the inspection task (that is, the relevant parameters of the pan-tilt head when shooting the template image) to obtain a snapshot image.

[0035] In one embodiment, before taking a snapshot, the process further includes obtaining a template image. Specifically, when the robot arrives at an inspection point along the inspection route, it responds to a snapshot command by taking a snapshot of the target object, obtaining a template image, and recording the location information and pan / tilt position (pan / tilt position parameters) at that time. The template image, location information, and pan / tilt position are then saved to the inspection task. This process is repeated until template images corresponding to all inspection points along the inspection route are obtained.

[0036] It can be understood that when shooting the template image, the robot is manually controlled to move to a position where it can accurately capture the image, observe the real-time image in the camera, control the pan-tilt head to aim at the target object, execute the camera capture, and save the captured image as the template image of the inspection point in the inspection task.

[0037] S102: performing template matching on the captured image and the template image in the inspection task, and in response to a successful match, calculating an offset distance between the center point of the captured image and the target sub-image.

[0038] In one embodiment, before performing template matching, it also includes: grayscale processing of the template image and the snapshot image. Specifically, the template image and the snapshot image can be grayscaled based on the weighted average method, that is, the weighted average of RGB is taken for the pixel value of each point in the image. By grayscale the template image and the snapshot image, the interference of color differences on the matching results can be reduced, and the accuracy of template matching can be improved; at the same time, the amount of data and computational complexity are reduced, and the efficiency of template matching is improved. In an optional embodiment, those skilled in the art can select a suitable method for grayscale according to actual needs, such as the average method.

[0039] It can be understood that template matching is to find the most similar area from the captured image based on the template image. In one embodiment, template matching can use normalized correlation matching. Specifically, the template image is superimposed on the captured image and translated. During the translation process, the area in the captured image that overlaps with the template image is defined as a sub-image. Traverse / search the entire captured image to obtain multiple sub-images. Among them, the search range is 1≤i≤Wn, 1≤j≤Hm, i is the horizontal index in the traversal search process, j is the vertical index in the traversal search process, W is the width of the captured image, H is the height of the captured image, n is the width of the sub-image, and m is the height of the sub-image.

[0040] For any subgraph, calculate the similarity between the template graph and the subgraph. Specifically, the similarity calculation expression is:

[0041]

[0042] Where D(i,j) is the similarity between the template image and the sub-image, M is the number of pixels in the vertical direction of the template image, N is the number of pixels in the horizontal direction of the template image, and S ij (m,n) is the sub-image of the captured image, and T(m,n) is the template image.

[0043] Normalize and get the similarity between the subgraph and the template graph. Specifically, the similarity calculation expression is:

[0044]

[0045] Where R(i,j) is the similarity between the template image and the sub-image, M is the number of pixels in the vertical direction of the template image, N is the number of pixels in the horizontal direction of the template image, and S ij (m,n) is the sub-image of the captured image, and T(m,n) is the template image.

[0046] It can be understood that the greater the similarity, the more similar the template image is to the snapshot image. When the similarity is 1, the template image is identical to the snapshot image.

[0047] Furthermore, the similarity of all sub-graphs is calculated, and the sub-graph with the largest similarity is used as the target sub-graph. It is determined whether the similarity of the target sub-graph is greater than the preset threshold. If so, it is determined that the match is successful; if not, it indicates that the difference between the current capture and deployment is too large, then it is determined that the match failed, and the robot is adjusted, returning to step S101 to re-shoot and match.

[0048] By reshooting snapshots whose similarity is lower than a preset threshold, the interference of invalid images can be avoided, the quality of the snapshots can be improved, the misjudgment and waste of computing resources can be reduced, and the reliability of template matching can be improved.

[0049] In this embodiment, the preset threshold is set to 0.6. In an optional embodiment, it can also be set to 0.7 or other values. Those skilled in the art can set it according to actual needs.

[0050] In other optional embodiments, the average value of the similarities of all sub-images may be calculated to determine whether the average value is greater than a preset threshold. If so, the match is determined to be successful; if not, the position is adjusted and the capture is performed again.

[0051] In addition, since light will affect the accuracy of template matching to a certain extent, the preset threshold can be adjusted according to different lighting conditions to improve the reliability and accuracy of template matching. Specifically, under conditions of good light (such as daytime, sunny days), since the similarity between the snapshot image (with good image quality) and the template image is generally high, the preset threshold can be appropriately increased (for example, set to 0.65); under conditions of poor light, the similarity between the snapshot image (with poor image quality) and the template image is generally low, then the preset threshold can be lowered (for example, set to 0.55) to avoid missing the target. By adjusting the preset threshold according to the lighting conditions, the reliability and accuracy of template matching can be improved, and the situation of missing targets can be reduced.

[0052] For the successfully matched snapshot, the offset distance between the center point of the target sub-image and the center point of the template image is calculated with the upper left corner of the template image as the coordinate axis origin. In this embodiment, the offset distance includes a first difference and a second difference.

[0053] Specifically, the abscissa of the center point of the target sub-image is subtracted from the abscissa of the center point of the template image to obtain a first difference, and the ordinate of the center point of the target sub-image is subtracted from the ordinate of the center point of the template image to obtain a second difference.

[0054] S103: Adjust the pan-tilt platform according to the offset distance and the current position parameters of the pan-tilt platform to take a snapshot.

[0055] Specifically, the horizontal field of view angle and the vertical field of view angle are determined according to the current focal length of the camera; wherein the horizontal field of view angle is negatively correlated with the focal length, and the vertical field of view angle is negatively correlated with the focal length.

[0056] In one embodiment, the calculation expression of the horizontal field of view angle is:

[0057]

[0058] Where A level is the horizontal field of view, W t is the width of the camera target surface (obtained from the camera manufacturer), and F is the current focal length.

[0059] In one embodiment, the vertical field of view angle is calculated as follows:

[0060]

[0061] Where A ver is the vertical field of view, H t is the camera target surface size height (obtained from the camera manufacturer), and F is the current focal length value.

[0062] By calculating the horizontal and vertical field of view angles based on the parameters provided by different manufacturers, the reliability of obtaining the horizontal and vertical field of view angles can be improved, and the accuracy of subsequent fine-tuning of the gimbal can be improved, thereby reducing the interference of gimbal errors on snapshots.

[0063] In one embodiment, the calculation expression of the current focal length is:

[0064] F=4.6793×V 0.843 ;

[0065] Where F is the focal length and V is the physical zoom value.

[0066] The physical zoom value is obtained by scaling the current optical zoom value. In one embodiment, the calculation formula of the physical zoom value is:

[0067]

[0068] Where V is the physical zoom value, and U is the current optical zoom value.

[0069] Furthermore, a horizontal angle is determined based on the horizontal field of view angle and the first difference; the horizontal angle is positively correlated with the horizontal field of view angle and negatively correlated with the first difference. In this embodiment, the horizontal angle is obtained by dividing the horizontal field of view angle by the first difference. Simultaneously, a vertical angle is determined based on the vertical field of view angle and the second difference; the vertical angle is positively correlated with the vertical field of view angle and negatively correlated with the second difference. In this embodiment, the vertical angle is obtained by dividing the vertical field of view angle by the second difference.

[0070] Furthermore, based on the obtained horizontal and vertical angles, addition and subtraction operations are performed with the current position of the gimbal to correct the position of the gimbal, and the target object is captured based on the corrected gimbal to obtain a captured image with a smaller error than the template image, and then the captured image is subjected to algorithmic recognition.

[0071] By correcting the gimbal position and then capturing the image again, the position error of the gimbal is reduced, thereby reducing the error between the captured image and the template image, and improving the accuracy and reliability of subsequent algorithm recognition of the captured image.

[0072] In other optional embodiments, the method of the present invention can also be as follows Figure 2As shown. In this embodiment, there are two processes: recording the template image and fine-tuning the pan / tilt head for template matching. Specifically, recording the template image includes: driving the robot to the specified position, manually controlling the pan / tilt head to aim the camera at the target, observing the camera image, and taking a snapshot after meeting the requirements of the template image (high clarity, target object located in the center of the image, etc.) to obtain the template image. The template image, the position when the template image was taken, and the pan / tilt head preset position are stored in the inspection task.

[0073] In this embodiment, the template matching pan-tilt fine-tuning includes: the robot autonomously navigates to the snapshot position according to the inspection route, calls the preset position specified by the pan-tilt, controls the camera to snapshot, grayscales the snapshot image and the template image and then performs template matching, determines whether the template matching is successful, and if the matching fails, the robot re-navigates to the snapshot position for a second snapshot, etc.; if the matching is successful, calculates the offset distance between the center point of the snapshot image and the template image, adjusts the pan-tilt according to the offset distance and the current pan-tilt position parameters, and then controls the camera to snapshot to obtain the snapshot image after optimization, and performs algorithm recognition on the snapshot image.

[0074] By correcting the pan / tilt position, it is possible to capture images even when there are errors in the navigation to the target position, thereby improving the accuracy of the algorithm's recognition results. Compared with manual inspection methods, the method of the present invention can reduce human intervention, record and analyze data in real time, and improve the efficiency of capture. Compared with traditional robot inspection capture, by re-capturing images with a similarity below a preset threshold, it can avoid problems such as target loss, target incompleteness, and target blur.

[0075] Figure 3 Schematically shows a structural block diagram of a robot capture system based on template matching and pan-tilt fine-tuning according to this embodiment.

[0076] In a second aspect, the present invention also provides a robot capture system based on template matching and pan / tilt fine-tuning. Figure 3 As shown, the robot capture system includes a processor and a memory, and the memory stores computer program instructions. When the computer program instructions are executed by the processor, a robot capture method based on template matching and pan-tilt fine-tuning according to the first aspect of the present invention is implemented.

[0077] The robot capture system also includes other components well known to those skilled in the art, such as a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.

[0078] In the present invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, the computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible or connectable to a device. Any application or module described in the present invention can be implemented using computer-readable / executable instructions that can be stored or otherwise retained by such a computer-readable medium.

[0079] In this specification, "multiple" means at least two, such as two, three, or more, unless otherwise specifically defined. Furthermore, the steps of the above method are divided for clarity of description only. During implementation, they can be combined into a single step, or some steps can be split into multiple steps, as long as they share the same logical relationship.

[0080] While several embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, variations, and alternatives will occur to those skilled in the art without departing from the concept and spirit of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention.

Claims

1. A robot capture method based on template matching and pan / tilt fine-tuning, characterized in that: include: In response to receiving an inspection task, according to the inspection task, driving to a capture position, calling a pan-tilt preset position to capture, and obtaining a capture image; the inspection task includes a template image, a capture position, and a pan-tilt preset position; Performing template matching on the captured image and the template image in the inspection task, and in response to a successful match, calculating an offset distance between the captured image and a center point of a target sub-image, the target sub-image representing the sub-image of all sub-images of the captured image that is most similar to the template image; The pan-tilt platform is adjusted according to the offset distance and the current position parameters of the pan-tilt platform to take snapshots.

2. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1 is characterized in that: The method for obtaining the template graph includes: In response to receiving the snapshot instruction, capturing the target object according to the snapshot instruction to obtain the template image; The template image, the capture position when capturing the template image, and the pan / tilt position parameters are saved in the inspection task.

3. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1 is characterized in that: Performing template matching on the captured image and the template image in the inspection task includes: Overlaying the template image on the snapshot image and translating the template image, and during the translation process, defining the area of ​​the snapshot image covered by the template image as a sub-image of the snapshot image; For any subgraph, calculating the similarity between the subgraph and the template graph; In response to the maximum similarity among all similarities being greater than a preset threshold, it is determined that the match is successful; in response to the maximum similarity among all similarities being less than or equal to the preset threshold, it is determined that the match is unsuccessful.

4. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 3 is characterized in that: The calculation expression of the similarity is: Where R(i,j) is the similarity, M is the number of pixels in the vertical direction of the template image, N is the number of pixels in the horizontal direction of the template image, and S ij (m,n) is the sub-image of the captured image, and T(m,n) is the template image.

5. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1, characterized in that: The template matching process further includes: in response to a matching failure, returning to the capture position and re-capturing.

6. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1, characterized in that: The offset distance includes a first difference and a second difference, and calculating the offset distance between the captured image and the center point of the target sub-image includes: Subtracting the abscissa of the center point of the template image from the abscissa of the center point of the target sub-image to obtain the first difference; The second difference is obtained by subtracting the vertical coordinate of the center point of the template image from the vertical coordinate of the center point of the target sub-image.

7. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1 is characterized in that: Adjusting the gimbal to capture the image according to the offset distance and the gimbal's current position parameter includes: Determine a horizontal field of view angle and a vertical field of view angle according to a current focal length of the camera; wherein the horizontal field of view angle is negatively correlated with the focal length, and the vertical field of view angle is negatively correlated with the focal length; determining a horizontal angle according to the horizontal field of view angle and the first difference; wherein the horizontal angle is positively correlated with the horizontal field of view angle and negatively correlated with the first difference; determining a vertical angle according to the vertical field of view angle and the second difference; wherein the vertical angle is positively correlated with the vertical field of view angle and negatively correlated with the second difference; Based on the horizontal angle and the vertical angle, the pan / tilt head is rotated according to the current position parameters to capture the image.

8. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 7, characterized in that: The calculation expression of the horizontal field angle is: The calculation expression of the vertical field angle is: Where A ver is the vertical field of view, H t is the camera target surface size height, F is the current focal length value, A level is the horizontal field of view, W t is the width of the camera target surface.

9. The robot capture method based on template matching and pan / tilt fine-tuning according to claim 1, characterized in that: Before matching, the method further includes: grayscale processing of the template image and the captured image.

10. A robot capture system based on template matching and pan / tilt fine-tuning, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a robot capture method based on template matching and pan-tilt fine-tuning according to any one of claims 1 to 9 is implemented.