Positioning method and device applied to known environment and storage medium

By selecting fixed feature points in known environments and matching the initial position of the monocular camera, the problem of insufficient robustness of the drone positioning in the repetitive scenery environment is solved, and the positioning effect of high precision and high robustness is achieved.

CN120125658AActive Publication Date: 2025-06-10TANGSHAN BAICHUAN INTELLIGENT MACHINE
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510187355.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing drone positioning methods are not robust enough in the environment of repeated scenes or similar scenes, have low positioning accuracy, and are prone to incorrect matching, resulting in pose information deviation.

Method used

By selecting fixed and unchanging feature points in a known environment, establishing a world coordinate system, and matching the initial position of the monocular camera and the preset threshold range to calculate the position of the drone.

Benefits of technology

It improves the positioning accuracy and robustness of the drone in known environments, and can accurately calculate the positioning even in a small number of feature points, reducing the impact of environmental changes on positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125658A_ABST
    Figure CN120125658A_ABST
Patent Text Reader

Abstract

The invention discloses a positioning method and device applied to a known environment and a storage medium, and relates to the technical field of space positioning. Comprising the following steps: acquiring world coordinates of each first feature point in a known environment; acquiring a first image shot by the monocular camera at the current moment; extracting a second feature point in the first image and acquiring a second pixel coordinate of the second feature point; calculating a first pixel coordinate of the first feature point based on the pose of the monocular camera at the previous moment and the world coordinate of the first feature point; matching the second feature point and the first feature point based on the second pixel coordinate of the second feature point, the first pixel coordinate of the first feature point and a preset threshold range; and calculating the pose of the monocular camera at the current moment based on the second pixel coordinates of the second feature points of the plurality of matching pairs and the world coordinates of the first feature points. The method is high in positioning precision, high in robustness, simple and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatial positioning, and particularly relates to a positioning method, device and storage medium applied to a known environment. Background Art

[0002] Unmanned aerial vehicles include unmanned aircraft and AGV carts, etc. During the process of performing tasks, real-time positioning is required, that is, it is necessary to determine its own pose in real time, that is, to determine its own position and attitude. A current positioning method is as follows: By setting a camera on the unmanned aerial vehicle, the pose of the unmanned aerial vehicle is determined according to the images captured by the camera in real time. However, the existing positioning method has defects such as insufficient robustness and low positioning accuracy when applied in some industrial scenarios. For example, in some environments with repetitive or similar scenes, the images captured by the camera at different positions are highly similar. When extracting feature points from the images and matching the feature points with the position points of the actual environment, incorrect matching is likely to occur, resulting in incorrect acquisition of the world coordinates of the feature points, and thus deviation of the pose information of the unmanned aerial vehicle calculated based on the feature points. Summary of the Invention

[0003] The purpose of the present invention is to provide a positioning method, device and storage medium applied to a known environment, which have high positioning accuracy, strong robustness, and are simple and reliable.

[0004] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0005] The present invention provides a positioning method applied to a known environment, including:

[0006] Step 101, obtaining the world coordinates of each first feature point in the known environment;

[0007] Step 102, obtaining a first image captured by a monocular camera at the current moment;

[0008] Step 103, extracting second feature points in the first image and obtaining second pixel coordinates of the second feature points, where the second pixel coordinates are actual measured values of the pixel coordinates of the second feature points in the first image; the second feature points are two-dimensional projection points of the first feature points in the first image;

[0009] Step 104, calculating first pixel coordinates of the first feature point based on the pose of the monocular camera at the previous moment and the world coordinates of the first feature point; the first pixel coordinates are theoretical calculated values of the pixel coordinates of the first feature point imaged under the pose of the monocular camera at the previous moment; the initial pose of the monocular camera is a known pose;

[0010] Step 105: Based on the second pixel coordinates of the second feature points, the first pixel coordinates of the first feature points, and a preset threshold range, match the second feature points and the first feature points to obtain a number of matching pairs, where each matching pair includes a second feature point and a first feature point;

[0011] Step 106: Based on the second pixel coordinates of the second feature points and the world coordinates of the first feature points in the number of matching pairs, calculate the pose of the monocular camera at the current moment;

[0012] Step 107: Repeat steps 102 - 106.

[0013] Optionally, the first feature points in step 104 are the first feature points located within a first range corresponding to the pose of the monocular camera at the previous moment, and different poses of the monocular camera correspond to different first ranges.

[0014] Optionally, in step 103, the extraction of the second feature points in the first image specifically includes: using a deep learning model to extract the second feature points in the first image.

[0015] Optionally, the training method of the deep learning model includes:

[0016] Step 201: Shoot a video of the known environment, extract frames from the video frame by frame using a frame extraction technique, label the second feature points in each extracted frame using image annotation software and generate corresponding labels, and divide the labeled dataset into a training set, a validation set, and a test set;

[0017] Step 202: Determine the architecture, number of layers, number of neurons, and connection method of the deep learning model to obtain the deep learning model;

[0018] Step 203: Use the training set to train the deep learning model, adjust the parameters of the deep learning model through the backpropagation algorithm to minimize the loss function, and optimize the parameters of the deep learning model through an optimization algorithm to ensure that the deep learning model converges and does not overfit;

[0019] Step 204: Use the validation set to evaluate the performance of the model, evaluate the deep learning model through accuracy, precision, recall, and F1 score, and according to the evaluation results, optimize the deep learning model by adjusting hyperparameters, modifying the model structure, or using regularization methods.

[0020] Optionally, in step 103, the extraction of the second feature points in the first image specifically includes:

[0021] Step 301, perform grayscale processing on the first image;

[0022] Step 302, perform smoothing processing on the grayscaled first image;

[0023] Step 303, use an edge detection method to perform edge detection on the smoothed first image to obtain an edge image;

[0024] Step 304, perform Hough transform on the edge image to extract edge line segments;

[0025] Step 305, extract the second feature points based on the edge line segments.

[0026] Optionally, when the second feature points are corner points, extracting the second feature points based on the edge line segments in step 305 specifically includes:

[0027] Extracting the intersection points of the edge line segments, and the intersection points are the second feature points.

[0028] Optionally, when there are square columns in the known environment, the first feature points include the right-angle points of the square columns;

[0029] Alternatively, when there are cylindrical columns in the known environment, the first feature points include the center points of the cross-sections of the cylindrical columns.

[0030] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of any one of the methods described above.

[0031] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of any one of the methods described above are implemented

[0032] In the technical solution of the present invention, for the scenario where a drone performs tasks in a known environment, a world coordinate system is established for the known environment, and fixed points in the known environment are selected as feature points. When performing tasks each time, the pose of the monocular camera on the drone is initialized, so that the initial pose of the monocular camera is known. A preset threshold range is set in advance according to the moving speed of the drone and the shooting frequency of the monocular camera. After taking the first image at the current moment, when determining the feature points in the real environment corresponding to the feature points in the first image, the pose of the monocular camera at the previous moment and the feature points that can be captured by the monocular camera in this pose are used as references, and are constrained by the preset threshold range, so that the obtained corresponding relationship is reliable. Moreover, by using this method, even when there are a small number of available feature points in the known environment, the pose of the monocular camera can be accurately calculated, and the positioning will not be affected by the changes of other objects in the known environment except the feature points, and both the robustness and the positioning accuracy are relatively high. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 FIG. is a schematic diagram of a flowchart of a positioning method applied to a known environment;

[0034] Figure 2 FIG. is a perspective view of a known environment;

[0035] Figure 3 FIG. is a schematic diagram of the first image captured by the monocular camera;

[0036] Figure 4 FIG. is a schematic diagram of marking the second feature points in the first image;

[0037] Figure 5 FIG. is a schematic diagram of matching the second feature points and the first feature points;

[0038] Figure 6 FIG. is a schematic diagram of extracting the second feature points on the first image based on edge segments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Here, exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of devices or methods consistent with some aspects of the present invention.

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other implementation manners obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope protected by the present invention.

[0041] Hereinafter, embodiments will be described with reference to the accompanying drawings. In addition, the embodiments shown below do not impose any limitation on the inventive concept described in the claims. Further, all the contents of the configurations shown in the following embodiments are not necessarily essential for the solution of the invention described in the claims.

[0042] See Figure 1 , the present invention provides a positioning method applicable to a known environment. From a program perspective, the execution subject of this process can be a program installed on a server or a positioning device; from a hardware perspective, the execution subject of this process can be a positioning device or a server, specifically, a positioning device installed on a drone, a computer, or a drone control platform, etc. As Figure 1 shown, this method may include the following steps:

[0043] Step 101, obtain the world coordinates of each first feature point in the known environment.

[0044] When performing certain tasks, a drone sails in a certain known environment. That is, some objects in this environment are fixed and unchanged, and the position information of these objects can be measured in advance. Furthermore, a world coordinate system origin can be defined customarily to establish the world coordinate system of this real environment. Select some fixed and unchanged points in the real environment as feature points. Here, the feature points in the real environment are called first feature points, and the world coordinates of each first feature point can be obtained through measurement. See Figure 2 , Figure 2 is a certain known environment. The drone sails in Passage 1. There are columns on both sides of Passage 1. These columns are generally fixed and unchanged. The right-angle points where the columns meet the ground can be selected as the first feature points. For ease of understanding, some of the first feature points are labeled, such as first feature points A, B, C, and D. After selecting a certain point in the real environment as the world coordinate system origin, the world coordinates of each right-angle point, that is, each first feature point, can be obtained through measurement. Obtain the world coordinates of each first feature point and store them for subsequent steps to call.

[0045] Step 102, obtain the first image captured by the monocular camera at the current moment.

[0046] A monocular camera is installed on the drone. During the flight of the drone, the monocular camera takes pictures of the surrounding environment at a certain frequency, and the taken images are called the first images. After the monocular camera takes the first images, the execution entity acquires the first images, that is, the monocular camera transmits the first images to the execution entity.

[0047] Step 103: Extract the second feature points in the first image and obtain the second pixel coordinates of the second feature points. The second pixel coordinates are the actual measured values of the pixel coordinates of the second feature points in the first image; the second feature points are the two-dimensional projection points of the first feature points in the first image.

[0048] After the execution entity acquires the first images taken by the monocular camera at the current moment, it processes the first images and extracts the second feature points in the first images. After the monocular camera takes pictures, the first feature points in the real environment will be imaged in the first images. Here, the images of the first feature points in the first images are called the second feature points. The second feature points are also the two-dimensional projection points of the first feature points in the first images. Refer to Figure 3 , Figure 3 is the first image taken by the monocular camera at a certain pose. The right-angle points of the columns on the first image are called the second feature points here. For ease of understanding, some of the second feature points are labeled, such as the second feature points a, b, and c. After the execution entity acquires the first images, it will extract the second feature points. After analyzing and processing the first images and the second feature points, the second pixel coordinates of the second feature points can be obtained, that is, the actual measured values of the pixel coordinates of the second feature points in the first images can be obtained.

[0049] Step 104: Based on the pose of the monocular camera at the previous moment and the world coordinates of the first feature points, calculate the first pixel coordinates of the first feature points. The first pixel coordinates are the theoretical calculated values of the pixel coordinates of the first feature points imaged in the pose of the monocular camera at the previous moment; the initial pose of the monocular camera is a known pose.

[0050] The initial pose of the monocular camera is known. Each time the UAV executes a task, it will first be initialized, that is, the UAV is placed at a certain position and the monocular camera is adjusted to a certain pose. The initial pose of the monocular camera can be obtained through pre-measurement or directly set. Each time the UAV starts to execute a task, the monocular camera is initialized to the initial pose. For the sake of easy understanding, take the next moment of the initial pose as the current moment as an example. At this time, the pose of the monocular camera at the previous moment is the initial pose and is known. According to the pose of the monocular camera at the previous moment and the world coordinates of the first feature point, the first pixel coordinates of the first feature point can be calculated. When the monocular camera is in the pose at the previous moment, there is a pixel coordinate system. The pixel coordinates obtained by theoretically converting the first feature point in the real environment to this pixel coordinate system are the first pixel coordinates. The first pixel coordinates can be calculated based on the physical parameters of the monocular camera, the world coordinates of the first feature point, and the pose of the monocular camera.

[0051] After defining the origin of the world coordinates of the real environment, the world coordinates of each first feature point in the real environment can be obtained through pre-measurement, and the initial pose of the monocular camera can be obtained. The initial pose of the monocular camera can be pre-set or obtained through pre-measurement. The first pixel coordinates (u exp_x , v exp_y ) of each first feature point can be calculated through the internal parameter matrix and distortion parameters of the monocular camera. The specific calculation process is as follows:

[0052] First, convert the world coordinates of the first feature point to the camera coordinates in the camera coordinate system. The calculation is carried out through the following formula:

[0053]

[0054] Among them, [X c Y c Z c T is the camera coordinates of the first feature point in the monocular camera coordinate system, X w Y w Z w T is the world coordinates of the first feature point; R is the rotation matrix used to describe the rotation of the monocular camera, and T is the translation matrix used to describe the position of the monocular camera in the world coordinate system. The two form a 3×4 matrix, that is, the initial pose of the camera. After obtaining the initial pose T init-x , T init-y , T init-z , R init-x , R init-y , R init-z ) of the monocular camera, the matrix R and the matrix T can be obtained. ​​

[0055] Secondly, normalize the camera coordinates of the first feature point:

[0056]

[0057] Here, P c can be regarded as a two-dimensional homogeneous coordinate, called the normalized coordinate. The pixel coordinate is regarded as the result of quantifying and measuring the point on the normalized plane.

[0058] Thirdly, according to the distortion parameters of the monocular camera, correct the radial distortion and tangential distortion of the points on the normalized plane:

[0059]

[0060] Finally, project the corrected points onto the pixel plane through the internal parameter matrix to obtain the first pixel coordinates (u exp_x , v exp_y ) of the point on the image.

[0061]

[0062] Step 105: Based on the second pixel coordinates of the second feature point, the first pixel coordinates of the first feature point, and a preset threshold range, match the second feature point and the first feature point to obtain a number of matching pairs. Each matching pair includes a second feature point and a first feature point.

[0063] The preset threshold range can be set according to the moving speed and photographing frequency of the monocular camera. Based on the second pixel coordinates and the first pixel coordinates, the position difference between the second feature point and the first feature point can be obtained. If the position difference is within the preset threshold range, then the second feature point and the first feature point are matched as a matching pair, and the second feature point is the two-dimensional projection point of the first feature point on the first image.

[0064] See Figure 5 , Figure 5 for a schematic diagram of matching the second feature point and the first feature point. Figure 5 is a first image, and there is a second feature point in the first image. After theoretically calculating the first pixel coordinates of the first feature point and marking the first feature point at the position of the first pixel coordinates on this first image, Figure 5 is obtained. Figure 5 The two points connected by the line in are the points that meet the preset threshold range, and the two points connected by the line form a matching pair. The second feature point a is matched with the first feature point A, the second feature point b is matched with the first feature point B, and the second feature point c is matched with the first feature point C.

[0065] In some cases, when performing matching, in addition to meeting the requirements of the preset threshold range, additional limiting conditions can be added for the second feature point and the first feature point, that is, the first feature point with the smallest position difference from the second feature point and meeting the preset threshold range, and only then can the second feature point and the first feature point form a matching pair. It should be noted that after the second feature point and the first feature point are successfully matched, they will not participate in the matching again.

[0066] Specifically, if a second feature point and a first feature point satisfy the following formula, then the second feature point and the first feature point are matched as a matching pair.

[0067]

[0068] Wherein, is the preset threshold, is the second pixel coordinate, is the first pixel coordinate.

[0069] That is, when the second pixel coordinate minus the first pixel coordinate [u cap_x , v cap_y is less than the set threshold u threshold , v threshold ), it can be determined that the second feature point is the two-dimensional projection point of the first feature point on the first image.

[0070] Step 106, based on the second pixel coordinates of the second feature points of the several matching pairs and the world coordinates of the first feature points, calculate the pose of the monocular camera at the current moment.

[0071] When there are more than four matching pairs, and the feature points of these matching pairs are coplanar but not collinear, that is, at least three of the four groups of feature points are not on the same straight line, the pose of the current monocular camera can be calculated. Specifically, the pose of the monocular camera at the current moment can be calculated by the solvePnP algorithm. Since the monocular camera is set on the unmanned aerial vehicle, there is a definite mapping relationship between the pose of the monocular camera and the pose of the unmanned aerial vehicle, and thus the pose of the unmanned aerial vehicle at the current moment can be obtained.

[0072] Step 107, repeat steps 102 - 106.

[0073] When the next moment arrives, the next moment becomes the current moment, and the current moment in step 106 becomes the previous moment, that is, the pose of the monocular camera at the previous moment has been calculated and is known. By executing steps 102 - 106 at each moment, the pose of the monocular camera can be obtained in real time. "The pose of the previous moment in step 102" is only the initial pose when calculating for the first time, which is pre-measured or set. In the subsequent calculation process, it is the pose obtained from the previous calculation.

[0074] This method is applicable to the scenario where a drone performs tasks in a known environment. A world coordinate system is established for the known environment, and fixed points in the known environment are selected as feature points. The initial pose of the monocular camera on the drone is initialized each time a task is executed, so that the initial pose of the monocular camera is known. A preset threshold range is set in advance according to the moving speed of the drone and the shooting frequency of the monocular camera. After taking the first image at the current moment, when determining the feature points in the real environment corresponding to the feature points in the first image, the pose of the monocular camera at the previous moment and the feature points that can be captured by the monocular camera at this pose are used as references and constrained by the preset threshold range, so that the obtained corresponding relationship is reliable. Moreover, by using this method, even when there are a small number of available feature points in the known environment, the pose of the monocular camera can be accurately calculated, and the positioning will not be affected by changes in other objects in the known environment except for the feature points, and both the robustness and the positioning accuracy are relatively high.

[0075] Optionally, the first feature points in step 104 are the first feature points located within a first range corresponding to the pose of the monocular camera at the previous moment, and different poses of the monocular camera correspond to different first ranges.

[0076] When calculating the first pixel coordinates of the first feature points, the first pixel coordinates of all the first feature points in the known environment can be calculated, or only the first feature points within the first range can be calculated to obtain the first pixel coordinates of the first feature points within the first range. The first range can be a range determined based on the field of view of the monocular camera. For example, it can be the field of view of the monocular camera, or a range obtained by appropriately expanding the edges on the basis of the field of view of the monocular camera. Refer to Figure 2 , assuming Figure 2 that the objects in are within a certain field of view of the monocular camera. Although point C is blocked by the column 2, point C also belongs to the first feature points within the first range. There is a definite mapping relationship between the pose of the monocular camera and the first feature points within the field of view of the monocular camera. The mapping relationship between different poses and the first range can be established in advance, and when determining the first range, it can be directly retrieved. It can also be calculated in real time according to the pose and the physical parameters of the monocular camera to determine the first range.

[0077] Selecting the method of only calculating the first feature points within the first range can reduce the amount of calculation.

[0078] Optionally, in step 103, the extracting the second feature points in the first image specifically includes: extracting the second feature points in the first image by using a deep learning model.

[0079] Optionally, the training method of the deep learning model includes:

[0080] Step 201: Shoot a video of the known environment, extract each frame of the video through frame extraction technology, use image annotation software to annotate the second feature points in each extracted frame image and generate corresponding labels, and divide the annotated data set into a training set, a validation set, and a test set.

[0081] See Figure 4 , after annotating the second feature points e and f, the generated label is the coordinate (x, y) of the second feature point. At this time, a rectangular box with a size of 10 * 10 pixels will be generated centered on the annotated second feature point in the image, and the detection of the second feature point will be converted into the detection of a small target rectangular box. After annotation, divide the annotated images into a training set, a validation set, and a test set.

[0082] Step 202: Determine the architecture, number of layers, number of neurons, and connection method of the deep learning model to obtain the deep learning model.

[0083] Step 203: Use the training set to train the deep learning model, adjust the parameters of the deep learning model through the backpropagation algorithm to minimize the loss function, and optimize the parameters of the deep learning model through the optimization algorithm to ensure that the deep learning model converges and does not overfit.

[0084] Step 204: Use the validation set to evaluate the performance of the model, evaluate the deep learning model through accuracy, precision, recall, and F1 score, and according to the evaluation results, optimize the deep learning model by adjusting hyperparameters, modifying the model structure, or using regularization methods.

[0085] At this time, the training of the deep learning model is completed and can be applied to the extraction of the second feature points.

[0086] Of course, step 205 can also be added to evaluate the performance of the deep learning model using the test set, analyze the prediction results of the model, calculate various performance metrics to determine the actual effect of the deep learning model, and after determining the actual effect of the deep learning model, put the deep learning model into application.

[0087] Of course, during the application process of the model, step 206 can be added to continuously monitor the performance of the deep learning model in the application environment, detect signs of degradation or performance decline of the deep learning model, and perform necessary updates and maintenance.

[0088] Optionally, in step 103, the extraction of the second feature points in the first image specifically includes:

[0089] Step 301: Perform grayscale processing on the first image.

[0090] Perform grayscale processing on the first image to convert the color first image into a grayscale image G(x, y).

[0091] Step 302: Perform smoothing processing on the grayscale-processed first image.

[0092] Specifically, the Gaussian smoothing method or other methods can be used to perform smoothing processing on the grayscale image G(x, y) to reduce noise interference. Here, the Gaussian smoothing method is used as an example for description:

[0093] Perform a convolution operation on the grayscale image using a two-dimensional Gaussian filter to obtain a smoothed image, that is, the filtered image G f (x, y):

[0094]

[0095] Among them, σ is the standard deviation of the Gaussian function.

[0096] Step 303: Use an edge detection method to perform edge detection on the smoothed first image to obtain an edge image.

[0097] First, calculate the gradient magnitude M(x, y) and direction θ(x, y) of the filtered image,

[0098]

[0099] Among them, G x (x, y) and G y (x, y) are the gradients of the filtered image in the x and y directions respectively.

[0100] Then perform non-maximum suppression on the gradient magnitude.

[0101] Finally, through double-threshold detection and edge connection, obtain the edge image E=(x, y).

[0102] Step 304: Perform Hough transform on the edge image to extract edge segments.

[0103] Perform Hough transform on the edge image E=(x, y) to map the edge points in the image space to the Hough parameter space. Through random sampling and endpoint detection of the edge points, convert each edge point from Cartesian coordinates to ρ and θ in the Hough parameter space through polar coordinate transformation. Through random sampling and accumulation of ρ and θ, find the edge segments in the image through voting and threshold judgment, and obtain the coordinate information of the two endpoints of each edge segment in the Cartesian coordinate system (x 1 , y 1 ) and (x 2 , y 2), and then convert it into polar coordinates (ρ, θ), and fuse line segments with similar polar coordinate information to reduce interference information. In the polar coordinate system, the equation of a line can be expressed as ρ = xcosθ + ysinθ, where ρ is the vertical distance from the origin to the line, and θ is the angle between the line and the positive direction of the x-axis.

[0104] Step 305: extract the second feature point based on the edge segment.

[0105] After the edge line segment is extracted, the second feature point is extracted based on the extracted edge line segment. For example, when the second feature point is the center of a circle, the extracted edge line segment is a circle, and the second feature point can be obtained by extracting the center of the circle. When the second feature point is a corner point, after the edge line segment is extracted, the second feature point can be obtained by calculating the intersection of the edge line segment.

[0106] Optionally, when the second feature point is a corner point, extracting the second feature point based on the edge segment in step 305 specifically includes:

[0107] The intersection points of the edge line segments are extracted, and the intersection points are the second feature points.

[0108] See also Figure 6 , Figure 6 The second feature point in is a right angle point. The edge segments extracted by Hough transform are straight line segments, such as straight line segments l1, l2, l3, l4, etc. The second feature point can be obtained by calculating the intersection of the straight line segments by combining the straight line equations. The intersection of the straight line segments here can refer to the intersection of the straight line segments or the extension of the straight line segments.

[0109] For example, by calculating the intersection point between the straight line segment l1 and the straight line segment l2, the intersection point between the straight line segment l3 and the straight line segment l4, etc., the second feature points a, b, etc. can be obtained, thereby extracting each second feature point in the first image.

[0110] Optionally, when a square pillar exists in the known environment, the first feature point includes a right-angle point of the square pillar.

[0111] See also Figure 2 When there is a square column in the known environment, four right-angle points at the intersection of the square column and the ground are selected as the first feature points.

[0112] Alternatively, when there is a cylindrical pillar in the known environment, the first feature point includes a center point of a cross section of the cylindrical pillar.

[0113] Of course, other points in the known environment that are easy to identify can also be selected as the first feature point, such as corner points of some polygons, or the center points of screws or nuts used to fix the fan, and the center point of the fan can all be used as the first feature point.

[0114] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of any one of the methods described above.

[0115] The present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps of any one of the methods described above are implemented.

Claims

1. A positioning method applied to a known environment, characterized in that: include: Step 101, obtaining the world coordinates of each first feature point in a known environment; Step 102, obtaining a first image captured by the monocular camera at the current moment; Step 103, extracting a second feature point in the first image and acquiring a second pixel coordinate of the second feature point, where the second pixel coordinate is an actual measurement value of the pixel coordinate of the second feature point in the first image; The second feature point is a two-dimensional projection point of the first feature point in the first image; Step 104: based on the posture of the monocular camera at the last moment and the world coordinates of the first feature point, calculate the first pixel coordinates of the first feature point; the first pixel coordinates are theoretically calculated values ​​of the pixel coordinates of the first feature point imaged at the posture of the monocular camera at the last moment; the initial posture of the monocular camera is a known posture; Step 105, matching the second feature point with the first feature point based on the second pixel coordinate of the second feature point, the first pixel coordinate of the first feature point and a preset threshold range to obtain a plurality of matching pairs, each of which includes a second feature point and a first feature point; Step 106, calculating the position and posture of the monocular camera at the current moment based on the second pixel coordinates of the second feature points of the plurality of matching pairs and the world coordinates of the first feature point; Step 107, repeat steps 102-106.

2. The positioning method applied to a known environment as claimed in claim 1, characterized in that: The first feature point in step 104 is a first feature point located in a first range corresponding to the position and posture of the monocular camera at a previous moment, and different positions and postures of the monocular camera correspond to different first ranges.

3. The positioning method applied to a known environment as claimed in claim 1, characterized in that: In step 103, extracting the second feature points in the first image specifically includes: extracting the second feature points in the first image using a deep learning model.

4. The positioning method applied to a known environment as claimed in claim 3, characterized in that: The training method of the deep learning model includes: Step 201, shooting a video of the known environment, extracting the video frame by frame using a frame extraction technique, annotating the second feature points in each extracted frame of the image using image annotation software and generating corresponding labels, and dividing the annotated data set into a training set, a validation set, and a test set; Step 202, determining the architecture, number of layers, number of neurons and connection mode of the deep learning model to obtain the deep learning model; Step 203, training the deep learning model using the training set, adjusting the parameters of the deep learning model by a back propagation algorithm to minimize the loss function, and optimizing the parameters of the deep learning model by an optimization algorithm to ensure that the deep learning model converges and does not overfit; Step 204, using the validation set to evaluate the performance of the model, evaluating the deep learning model through accuracy, precision, recall and F1 score, and based on the evaluation results, tuning the deep learning model by adjusting hyperparameters, modifying the model structure or using regularization methods.

5. The positioning method applied to a known environment as claimed in claim 1, characterized in that: In step 103, extracting the second feature point in the first image specifically includes: Step 301, grayscale processing is performed on the first image; Step 302, performing smoothing processing on the first image after grayscale processing; Step 303, using an edge detection method to perform edge detection on the first image after smoothing to obtain an edge image; Step 304, performing Hough transform on the edge image to extract edge segments; Step 305: extract the second feature point based on the edge segment.

6. The positioning method applied to a known environment as claimed in claim 5, characterized in that: When the second feature point is a corner point, extracting the second feature point based on the edge line segment in step 305 specifically includes: The intersection points of the edge line segments are extracted, and the intersection points are the second feature points.

7. The positioning method applied to a known environment as claimed in claim 1, characterized in that: When there is a square pillar in the known environment, the first feature point includes a right angle point of the square pillar; Alternatively, when there is a cylindrical pillar in the known environment, the first feature point includes a center point of a cross section of the cylindrical pillar.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method based on artificial marker and stereoscopic vision

    CN112509125A

  • Monocular camera pose estimation method and system

    CN113256711A

  • Transparent object positioning method and device based on monocular color and storage medium

    CN115830103A

  • Box workpiece pose measurement method based on point features

    CN116091603A

  • Three-dimensional reconstruction method and apparatus for monocular endoscope image, and terminal device

    WO2021115071A1