Road target visual positioning system and method based on urban street scene
Through a visual positioning method based on urban street scenes, feature matching and deep learning networks are used to solve the pose parameters. Combined with GPS coordinates and epipolar geometry, the problem of low accuracy of traditional positioning technology in urban environments is solved, and high-precision and stable visual positioning is achieved.
Patent Information
- Application Number
- CN202510970558.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Traditional positioning technology is affected by reflection, refraction and electromagnetic interference from tall buildings in urban environments, resulting in low positioning accuracy. Existing visual positioning systems have high computing resource and time costs when processing image data and eliminating interference from moving objects, and rely on large-scale image databases, making it difficult to provide high-precision positioning in complex urban environments.
A visual positioning method based on urban street scenes is adopted. By obtaining positioning pictures and query pictures for feature matching, deep learning networks or visual algorithms are used to solve the pose parameters. The absolute coordinate position of the query picture is calculated by combining GPS coordinates and epipolar geometry to enhance positioning accuracy and stability.
It improves the accuracy and reliability of visual positioning in complex urban environments, solves the problem of limited accuracy of traditional methods when the shooting interval is large or the distance is far, and ensures that targets can be accurately identified and located even when there are large spatial or temporal gaps between images.
Smart Images

Figure CN120472006B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision, and more particularly to a road target visual positioning system and method based on urban street scenes. Background Art
[0002] In complex urban environments, traditional positioning technologies such as GPS (Global Positioning System) and RTK (Real-Time Kinematic) struggle to provide high-precision location information due to factors such as reflection and refraction from tall buildings and electromagnetic interference. Particularly in dense urban canyons, the instability of satellite signals causes positioning drift, making these traditional methods a significant challenge in applications requiring precise location information, such as autonomous driving and smart city management. Furthermore, existing positioning systems based on radio signals cannot effectively address the multipath effect in urban environments, further limiting their performance and reliability in real-world applications.
[0003] To address these issues, visual positioning systems have been proposed as an alternative, using relatively fixed visual features in urban environments for position matching and positioning. However, existing technologies have encountered many difficulties in their implementation, such as how to efficiently process large amounts of image data and how to accurately eliminate the interference of moving objects such as pedestrians and vehicles on the positioning results. Despite this, such systems still rely on pre-established large-scale image databases and require complex algorithms to ensure rapid matching between real-time images and reference images in the database, which places high demands on computing resources and time costs.
[0004] Therefore, it is necessary to design a new method to achieve rapid matching of images of long-time scale-invariant elements such as buildings and environments, and to solve the accuracy limitation problem of traditional visual positioning methods based on image feature matching when the interval between positioning images is large or the distance between the query image and the positioning image is far. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a road target visual positioning system and method based on urban street scenes.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for visually locating road objects based on urban street scenes, comprising:
[0007] Get the positioning image and multiple query images;
[0008] Perform feature matching and extraction on each query image and the positioning image to obtain a set of feature points corresponding to each query image;
[0009] Solve the pose parameters based on deep learning networks or visual algorithms to obtain the solution results;
[0010] Determine the movement direction and distance between adjacent query images in multiple query images, and calculate the absolute coordinate position of the query image by combining the solution results, the GPS coordinates of the positioning image and the epipolar geometry relationship.
[0011] A further technical solution is: the positioning picture includes a basic picture with GPS coordinates and direction angles.
[0012] A further technical solution is: solving the posture parameters based on the deep learning network or visual algorithm to obtain the solution result includes:
[0013] Calculating a fundamental matrix for the set of feature points corresponding to each query image, constructing an intrinsic parameter matrix based on camera parameters, and calculating an essential matrix using the fundamental matrix and the intrinsic parameter matrix; wherein the essential matrix includes relative rotation and translation information;
[0014] The pose parameters of the essential matrix corresponding to each query image are solved to obtain the solution.
[0015] A further technical solution is: calculating a basic matrix for the feature point set corresponding to each query image, constructing an intrinsic parameter matrix in combination with camera parameters, and calculating an essential matrix using the basic matrix and the intrinsic parameter matrix, including:
[0016] Randomly selecting feature points from the feature point set corresponding to each query image, and calculating a basic matrix based on the randomly selected feature points;
[0017] Construct the intrinsic parameter matrix using known camera parameters;
[0018] The fundamental matrix is converted into an intrinsic matrix using the intrinsic parameter matrix.
[0019] A further technical solution is: randomly selecting feature points from the feature point set corresponding to each query image, and calculating a basic matrix based on the randomly selected feature points, including:
[0020] Feature points are randomly selected from the feature point set corresponding to each query image, and a basic matrix is calculated using an eight-point method algorithm based on the randomly selected feature points.
[0021] A further technical solution is: solving the posture parameters based on the deep learning network or visual algorithm to obtain the solution result includes:
[0022] When a deep learning network is used to solve the pose parameters, the query image is input into the pose solving model to solve the pose parameters to obtain a solution result;
[0023] Among them, the posture solution model is obtained by training a neural network model by using several pictures with position, orientation, and camera parameter information as a sample set; the loss function of the posture solution model includes a loss function that combines relative position and posture errors, and balances the importance of relative position and posture errors through hyperparameters.
[0024] A further technical solution is: determining the movement direction and distance between adjacent query images in a plurality of query images, and calculating the absolute coordinate position of the query image in combination with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship, including:
[0025] Using sensor data from the vehicle-mounted device to determine the movement direction and distance between adjacent query images in multiple query images to obtain a direction vector;
[0026] The absolute coordinate position of the query image is calculated based on the direction vector combined with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship.
[0027] A further technical solution is: the absolute coordinate position of the query image is calculated based on the direction vector in combination with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship, including:
[0028] According to the direction vector and the solution result, the relative direction and distance of the query image relative to the positioning image are calculated based on the epipolar geometric relationship using the triangulation principle to determine the absolute coordinate position of the query image.
[0029] A further technical solution is: calculating the relative direction and distance of the query image relative to the positioning image based on the direction vector and the solution result using the principle of triangulation based on the epipolar geometric relationship to determine the absolute coordinate position of the query image, including:
[0030] Calculating the relative direction between the adjacent query image and the positioning image using the epipolar geometry relationship according to the direction vector and the solution result;
[0031] Considering the adjacent query image and the positioning image as triangle vertices, and calculating the length between the adjacent query image and the positioning image according to the direction vector based on the cosine theorem to obtain a relative distance;
[0032] The absolute coordinate position of the query image is determined in combination with the relative direction and the relative distance.
[0033] The present invention also provides a road target visual positioning system based on urban street scenes, comprising:
[0034] An acquisition unit, used to acquire a positioning image and multiple query images;
[0035] An extraction unit, configured to perform feature matching and extraction on each query image and the positioning image, respectively, to obtain a set of feature points corresponding to each query image;
[0036] A solving unit is used to solve the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result;
[0037] The absolute coordinate calculation unit is used to determine the corresponding moving direction and distance between adjacent query images in multiple query images, and calculate the absolute coordinate position of the query image in combination with the GPS coordinates and epipolar geometry of the positioning image.
[0038] Compared with the prior art, the present invention has the following beneficial effects: the present invention obtains a positioning picture and multiple query pictures, performs feature matching and extraction on each query picture and the positioning picture, obtains a set of feature points, and then solves the posture parameters of each query picture; at the same time, determines the movement direction and distance between adjacent query pictures, and calculates the absolute coordinate position of the query picture using the GPS coordinates and epipolar geometry of the positioning picture; this method solves the problem of limited accuracy of traditional visual positioning methods based on picture feature matching when facing large shooting intervals or long distances, enhances positioning accuracy by integrating information from multiple perspectives, ensures that targets can be accurately identified and positioned even when there are large spatial or temporal gaps between images, effectively improves the stability and reliability of visual positioning in complex urban environments, and improves positioning errors caused by rough GPS information when the density of positioning pictures is low.
[0039] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 A schematic diagram of a flow chart of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention;
[0042] Figure 2A schematic diagram of a sub-process of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention;
[0043] Figure 3 A schematic diagram of a sub-process of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention;
[0044] Figure 4 A schematic diagram of a sub-process of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention;
[0045] Figure 5 A schematic diagram of a sub-process of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention;
[0046] Figure 6 A schematic diagram of matching a positioning image and a query image provided by an embodiment of the present invention;
[0047] Figure 7 Schematic diagram of feature point matching between the positioning image and the query image provided by an embodiment of the present invention;
[0048] Figure 8 A schematic diagram of the positional relationship between the positioning image and the query image provided by an embodiment of the present invention;
[0049] Figure 9 A schematic block diagram of a road target visual positioning system based on urban street scenes provided by an embodiment of the present invention;
[0050] Figure 10 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0052] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0053] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0054] It should be further understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0055] See also Figure 1 , Figure 1 This is a schematic flow chart of a method for visually localizing road objects based on urban street scenes, provided by an embodiment of the present invention. This method is applied to a server. The server interacts with a camera. The method combines GPS coordinates and azimuth angles in a base image (i.e., the positioning image) with feature matching between multiple query images to calculate the relative direction and absolute coordinate position of each query image relative to the positioning image. The specific steps include: obtaining a positioning image and a query image; matching and extracting a set of feature points; calculating an intrinsic matrix using a fundamental matrix and a camera intrinsic parameter matrix; decomposing the intrinsic matrix to obtain pose parameters; determining the movement direction and distance between adjacent query images, and calculating the absolute coordinate position of the query image using triangulation principles based on epipolar geometry. This method is particularly suitable for addressing the accuracy issues of traditional visual localization techniques based on image feature matching when the images are taken far apart or at long distances. By rapidly matching long-timescale invariant elements, such as buildings and the surrounding environment, it improves positioning accuracy and reliability. Therefore, positioning accuracy can be effectively improved even when the positioning images are taken far apart or when the query image is far from the positioning image.
[0056] Figure 1 FIG. 1 is a flow chart of a method for visually locating road objects based on urban street scenes provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S150.
[0057] S110: Obtain a positioning image and multiple query images.
[0058] In this embodiment, the positioning picture includes a basic picture with GPS coordinates and direction angles.
[0059] Location images are basic images with precise GPS coordinates and azimuth information. These images are typically city street scenes taken at known locations, and each image is accompanied by accurate geographic information (such as latitude and longitude) and azimuth information (such as the direction the camera is facing). These images can serve as benchmarks or reference points for subsequently calculating the absolute coordinates of query images at unknown locations.
[0060] The role of localization images is crucial because they provide a relatively fixed, temporally persistent spatial reference frame, which allows the location of query images to be determined by matching features with the localization images even when the environment changes.
[0061] Query images are images that don't carry GPS coordinates but whose locations need to be determined. These images may be captured by vehicles or other mobile devices while in motion. To improve positioning accuracy, multiple query images are typically captured consecutively, leveraging the spatial relationships between them to aid positioning.
[0062] By analyzing and processing the query image and comparing it with the localization image, the common feature points between the two sets of images are identified. This process relies on efficient feature extraction algorithms (such as SIFT and ORB) and calculates geometric transformation models such as the fundamental matrix and the essential matrix to determine the exact position of the query image relative to the localization image.
[0063] In practice, the quality and accuracy of the positioning images must be ensured, including clear image content and accurate GPS coordinates and azimuth angles. Query images must cover as many different viewing angles and distances as possible of the target area to increase the success rate and reliability of feature matching.
[0064] To improve positioning accuracy, other sensor data, such as the direction and acceleration information provided by the on-board nine-axis gyroscope, can also be used in combination to more accurately estimate the actual movement between adjacent query images, thereby optimizing the final positioning results.
[0065] In summary, step S110 involves not only acquiring a positioning image with precise geographic information but also collecting a series of high-quality query images. Together, these two elements form the foundation of a precise vision-based positioning method. The effective execution of this step is directly related to the performance of the entire positioning process.
[0066] S120 , performing feature matching and extraction on each query image and the positioning image respectively, to obtain a feature point set corresponding to each query image.
[0067] In this embodiment, a feature point set refers to a set of significant feature points extracted from the query image and the positioning image using a specific feature extraction algorithm (such as SIFT or ORB). The corresponding feature points between the two images are then found using a feature matching algorithm. These feature points are typically located in areas of the image with unique texture or structure, such as edges and corners.
[0068] For each query image Q and its corresponding localization image A, we first independently extract feature points from each image using a selected feature extraction algorithm (e.g., SIFT, ORB, etc.). These feature points contain not only location information but also descriptors for subsequent feature matching.
[0069] A matching algorithm is used to pair the feature points of the query image Q with the feature points of its corresponding positioning image A. This process typically involves calculating the similarity or distance between the two sets of feature points and selecting the matching pairs that meet certain conditions (such as minimizing the Euclidean distance between descriptors) as the final feature point matching results.
[0070] To improve the accuracy and robustness of the matching results, algorithms such as Random Sample Consensus (RANSAC) can be used to remove incorrectly matched feature point pairs (i.e., outliers). This step is crucial for the subsequent calculation of the fundamental matrix F and other geometric transformation models.
[0071] In actual operation, a threshold M is set to ensure that the number of feature point matches obtained exceeds the threshold. This is because only when a sufficient number of valid feature points are matched can the calculated geometric transformation model (such as the basic matrix F) be guaranteed to have sufficient accuracy and reliability.
[0072] The obtained feature point set This is not only the foundation for constructing the fundamental matrix F but also the prerequisite for subsequently calculating the essential matrix E using the eight-point method or other algorithms. Furthermore, analyzing the feature point set further allows for understanding the spatial relationship between the query image and the positioning image, providing critical data support for accurately calculating the query image's position.
[0073] In summary, in step S120, obtaining the corresponding feature point set for each query image by performing feature matching and extraction on the query image and the positioning image is an essential and important step in the entire visual positioning method based on image feature matching. The quality of this step directly affects the subsequent positioning accuracy.
[0074] S130: Solve the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result.
[0075] In one embodiment, see Figure 2, the above-mentioned step S130 may include steps S131~S132.
[0076] S131, calculating a fundamental matrix for the feature point set corresponding to each query image, constructing an intrinsic parameter matrix in combination with camera parameters, and calculating an essential matrix using the fundamental matrix and the intrinsic parameter matrix; wherein the essential matrix includes relative rotation and translation information;
[0077] In one embodiment, see Figure 3 , the above-mentioned step S131 may include steps S1311~S1313.
[0078] S1311. Randomly select feature points from the feature point set corresponding to each query image, and calculate a basic matrix based on the randomly selected feature points.
[0079] In this embodiment, the basic matrix refers to a matrix representing the epipolar geometric relationship satisfied by the corresponding feature points between the query image and the positioning image, which is calculated by randomly selecting at least 8 pairs of feature points from the matching feature point sets of the query image and the positioning image, and using an algorithm such as the eight-point method.
[0080] Specifically, feature points are randomly selected from the feature point set corresponding to each query image, and the basic matrix is calculated using the eight-point method algorithm based on the randomly selected feature points.
[0081] In step S1311, a set of feature points is first randomly selected from the set of feature points corresponding to each query image. The purpose of this step is to calculate the fundamental matrix F between the two images (the positioning image and the query image) using these feature points. Since some mismatched feature points (i.e., outliers) may exist in practical applications, random sampling can help to more accurately estimate the fundamental matrix. The specific steps are as follows:
[0082] From the matched feature point set Randomly extract several pairs of feature points (usually at least 8 pairs of feature points are required to meet the requirements of the eight-point method algorithm).
[0083] Using the selected feature point pairs, the fundamental matrix FF is calculated using algorithms such as the eight-point method. This process not only helps to remove possible outliers but also improves the robustness of the calculation results.
[0084] S1312. Construct an intrinsic parameter matrix using known camera parameters.
[0085] In this embodiment, the intrinsic parameter matrix refers to a matrix describing the internal geometric characteristics of the camera, which is constructed based on known camera parameters (such as focal length, principal point coordinates, etc.), and is specifically a 3x3 matrix determined according to the camera's focal length and image size.
[0086] Step S1312 involves constructing an intrinsic parameter matrix K based on the known camera parameters. The intrinsic parameter matrix describes the internal geometric properties of the camera, including information such as focal length and principal point coordinates. For a given camera, its intrinsic parameter matrix K can be expressed as: ,in = is the camera's focal length, w and h are the image's pixel width and height, and w / 2 and h / 2 are the coordinates of the image's center. These parameters are typically provided by the camera manufacturer or obtained through calibration.
[0087] S1313: Use the intrinsic parameter matrix to convert the basic matrix into an essential matrix.
[0088] In this embodiment, the intrinsic matrix refers to a matrix that is converted by combining the basic matrix with the intrinsic parameter matrix of the camera, and contains the relative rotation and translation information between the shooting positions of the two pictures. This matrix realizes the conversion from the basic matrix to the matrix containing the actual motion parameters (rotation and translation) through the epipolar geometry constraint, thereby providing a basis for accurately calculating the position of the query picture relative to the positioning picture.
[0089] Finally, in step S1313, the previously obtained fundamental matrix F and the intrinsic parameter matrix K are used to calculate the essential matrix E. The relationship between the essential matrix E and the fundamental matrix F can be expressed by the following formula: Where E contains the relative rotation and translation between the two image locations, specifically E = t × R, where t is the translation vector and R is the rotation matrix. By decomposing the essential matrix, we can extract these two important motion parameters and determine the orientation and displacement of the query image relative to the positioning image.
[0090] In summary, step S131 achieves the goal of accurately locating the query image by calculating the basic matrix, constructing the internal reference matrix, and finally calculating the essential matrix. This series of steps provides a solid foundation for subsequent vision-based positioning technology.
[0091] S132. Solve the pose parameters of the essential matrix corresponding to each query image to obtain a solution result.
[0092] In this embodiment, the decomposition result refers to four possible combinations of rotation matrices R and translation vectors t representing relative camera motion, obtained by performing singular value decomposition (SVD) on the essential matrix E between the query image and the localization image. The correct combination is then selected using geometric constraints as the final solution, accurately describing the spatial relationship between the query image and the localization image. This process provides key parameters for the subsequent calculation of the absolute coordinates of the query image.
[0093] Decompose the essential matrix corresponding to each query image to obtain the relative rotation matrix and translation vector to form the solution.
[0094] In this embodiment, step S132 involves solving the pose parameters of the essential matrix corresponding to each query image to obtain an accurate relative rotation matrix and translation vector to form the final solution. Specifically, this process mainly includes:
[0095] Once the essential matrix E between the query and localization images is obtained through the above method, the next step is to decompose it to extract the relative rotation R and translation t between the two image positions. The relationship between the essential matrix E and the camera motion parameters can be expressed as follows: E = t × R, where t × represents the antisymmetric matrix of the translation vector t, and R is the 3x3 rotation matrix that describes the camera's rotation state.
[0096] There are many methods to decompose the essential matrix E, including but not limited to the singular value decomposition (SVD) method. Using SVD decomposition, the E matrix can be decomposed into the product of three matrices: E=UΣV T Where U and V are orthogonal matrices, and Σ is a diagonal matrix with a specific form (usually with diagonal elements σ, -σ, 0). Based on this decomposition, four possible combinations of (R, t) can be calculated, because the mapping from E to (R, t) is ambiguous.
[0097] According to the results of SVD decomposition, the rotation matrix R and translation vector t can be calculated as follows:
[0098] R=UV T Or R=Udiag(1,1,-1)V T ; t=u3 or t=-u3, where u3 is the last column of the U matrix;
[0099] Since the decomposition process produces two valid sets of (R, t) combinations, additional geometric constraints or prior knowledge are needed to determine which set is correct. This usually involves checking which solution complies with the physical constraints of the actual scene, such as objects cannot be upside down relative to the camera.
[0100] After completing the above steps, the relative rotation matrix R and translation vector t of the query image relative to the positioning image are obtained, which constitute the solution. This result not only provides the spatial positional relationship of the query image Q relative to the positioning image A with known GPS coordinates, but also lays the foundation for further calculation of the absolute coordinates of the query image. Combining the inter-frame motion information provided by the on-board equipment with geometric principles such as the triangle cosine theorem, the absolute position coordinates of the query image can be accurately inferred, thus achieving high-precision application of visual positioning technology.
[0101] In addition, in another embodiment, when a deep learning network is used to solve the pose parameters, the query image is input into a pose solving model to solve the pose parameters to obtain a solution result;
[0102] Among them, the posture solution model is obtained by training a neural network model by using several pictures with position, orientation, and camera parameter information as a sample set; the loss function of the posture solution model includes a loss function that combines relative position and posture errors, and balances the importance of relative position and posture errors through hyperparameters.
[0103] The problem of estimating the relative pose between images is transformed into a regression problem, which uses a neural network to directly extract features from the input images and predict their relative position and orientation. This method does not rely on specific camera parameters and therefore has wider applicability.
[0104] To train such a neural network model, a batch of image pairs or groups containing relative pose information is required. Each group consists of at least two images, and each image is accompanied by information such as its position in space, orientation, and camera parameters. This data forms the basis for training the model, ensuring that it can learn the relative relationships between objects from different perspectives.
[0105] The design of the loss function is crucial for training the model. A specific form of loss function is used here: ; and are the predicted value and true value of the relative position, and They are the predicted value and true value of relative posture (rotation information), It is a hyperparameter used to balance the position dimension and the angle dimension, and to balance the magnitude difference between the position error and the angle error, because usually the units of measurement of position and angle are different, which directly affects their weight distribution in the total loss.
[0106] During training, the model receives paired input images and attempts to minimize the loss function defined above. This means the model must not only learn to identify the relative position changes between the two images, but also accurately capture the rotational relationship between them. By continuously adjusting the network weights to reduce the gap between the predicted and true values, the model gradually learns how to infer relative pose information from the images.
[0107] This approach is particularly suitable for application scenarios where it is difficult to obtain accurate camera parameters. It provides a flexible and efficient way to understand and process visual information in the environment, thereby enhancing the system's adaptability to unknown environments. In addition, since it does not rely on a specific camera model, this technology has a wider range of applications.
[0108] S140: Determine the movement direction and distance corresponding to adjacent query images in the plurality of query images, and calculate the absolute coordinate position of the query image in combination with the GPS coordinates and epipolar geometry of the positioning image.
[0109] In this embodiment, the movement direction refers to the change in the relative angle between the cameras when capturing two adjacent query images (e.g., Q1 and Q2) while the vehicle is moving. This direction can be accurately measured using data from the vehicle's nine-axis gyroscope, including the vehicle's direction of travel and possible rotation angles.
[0110] Distance refers to the distance the vehicle traveled between the capture times of two adjacent query images. This information can also be obtained from the vehicle's sensor data, such as the displacement provided by the odometer or GPS data.
[0111] Absolute coordinate location refers to calculating the exact geographic coordinates of a query image relative to a positioning image using visual positioning technology combined with vehicle sensor data, based on the location information of a known positioning image (a base image with precise GPS coordinates and heading angles). This involves using triangulation principles, combining the epipolar geometry between the query and positioning images, and the spatial displacement between the two query images, to ultimately determine the absolute coordinates of the query image.
[0112] In one embodiment, see Figure 4 , the above-mentioned step S140 may include steps S141~S142.
[0113] S141 , using sensor data from the vehicle-mounted device to determine a moving direction and distance between adjacent query images in a plurality of query images to obtain a direction vector.
[0114] In this embodiment, the direction vector refers to a mathematical expression that represents the direction and magnitude of the vehicle-mounted device's movement from one query image to the next adjacent query image. Specifically, it contains two key elements:
[0115] Direction: describes the direction from the first query image to the second query image, usually expressed as an angle or unit vector.
[0116] Size: The actual distance the vehicle-mounted device moves between two frames of images, reflecting the specific length of the movement.
[0117] This direction vector is calculated based on real-time data from sensors built into the vehicle's equipment (such as a nine-axis gyroscope and accelerometer). It is then used to accurately calculate the absolute coordinates of the query image relative to the positioning image. This allows the query image's spatial position to be accurately calculated, even without direct GPS coverage, based on the relative movement between consecutive query images and information from the known positioning image.
[0118] S142. Calculate the absolute coordinate position of the query image based on the direction vector, the solution result, the GPS coordinates of the positioning image, and the epipolar geometry relationship.
[0119] In this embodiment, the relative direction and distance of the query image relative to the positioning image are calculated based on the direction vector and the solution result and the epipolar geometry relationship using the triangulation principle to determine the absolute coordinate position of the query image.
[0120] In one embodiment, see Figure 5 The above-mentioned step S142 may include steps S1421~S1423.
[0121] S1421. Calculate the relative direction between the adjacent query image and positioning image using the epipolar geometry relationship according to the direction vector and the solution result.
[0122] In this embodiment, the direction vector provided by the vehicle-mounted device (i.e., the direction and distance from Q1 to Q2) is first used in conjunction with the basic principles of epipolar geometry to determine the relative direction of the query images (e.g., Q1 and Q2) relative to the positioning image D. Epipolar geometry is an important concept based on two-view geometry, which describes the camera motion information by the positional relationship of feature matching points in different images. Specifically, here, by matching known feature points ( ) and the basic matrix F, we can get the epipolar line between the query image and the positioning image, and further infer the relative direction between them.
[0123] S1422: Consider the adjacent query image and the positioning image as triangle vertices, and calculate the length between the adjacent query image and the positioning image according to the direction vector based on the cosine theorem to obtain a relative distance.
[0124] In this embodiment, if Figure 8 As shown, the query images Q1, Q2 and the positioning image D are regarded as the three vertices of a triangle, and the distance between these vertices is calculated by the cosine theorem using the direction vector and angle information obtained previously. The specific operation is as follows:
[0125] Known conditions: the distance between Q1 and Q2 (provided by on-board sensors), and the direction information of Q1D and Q2D (calculated by epipolar geometry).
[0126] Apply the cosine theorem: Let Q1D and Q2D be a and b respectively, and the distance between Q1Q2 be c, then the formula To solve the lengths of a and b, where C is the angle between Q1D and Q2D, we can get the exact distance between the query image and the positioning image.
[0127] S1423. Determine the absolute coordinate position of the query image in combination with the relative direction and the relative distance.
[0128] In this embodiment, the final step integrates the results of all previous steps—namely, the relative direction and distance of the query image relative to the positioning image—to determine the absolute coordinate position of the query image. Since the GPS coordinates of positioning image D are known, the geographic coordinates of query images Q1 and Q2 can be precisely calculated based on their coordinates and the relative direction and distance. This process involves simple vector operations, converting relative positions into actual coordinate values in a geographic coordinate system. This method allows precise positioning of any location along the vehicle's travel path, even in the absence of a direct GPS signal.
[0129] The method of this embodiment uses visual positioning technology based on image feature matching, whose accuracy is limited by multiple factors, including the shooting distance of the positioning image (the base image with precise GPS coordinates and orientation angles) and the distance between the query image (the image with unknown GPS coordinates and a location to be determined) and the positioning image. To improve the accuracy of visual positioning, the method of this embodiment utilizes the visual positioning technology of a vehicle-mounted dynamic device. This technology relies on the spatial movement between adjacent image frames Q1 and Q2 when the vehicle-mounted device is in motion, combined with parallax and epipolar geometry calculations with the positioning image, to obtain precise positioning information for the query image.
[0130] In this embodiment, the solution result refers to the direction information of the query image Q relative to the positioning image, which is the direction of the translation vector t obtained by decomposing the essential matrix E. The essential matrix E contains the relative rotation and translation information between the shooting positions of the two images, that is, E=t×R. By decomposing E, the relative pose parameters of the camera motion can be obtained, including the rotation matrix R and the translation vector t. The t direction information here represents the direction information of the query image Q relative to the positioning image. However, due to the lack of information about the actual distance, only the direction can be determined at this time, and the absolute position cannot be directly obtained.
[0131] The direction vector involves using onboard equipment (such as a nine-axis gyroscope) to obtain the direction and displacement between adjacent frames (e.g., Q1 and Q2). This not only obtains the direction vector T but also the exact distance between the two frames. Combining this with epipolar geometry, the relative orientations of Q1 and Q2 relative to the positioning image D can be determined. With this information, combined with the known GPS coordinates of positioning image D, the absolute coordinates of Q1 and Q2 can be calculated using methods such as triangulation.
[0132] The relationship between the solution and the direction vector is that they together form the basis for accurate positioning:
[0133] First, the directional information of the translation vector t obtained by decomposing the essential matrix provides a preliminary basis for establishing the relative orientation relationship between the query image and the positioning image. Then, using the real-world motion data (direction and distance) between consecutive frames provided by the on-board equipment and epipolar geometry, the precise spatial relationship between multiple query images and the positioning image is further determined.
[0134] Finally, by combining the above two types of direction information and using geometric principles such as the triangle cosine theorem, the conversion from relative direction to absolute coordinates is achieved based on the known angle information and distance information, thereby completing the precise positioning of the query image.
[0135] Therefore, these two types of direction information complement each other. The solution result provides basic direction guidance, and the direction vector supplements the actual distance and more detailed direction data, which together contribute to the final precise positioning.
[0136] For example, Figure 6 and Figure 7 The figure shows the matching feature points between two query images Q1 and Q2 and the localization image. The left image is the localization image, and the right image is the query image. Blue lines connect the same image feature points, visually presenting the image correspondences obtained through feature matching. Based on these image correspondences, the relative orientation between the localization image and the query image can be determined based on the epipolar relationship, thereby achieving high-precision visual localization.
[0137] Specifically, for the localization image A and the query image Q1, we first extract and match feature points using algorithms such as SIFT and ORB to obtain a set of corresponding feature points P_matches between the two images. The number of these feature points must exceed a certain threshold M. For example, suppose we use the SIFT algorithm to extract 500 key points from both the localization image A and the query image Q1, and ultimately match 300 common feature points, which meets the set threshold requirement.
[0138] Based on the above matched feature point set , you can use algorithms such as the eight-point method to calculate the basic matrix F between the two images. At the same time, you can also use methods such as random sampling consensus (RANSAC) to We randomly select feature points that meet the algorithm's number requirements to eliminate outliers and improve the robustness of the fundamental matrix calculation. For example, after applying the RANSAC algorithm, we can further filter out 280 reliable matching points from 300 matching points for calculating the fundamental matrix F. This represents the epipolar relationship between corresponding feature points in the two images.
[0139] When the camera parameters such as focal length, aperture, sensor size, etc. are known, the camera intrinsic parameter matrix K can be calculated. For example, assuming a typical camera setting, the focal length in the x and y directions are both = =1.2 mm, the image resolution is w=1920 pixels, h=1080 pixels, then the camera intrinsic parameter matrix K can be expressed as: ;
[0140] Through the camera's intrinsic parameter matrix K, we can get the essential matrix E between the query image and the positioning image, that is, . This matrix contains the relative rotation and translation information between the shooting positions of the two pictures. For example, by decomposing the essential matrix E, the relative posture parameters of the camera motion can be obtained, namely the rotation matrix R and the translation vector t. Assume that the decomposition results are R=[r1, r2, r3] and t=[tx, ty, tz], where r1, r2, r3 are the three column vectors of the rotation matrix, and tx, ty, tz are the components of the translation vector, which provide the direction information of the query image relative to the positioning image.
[0141] Although the specific coordinates of the query image cannot be directly determined based on the directional information, the movement information provided by the onboard device can be used to determine the movement direction and displacement distance between the two frames Q1 and Q2, namely the direction vector T. Combining this information with the principles of epipolar geometry, the absolute coordinates of Q1 and Q2 can be calculated by locating the GPS coordinate points of the image. For example, if the GPS coordinates of the located image D are known, and points Q1, Q2, and D form a triangle, the distance between Q1D and Q2D can be solved using the law of cosines and the known angle information and the distance information between Q1 and Q2, thereby determining the absolute positions of Q1 and Q2.
[0142] It can be seen that in the complex and changing environment of the city, the goal of the method of this embodiment is to accurately identify and filter out those dynamically changing factors in the image, such as moving vehicles, flowing crowds and other short-lived elements, and focus on retaining and matching those long-term stable and unchanging urban buildings and environmental features.
[0143] For example, in a busy city neighborhood, a large number of pedestrians and vehicles come and go every day. These factors can make image-based positioning extremely difficult because they significantly change the content of photos taken at different times at the same location. However, the method of this embodiment can effectively focus on relatively fixed building facades, road layouts, or specific landmarks, thereby greatly improving the reliability and accuracy of visual positioning in such complex environments. This process involves using advanced feature extraction algorithms (such as SIFT and ORB) to select feature points that are not easily changed over time, and calculating epipolar geometry to determine the position of the query image relative to the base image with known GPS coordinates. Ultimately, high-precision positioning can be achieved even in urban environments full of moving objects.
[0144] The above-mentioned visual positioning method for road targets based on urban street scenes obtains a positioning image and multiple query images, and performs feature matching and extraction on each query image and the positioning image to obtain a set of feature points, and then solves the pose parameters of each query image; at the same time, it determines the movement direction and distance between adjacent query images, and uses the GPS coordinates and epipolar geometry relationship of the positioning image to calculate the absolute coordinate position of the query image; this method solves the problem of limited accuracy of traditional visual positioning methods based on image feature matching when facing large shooting intervals or long distances. It enhances positioning accuracy by integrating information from multiple perspectives, ensuring that targets can be accurately identified and positioned even when there are large spatial or temporal gaps between images, effectively improving the stability and reliability of visual positioning in complex urban environments, and improving the positioning error caused by rough GPS information when the density of positioning images is low.
[0145] Figure 9 FIG is a schematic block diagram of a road target visual positioning system 300 based on urban street scenes provided by an embodiment of the present invention. Figure 9 As shown, corresponding to the above-mentioned road target visual positioning method based on urban street scenes, the present invention also provides a road target visual positioning system 300 based on urban street scenes. The road target visual positioning system 300 based on urban street scenes includes a unit for executing the above-mentioned road target visual positioning method based on urban street scenes, and the system can be configured in a server. Specifically, please refer to Figure 8 The road target visual positioning system 300 based on urban street scenes includes an acquisition unit 301, an extraction unit 302, a solution unit 303 and an absolute coordinate calculation unit 304.
[0146] An acquisition unit 301 is used to acquire a positioning image and multiple query images; an extraction unit 302 is used to perform feature matching and extraction on each query image with the positioning image to obtain a set of feature points corresponding to each query image; a solution unit 303 is used to solve the posture parameters based on a deep learning network or a visual algorithm to obtain a solution result; an absolute coordinate calculation unit 304 is used to determine the corresponding moving direction and distance between adjacent query images in multiple query images, and calculate the absolute coordinate position of the query image in combination with the GPS coordinates of the positioning image and the epipolar geometry relationship.
[0147] In one embodiment, the solving unit 303 includes:
[0148] The essential matrix calculation subunit is used to calculate the basic matrix for the feature point set corresponding to each query image, construct the intrinsic parameter matrix in combination with the camera parameters, and calculate the essential matrix using the basic matrix and the intrinsic parameter matrix; the parameter solution subunit is used to solve the posture parameters of the essential matrix corresponding to each query image to obtain the solution result.
[0149] In one embodiment, the essential matrix calculation subunit includes:
[0150] A basic matrix calculation module is used to randomly select feature points from the feature point set corresponding to each query image and calculate the basic matrix based on the randomly selected feature points; an intrinsic parameter matrix construction module is used to construct the intrinsic parameter matrix using known camera parameters; and a conversion module is used to convert the basic matrix into an essential matrix using the intrinsic parameter matrix.
[0151] In one embodiment, the basic matrix calculation module is configured to randomly select feature points from a feature point set corresponding to each query image, and calculate the basic matrix using an eight-point method algorithm based on the randomly selected feature points.
[0152] In one embodiment, the decomposition unit is used to decompose the essential matrix corresponding to each query image to obtain a relative rotation matrix and a translation vector to form a solution result.
[0153] In one embodiment, the absolute coordinate calculation unit 304 includes:
[0154] A direction vector determination subunit is used to use the sensor data of the vehicle-mounted equipment to determine the movement direction and distance between adjacent query images in multiple query images to obtain a direction vector; a position calculation subunit is used to calculate the absolute coordinate position of the query image based on the direction vector combined with the solution result, the GPS coordinates of the positioning image and the epipolar geometric relationship.
[0155] In one embodiment, the position calculation subunit is used to calculate the relative direction and distance of the query image relative to the positioning image based on the epipolar geometric relationship according to the direction vector and the solution result using the triangulation principle to determine the absolute coordinate position of the query image.
[0156] In one embodiment, the position calculation subunit includes:
[0157] A relative direction calculation module is used to calculate the relative direction between the adjacent query images and the positioning images using the epipolar geometry relationship based on the direction vector and the solution result; a relative distance calculation module is used to regard the adjacent query images and the positioning images as triangle vertices, and calculate the length between the adjacent query images and the positioning images based on the cosine theorem according to the direction vector to obtain the relative distance; a coordinate determination module is used to determine the absolute coordinate position of the query image in combination with the relative direction and the relative distance.
[0158] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned road target visual positioning system 300 based on urban street scenes and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.
[0159] The above-mentioned road target visual positioning system 300 based on urban street scenes can be implemented in the form of a computer program. The computer program can be used in Figure 9 Runs on the computer equipment shown.
[0160] See also Figure 10 , Figure 10 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.
[0161] See Figure 10 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0162] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a road target visual positioning method based on urban street scenes.
[0163] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0164] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a road target visual positioning method based on urban street scenes.
[0165] The network interface 505 is used to communicate with other devices through the network. Figure 10 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0166] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:
[0167] Obtain a positioning image and multiple query images; perform feature matching and extraction on each query image with the positioning image to obtain a set of feature points corresponding to each query image; solve the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result; determine the corresponding movement direction and distance between adjacent query images in the multiple query images, and calculate the absolute coordinate position of the query image in combination with the GPS coordinates and epipolar geometry of the positioning image.
[0168] The positioning picture includes a basic picture with GPS coordinates and direction angles.
[0169] In one embodiment, when the processor 502 implements the step of solving the pose parameters based on the deep learning network or the visual algorithm to obtain the solution result, it specifically implements the following steps:
[0170] A basic matrix is calculated for the set of feature points corresponding to each query image, an intrinsic parameter matrix is constructed in combination with camera parameters, and an essential matrix is calculated using the basic matrix and the intrinsic parameter matrix; wherein the essential matrix includes relative rotation and translation information; and the pose parameters of the essential matrix corresponding to each query image are solved to obtain a solution result.
[0171] In one embodiment, when the processor 502 implements the step of randomly selecting feature points from the feature point set corresponding to each query image and calculating the basic matrix based on the randomly selected feature points, the processor 502 specifically implements the following steps:
[0172] Randomly select feature points from the feature point set corresponding to each query image, and calculate the basic matrix based on the randomly selected feature points
[0173] In one embodiment, when the processor 502 implements the step of solving the pose parameters based on the deep learning network or the visual algorithm to obtain the solution result, it specifically implements the following steps:
[0174] When a deep learning network is used to solve the pose parameters, the query image is input into the pose solving model to solve the pose parameters to obtain a solution result;
[0175] Among them, the posture solution model is obtained by training a neural network model by using several pictures with position, orientation, and camera parameter information as a sample set; the loss function of the posture solution model includes a loss function that combines relative position and posture errors, and balances the importance of relative position and posture errors through hyperparameters.
[0176] In one embodiment, the processor 502 implements the following steps when implementing the step of determining the movement direction and distance between adjacent query images in the plurality of query images and calculating the absolute coordinate position of the query image in combination with the GPS coordinates of the positioning image and the epipolar geometry relationship:
[0177] The sensor data of the vehicle-mounted device is used to determine the movement direction and distance between adjacent query images in multiple query images to obtain a direction vector; based on the direction vector combined with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship, the absolute coordinate position of the query image is calculated.
[0178] In one embodiment, when the processor 502 implements the step of calculating the absolute coordinate position of the query image based on the direction vector combined with the solution result, the GPS coordinates of the positioning image, and the epipolar geometry relationship, it specifically implements the following steps:
[0179] According to the direction vector and the solution result, the relative direction and distance of the query image relative to the positioning image are calculated based on the epipolar geometric relationship using the triangulation principle to determine the absolute coordinate position of the query image.
[0180] In one embodiment, when the processor 502 implements the step of calculating the relative direction and distance of the query image relative to the positioning image based on the epipolar geometric relationship according to the direction vector and the solution result to determine the absolute coordinate position of the query image, the processor 502 specifically implements the following steps:
[0181] The relative direction between the adjacent query image and the positioning image is calculated using the epipolar geometry relationship according to the direction vector and the solution result; the adjacent query image and the positioning image are regarded as triangle vertices, and the length between the adjacent query image and the positioning image is calculated based on the cosine theorem according to the direction vector to obtain the relative distance; the absolute coordinate position of the query image is determined in combination with the relative direction and the relative distance.
[0182] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0183] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0184] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:
[0185] Obtain a positioning image and multiple query images; perform feature matching and extraction on each query image with the positioning image to obtain a set of feature points corresponding to each query image; solve the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result; determine the corresponding movement direction and distance between adjacent query images in the multiple query images, and calculate the absolute coordinate position of the query image in combination with the GPS coordinates and epipolar geometry of the positioning image.
[0186] The positioning picture includes a basic picture with GPS coordinates and direction angles.
[0187] The intrinsic matrix includes relative rotation and translation information.
[0188] In one embodiment, when the processor executes the computer program to implement the step of solving the pose parameters based on the deep learning network or the visual algorithm to obtain the solution result, the processor specifically implements the following steps:
[0189] A basic matrix is calculated for the set of feature points corresponding to each query image, an intrinsic parameter matrix is constructed in combination with camera parameters, and an essential matrix is calculated using the basic matrix and the intrinsic parameter matrix; wherein the essential matrix includes relative rotation and translation information; and the pose parameters of the essential matrix corresponding to each query image are solved to obtain a solution result.
[0190] In one embodiment, when the processor executes the computer program to implement the step of randomly selecting feature points from the set of feature points corresponding to each query image and calculating the basic matrix based on the randomly selected feature points, the processor specifically implements the following steps:
[0191] Feature points are randomly selected from the feature point set corresponding to each query image, and a basic matrix is calculated using an eight-point method algorithm based on the randomly selected feature points.
[0192] In one embodiment, when the processor executes the computer program to implement the step of solving the pose parameters based on the deep learning network or the visual algorithm to obtain the solution result, the processor specifically implements the following steps:
[0193] When a deep learning network is used to solve the pose parameters, the query image is input into the pose solving model to solve the pose parameters to obtain a solution result;
[0194] Among them, the posture solution model is obtained by training a neural network model by using several pictures with position, orientation, and camera parameter information as a sample set; the loss function of the posture solution model includes a loss function that combines relative position and posture errors, and balances the importance of relative position and posture errors through hyperparameters.
[0195] In one embodiment, when the processor executes the computer program to implement the step of determining the movement direction and distance corresponding to adjacent query images among the plurality of query images, and calculating the absolute coordinate position of the query image in combination with the GPS coordinates of the positioning image and the epipolar geometry relationship, the processor specifically implements the following steps:
[0196] The sensor data of the vehicle-mounted device is used to determine the movement direction and distance between adjacent query images in multiple query images to obtain a direction vector; based on the direction vector combined with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship, the absolute coordinate position of the query image is calculated.
[0197] In one embodiment, when the processor executes the computer program to implement the step of calculating the absolute coordinate position of the query image based on the direction vector combined with the solution result, the GPS coordinates of the positioning image, and the epipolar geometry relationship, the processor specifically implements the following steps:
[0198] According to the direction vector and the solution result, the relative direction and distance of the query image relative to the positioning image are calculated based on the epipolar geometric relationship using the triangulation principle to determine the absolute coordinate position of the query image.
[0199] In one embodiment, when the processor executes the computer program to implement the step of calculating the relative direction and distance of the query image relative to the positioning image based on the epipolar geometric relationship using the triangulation principle according to the direction vector and the solution result to determine the absolute coordinate position of the query image, the processor specifically implements the following steps:
[0200] The relative direction between the adjacent query image and the positioning image is calculated using the epipolar geometry relationship according to the direction vector and the solution result; the adjacent query image and the positioning image are regarded as triangle vertices, and the length between the adjacent query image and the positioning image is calculated based on the cosine theorem according to the direction vector to obtain the relative distance; the absolute coordinate position of the query image is determined in combination with the relative direction and the relative distance.
[0201] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0202] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0203] In the several embodiments provided herein, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0204] The steps in the method of the embodiment of the present invention may be adjusted in order, combined, or deleted as needed. The units in the system of the embodiment of the present invention may be combined, divided, or deleted as needed. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0205] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (such as a personal computer, terminal, or network device) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0206] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A road target visual positioning method based on urban street scenes, characterized in that: include: Get the positioning image and multiple query images; Perform feature matching and extraction on each query image and the positioning image to obtain a set of feature points corresponding to each query image; Solving the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result; the solution result is the relative pose between the query image and the positioning image; Determine the movement direction and distance between adjacent query images in the plurality of query images, and calculate the absolute coordinate position of the query image by combining the solution result, the GPS coordinates of the positioning image, and the epipolar geometry relationship; Determining the movement direction and distance between adjacent query images in the plurality of query images, and calculating the absolute coordinate position of the query image by combining the solution result, the GPS coordinates of the positioning image, and the epipolar geometry, including: Calculating the relative direction between the adjacent query images and the positioning image using the epipolar geometry relationship according to the direction vector and the solution result; wherein the direction vector refers to the movement direction and distance between adjacent query images in the plurality of query images; Considering the adjacent query image and the positioning image as triangle vertices, and calculating the length between the adjacent query image and the positioning image according to the direction vector based on the cosine theorem to obtain a relative distance; The absolute coordinate position of the query image is determined in combination with the relative direction and the relative distance.
2. The method for visually locating road objects based on urban street scenes according to claim 1, characterized in that: The positioning picture includes a basic picture with GPS coordinates and direction angles.
3. The method for visually locating road objects based on urban street scenes according to claim 1, characterized in that: Solving the pose parameters based on the deep learning network or visual algorithm to obtain a solution result includes: Calculating a fundamental matrix between the positioning image and the query image based on the set of feature points corresponding to each query image, constructing an intrinsic parameter matrix in combination with camera parameters, and calculating an intrinsic matrix using the fundamental matrix and the intrinsic parameter matrix; wherein the intrinsic matrix includes relative rotation and translation information; The pose parameters of the essential matrix corresponding to each query image are solved to obtain the solution.
4. The method for visually locating road objects based on urban street scenes according to claim 3, characterized in that: The step of calculating a basic matrix for a set of feature points corresponding to each query image, constructing an intrinsic parameter matrix in combination with camera parameters, and calculating an essential matrix using the basic matrix and the intrinsic parameter matrix includes: Randomly selecting feature points from the feature point set corresponding to each query image, and calculating a basic matrix based on the randomly selected feature points; Construct the intrinsic parameter matrix using known camera parameters; The fundamental matrix is converted into an intrinsic matrix using the intrinsic parameter matrix.
5. The method for visually locating road objects based on urban street scenes according to claim 4, characterized in that: The randomly selecting feature points from the feature point set corresponding to each query image, and calculating a basic matrix based on the randomly selected feature points, includes: Feature points are randomly selected from the feature point set corresponding to each query image, and a basic matrix is calculated using an eight-point method algorithm based on the randomly selected feature points.
6. The method for visually locating road objects based on urban street scenes according to claim 1, characterized in that: Solving the pose parameters based on the deep learning network or visual algorithm to obtain a solution result includes: When a deep learning network is used to solve the pose parameters, the query image and the positioning image are input into the pose solving model to solve the pose parameters to obtain a solution result; The pose solution model is obtained by training a neural network model using a number of images with position, orientation, and camera parameter information as a sample set, and the sample set includes a group of images formed by the query image and the positioning image; the loss function of the pose solution model includes a loss function that combines relative position and pose errors, and balances the importance of relative position and pose errors through hyperparameters.
7. The method for visually locating road objects based on urban street scenes according to claim 1, characterized in that: The process of determining the direction vector includes: The sensor data of the vehicle-mounted device is used to determine the movement direction and distance between adjacent query images in multiple query images to obtain a direction vector.
8. A road target visual positioning system based on urban street scenes, characterized by: The system uses the road target visual positioning method based on urban street scenes as claimed in claim 1, including: An acquisition unit, used to acquire a positioning image and multiple query images; An extraction unit, configured to perform feature matching and extraction on each query image and the positioning image, respectively, to obtain a set of feature points corresponding to each query image; A solving unit is used to solve the pose parameters based on a deep learning network or a visual algorithm to obtain a solution result; The absolute coordinate calculation unit is used to determine the corresponding moving direction and distance between adjacent query images in multiple query images, and calculate the absolute coordinate position of the query image in combination with the solution result, the GPS coordinates of the positioning image and the epipolar geometry relationship.
Citation Information
Patent Citations
Unmanned aerial vehicle autonomous positioning method and system based on monocular vision inertial navigation fusion
CN114693754A
Computer vision positioning method and device
TWI746417B