A Visual Localization Method and System Based on Scene Recognition and Deep Learning

By adopting visual positioning methods based on scene recognition and deep learning in robot navigation, the problem of traditional visual positioning technology degradation in complex environments is solved, and higher positioning robustness and accuracy are achieved.

CN118840529BActive Publication Date: 2025-06-17ADVANCED TECH RES INST OF BEIJING UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411310480.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-06-17
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

The visual positioning technology in existing robot navigation is not performing well in complex environments, mainly because the traditional method relies on the matching of local feature points, is susceptible to the influence of dynamic environment and lighting changes, and lacks dynamic adjustment and adaptive update mechanisms, resulting in a decrease in positioning accuracy.

Method used

The visual positioning method based on scene recognition and deep learning is adopted. By obtaining the image information and displacement rotation information of the robot environment, scene recognition is performed, matching image feature points are obtained, initial position pose is calculated, and deep learning model is input for correction, and the model parameters are adaptively updated to improve positioning accuracy.

Benefits of technology

The robustness and accuracy of visual positioning are improved, especially in the case of dynamic environments, lighting changes and occlusion. Through the semantic information provided by scene recognition, the system can more effectively select and match feature points, enhancing the stability and accuracy of positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118840529B_ABST
    Figure CN118840529B_ABST
Patent Text Reader

Abstract

The present invention provides a visual positioning method and system based on scene recognition and deep learning, belonging to the technical fields of computer vision and deep learning, including: acquiring image information of the environment where the robot is located and the displacement and rotation information of the robot; inputting the acquired information into a deep learning model for scene recognition to obtain the recognized scene image; obtaining the well-matched image feature points in the recognized scene image; selecting a pair of images from the feature matching results for the well-matched feature points, and calculating the initial pose of the camera on the robot; inputting the image feature points and the initial pose into the deep learning model to output the corrected pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and deep learning, and particularly relates to a vision positioning method and system based on scene recognition and deep learning. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Currently, robot navigation mainly adopts visual slam technology, visual and lidar fusion positioning technology, visual and integrated navigation fusion positioning technology, etc.

[0004] Visual slam positioning such as ORB-SLAM, PTAM, etc. realizes positioning by extracting and matching feature points in images. Feature extraction mainly uses algorithms such as ORB, SIFT, SURF, etc., and these algorithms are very sensitive to illumination, dynamic environment, and feature point richness, and perform poorly in environments with insufficient illumination and scarce features.

[0005] Visual and lidar fusion positioning technology uses the environmental point cloud data obtained by lidar to assist visual positioning, but it performs poorly in objects with low reflectivity (such as glass, mirrors) and complex environments (such as fog, smoke).

[0006] Visual and integrated navigation fusion positioning technology, in an environment where satellite navigation signals are blocked (such as tunnels, dense forests, buildings), GPS cannot provide reliable positioning, which will cause the positioning errors of the IMU and visual sensors to accumulate over time, and the positioning accuracy will decrease after long-term use.

[0007] Through analysis, the current visual positioning technology in robot navigation performs poorly in complex environments, such as dynamic environments, illumination changes, and occlusion situations. The main reasons are as follows:

[0008] Traditional visual positioning methods mainly rely on the matching of local feature points and are easily affected by dynamic objects and environmental changes, resulting in a decrease in positioning accuracy and robustness.

[0009] However, the current scene recognition is mainly based on the results of deep learning. The problems existing in the above solutions are:

[0010] Existing scene recognition and positioning systems usually use predefined models and parameters for scene recognition and positioning, lack a dynamic adjustment and adaptive update mechanism, and are difficult to adaptively update according to environmental changes and new data, resulting in a gradual decrease in the accuracy of the system during long-term use. Summary of the Invention

[0011] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a visual positioning method based on scene recognition and deep learning, which can improve the positioning accuracy.

[0012] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0013] In the first aspect, a visual positioning method based on scene recognition and deep learning is disclosed, including:

[0014] Obtain image information in the environment where the robot is located and the displacement and rotation information of the robot;

[0015] Perform scene recognition on the obtained information to obtain the recognized scene image;

[0016] In the recognized scene image, obtain the well-matched image feature points;

[0017] For the well-matched feature points, select a pair of images from the feature matching results, and calculate the initial pose of the camera on the robot;

[0018] Input the image feature points and the initial pose into the deep learning model, output the corrected pose, and perform visual positioning based on the corrected pose;

[0019] Among them, the deep learning model continuously adjusts parameters according to the collected information and adaptively updates the deep learning model.

[0020] As a further technical solution, the deep learning model is a trained ResNet deep learning model. When the ResNet deep learning model is trained:

[0021] Obtain multi-scene image information, input it into the ResNet deep learning model, adjust the output dimension of the last fully connected layer to match the number of scene categories, use the backpropagation algorithm and the optimizer to train the model, and adjust the model parameters to minimize the classification error.

[0022] As a further technical solution, when inputting the obtained information into the deep learning model for scene recognition, match the recognition result with the reference scene in the database, determine the current scene position, and add annotation information.

[0023] As a further technical solution, in the recognized scene image, to obtain the well-matched image feature points, the ORB method is specifically used to extract the feature points.

[0024] As a further technical solution, the PnP algorithm is used to estimate the camera pose when calculating the initial pose of the camera on the robot.

[0025] PnP, corresponding to Perspective-n-Point in English, Chinese: It is a method for solving the correspondence between 3D and 2D points.

[0026] As a further technical solution, training the ResNet deep learning model includes:

[0027] Input an image dataset containing true pose labels, where each image should have corresponding 3D points, their 2D projection points in the image, and the pose of the initial PnP estimate;

[0028] The model formula is expressed as follows:

[0029] ;

[0030] ;

[0031] Where: is the input image, is the initial pose estimate, is the corrected pose, , , , are the weights and biases of the neural network.

[0032] In the second aspect, a visual positioning system based on scene recognition and deep learning is disclosed, including:

[0033] An information acquisition module, configured to: acquire image information and the displacement and rotation information of the robot in the environment where the robot is located;

[0034] A scene image recognition module, configured to: perform scene recognition on the acquired information to obtain the recognized scene image;

[0035] An image feature point acquisition module, configured to: obtain well-matched image feature points in the recognized scene image;

[0036] An initial pose calculation module, configured to: select a pair of images from the feature matching results for the well-matched feature points and calculate the initial pose of the camera on the robot;

[0037] A pose determination module, configured to: input the image feature points and the initial pose into the neural network model, output the corrected pose, and perform visual positioning based on the corrected pose;

[0038] Among them, the deep learning model continuously adjusts parameters according to the acquired information and adaptively updates the deep learning model.

[0039] The above one or more technical solutions have the following beneficial effects:

[0040] The technical solution of the present invention uses a pre-trained ResNet deep learning model for scene recognition. By matching with reference images in the database, the current scene is recognized, thereby improving the robustness and accuracy of positioning. Scene recognition can help determine the specific location or area where the camera is located and provide context information. This is particularly useful in complex environments, such as dynamic environments, lighting changes, and occlusion situations. Through the semantic information provided by scene recognition, the visual positioning system can more effectively select and match feature points, improving the positioning accuracy.

[0041] According to environmental changes and new data, i.e., the acquired image data, the deep learning model continuously adjusts its parameters according to the newly acquired data, optimizes the model, adaptively updates the parameters of the scene recognition model and the feature extraction algorithm, including position and pose information, continuously optimizes the ResNet deep learning model, performs PnP initial pose calculation, corrects the pose of the deep learning model, and conducts BA optimization for positioning to improve the positioning accuracy.

[0042] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0044] Figure 1 It is a data processing flow chart of an embodiment of the present invention;

[0045] Figure 2 It is a ResNet model structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0049] Embodiment 1

[0050] With the development of deep learning technology, scene recognition and feature extraction methods based on deep learning can significantly improve the performance of visual positioning systems. This embodiment discloses a visual positioning method based on scene recognition and deep learning. See Appendix Figure 1 as shown, including:

[0051] Step 1: Dataset collection.

[0052] In this embodiment, two high-resolution cameras and a high-precision inertial navigation device IMU are used to capture environmental image information and displacement and rotation information. The data of the camera device is connected to the built-in camera data acquisition card through the GSML2 interface. The camera data acquisition card is connected to the computer through PCI. The computer processes the received data and stores the trained model. The IMU device is connected to the computer through the USB interface. The camera and IMU devices are started. The computer reads the camera and IMU data, performs denoising processing on the image data using Gaussian filtering, normalizes the image, and adjusts the image size to meet the input requirements of the neural network model.

[0053] Step 2: Use the ResNet deep learning model for scene recognition.

[0054] The model is pre-trained, multi-scene image information is collected and input into the ResNet deep learning model. The output dimension of the last fully connected layer is adjusted to match the number of scene categories. The model is trained using the backpropagation algorithm and optimizer, and the model parameters are adjusted to minimize the classification error. In this embodiment, the multi-scene image information includes image information in scenarios such as urban streets, buildings, traffic lights, vehicles, and pedestrians.

[0055] The image information preprocessed in Step 1 is input into the ResNet deep learning model for scene recognition. The recognition result is matched with the reference scene in the database to determine the current scene location, and annotation information is added. By storing the recognition result and location annotation information of the current scene in the database, location information is provided for the subsequently recognized scenes to assist in improving the positioning accuracy.

[0056] The above database is used to store the reference information required in the scene recognition process. It has an input-output relationship with the ResNet deep learning model. After the ResNet deep learning model recognizes the scene category of the current image, the result is sent to the database for matching and verification. The recognition result of the ResNet deep learning model can be used as feedback information to help the database update its content. Among them, when the recognized scene type matches the scene in the database, since the database stores the type, description, and location information of the scene, etc., the current scene location can be determined.

[0057] In one embodiment, the structure of the above ResNet deep learning model is shown in AppendixFigure 2 As shown. When specifically performing data processing, the preprocessed image is input into the input layer of the ResNet deep learning model. The model calculates through multiple convolutional layers, residual blocks, and pooling layers to extract the eigenvalue of the input image. Through the final fully connected layer, the probability distribution of each scenario is output. According to the probability distribution output by the model, the category with the highest probability is selected as the final scenario recognition result, and the recognition result is recorded in the database. In the training phase, the performance of the model is measured by calculating the loss between the prediction result and the actual label. The gradient is calculated through the backpropagation algorithm, and the model weights are updated using the Adam optimization algorithm to enable the model to gradually learn and improve accuracy.

[0058] Step 3: Extract feature points.

[0059] In the recognized scene image, local feature points are extracted. In this embodiment, the traditional ORB (Oriented FAST and Rotated BRIEF) method is used to extract feature points. Local feature points are points in the image with significant local features. By matching and pose estimating these feature points, the displacement and rotation information of the image can be obtained.

[0060] The ORB feature point consists of two parts: a key point and a descriptor, which adds directionality to the FAST (Features from Accelerated Segment Test) corner point. The process of extracting FAST corner points is as follows: A pixel point p extracted from the scene image is compared with 16 other pixel points within a circle with a radius of 3 around it. If the brightness of this pixel differs from the brightness of 12 of them by more than a threshold, and the threshold is specifically 20% of, then this pixel point will be selected as a FAST corner point.

[0061] The FAST corner point lacks rotational invariance and scale invariance. ORB performs FAST corner point detection on images with different resolutions on each layer by constructing a pyramid and adds scale description. For the problem of rotational invariance, the gray centroid method is used to solve it. The specific steps are as follows:

[0062] For a small image block B extracted from the scene image, define the moment of this image block as follows:

[0063] ,

[0064] where x and y represent the image pixel coordinates, p and q represent the order of the moment, is the gray value of x and y, , is the weight, is the weighted pixel grayscale value.

[0065] The centroid C of the image block is obtained through moments as follows:

[0066]

[0067] where, is the weighted sum of the x coordinates in the image block, is the weighted sum of the y coordinates in the image block, is the total intensity of the image block area.

[0068] The direction of the feature point is defined as the direction from the geometric center O to the centroid C of the image block:

[0069]

[0070] represents the direction of the feature point.

[0071] By the above method, the FAST corner points are described with scale and rotation, thus greatly improving the robustness of representation between different images.

[0072] The descriptor of the ORB feature point adopts the BRIEF descriptor and is combined with the key point direction information calculated in the FAST corner point extraction stage, making this descriptor also have good rotational invariance. The original BRIEF descriptor is a high-dimensional vector composed of many 0s and 1s, where 0 and 1 represent the size relationship between two random pixels near the key point. The comparison process of this descriptor is fast and intuitive, and it is convenient to store using binary representation, so it is suitable for real-time image matching.

[0073] The extracted feature points are feature-matched using the Hamming distance.

[0074] In an embodiment, for two binary feature descriptors and , their Hamming distance d can be calculated by the following formula:

[0075]

[0076] where: and are two feature descriptors, represents the exclusive OR operation, and n is the length of the descriptor. , represent two different i-th feature descriptors.

[0077] Step Four: Pose Estimation.

[0078] In this embodiment, the traditional PNP algorithm is used to estimate the camera pose. In step 3, the well-matched feature points have been obtained. Select a pair of images from the feature matching results and calculate the camera pose. The camera pose includes: R represents the rotation matrix, and t represents the translation vector.

[0079] Suppose the coordinates of a certain spatial point are , representing the spatial point coordinates, and its projected pixel coordinates are , being the x-axis coordinate, being the y-axis coordinate. The relationship between the pixel position and the spatial point position is as follows:

[0080]

[0081] where K represents the internal parameter matrix of the camera, , being the scale factor of a certain spatial point i. , are the focal lengths of the camera in the x and y directions respectively, , being the camera optical center.

[0082] Represented by the Lie group T , the above formula is written in matrix form as:

[0083]

[0084] representing the spatial point coordinates.

[0085] Solve PNP based on minimizing the reprojection error. Define the reprojection error as:

[0086]

[0087] where, is the actual 2D point, is the projected 2D point, and its calculation formula is:

[0088]

[0089] where, , , , , , , , are the numbers in the rotation matrix respectively, where is the in the internal parameter matrix K of the camera, is in the internal parameter matrix K of the camera , is in the internal parameter matrix K of the camera , is in the internal parameter matrix K of the camera , 、 、 are numbers in the translation vector, and s is the scale factor.

[0090] The total reprojection error is the sum of the errors of all points:

[0091]

[0092] represents the number of points.

[0093] Convert the rotation matrix R to Rodrigues parameters r for optimization. The Rodrigues rotation formula represents the rotation matrix R as a vector in axis-angle form, and the Rodrigues parameter r is the product of the angle rotated about the unit vector axis .

[0094]

[0095] where represents the vector representing the rotation matrix R in axis-angle form.

[0096] The Jacobian matrix J is the partial derivative of the reprojection error with respect to the optimization variables:

[0097]

[0098] where e represents the reprojection error, r represents the Rodrigues parameters, t represents the translation vector, are the partial derivatives of the reprojection error with respect to the Rodrigues parameters and the translation vector.

[0099] Use the Levenberg-Marquardt algorithm to update the optimization variables:

[0100]

[0101] where, is the damping factor, I is the identity matrix, e is the current reprojection error, r represents the Rodrigues parameters, where k + 1 means the iteration number, and J is the Jacobian matrix.

[0102] Continuously iterate to update the optimization variables until convergence, that is, the reprojection error no longer decreases significantly. After optimization, convert the Rodrigues parameters r back to the rotation matrix R, and output the final rotation matrix R and the translation vector t.

[0103] In this embodiment, through continuous iterative optimization, the reprojection error is gradually reduced, so that the final rotation matrix R and translation vector t reach high precision. Especially when dealing with non-linear optimization problems, this iterative method can effectively approximate the global optimal solution.

[0104] Step Five: Use the ResNet deep learning model to correct the pose estimation.

[0105] Load the pre-trained ResNet deep learning model, and input an image dataset containing real pose labels. Each image should have corresponding 3D coordinate points, their 2D projection points in the image, and the pose of the initial PnP estimation. Input the extracted image features and the initial pose, i.e., the extracted feature points and the calculated camera pose, into the model, and the model outputs the corrected pose.

[0106] Method for obtaining 2D projection points: Use a calibration board to take images at different angles, and the opencv tool can be used to calculate the 2D projection points. The initial PNP pose is the result calculated in Step Four.

[0107] The model formula is expressed as follows:

[0108] (Connecting image features and initial pose)

[0109] (Corrected pose)

[0110] Where: is the input image, is the initial pose estimation, is the corrected pose, , , , are the weights and biases of the neural network.

[0111] Define the loss function, and use the mean squared error loss (MSE Loss) to measure the difference between the corrected pose and the real pose:

[0112]

[0113] Where, represents the mean squared error loss, and are the real poses of the i-th sample, and are the corrected poses.

[0114] Initialize the model parameters and the optimizer, and perform iterative training on the training set. Update the model parameters through backpropagation to minimize the loss function. The training process is as follows:

[0115]

[0116] Among them, are model parameters, is the learning rate, represents the mean squared error loss.

[0117] Load the pre-trained model, input the image and the initial PNP estimated pose, and obtain the corrected pose estimate:

[0118]

[0119] Among them, is the input image, is the initial pose estimate, is the corrected pose estimate, is the deep learning network model.

[0120] Step Six: BA graph optimization.

[0121] Finally, use the BA graph optimization method to optimize the global pose.

[0122] Through the above steps one to six, the pose can be obtained. The initial pose is obtained through preliminary PnP calculation. Step Five performs a correction process based on the initial pose, and this Step Six performs a global optimization process on the above pose.

[0123] Use the variable x to represent the camera pose and the landmark to be optimized:

[0124]

[0125]

[0126] Among them, represents the pose variable. represents the increment of each pose, the reprojection error increment.

[0127] Then construct . Its meaning is: starting from a certain initial value, continuously find the descent direction to find the optimal solution of the objective function, and continuously solve the increment in the increment equation . That is, find a to minimize the reprojection error corresponding to each ; find a to minimize the reprojection error corresponding to each .

[0128] For the error function e(x + ) Perform linearization:

[0129] m

[0130] Among them, represents the partial derivative of the camera pose in the current state, represents the partial derivative of the camera position in the current state, the m - reprojection error, represents the rotation increment, represents the error.

[0131] The BA optimization considers not a single pose point and landmark point, but m poses and n spatial coordinates, that is, summing the above formula. So here we re - define the error function:

[0132]

[0133] Among them, represents the partial derivative of the entire cost function with respect to the camera pose in the current state, represents the partial derivative of this function with respect to the landmark point position, ) represents the error function, represents the number of poses, represents the number of spatial coordinates, represents the rotation increment.

[0134] To simplify the calculation, now define:

[0135]

[0136]

[0137]

[0138] E, F, J are Jacobian matrices.

[0139] Then the error function can be modified to

[0140]

[0141] Expanding gives:

[0142]

[0143] Taking the derivative of the above formula and setting it equal to 0, we get the gradient of the error function:

[0144]

[0145] x = 0

[0146] x = -e

[0147] Let

[0148]

[0149]

[0150] Therefore, the gradient of the error function is as follows:

[0151]

[0152] At this time, only by iteratively updating and calculating the gradient according to the above formula for multiple rounds can we obtain the optimal solution that minimizes the error function with the smallest value.

[0153] Matrices E, F, and J represent the partial derivatives of the error function with respect to the camera pose and the positions of the landmark points. By constructing the gradient expression of the error function and taking the derivative to obtain the gradient, this gradient indicates how the current parameters should be updated in the parameter space to better reduce the error. The final projection error is output, and the translation vector t and rotation matrix R are updated according to the projection error.

[0154] After the above steps, according to environmental changes and new data, i.e., the acquired image data, the deep learning model continuously adjusts the parameters according to the newly acquired data, optimizes the model, adaptively updates the parameters of the scene recognition model and feature extraction algorithm, including position and pose information, continuously optimizes the neural network model, calculates the initial pose of PnP, corrects the pose of the deep learning model, and performs BA optimization for positioning to improve the positioning accuracy.

[0155] The technical solution of this embodiment utilizes scene recognition combined with visual positioning: specifically, it uses the pre-trained deep learning model ResNet for scene recognition, and identifies the current scene by matching with the reference images in the database, thereby improving the robustness and accuracy of positioning. Scene recognition can help determine the specific location or area where the camera is located and provide context information. This is particularly useful in complex environments, such as dynamic environments, lighting changes, and occlusion situations. Through the semantic information provided by scene recognition, the visual positioning system can more effectively select and match feature points to improve the positioning accuracy.

[0156] In this embodiment, the context information refers to the additional background information that can help determine the environment where the camera is located, including not only visual features and objects, but also scene categories such as streets, campuses, indoors, etc., feature objects in the scene (such as trees, cars), lighting and weather conditions, geographical information, etc.

[0157] In this embodiment, semantic information helps the system understand the actual use or function of the scene, so as to further accurately locate. By identifying the object information in the image, the system can determine its location.

[0158] The technical solution of this embodiment combines multi-view geometric positioning with deep learning for auxiliary correction: The PnP algorithm may be affected by noise and lack of matches during the initial estimation process, resulting in errors. The deep learning model can use more context information and data-driven methods to correct the initial pose and reduce errors. Use the traditional PnP algorithm for initial pose estimation, and improve the positioning accuracy by minimizing the reprojection error. Introduce a deep learning model to correct the PNP pose estimation result, further improving the positioning accuracy and robustness.

[0159] The technical solution of this embodiment includes adaptive update and optimization: According to environmental changes and new data, adaptively update the parameters of the scene recognition model and feature extraction algorithm, and use the BA graph optimization algorithm to optimize the global pose, further improving the positioning accuracy.

[0160] Embodiment 2

[0161] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the above method.

[0162] Embodiment 3

[0163] The purpose of this embodiment is to provide a computer-readable storage medium.

[0164] A computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it executes the steps of the above method.

[0165] Embodiment 4

[0166] The purpose of this embodiment is to provide a visual positioning system based on scene recognition and deep learning, including:

[0167] An information acquisition module, configured to: acquire image information of the environment where the robot is located and the displacement and rotation information of the robot;

[0168] A scene image recognition module, configured to: perform scene recognition on the acquired information to obtain the recognized scene image;

[0169] An image feature point acquisition module, configured to: obtain well-matched image feature points in the recognized scene image;

[0170] An initial pose calculation module, configured to: for the matched feature points, select a pair of images from the feature matching results, and calculate the initial pose of the camera on the robot;

[0171] A pose determination module, configured to: input the image feature points and the initial pose into a deep learning model, and output the corrected pose.

[0172] Embodiment 5

[0173] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any one of the above embodiments.

[0174] Each step involved in the device of the above embodiments corresponds to the first method embodiment. For the specific implementation, please refer to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0175] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps of them can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0176] Although the specific implementation of the present invention has been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A visual positioning method based on scene recognition and deep learning, characterized by: include: Obtain image information of the robot's environment and the robot's displacement and rotation information; Perform scene recognition based on the acquired information to obtain a recognized scene image; Use the ResNet deep learning model for scene recognition; Obtain matched image feature points in the recognized scene image; For the matched feature points, select a pair of images from the feature matching results and calculate the initial pose of the camera on the robot; Input the image feature points and initial pose into the deep learning model, output the corrected pose, and perform visual positioning based on the corrected pose; Among them, the deep learning model continuously adjusts parameters according to the information collected and updated the deep learning model adaptively; When calculating the initial pose of the camera on the robot, the PnP algorithm is used to estimate the camera pose; The ResNet deep learning model formula is as follows: ; ; in: is the input image, is the initial pose estimate, is the corrected pose, , , , are the weights and biases of the neural network; The mean square error loss is used to measure the difference between the corrected pose and the true pose: ; in, represents the mean square error loss, and is the true pose of the i-th sample, and is the corrected pose; The BA graph optimization method is used to optimize the global pose; Redefine the error function: ; in, Represents the partial derivative of the entire cost function with respect to the camera posture in the current state, represents the partial derivative of the function with respect to the landmark position, ) represents the error function, Indicates the number of poses, Represents the number of spatial coordinates, Indicates the rotation increment.

2. The visual positioning method based on scene recognition and deep learning as claimed in claim 1, characterized in that: The deep learning model is a trained ResNet deep learning model. During training, the ResNet deep learning model: Obtain multi-scene image information, input it into the ResNet deep learning model, adjust the output dimension of the last fully connected layer to match the number of scene categories, train the model using the backpropagation algorithm and optimizer, and adjust the model parameters to minimize the classification error.

3. The visual positioning method based on scene recognition and deep learning as claimed in claim 1, characterized in that: When the acquired information is input into the deep learning model for scene recognition, the recognition results are matched with the reference scenes in the database to determine the current scene location and add annotation information.

4. The visual positioning method based on scene recognition and deep learning as claimed in claim 1, characterized in that: For the recognized scene image, matched image feature points are obtained, and the ORB method is specifically used to extract the feature points.

5. The visual positioning method based on scene recognition and deep learning as claimed in claim 2, characterized in that: Training the ResNet deep learning model includes: Input an image dataset containing real pose labels. Each image should have a corresponding 3D point and its 2D projection point in the image, as well as the initial PnP estimated pose.

6. The visual positioning method based on scene recognition and deep learning as claimed in claim 2, characterized in that: After obtaining the image information of the robot's environment and the displacement and rotation information of the robot, preprocessing is performed and the preprocessed image is input into the input layer of the ResNet deep learning model. The model calculates through multiple convolutional layers, residual blocks, and pooling layers to extract the feature values ​​of the input image, and outputs the probability distribution of each scene through the final fully connected layer.

7. The visual positioning method based on scene recognition and deep learning as claimed in claim 6, characterized in that: According to the output probability distribution, the category with the highest probability is selected as the final scene recognition result, and the recognition result is recorded in the database.

8. A visual positioning system based on scene recognition and deep learning, characterized by: include: The information acquisition module is configured to: acquire image information of the environment in which the robot is located and displacement and rotation information of the robot; The scene image recognition module is configured to: perform scene recognition on the acquired information to obtain a recognized scene image; Use the ResNet deep learning model for scene recognition; The image feature point acquisition module is configured to: obtain matched image feature points in the recognized scene image; The initial pose calculation module is configured to: for the matched feature points, select a pair of images from the feature matching results, and calculate the initial pose of the camera on the robot; The posture determination module is configured to: input the image feature points and the initial posture into the deep learning model, output the corrected posture, and perform visual positioning based on the corrected posture; Among them, the deep learning model continuously adjusts parameters according to the information collected and updated the deep learning model adaptively; When calculating the initial pose of the camera on the robot, the PnP algorithm is used to estimate the camera pose; The ResNet deep learning model formula is as follows: ; ; in: is the input image, is the initial pose estimate, is the corrected pose, , , , are the weights and biases of the neural network; The mean square error loss is used to measure the difference between the corrected pose and the true pose: ; in, represents the mean square error loss, and is the true pose of the i-th sample, and is the corrected pose; The BA graph optimization method is used to optimize the global pose; Redefine the error function: ; in, Represents the partial derivative of the entire cost function with respect to the camera posture in the current state, represents the partial derivative of the function with respect to the landmark position, ) represents the error function, Indicates the number of poses, Represents the number of spatial coordinates, Indicates the rotation increment.

Citation Information

Patent Citations

  • Pose information estimation method and apparatus, and mobile device

    CN106780608A

  • Location recognition and relative positioning method and system adapting to changes of visual characteristics

    CN107967457A

  • Indoor visual repositioning method and system

    CN111144349A

  • Indoor visual positioning method and device and electronic equipment

    CN112686962A

  • Visual positioning method and device and electronic equipment

    CN113657283A