Visual repositioning method, device, equipment and medium
By using neural network models to obtain the global and local features of real-time images in the visual relocation method, the candidate images and poses are determined, and the shortcomings of traditional methods in complex environments and high computing power requirements are solved, and efficient and real-time visual relocation is achieved.
Patent Information
- Application Number
- CN202510255954.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional visual relocation method is difficult to adapt to when facing scenes such as lighting changes and complex environments, and has high computing density and high computing power requirements, making it difficult to meet real-time requirements.
By inputting the real-time image at the current moment to the trained neural network model, the global and local features of the real-time image are obtained, and candidate images and visual relocation positions are determined based on these features, reducing computing needs and improving real-time performance.
While ensuring repositioning accuracy, it reduces computing requirements, adapts to application scenarios with high real-time requirements and limited computing power, and improves matching efficiency and accuracy.
Smart Images

Figure CN120182376A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual relocalization technology, and particularly relates to a visual relocalization method, device, equipment, and medium. Background Art
[0002] Visual relocalization technology has developed rapidly in recent years, especially in fields such as autonomous driving and augmented reality, where it has important application value.
[0003] Traditional visual relocalization methods use feature extraction methods based on gradient changes. Such methods are difficult to adapt to scenarios such as illumination changes and complex environments. At the same time, the computational density of such methods is relatively high, requiring high computing power and making it difficult to meet the real-time requirements of end-side devices.
[0004] In view of this, this application is specifically proposed. Summary of the Invention
[0005] The following gives a brief overview of one or more aspects to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects, and neither is it intended to identify key or decisive elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0006] This application provides a visual relocalization method, device, equipment, and medium, which reduces the computational requirements while ensuring accuracy, so as to better adapt to application scenarios with high real-time requirements and limited computing power.
[0007] In a first aspect, this application provides a visual relocalization method, including the following steps:
[0008] Input the real-time image collected at the current moment into a trained neural network model to obtain the global feature and local feature of the real-time image;
[0009] Determine candidate images according to the global feature of the real-time image and the global features of each reference image in the map database;
[0010] Determine the visual relocalization pose at the current moment according to the local feature of the real-time image and the local feature of the candidate image.
[0011] Further, the neural network model includes a basic network module, a convolutional network module, a depth-to-space module, and a grid sampling module; the local feature includes local key points and local key point descriptors;
[0012] The step of inputting the real-time image collected at the current moment into a trained neural network model to obtain the global feature and local feature of the real-time image includes:
[0013] Input the real-time image collected at the current moment into the basic network module to obtain global features;
[0014] Input the global features into the convolutional network module to obtain intermediate results, and input the intermediate results into the depth-to-space module to obtain local key points;
[0015] Input the local key points into the grid sampling module to obtain the local key point descriptors.
[0016] Furthermore, the neural network model is trained in a self-supervised manner using knowledge distillation.
[0017] Furthermore, the determination of candidate images based on the global features of the real-time image and the global features of each reference image in the map database includes:
[0018] Calculate the cosine similarity between the global features of the real-time image and the global features of each reference image in the map database;
[0019] Determine a preset number of candidate images from the map database according to the magnitude of the cosine similarity.
[0020] Furthermore, the determination of the visual relocalization pose at the current moment based on the local features of the real-time image and the local features of the candidate images includes:
[0021] Determine initial key point matching pairs through a key point matcher according to the local features of the real-time image and the local features of the candidate images;
[0022] Verify the initial key point matching pairs based on the Random Sample Consensus (RANSAC) algorithm to obtain the final key point matching pairs;
[0023] Determine the visual relocalization pose at the current moment according to the final key point matching pairs.
[0024] Furthermore, the local features include local key points and local key point descriptors. The determination of initial key point matching pairs through a key point matcher according to the local features of the real-time image and the local features of the candidate images includes:
[0025] Calculate the similarity between the local key point descriptors of the real-time image and the local key point descriptors of each candidate image through a key point matcher;
[0026] Determine the two local key points corresponding to the highest similarity as the initial key point matching pairs.
[0027] Further, determining the visual relocalization pose at the current moment according to the final key-point matching pairs includes:
[0028] Determining the pose of the key points of the candidate image in the final key-point matching pairs as the visual relocalization pose of the corresponding key points in the real-time image at the current moment.
[0029] In a second aspect, the present application further provides a visual relocalization device, including:
[0030] An input module, configured to input a real-time image collected at the current moment into a trained neural network model to obtain the global feature and the local feature of the real-time image;
[0031] A first determination module, configured to determine candidate images according to the global feature of the real-time image and the global features of each reference image in the map database;
[0032] A second determination module, configured to determine the visual relocalization pose at the current moment according to the local feature of the real-time image and the local features of the candidate images.
[0033] In a third aspect, the present application further provides an electronic device, where the electronic device includes:
[0034] One or more processors;
[0035] A storage device, configured to store one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the visual relocalization method as described above.
[0037] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the visual relocalization method as described above is implemented.
[0038] The visual relocalization method disclosed in the present application inputs a real-time image collected at the current moment into a trained neural network model to obtain the global feature and the local feature of the real-time image, improves the encoding speed and accuracy of the global feature and the local feature, and ensures the relocalization accuracy; further determines candidate images according to the global feature of the real-time image and the global features of each reference image in the map database, narrows the matching range, and further improves the matching efficiency; determines the visual relocalization pose at the current moment according to the local feature of the real-time image and the local features of the candidate images, reduces the calculation requirement while ensuring the accuracy, so as to better adapt to application scenarios with high real-time requirements and limited computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of a visual relocalization method provided by an embodiment of the present application;
[0041] Figure 2 It is a schematic structural diagram of a neural network model provided by an embodiment of the present application;
[0042] Figure 3 It is a schematic structural diagram of a visual relocalization device provided by an embodiment of the present application;
[0043] Figure 4 It is a schematic structural diagram of an electronic device in an embodiment of the present application. Specific Embodiments
[0044] The following will further elaborate on the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only the parts related to the invention are shown in the accompanying drawings.
[0045] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will detail the present application with reference to the accompanying drawings and embodiments.
[0046] Figure 1 It is a schematic flowchart of a visual relocalization method proposed by the present application. As Figure 1 shown, the visual relocalization method includes the following steps:
[0047] S110. Input the real-time image collected at the current moment into the trained neural network model to obtain the global feature and local feature of the real-time image.
[0048] Among them, the real-time image can be collected by an in-vehicle vision sensor. Specifically, during the vehicle driving process, the in-vehicle vision sensor captures the surrounding environment of the vehicle at a certain frequency to obtain the real-time image.
[0049] In some embodiments, referring to a schematic structural diagram of a neural network model as Figure 2 shown, it includes a basic network module 210, a convolutional network module 220, a depth-to-space module 230, and a network sampling module 240.
[0050] The local features include local key points and local key point descriptors.
[0051] Inputting the real-time image collected at the current moment into the trained neural network model to obtain the global features and local features of the real-time image includes:
[0052] Inputting the real-time image collected at the current moment into the basic network module to obtain global features; inputting the global features into the convolutional network module to obtain an intermediate result, and inputting the intermediate result into the depth-to-space module to obtain local key points; inputting the local key points into the grid sampling module to obtain the local key point descriptors.
[0053] Obtaining the global features and local features of the real-time image through the neural network model is beneficial to improving the running efficiency and relocalization accuracy of the entire relocalization process.
[0054] In some embodiments, the sample data for training the neural network model can be obtained in the following manner:
[0055] Using the publicly available landmark data, which itself has global features, and further using the SuperPoint algorithm to process the landmark data to obtain local key points and local key point descriptors (i.e., local features). SuperPoint is an algorithm used in computer vision to detect local key points in an image and generate corresponding descriptors. Thus, the purpose of not manually annotating data is achieved. Generally speaking, the neural network model is trained in a self-supervised manner of knowledge distillation. The neural network model is constructed in a fully convolutional manner, which can improve the generality of the model, support the quantization process well, and the quantization result is relatively robust.
[0056] Among them, the depth-to-space module 230 is also called the depth2space module, and is also called the inverse operation of deconvolution. It is a transformation that rearranges the depth dimension information of the feature map to the spatial dimension. For example, assuming that the input feature map has a certain height, width, and depth, in the "depth2space" operation, the depth dimension information will be distributed to the height and width dimensions according to certain rules, so that the depth of the output feature map decreases, while the height and width increase accordingly. Generally speaking, the depth-to-space module 230 is used to adjust the dimension structure of the feature map without adding too much computational effort to meet the subsequent processing requirements.
[0057] The network sampling module 240 is also called the grid sample module and is used to sample the feature map.
[0058] S120. Determine candidate images based on the global features of the real-time image and the global features of each reference image in the map database.
[0059] It can be understood that each reference image in the map database stores corresponding global features, local features, and poses (the pose can specifically be determined based on the dead reckoning method). By screening each reference image in the map database according to the global features of the real-time image, candidate images with a smaller range or fewer numbers can be obtained, thereby reducing the matching range of subsequent local features and improving the overall processing efficiency and accuracy.
[0060] In some embodiments, the determining candidate images based on the global features of the real-time image and the global features of each reference image in the map database includes:
[0061] Calculate the cosine similarity between the global features of the real-time image and the global features of each reference image in the map database; determine a preset number of candidate images from the map database according to the magnitude of the cosine similarity.
[0062] The greater the cosine similarity of the global features of two images, the more similar the global features of the two figures are.
[0063] Specifically, sort according to the magnitude of the cosine similarity, from large to small, and determine the reference images corresponding to the preset number of cosine similarities ranked in the front as candidate images. That is, the candidate images are images with relatively similar global features to the real-time image.
[0064] S130. Determine the visual relocalization pose at the current moment based on the local features of the real-time image and the local features of the candidate images.
[0065] In some embodiments, the determining the visual relocalization pose at the current moment based on the local features of the real-time image and the local features of the candidate images includes: determining initial key-point matching pairs through a key-point matcher according to the local features of the real-time image and the local features of the candidate images; verifying the initial key-point matching pairs based on the random sample consensus algorithm to obtain final key-point matching pairs; determining the visual relocalization pose at the current moment according to the final key-point matching pairs.
[0066] Among them, the local features include local key points and local key-point descriptors, and the determining initial key-point matching pairs through a key-point matcher according to the local features of the real-time image and the local features of the candidate images includes:
[0067] Calculate the similarity between the local key-point descriptors of the real-time image and those of each candidate image through a key-point matcher; determine the two local key-points corresponding to the highest similarity as the initial key-point matching pairs. Optionally, when calculating the similarity between the local key-point descriptors of the real-time image and those of the candidate image, Euclidean distance, Hamming distance, etc. can be used. Generally, Hamming distance is used to represent the similarity for binary descriptors, and the smaller the distance, the more similar the two key-points are. For non-binary descriptors, Euclidean distance is generally used. By setting a distance threshold, when the distance between two key-point descriptors is less than this distance threshold, these two key-points are considered a matching pair. In addition, some strategies can also be adopted to improve the matching accuracy, such as the nearest neighbor - second nearest neighbor ratio test, that is, calculate the nearest neighbor distance and the second nearest neighbor distance between each key-point and the key-points in other images, and when the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than a certain threshold, it is determined as a matching pair.
[0068] Verify the initial key-point matching pairs based on the Random Sample Consensus (RANSAC) algorithm to obtain the final key-point matching pairs. Specifically: The RANSAC algorithm estimates the transformation model by continuously randomly selecting some matching pairs, and determines whether other matching pairs are "inliers" according to this transformation model. If they are "inliers", they are retained, otherwise they are discarded, thereby improving the final matching accuracy.
[0069] Further, determining the visual relocalization pose at the current moment according to the final key-point matching pairs includes:
[0070] Determine the pose of the key-points of the candidate image in the final key-point matching pairs as the visual relocalization pose at the current moment of the corresponding key-points in the real-time image. Among them, the candidate image is an image in the map database, and this image correspondingly stores global features, local features, and key-point poses.
[0071] The map database is generated during offline map building. The process of visual relocalization can be summarized as: when obtaining a real-time image, determine whether the real-time image already exists in the map database. If it exists, further determine the poses of the key-points in this real-time image. Specifically: Input the real-time image into a neural network model to obtain its global features and local features; further screen candidate images from the map database according to the global features (where each image in the map database stores the corresponding global features, local features, and key-point poses), aiming to narrow the subsequent matching range, thereby improving the matching efficiency and accuracy; after obtaining the candidate images, further perform fine-grained matching based on the local features to determine the candidate image most similar to the real-time image, and then determine the key-point pose of this candidate image as the key-point pose of the real-time image to achieve the purpose of visual relocalization.
[0072] The visual relocalization method disclosed in this application obtains global features and local features through a neural network model, improving the inference speed and accuracy. By using a self-supervised method of knowledge distillation to train the neural network model, there is no need for additional manual data annotation, reducing the implementation difficulty and cost. By constructing the network in a fully convolutional manner, the network is more general, supports the quantization process well, and the quantization process is relatively robust. In summary, the solution of this application can run on edge terminals with insufficient computing resources, such as home memory parking terminals.
[0073] Figure 3 FIG. is a schematic structural diagram of a visual relocalization device provided by an embodiment of the present disclosure, specifically including: an input module 310, a first determination module 320, and a second determination module 330.
[0074] Among them, the input module 310 is configured to input the real-time image collected at the current moment into the trained neural network model to obtain the global feature and local feature of the real-time image; the first determination module 320 is configured to determine candidate images according to the global feature of the real-time image and the global features of each reference image in the map database; the second determination module 330 is configured to determine the visual relocalization pose at the current moment according to the local feature of the real-time image and the local feature of the candidate image.
[0075] Further, the neural network model includes a basic network module, a convolutional network module, a depth-to-space module, and a grid sampling module; the local feature includes local key points and local key point descriptors.
[0076] The input module 310 is specifically configured to: input the real-time image collected at the current moment into the basic network module to obtain a global feature; input the global feature into the convolutional network module to obtain an intermediate result, and input the intermediate result into the depth-to-space module to obtain local key points; input the local key points into the grid sampling module to obtain the local key point descriptors.
[0077] The neural network model is trained in a self-supervised manner of knowledge distillation.
[0078] Further, the first determination module 320 includes: a first calculation unit, configured to calculate the cosine similarity between the global feature of the real-time image and the global features of each reference image in the map database; a first determination unit, configured to determine a preset number of candidate images from the map database according to the magnitude of the cosine similarity.
[0079] Further, the second determination module 330 includes: a second determination unit, configured to determine an initial key point matching pair through a key point matcher according to local features of the real-time image and local features of the candidate image; a verification unit, configured to verify the initial key point matching pair based on the random sample consensus algorithm to obtain a final key point matching pair; and a third determination unit, configured to determine a visual relocalization pose at the current moment according to the final key point matching pair.
[0080] Further, the second determination unit is specifically configured to calculate the similarity between the local key point descriptors of the real-time image and the local key point descriptors of each candidate image through a key point matcher; and determine the two local key points corresponding to the highest similarity as the initial key point matching pair.
[0081] Further, the third determination unit is specifically configured to: determine the pose of the key point of the candidate image in the final key point matching pair as the visual relocalization pose of the corresponding key point in the real-time image at the current moment.
[0082] The visual relocalization device disclosed in this application inputs the real-time image collected at the current moment into a trained neural network model to obtain the global features and local features of the real-time image, which improves the encoding speed and accuracy of the global features and local features and ensures the relocalization accuracy; furthermore, candidate images are determined according to the global features of the real-time image and the global features of each reference image in the map database, narrowing the matching range and thus improving the matching efficiency; the visual relocalization pose at the current moment is determined according to the local features of the real-time image and the local features of the candidate image, reducing the computational requirements while ensuring the accuracy, so as to better adapt to application scenarios with high real-time requirements and limited computing power.
[0083] Figure 4 It is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. Specifically refer to Figure 4 below, which shows a schematic structural diagram of the electronic device 500 suitable for implementing the embodiment of the present disclosure. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0084] Such as Figure 4As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device 508 into a random access memory (RAM) to implement the method of the embodiments as described in the present disclosure. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O interface 505 is also connected to the bus 504. An input device 506, an output device 507, a storage device 508, and a communication device 509 are all connected to the I / O interface 505.
[0085] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart, so as to implement the visual relocalization method as described above. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiments of the present disclosure are executed.
[0086] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0087] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method steps in this application.
[0088] Optionally, when the above-mentioned one or more programs are executed by the electronic device, the electronic device can also execute the other steps described in the above embodiments.
[0089] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0090] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0091] Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only for helping to understand the method and its core idea of the present application. The above are only the preferred implementation manners of the present application. It should be noted that due to the limitation of literal expression and objectively existing infinite specific structures, for those of ordinary skill in the art, without departing from the principle of the present application, several improvements, refinements, or changes can be made, or the above technical features can be combined in an appropriate manner; these improvements, refinements, changes, or combinations, or directly applying the inventive concept and technical solution to other occasions without improvement, shall all be regarded as the protection scope of the present application.
Claims
1. A visual relocalization method, characterized in that: include: Inputting the real-time image collected at the current moment into the trained neural network model to obtain the global features and local features of the real-time image; Determining a candidate image according to the global features of the real-time image and the global features of each reference image in a map database; The visual relocation posture at the current moment is determined according to the local features of the real-time image and the local features of the candidate image.
2. The visual relocation method according to claim 1, characterized in that: The neural network model includes a basic network module, a convolutional network module, a depth-to-space module, and a grid sampling module; the local features include local key points and local key point descriptors; The step of inputting the real-time image collected at the current moment into the trained neural network model to obtain the global features and local features of the real-time image includes: Inputting the real-time image collected at the current moment into the basic network module to obtain global features; Inputting the global features into the convolutional network module to obtain an intermediate result, and inputting the intermediate result into the depth-to-space module to obtain a local key point; The local key points are input into the grid sampling module to obtain the local key point descriptors.
3. The visual relocation method according to claim 1, characterized in that: The neural network model is trained using a self-supervised approach using knowledge distillation.
4. The visual relocation method according to claim 1, characterized in that: The determining of the candidate image according to the global features of the real-time image and the global features of each reference image in the map database comprises: Calculating the cosine similarity between the global features of the real-time image and the global features of each reference image in the map database; A preset number of candidate images are determined from the map database according to the magnitude of the cosine similarity.
5. The visual relocation method according to claim 1, characterized in that: The determining the visual relocation posture at the current moment according to the local features of the real-time image and the local features of the candidate image comprises: Determining an initial key point matching pair through a key point matcher according to the local features of the real-time image and the local features of the candidate image; Verifying the initial key point matching pair based on a random sampling consistency algorithm to obtain a final key point matching pair; The visual relocalization pose at the current moment is determined based on the final key point matching pair.
6. The visual relocation method according to claim 5, characterized in that: The local features include local key points and local key point descriptors, and determining an initial key point matching pair through a key point matcher according to the local features of the real-time image and the local features of the candidate image, comprises: Calculating the similarity between the local key point descriptors of the real-time image and the local key point descriptors of each candidate image by a key point matcher; The two local key points corresponding to the highest similarity are determined as the initial key point matching pair.
7. The visual relocation method according to claim 5, characterized in that: Determining the visual relocation pose at the current moment according to the final key point matching pair includes: The pose of the key point of the candidate image in the final key point matching pair is determined as the visual relocation pose of the corresponding key point in the real-time image at the current moment.
8. A visual repositioning device, characterized in that: include: An input module, used to input the real-time image collected at the current moment into the trained neural network model to obtain the global features and local features of the real-time image; A first determination module, used to determine a candidate image according to the global features of the real-time image and the global features of each reference image in a map database; The second determination module is used to determine the visual relocation posture at the current moment according to the local features of the real-time image and the local features of the candidate image.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the visual relocalization method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the visual repositioning method according to any one of claims 1 to 7 is implemented.