Positioning method and apparatus, electronic device, and computer-readable storage medium

By using civilian-grade IMU sensor data and visual positioning timing models in combination with cloud services, the problem of high cost of high-precision positioning is solved, and a low-cost, high-precision positioning solution is provided, which is suitable for AR navigation and other scenarios with high precision requirements.

CN114460617BActive Publication Date: 2025-10-17ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011211374.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-03
Publication Date
2025-10-17
Estimated Expiration
2040-11-03

AI Technical Summary

Technical Problem

Existing high-precision positioning solutions rely on high-cost high-precision GNSS+INS integrated navigation systems, resulting in poor applicability. A low-cost and high-precision positioning solution is needed.

Method used

It uses civilian-grade IMU sensor data, pre-trained visual positioning time series model and cloud-based visual positioning service, combines the images collected by the camera and the visual positioning pose for positioning, and predicts and corrects the camera pose estimation data through the visual positioning time series model.

Benefits of technology

It achieves low-cost high-precision positioning and is suitable for a variety of scenarios, including AR navigation and elevated roads, reducing the demand for sensors and chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114460617B_ABST
    Figure CN114460617B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a positioning method and device, electronic equipment and computer readable storage medium, the method comprises: based on the initial frame image and GNSS position data collected by the camera, requesting the camera initial pose estimation data at time t0 returned by the visual positioning service; obtaining the IMU sensor data at time t N from the Nth frame image collected by the camera after the initial frame image, inputting the IMU sensor data at time t N and the camera pose estimation data at time t N‑1 into the pre-trained positioning time sequence model, and predicting the camera pose estimation data at time t N ; if the visual positioning correction condition is met when the Nth frame image is collected, then based on the position data at time t N and the Nth frame image, requesting the visual positioning pose at time t N returned by the visual positioning service; based on the visual positioning pose at time t N , correcting the camera pose estimation data at time t N output by the visual positioning time sequence model, and obtaining the positioning result at time t N .
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present disclosure relates to the technical field of computer vision, in particular to a positioning method and device, electronic equipment and computer readable storage medium. BACKGROUND

[0002] With the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter level and below) positioning results of the positioned object (such as a smart phone, a car, a robot, etc.). In order to obtain high-precision positioning results, the prior art provides a positioning scheme based on a high-precision GNSS+INS integrated navigation system, which obtains the positioning result (latitude and longitude position and attitude) of the positioned object at each time by means of the high-precision GNSS+INS integrated navigation system. The inventors found in the process of studying the scheme that, although the scheme can obtain high-precision positioning results, the sensors and chips used by the high-precision GNSS+INS integrated navigation system are expensive, and the high cost leads to poor applicability of the scheme. Therefore, it is urgent to provide a more universal positioning scheme that can guarantee positioning accuracy and is low in cost. SUMMARY

[0003] The embodiment of the present disclosure provides a positioning method, device, electronic equipment and computer readable storage medium.

[0004] In a first aspect, the embodiment of the present disclosure provides a positioning method.

[0005] Specifically, the positioning method comprises:

[0006] requesting a vision positioning service to return camera initial pose estimation data of a camera at a t0 moment based on an initial frame image collected by the camera and GNSS position data output by a GNSS receiver;

[0007] obtaining IMU sensor data at a tN moment for an Nth frame image collected by the camera after the initial frame image, inputting the IMU sensor data at the tN moment and camera pose estimation data at a tN-1 moment into a pre-trained vision positioning time sequence model, and predicting camera pose estimation data at the tN moment, wherein N is an integer greater than or equal to 1; N N N-1 N

[0008] if a set vision positioning correction condition is met when the Nth frame image is collected, requesting the vision positioning service to return vision positioning pose at the tN moment based on position data at the tN moment and the Nth frame image; N N

[0009] based on the vision positioning pose at the tN moment, obtaining a final positioning result of the camera.​​​​​​N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

[0010] In combination with the first aspect, in a first implementation of the first aspect of the embodiment of the present disclosure, requesting the visual positioning service to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver includes:

[0011] Obtain the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, and send a positioning request to the visual positioning service, wherein the positioning request carries the initial frame image and GNSS position data, so that it calculates and returns the camera's initial pose estimation data at time t0 based on the initial frame image and GNSS position data.

[0012] In combination with the first aspect and the first implementation of the first aspect, in the second implementation of the first aspect of the embodiment of the present disclosure, if the set visual positioning correction condition is met when the N-th frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, including:

[0013] If the set visual positioning correction condition is met when the Nth frame image is collected, the Nth frame image collected by the camera and t N Location data at the moment;

[0014] Send a positioning request to the visual positioning service, wherein the positioning request carries the Nth frame image and t N The position data at the moment, so that the visual positioning service is based on the Nth frame image and t N Calculate and return the position data at time t N Visual positioning pose at the moment.

[0015] In combination with the first aspect, the first implementation method of the first aspect, the second implementation method of the first aspect and the third implementation method of the first aspect, in the fourth implementation method of the first aspect of the present disclosure, the visual positioning correction condition includes: performing a correction every M frames of image, or the scene of the captured image is a set scene for visual positioning correction.

[0016] In a fifth implementation manner of the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, or the fifth implementation manner of the first aspect, the position data at the t N time is position data in the camera pose estimation data at the t N time.

[0017] In a sixth implementation manner of the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, or the sixth implementation manner of the first aspect, the visual positioning pose at the t N time is corrected based on the visual positioning timing model output of the camera pose estimation data at the t N time, to obtain a positioning result at the t N time, including:

[0018] In a seventh implementation manner of the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, or the seventh implementation manner of the first aspect, the visual positioning pose at the t N time is asynchronously corrected to obtain an asynchronously corrected visual positioning pose at the t N time.

[0019] In an eighth implementation manner of the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, the seventh implementation manner of the first aspect, or the eighth implementation manner of the first aspect, the asynchronously corrected visual positioning pose at the t N time is corrected based on the visual positioning timing model output of the camera pose estimation data at the t N time, to obtain a positioning result at the t N time.

[0020] In a ninth implementation manner of the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, the seventh implementation manner of the first aspect, the eighth implementation manner of the first aspect, or the ninth implementation manner of the first aspect, the visual positioning pose at the t N time is asynchronously corrected to obtain an asynchronously corrected visual positioning pose at the t N time, including:

[0021] After sending the positioning request to the visual positioning service, a positioning request sending time t T is recorded, and a state quantity of the visual positioning timing model at the t T time is backed up.

[0022] Speed values and IMU sensor data at each time after the t T time are recorded until the t T+P time, at which the visual positioning pose at the t N time is received.

[0023] The t N The visual positioning pose at the t T The state quantity of the visual positioning time sequence model at the t

[0024] The state quantity of the visual positioning time sequence model at the t T The state quantity of the visual positioning time sequence model at the t N The visual positioning pose at the t

[0025] In a second aspect, the embodiments of the present disclosure provide a positioning method.

[0026] Specifically, the positioning method comprises:

[0027] In response to receiving a positioning request sent by a terminal, obtaining position data and an image carried by the positioning request;

[0028] Calculating camera pose estimation data based on the position data and the image;

[0029] Sending the camera pose estimation data to the terminal as an input of a pre-trained visual positioning time sequence model running on the terminal or for correcting a result output by the visual positioning time sequence model.

[0030] With reference to the second aspect, in a first implementation of the second aspect, when the image is an initial frame image, the position data is GNSS position data; when the image is an Nth frame image collected after the initial frame image, the position data is position data in camera pose estimation data at a t N The position data in the camera pose estimation data at the t N The camera pose estimation data at the t N The camera pose estimation data at the t N-1 The camera pose estimation data at the t

[0031] With reference to the second aspect and the first implementation of the second aspect, in a second implementation of the second aspect, the camera pose estimation data is calculated based on the position data and the image, comprising:

[0032] Determining a candidate reference image with a distance to the position data within a preset range from a reference image database, wherein the reference image database stores images with pre-calculated pose data;

[0033] calculate image similarity between the image and the candidate reference image, and determine the candidate reference image satisfying a preset condition as the reference image corresponding to the image;

[0034] calculate relative pose data between the image and the reference image;

[0035] obtain pose data of the reference object, and calculate pose data of the image based on the pose data of the reference object and the relative pose data, as camera pose estimation data.

[0036] In a third implementation manner of the second aspect, the second implementation manner of the second aspect, or the second implementation manner of the second aspect, the calculating of the relative pose data between the image and the reference image is implemented as:

[0037] input the image and the reference image into a pre-trained relative pose data calculation model to calculate the relative pose data between the image and the reference image.

[0038] In a third aspect, a positioning method is provided in the embodiments of the present disclosure.

[0039] Specifically, the positioning method comprises:

[0040] The terminal requests a visual positioning service to return camera initial pose estimation data at a t0 moment of the camera based on an initial frame image collected by a camera and GNSS position data output by a GNSS receiver;

[0041] The cloud calculates camera initial pose estimation data at the t0 moment based on the GNSS position data and the initial frame image, and sends the camera initial pose estimation data at the t0 moment to the terminal in response to receiving the positioning request sent by the terminal;

[0042] The terminal obtains IMU sensor data at a tN moment of the camera for an Nth frame image collected by the camera after the initial frame image, inputs the IMU sensor data at the tN moment and the camera pose estimation data at the t0 moment into a pre-trained visual positioning time sequence model, and predicts camera pose estimation data at the tN moment. N N N-1 N wherein N is an integer greater than or equal to 1, if a preset visual positioning correction condition is met when the Nth frame image is collected, the terminal requests a visual positioning service to return visual positioning pose at the tN moment based on the position data at the t0 moment and the Nth frame image, and corrects the camera pose estimation data at the tN moment output by the visual positioning time sequence model based on the visual positioning pose at the tN moment. N N N N ​​​​​​The camera pose estimation data at the time t0 is corrected to obtain t N The positioning result at the time t0.

[0043] With reference to the third aspect, in a first implementation manner of the third aspect, the cloud end calculates the initial camera pose estimation data at the time t0 based on the GNSS position data and the initial frame image, including:

[0044] determining a candidate reference image from the reference image database according to the GNSS position data, wherein the distance between the GNSS position data and the candidate reference image is within a preset range, and the reference image database stores images with pre-calculated pose data;

[0045] calculating the image similarity between the initial frame image and the candidate reference image, and determining the candidate reference image satisfying a preset condition as the reference image corresponding to the initial frame image;

[0046] calculating the relative pose data between the initial frame image and the reference image;

[0047] obtaining the pose data of the reference object, and calculating the pose data of the initial frame image based on the pose data of the reference object and the relative pose data, and taking the pose data of the initial frame image as the initial camera pose estimation data at the time t0.

[0048] With reference to the third aspect and the first implementation manner of the third aspect, in a second implementation manner of the third aspect, the calculation of the relative pose data between the initial frame image and the reference image is implemented as:

[0049] inputting the initial frame image and the reference image into a pre-trained relative pose data calculation model to calculate the relative pose data between the initial frame image and the reference image.

[0050] With reference to the third aspect, the first implementation manner of the third aspect, and the second implementation manner of the third aspect, in a third implementation manner of the third aspect, the t N The position data at the time t0 is the position data in the camera pose estimation data at the time t N The position data at the time t0 is the position data in the camera pose estimation data at the time t

[0051] With reference to the third aspect, the first implementation manner of the third aspect, the second implementation manner of the third aspect, and the third implementation manner of the third aspect, in a fourth implementation manner of the third aspect, the method further includes:

[0052] The terminal performs asynchronous correction on the visual positioning pose at the time t N The time t0 is the time t N The time t0 is the time t

[0053] In combination with the third aspect, the first implementation of the third aspect, the second implementation of the third aspect, the third implementation of the third aspect and the fourth implementation of the third aspect, in a fifth implementation of the third aspect of the present disclosure, the terminal is for the t N The visual positioning pose at the moment is asynchronously corrected to obtain the asynchronously corrected pose t N The visual positioning pose at the moment, including:

[0054] After sending a positioning request to the visual positioning service, record the time t when the positioning request is sent T , and back up t T The state quantity of the visual positioning timing model at the moment;

[0055] Record T The speed value and IMU sensor data at each moment after time t T+P Received at time t N Visual positioning posture at the moment;

[0056] Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment;

[0057] The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning pose at the moment.

[0058] In a fourth aspect, an embodiment of the present disclosure provides a positioning device.

[0059] Specifically, the positioning device includes:

[0060] A first request module is configured to request the visual positioning service to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver;

[0061] The prediction module is configured to obtain t N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, where N is an integer greater than or equal to 1;

[0062] The second request module is configured to, if the set visual positioning correction condition is met when the Nth frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N Visual positioning posture at the moment;

[0063] The correction module is configured to N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

[0064] In a fifth aspect, a cloud server is provided in an embodiment of the present disclosure.

[0065] Specifically, the cloud server includes:

[0066] an acquisition module, configured to, in response to receiving a positioning request sent by a terminal, acquire the position data and image carried in the positioning request;

[0067] A calculation module is configured to calculate camera pose estimation data based on the position data and the image;

[0068] A sending module is configured to send the camera pose estimation data to the terminal as an input to a pre-trained visual positioning timing model running on the terminal or to correct a result output by the visual positioning timing model.

[0069] In a sixth aspect, a positioning device is provided in an embodiment of the present disclosure.

[0070] Specifically, the positioning device includes:

[0071] The terminal is configured to request the visual positioning service to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, and obtain t N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, wherein N is an integer greater than or equal to 1, if the set visual positioning correction condition is met when the N-th frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the tN the visual positioning pose at the time t N the visual positioning pose at the time t N the camera pose estimation data at the time t N the positioning result at the time t

[0072] The cloud is configured to, in response to receiving the positioning request sent by the terminal, calculate the camera initial pose estimation data at the time t0 based on the GNSS position data and the initial frame image, and send the camera initial pose estimation data to the terminal.

[0073] In a seventh aspect, an electronic device is provided, which includes a memory and a processor, the memory is configured to store one or more computer instructions supporting a positioning apparatus to perform the positioning method described above, and the processor is configured to execute the computer instructions stored in the memory. The positioning apparatus can further include a communication interface configured to enable the positioning apparatus to communicate with other devices or a communication network.

[0074] In an eighth aspect, a computer readable storage medium is provided, which is configured to store computer instructions for a positioning apparatus, and the computer instructions include computer instructions for performing the positioning method described above.

[0075] In a ninth aspect, a navigation service is provided, in which a positioning result of a navigated object is obtained based on the positioning method described above, and navigation guidance services for a corresponding scene are provided for the navigated object based on the positioning result.

[0076] In a first implementation manner of the ninth aspect, the corresponding scene is one or a combination of AR navigation, elevated navigation, and main and auxiliary road navigation.

[0077] The technical solution provided by the embodiments of the present disclosure can include the following beneficial effects:

[0078] The technical solution described above only needs IMU sensor data obtained by a civilian IMU, a pre-trained visual positioning time sequence model, and a visual positioning pose provided by a visual positioning service from a cloud, and does not need expensive sensors and chips, compared with the prior art. Furthermore, the technical solution combines images collected by a camera and a visual positioning pose provided by a visual positioning service in the positioning process, which guarantees the accuracy of the final positioning result, and is a low-cost positioning solution with stronger universality.

[0079] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0080] Other features, objects, and advantages of the embodiments of the present disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. In the drawings:

[0081] Figure 1 A flowchart of a positioning method according to an embodiment of the present disclosure is shown;

[0082] Figure 2 A motion model diagram according to an embodiment of the present disclosure is shown;

[0083] Figure 3 A correction timing diagram of a visual positioning pose for t N

[0084] Figure 4 A flowchart of a positioning method according to another embodiment of the present disclosure is shown;

[0085] Figure 5 A flowchart of a positioning method according to still another embodiment of the present disclosure is shown;

[0086] Figure 6 A structural block diagram of a positioning device according to an embodiment of the present disclosure is shown;

[0087] Figure 7 A structural block diagram of a cloud server according to an embodiment of the present disclosure is shown;

[0088] Figure 8 A structural block diagram of a positioning device according to another embodiment of the present disclosure is shown;

[0089] Figure 9 A structural diagram of a computer system suitable for implementing a positioning method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0090] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so as to make those skilled in the art readily implement them. Also, portions irrelevant to the description of the exemplary embodiments are omitted in the accompanying drawings for the sake of clarity.

[0091] In the embodiments of the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate that there are features, numbers, steps, actions, parts, or combinations thereof, disclosed in the specification, and do not exclude the possibility that one or more other features, numbers, steps, actions, parts, or combinations thereof exist or are added.

[0092] ​It should be further noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other in the case of no conflict. The embodiments of the present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0093] The technical solution provided by the embodiments of the present disclosure only needs IMU sensor data obtained by a civilian IMU, a pre-trained visual positioning time sequence model, and a visual positioning pose provided by a visual positioning service from the cloud, compared with the prior art, without the need for high-end sensors and chips, further, the technical solution combines the image collected by the camera and the visual positioning pose provided by the visual positioning service in the positioning process, ensuring the accuracy of the final positioning result, and is a low-cost positioning solution with stronger universality.

[0094] Figure 1 A flowchart of a positioning method according to an embodiment of the present disclosure is shown as follows, Figure 1 The positioning method includes the following steps S101-S104:

[0095] In step S101, based on an initial frame image collected by a camera and GNSS position data output by a GNSS receiver, a visual positioning service is requested to return camera initial pose estimation data of t0 moment of the camera; wherein the camera can be a camera, the GNSS receiver can be a GPS receiver, the camera and the GPS receiver can be integrated in the same terminal, such as a mobile phone or a driving recorder, or can be independent components but mounted on the same object to be positioned, such as a car, a robot, etc.

[0096] In step S102, for the Nth frame image collected by the camera after the initial frame image, IMU sensor data of t N moment is obtained, the IMU sensor data of t N moment and the camera pose estimation data of t N-1 moment are input into a pre-trained visual positioning time sequence model, and camera pose estimation data of t N moment is predicted, wherein N is an integer greater than or equal to 1;

[0097] In step S103, if the visual positioning correction condition is met when the Nth frame image is collected, the visual positioning service is requested to return the visual positioning pose of t N moment based on the position data of t N moment and the Nth frame image;

[0098] In step S104, the visual positioning pose of t N moment is used to correct the camera pose estimation data of t NThe camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

[0099] As mentioned above, with the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter-level and below) positioning results of the located objects (such as smart phones, cars, robots, etc.). In order to obtain high-precision positioning results, the existing technology provides a positioning solution based on a high-precision GNSS+INS combined navigation system. This solution uses the high-precision GNSS+INS combined navigation system to obtain the positioning results (latitude and longitude positions and postures) of the located objects at every moment. In the process of studying this solution, the inventors found that although this solution can obtain high-precision positioning results, the sensors and chips used in the high-precision GNSS+INS combined navigation system are expensive. The excessively high cost makes the applicability of this solution not strong. Therefore, there is an urgent need to provide a more universal positioning solution that can guarantee positioning accuracy and is low-cost.

[0100] Taking the above issues into consideration, in this embodiment, a positioning method is proposed. This method performs positioning based solely on IMU sensor data acquired by a civilian-grade IMU, a pre-trained visual positioning time series model, and the visual positioning pose provided by a cloud-based visual positioning service. Compared with existing technologies, this method does not require expensive sensors and chips. Furthermore, this technical solution combines the images captured by the camera and the visual positioning pose provided by the visual positioning service during the positioning process, ensuring the accuracy of the final positioning result. It is a low-cost and more universal positioning solution. This positioning solution can be applied to AR navigation scenarios. When the user starts the AR navigation service, this solution begins to run. It can also be applied to other scenarios that require higher-precision positioning results, such as elevated roads, main and auxiliary roads, etc.

[0101] In one embodiment of the present disclosure, the positioning method may be applicable to a terminal computer, a computing device, an electronic device, a server, a service cluster, etc. that can perform positioning processing.

[0102] In one embodiment of the present disclosure, the positioning method may be performed based on an image sequence including multiple frames of images.

[0103] In one embodiment of the present disclosure, the initial frame image refers to the 0th frame in the image sequence including multiple frame images captured by a camera.

[0104] In an embodiment of the present disclosure, the acquisition time of the initial frame image and the acquisition time of the GNSS position data can be the same or different, as long as the GNSS position data is the GNSS position data closest to the position of the initial frame image in the position data output by the GNSS receiver. The GNSS position data can be, for example, GPS position data, Beidou position data, etc.

[0105] In an embodiment of the present disclosure, the camera initial pose estimation data at time t0 refers to the camera pose estimation data at the time corresponding to the initial frame image determined by the cloud visual positioning service. In this embodiment, the camera pose estimation data at time t0 determined by the cloud visual positioning service can be used as the camera initial pose estimation data at time t0 corresponding to the initial frame image, and can also be used as the initial state data of the visual positioning time sequence model for subsequent prediction of the camera pose.

[0106] In an embodiment of the present disclosure, the pose estimation data can include position data such as latitude and longitude, and attitude data such as attitude yaw angle.

[0107] In an embodiment of the present disclosure, the IMU sensor data refers to the data detected by the IMU (Inertial Measurement Unit) sensor. For example, for a 6-axis IMU sensor composed of 3 accelerometers and 3 gyroscopes, the detected data can include 3-direction acceleration and 3-direction angular velocity.

[0108] In an embodiment of the present disclosure, the positioning time sequence model refers to a model that can be pre-trained to predict the visual positioning data at the next time. For example, when N = 1, after obtaining the camera initial pose estimation data at time t0, the camera pose estimation data at time t1 corresponding to the first frame image can be predicted according to the camera initial pose estimation data at time t0 and the acquired IMU sensor data at time t1. Then, when N = 2, the camera pose estimation data at time t2 corresponding to the second frame image can be predicted according to the camera pose estimation data at time t1 and the acquired IMU sensor data at time t2, and so on, so that the camera pose estimation data at the time corresponding to all images can be obtained. That is, for the Nth frame image acquired by the camera after the initial frame image, the IMU sensor data at time tN-1 is acquired, the IMU sensor data at time tN-1 and the camera pose estimation data at time tN-1 are input into the pre-trained visual positioning time sequence model, and the camera pose estimation data at time tN corresponding to the Nth frame image can be predicted. N N N-1 N wherein N is an integer greater than or equal to 1.

[0109] ​​​In an embodiment of the present disclosure, the visual positioning time sequence model can be implemented based on a Kalman filter, in which the visual positioning time sequence model comprises a prediction part based on a motion model and a correction part based on observation. Figure 2 As shown in the motion model according to an embodiment of the present disclosure, Figure 2 in which the motion model can satisfy the following update relationship:

[0110] R N+1 = R N exp((ω N dt) × )

[0111] V N+1 = V N +(R N a N +g)dt

[0112] P N+1 = P N +V N dt

[0113] wherein ω N and a N represent the angular velocity and acceleration in the IMU sensor data, respectively.

[0114] Based on the above visual positioning time sequence model, the state at time t N can be predicted from the state at time t N+1 .

[0115] In this embodiment, for the images acquired after the initial frame image, the camera pose estimation data at the corresponding time predicted by the visual positioning time sequence model can be directly used as the positioning result at the corresponding time of the images, but if it is determined that the visual positioning correction condition is met when the Nth frame image is acquired, the visual positioning service is requested to return the visual positioning pose at time t N based on the position data at time t N and the Nth frame image, and then the camera pose estimation data at time t N output by the visual positioning time sequence model is corrected based on the visual positioning pose at time t N , to obtain the corrected positioning result at time t N . The position data at time t N refers to the position data in the camera pose estimation data at time t N .

[0116] The above embodiment obtains accurate camera pose estimation data based on the spatiotemporal continuity assumption and in combination with historical time series information by means of a visual positioning time series model, and thus can better improve the problem of unstable single-frame prediction in visual positioning, thereby effectively improving the accuracy of visual positioning. Meanwhile, since the Kalman filter algorithm is very efficient, the calculation cost of the terminal can be basically ignored, that is, the calculation consumption can be greatly saved.

[0117] In an embodiment of the present disclosure, the step of requesting the visual positioning service to return the initial camera pose estimation data of the camera at the t0 moment based on the initial frame image collected by the camera and the GNSS position data output by the GNSS receiver in the step S101 can include the following steps:

[0118] The initial frame image collected by the camera and the GNSS position data output by the GNSS receiver are acquired, and a positioning request is sent to the visual positioning service, wherein the positioning request carries the initial frame image and the GNSS position data, so that the visual positioning service calculates and returns the initial camera pose estimation data at the t0 moment based on the initial frame image and the GNSS position data.

[0119] Since there is no historical time series information at the t0 moment, it is difficult to make a prediction, and therefore in this embodiment, the initial camera pose estimation data at the t0 moment is obtained by means of the visual positioning service of the cloud. Specifically, the terminal acquires the GNSS position data and the initial frame image at the t0 moment, and sends the GNSS position data and the initial frame image at the t0 moment to the visual positioning service of the cloud in the positioning request, so as to obtain the camera pose estimation data at the t0 moment calculated by the visual positioning service of the cloud based on the GNSS position data and the initial frame image at the t0 moment. The camera pose estimation data at the t0 moment obtained by the visual positioning service of the cloud can be used as the initial camera pose estimation data at the t0 moment.

[0120] Similarly, the step S103, that is, if the set visual positioning correction condition is met when the Nth frame image is collected, the step of requesting the visual positioning service to return the visual positioning pose at the t N moment based on the position data at the t N moment and the Nth frame image can include the following steps:

[0121] If the set visual positioning correction condition is met when the Nth frame image is collected, the Nth frame image collected by the camera and the position data at the t N moment are acquired.

[0122] A positioning request is sent to the visual positioning service, wherein the positioning request carries the Nth frame image and the position data at the t Nthe position data at the time t, so that the visual positioning service calculates the visual positioning pose based on the Nth frame of image and the position data at the time t N the position data at the time t, so that the visual positioning service calculates the visual positioning pose based on the Nth frame of image and the position data at the time t N the visual positioning pose at the time t.

[0123] In view of the estimation bias of the visual positioning timing model, in order to improve the accuracy of the camera pose estimation data, the camera pose estimation data generated by the visual positioning timing model needs to be corrected by referring to the cloud camera pose estimation data. However, in view of the control of the cloud computing cost, the cloud camera pose estimation data is not required to be referred to for each frame of image. Instead, it is first determined whether the visual positioning correction condition is met when the Nth frame of image is collected. If the visual positioning correction condition is met, the Nth frame of image collected by the camera and the position data at the time t N the position data at the time t, so that the visual positioning service calculates the visual positioning pose based on the Nth frame of image and the position data at the time t N the position data at the time t, so that the visual positioning service calculates the visual positioning pose based on the Nth frame of image and the position data at the time t N the position data at the time t, so that the visual positioning service calculates the visual positioning pose based on the Nth frame of image and the position data at the time t N the visual positioning pose at the time t. N the visual positioning pose at the time t. N the camera pose estimation data at the time t is corrected to obtain the t N the final positioning result at the time t. The visual positioning correction condition can include that the correction is performed once every M frames of image, or the scene of the collected image is a scene in which the visual positioning correction is performed, such as a scene in which the GNSS positioning signal is blocked to cause the GNSS position to be inaccurate, a scene in which the IMU data has an error due to a fork or a turn, a scene in which the uncertainty of the currently used visual positioning timing model exceeds a preset threshold (i.e., the error of the visual positioning timing model is large), or a scene in which the communication network is interrupted or the data transmission is unstable.

[0124] In another embodiment of the present disclosure, when the Nth frame of image is collected, the position data at the time t N the time t, the position data in the camera pose estimation data at the time t is used as the position data at the time t N the time t, the position data in the camera pose estimation data at the time t is used as the position data at the time t N the time t, the GNSS position data output by the GNSS receiver at the time t is used as the position data at the time t N the time t, the GNSS position data output by the GNSS receiver at the time t is used as the position data at the time t N the time t, the GNSS position data output by the GNSS receiver at the time t is used as the position data at the time t Nvisual positioning pose at the time tN. At this time, the step S103, i.e., if the set visual positioning correction condition is met when the Nth image is collected, the tN N visual positioning pose at the time tN. At this time, the step S103, i.e., if the set visual positioning correction condition is met when the Nth image is collected, the tN N The step of obtaining the visual positioning pose at the time tNmay include the following steps:

[0125] If the set visual positioning correction condition is met when the Nth image is collected, the Nth image collected by the camera and the GNSS position data at the time tN N GNSS position data output by the GNSS receiver at the time tN;

[0126] sending a positioning request to the visual positioning service, wherein the positioning request carries the Nth image and the GNSS position data at the time tN N GNSS position data output by the GNSS receiver at the time tN, so that the visual positioning service calculates and returns the visual positioning pose at the time tNbased on the Nth image and the GNSS position data at the time tN N GNSS position data output by the GNSS receiver at the time tN, so that the visual positioning service calculates and returns the visual positioning pose at the time tNbased on the Nth image and the GNSS position data at the time tN N visual positioning pose at the time tN.

[0127] In actual applications, the GNSS position data corresponding to some image frames may have unpredictable drift, which may cause the cloud to be unable to calculate or obtain accurate camera pose estimation data. At this time, the GNSS position data needs to be corrected, and the corrected GNSS position data is sent to the cloud to enable the cloud to calculate camera pose estimation data based on the corrected GNSS position data, thereby avoiding visual positioning errors caused by GNSS positioning errors. Considering that the visual positioning time sequence model based on the Kalman filter described above can be used to predict the pose data at the current time from the pose data at the previous time through a motion model, when the GNSS position data needs to be corrected, the position data in the pose data at the current time predicted by the visual positioning time sequence model is used as the correction data of the GNSS position data that has drifted, and the t N GNSS position data in the camera pose estimation data at the time tNis used as the t N GNSS position data at the time tN.

[0128] In this embodiment, the method further includes the following steps:

[0129] In response to receiving the retransmission position data request sent by the visual positioning service, the t N GNSS position data in the camera pose estimation data at the time tNis used as the t N visual positioning pose at the time tN.

[0130] The retransmission position data request refers to a request sent by the cloud visual positioning service to the terminal for retransmission of position data due to the fact that the GNSS position data previously sent to the cloud visual positioning service cannot be used or cannot be used to obtain accurate camera pose estimation data.

[0131] In an embodiment of the present disclosure, the collection time of the Nth image and the collection time of the GNSS position data can be the same or different, as long as the GNSS position data is the GNSS position data closest to the position of the Nth image in the position data output by the GNSS receiver.

[0132] In an embodiment of the present disclosure, the step S104, i.e., correcting the camera pose estimation data at the time t according to the visual positioning pose at the time t output by the visual positioning time sequence model, to obtain the positioning result at the time t, can include the following steps: N N N

[0133] N N

[0134] N N N

[0135] In actual applications, due to network transmission, cloud computing and other factors, there is a certain time cost for the terminal to call the cloud camera pose estimation data calculation service, which causes an asynchronous phenomenon between the cloud and the terminal, that is, when the terminal sends a positioning request and receives the camera pose estimation data sent by the cloud after a certain time, the camera pose estimation data received at this time is not the cloud camera pose estimation data corresponding to the time when the terminal sends the positioning request. Therefore, in this embodiment, the camera pose estimation data at the time t corresponding to the Nth image needs to be asynchronously corrected, and then the camera pose estimation data at the time t after asynchronous correction is used to correct the camera pose estimation data at the time t output by the visual positioning time sequence model, to obtain the positioning result at the time t. N N N N

[0136] ​​​​​​​​​​​​​​In one embodiment of the present disclosure, the t N The visual positioning pose at the moment is asynchronously corrected to obtain the asynchronously corrected pose t N The steps of visually locating the pose at the moment may include the following steps:

[0137] After sending a positioning request to the visual positioning service, record the time t when the positioning request is sent T , and back up t T The state quantity of the visual positioning timing model at the moment;

[0138] Record T The speed value and IMU sensor data at each moment after time t T+P Received at time t N Visual positioning posture at the moment;

[0139] Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment;

[0140] The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning pose at the moment.

[0141] In this embodiment, the state quantity of the visual positioning time series model when sending the positioning request is corrected, and the state quantity of the corrected visual positioning time series model is used for prediction to achieve the t N Correction of the visual positioning pose at the moment. Specifically, Figure 3 As shown, Figure 3 The serial number in the table represents the execution order. ① The terminal sends a visual positioning request to the cloud visual positioning service and records the time when the positioning request is sent, which is recorded as t T , and back up the state of the visual positioning time series model at this time; ② Since t T From the moment t T The speed value and IMU sensor data at each moment after time t, until reaching ③ t T+P The Nth frame of image sent by the cloud is received at time t N The visual positioning pose at the moment, that is, the request result stops; ④ Using the received N-th frame image corresponding to t N The visual positioning pose at time t T ⑤ The state quantity of the visual positioning time series model at the moment is corrected; t TThe state quantity of the visual positioning time sequence model at the time tN, the previously recorded speed value at each time, and the IMU sensor data are input into the visual positioning time sequence model, and the camera pose estimation data at the last time obtained through prediction and deduction is taken as the corrected camera pose estimation data corresponding to the Nth frame of image at the time tN N The camera pose estimation data at the time tN is obtained, and then the speed value and the IMU sensor data are combined to continue to predict and deduce the camera pose estimation data at the subsequent time.

[0142] In an embodiment of the present disclosure, the tN-1th frame of image is used to correct the state quantity of the visual positioning time sequence model at the time tN. N The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the visual positioning time sequence model to correct the state quantity of the visual positioning time sequence model at the time tN. T The step of correcting the state quantity of the visual positioning time sequence model at the time tN can be implemented as:

[0143] The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the visual positioning time sequence model to correct the state quantity of the visual positioning time sequence model at the time tN. N The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the visual positioning time sequence model to correct the state quantity of the visual positioning time sequence model at the time tN. T The state quantity of the visual positioning time sequence model at the time tN.

[0144] As mentioned above, the visual positioning time sequence model based on the Kalman filter includes a prediction part based on a motion model and a correction part based on observation, and therefore, in this embodiment, the visual positioning pose at the time tN-1 is taken as the observation quantity and input into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time tN is obtained. T The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time tN is obtained. T The state quantity of the visual positioning time sequence model at the time tN.

[0145] Similarly, in an embodiment of the present disclosure, the step S104 of correcting the camera pose estimation data at the time tN-1 output by the visual positioning time sequence model based on the visual positioning pose at the time tN-1 can be implemented as: N The camera pose estimation data at the time tN-1 output by the visual positioning time sequence model based on the visual positioning pose at the time tN-1 is corrected. N The step of correcting the camera pose estimation data at the time tN-1 output by the visual positioning time sequence model based on the visual positioning pose at the time tN-1 can be implemented as:

[0146] The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the visual positioning time sequence model to correct the state quantity of the visual positioning time sequence model at the time tN. N The visual positioning pose at the time tN-1 is taken as the observation quantity and input into the visual positioning time sequence model to correct the state quantity of the visual positioning time sequence model at the time tN. N The positioning result at the time tN.

[0147] Figure 4 A flowchart of a positioning method according to another embodiment of the present disclosure is shown as follows. Figure 4 As shown in the figure, the positioning method includes the following steps S401-S403:

[0148] In step S401, in response to receiving a positioning request sent by a terminal, position data and an image carried by the positioning request are obtained.

[0149] In step S402, camera pose estimation data is calculated based on the position data and the image;

[0150] In step S403, the camera pose estimation data is sent to the terminal as an input of a pre-trained visual positioning time sequence model running on the terminal or for correcting a result output by the visual positioning time sequence model.

[0151] It is mentioned above that, with the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter level and below) positioning results of a positioned object (such as a smart phone, a car, a robot, etc.). In order to obtain high-precision positioning results, the prior art provides a positioning scheme based on a high-precision GNSS+INS integrated navigation system, which obtains the positioning result (latitude and longitude position and attitude) of the positioned object at each moment by means of the high-precision GNSS+INS integrated navigation system. The inventors find in the research of the scheme that, although the scheme can obtain high-precision positioning results, the sensors and chips used by the high-precision GNSS+INS integrated navigation system are expensive, and the high cost leads to poor applicability of the scheme. Therefore, it is urgent to provide a more universal positioning scheme that can guarantee positioning accuracy and is low in cost.

[0152] In view of the above problems, in this embodiment, a positioning method is proposed, which obtains a visual positioning pose based on only IMU sensor data obtained by a civilian IMU and a visual positioning time sequence model pre-trained by a terminal and sends the visual positioning pose to the terminal, so that the terminal performs positioning. Compared with the prior art, the method does not need expensive sensors and chips, and further, the technical scheme combines images collected by a camera in the positioning process, guarantees the accuracy of the final positioning result, and is a low-cost positioning scheme with stronger universality.

[0153] In an embodiment of the present disclosure, the step S402, i.e., the step of calculating camera pose estimation data based on the position data and the image, can include the following steps:

[0154] determining a candidate reference image with a distance to the position data within a preset range from a reference image database according to the position data, wherein the reference image database stores images with pre-calculated pose data;

[0155] calculating an image similarity between the image and the candidate reference image, and determining the candidate reference image with an image similarity satisfying a preset condition as a reference image corresponding to the image;

[0156] calculating relative pose data between the image and the reference image;

[0157] Obtaining pose data of the reference object, and calculating pose data of the image based on the pose data of the reference object and the relative pose data, which is taken as camera pose estimation data.

[0158] In this embodiment, the camera pose estimation data corresponding to the time point of the terminal-sent image is determined based on the comparison between the reference image with pre-calculated pose data and the terminal-sent image. Specifically, first, one or more candidate reference images with a distance to the location data within a preset range are determined from the reference image database according to the terminal-sent location data, that is, one or more images with a closer distance to the location data are taken as the candidate reference images, wherein the reference image database refers to a database storing images with pre-calculated accurate pose data that can be used as a location reference in the future, and the preset range can be set according to the actual application needs; then, the image similarity between the terminal-sent image and the candidate reference image is calculated, and the candidate reference image with an image similarity satisfying a preset condition, such as being greater than a certain preset similarity threshold, is determined as the reference image corresponding to the image; then, the relative pose data between the image and the reference image is calculated; finally, the pre-calculated pose data of the reference object is obtained, and the pose data of the image is calculated based on the pose data of the reference object and the previously calculated relative pose data, such as adding the pose data of the reference object and the relative pose data, so that the pose data of the image can be obtained, which can be taken as the camera pose estimation data corresponding to the time point of the image. For example, if the image is an initial frame image, the obtained camera pose estimation data is the camera initial pose estimation data, which can be taken as the initial input of the pre-trained visual positioning time sequence model running on the terminal; if the image is the Nth frame image, the obtained camera pose estimation data is the camera pose estimation data at time point t N

[0159] In an embodiment of the present disclosure, the step of calculating the relative pose data between the image and the reference image can be implemented as:

[0160] The image and the reference image are input into a pre-trained relative pose data calculation model to calculate the relative pose data between the image and the reference image.

[0161] In this embodiment, the calculation of the relative pose data is obtained by means of a pre-trained deep neural network model or other relative pose model, wherein the pre-training process of the relative pose model belongs to the technology that should be mastered by those skilled in the art, and the present disclosure does not repeat it here.

[0162] ​In an embodiment of the present disclosure, when the image is an initial frame image, the position data is GNSS position data, i.e., the camera initial pose estimation data is calculated using the GNSS position data; when the image is an Nth frame image acquired after the initial frame image, the position data is the position data in the camera pose estimation data at time t N , i.e., the camera pose estimation data is calculated using the position data in the camera pose estimation data at time t N , wherein the camera pose estimation data at time t N is predicted by the visual positioning time-series model according to the IMU sensing data at time t N and the camera pose estimation data at time t N-1 .

[0163] In another embodiment of the present disclosure, when the image is an Nth frame image acquired after the initial frame image, the position data can also not directly use the position data in the camera pose estimation data at time t N , but similar to the calculation process of the camera initial pose estimation data, the camera pose estimation data is first calculated using the GNSS position data at time t N . At this time, the method can further include the following steps:

[0164] When no candidate reference image within a preset distance range from the GNSS position data can be determined from the reference image database according to the GNSS position data, a retransmission position data request is sent to the terminal, so that the terminal sends the position data in the camera pose estimation data at time t N predicted by the visual positioning time-series model to the visual positioning service;

[0165] A candidate reference image within a preset distance range from the position data is determined from the reference image database according to the received position data.

[0166] As mentioned above, in actual applications, the GNSS position data corresponding to some image frames can have unpredictable drift, so that the cloud cannot calculate or obtain accurate camera pose estimation data. At this time, the GNSS position data needs to be corrected, so that the cloud can calculate camera pose estimation data according to the corrected position data, thereby avoiding visual positioning errors caused by GNSS position positioning errors. Therefore, further, when no candidate reference image within a preset distance range from the GNSS position data can be determined from the reference image database according to the GNSS position data sent by the terminal, a retransmission position data request is sent to the terminal, so that the terminal sends the position data in the camera pose estimation data at time t NThe position data in the camera pose estimation data at the moment is sent to the cloud; after receiving the new position data, the cloud determines the candidate reference image within the preset range from the reference image database according to the new position data and the distance between the position data.

[0167] The following is a detailed description of the scheme provided by the present disclosure, taking the example of making a correction every five frames, wherein the terminal sends the initial frame image collected by the camera, i.e., the 0th frame image, and the GNSS position data output by the GNSS receiver to the visual positioning service, so that the visual positioning service calculates the camera pose estimation data at the t0 moment based on the 0th frame image and the GNSS position data as the initial camera pose estimation data of the terminal; for the 1st frame image collected by the camera, the terminal obtains the IMU sensing data at the t1 moment, and inputs the IMU sensing data at the t1 moment and the camera pose estimation data at the t0 moment into the pre-trained visual positioning time sequence model to predict the camera pose estimation data at the t1 moment; for the 2nd frame image collected by the camera, the difference compared with the t1 moment is that the camera pose estimation data input into the visual positioning time sequence model is the camera pose estimation data at the t1 moment predicted by the visual positioning time sequence model, and the processing principles of the 3rd-5th frame images are the same as those of the 2nd frame image, which will not be described here; for the 6th frame image collected by the camera, because the visual positioning correction condition is met, the camera pose estimation data at the t6 moment predicted by the visual positioning time sequence model needs to be obtained according to the same processing principle as that of the 2nd-5th frame images, and the position data in the 6th frame image and the camera pose estimation data at the t6 moment are sent to the visual positioning service to request the visual positioning service to return the visual positioning pose at the t6 moment, the camera pose estimation data at the t6 moment predicted by the visual positioning time sequence model is corrected using the visual positioning pose at the t6 moment returned by the visual positioning service to obtain the positioning result at the t6 moment (i.e., the final camera pose estimation data at the t6 moment); for the 7th frame image collected by the camera, the corrected positioning result at the t6 moment and the obtained IMU sensing data at the t7 moment are input into the pre-trained visual positioning time sequence model to predict the camera pose estimation data at the t7 moment, and so on for the subsequent frames, which will not be described here.

[0168] Figure 4 The technical terms and technical features involved in the embodiments shown and related embodiments are the same as or similar to those in the embodiments shown and related embodiments Figures 1-3 The technical terms and technical features mentioned in the embodiments shown and related embodiments are the same as or similar to those in the embodiments shown and related embodiments Figure 4 The explanations and descriptions of the technical terms and technical features involved in the embodiments shown and related embodiments can refer to the explanations and descriptions of the embodiments shown and related embodiments Figures 1-3 The explanations and descriptions of the embodiments shown and related embodiments can refer to the explanations and descriptions of the embodiments shown and related embodiments

[0169] Figure 5A flow chart of a positioning method according to still another embodiment of the present disclosure is shown in FIG. 7, which includes the following steps: Figure 5

[0170] The terminal requests a visual positioning service to return camera initial pose estimation data at time t0 based on initial frame images captured by the camera and GNSS position data output by a GNSS receiver;

[0171] The cloud calculates camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame images in response to receiving the positioning request sent by the terminal, and sends the camera initial pose estimation data at time t0 to the terminal;

[0172] The terminal acquires IMU sensor data at time tN for an Nth frame image captured by the camera after the initial frame image, inputs the IMU sensor data at time tN and camera pose estimation data at time tN into a pre-trained visual positioning time sequence model, and predicts camera pose estimation data at time tN output by the visual positioning time sequence model, wherein N is an integer greater than or equal to 1, and if a set visual positioning correction condition is met when the Nth frame image is captured, the terminal requests a visual positioning service to return visual positioning pose at time tN based on position data at time tN and the Nth frame image, corrects the camera pose estimation data at time tN output by the visual positioning time sequence model based on the visual positioning pose at time tN, and obtains positioning results at time tN. N N N-1 N N N N N N N N N

[0173] ​​​​​​​​​​​​​With the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) require high-precision (meter level and below) positioning results of a positioned object (such as a smart phone, a car, a robot, etc.). In order to obtain high-precision positioning results, the prior art provides a positioning scheme based on a high-precision GNSS+INS integrated navigation system, which obtains the positioning results (latitude and longitude position and attitude) of the positioned object at each moment by means of the high-precision GNSS+INS integrated navigation system. The inventors found in the research of the scheme that, although the scheme can obtain high-precision positioning results, the sensors and chips used by the high-precision GNSS+INS integrated navigation system are expensive, and the high cost leads to poor applicability of the scheme. Therefore, there is an urgent need to provide a more universal positioning scheme that can guarantee positioning accuracy and is low in cost.

[0174] In view of the above problems, in this embodiment, a positioning method is proposed, which only uses IMU sensor data obtained by a civilian-grade IMU, a visual positioning time sequence model pre-trained by a terminal, and visual positioning poses provided by a visual positioning service from a cloud, compared with the prior art, the method does not need expensive sensors and chips, further, the technical scheme combines images collected by a camera and visual positioning poses provided by a visual positioning service in the positioning process, which guarantees the accuracy of the final positioning result, and is a low-cost positioning scheme with stronger universality.

[0175] In an embodiment of the present disclosure, the positioning method can be applied to a positioning system including a terminal and a cloud.

[0176] In an embodiment of the present disclosure, the cloud calculates camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame image, including:

[0177] determining a candidate reference image with a distance to the GNSS position data within a preset range from a reference image database according to the GNSS position data, wherein the reference image database stores images with pre-calculated pose data;

[0178] calculating an image similarity between the initial frame image and the candidate reference image, and determining the candidate reference image satisfying a preset condition as the reference image corresponding to the initial frame image;

[0179] calculating relative pose data between the initial frame image and the reference image;

[0180] obtaining pose data of the reference object, and calculating pose data of the initial frame image based on the pose data of the reference object and the relative pose data, and taking the pose data as the camera initial pose estimation data at time t0.

[0181] In an embodiment of the present disclosure, the calculating the relative pose data between the initial frame image and the reference image is implemented as:

[0182] inputting the initial frame image and the reference image into a pre-trained relative pose data calculation model to calculate the relative pose data between the initial frame image and the reference image.

[0183] In an embodiment of the present disclosure, the t N time position data is position data in the camera pose estimation data at the t N time.

[0184] In an embodiment of the present disclosure, further comprising:

[0185] when the cloud cannot determine the candidate reference image with the distance to the GNSS position data within the preset range from the reference image database according to the GNSS position data, sending a retransmission position data request to the terminal, and determining the candidate reference image with the distance to the position data within the preset range from the reference image database according to the received position data;

[0186] the terminal sends the position data in the camera pose estimation data at the t N time to the cloud in response to receiving the retransmission position data request.

[0187] In an embodiment of the present disclosure, further comprising:

[0188] the terminal performs asynchronous correction on the visual positioning pose at the t N time to obtain the asynchronous corrected visual positioning pose at the t N time.

[0189] In an embodiment of the present disclosure, the terminal performs asynchronous correction on the visual positioning pose at the t N time to obtain the asynchronous corrected visual positioning pose at the t N time, comprising:

[0190] after sending the positioning request to the visual positioning service, recording the positioning request sending time t T , and backing up the state quantity of the visual positioning timing model at the t T time;

[0191] recording the speed value and the IMU sensor data at each time after the t T time until receiving the visual positioning pose at the t T+P time; N time;

[0192] Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment;

[0193] The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning pose at the moment.

[0194] Figure 5 The technical terms and technical features involved in the embodiments shown and related Figures 1-4 The technical terms and technical features mentioned in the embodiments shown and related are the same or similar. Figure 5 The explanation and description of the technical terms and technical features involved in the embodiments shown and related can refer to the above Figures 1-4 The explanations of the illustrated and related embodiments will not be repeated here.

[0195] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.

[0196] Figure 6 FIG1 shows a structural block diagram of a positioning device according to an embodiment of the present disclosure. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 6 As shown, the positioning device includes:

[0197] The first request module 601 is configured to request the visual positioning service to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver; wherein the camera can be a camera, and the GNSS receiver can be a GPS receiver. The camera and GPS receiver can be integrated into the same terminal, such as a mobile phone or a driving recorder, or they can be independent components but mounted on the same object to be positioned, such as a car, a robot, etc.

[0198] The prediction module 602 is configured to obtain t N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, where N is an integer greater than or equal to 1;

[0199] The second request module 603 is configured to, if the set visual positioning correction condition is met when the Nth frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N Visual positioning posture at the moment;

[0200] The correction module 604 is configured to N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

[0201] As mentioned above, with the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter-level and below) positioning results of the located objects (such as smart phones, cars, robots, etc.). In order to obtain high-precision positioning results, the existing technology provides a positioning solution based on a high-precision GNSS+INS combined navigation system. This solution uses the high-precision GNSS+INS combined navigation system to obtain the positioning results (latitude and longitude positions and postures) of the located objects at every moment. In the process of studying this solution, the inventors found that although this solution can obtain high-precision positioning results, the sensors and chips used in the high-precision GNSS+INS combined navigation system are expensive. The excessively high cost makes the applicability of this solution not strong. Therefore, there is an urgent need to provide a more universal positioning solution that can guarantee positioning accuracy and is low-cost.

[0202] Taking the above issues into consideration, in this embodiment, a positioning device is proposed. This device performs positioning based solely on IMU sensor data acquired by a civilian-grade IMU, a pre-trained visual positioning time series model, and the visual positioning pose provided by a cloud-based visual positioning service. Compared with existing technologies, this device does not require expensive sensors and chips. Furthermore, this technical solution combines the images captured by the camera and the visual positioning pose provided by the visual positioning service during the positioning process, ensuring the accuracy of the final positioning result. It is a low-cost and more universal positioning solution. This positioning solution can be applied to AR navigation scenarios. When the user starts the AR navigation service, this solution begins to run. It can also be applied to other scenarios that require higher-precision positioning results, such as elevated roads, main and auxiliary roads, etc.

[0203] In one embodiment of the present disclosure, the positioning device may be implemented as a terminal computer, a computing device, an electronic device, a server, a service cluster, etc. that can perform positioning processing.

[0204] In one embodiment of the present disclosure, the positioning device may be implemented based on an image sequence including multiple frames of images.

[0205] In an embodiment of the present disclosure, the initial frame image refers to the 0th frame in the image sequence including multiple frames obtained by a camera.

[0206] In an embodiment of the present disclosure, the acquisition time of the initial frame image and the acquisition time of the GNSS position data can be the same or different, as long as the GNSS position data is the GNSS position data closest to the position of the initial frame image in the position data output by the GNSS receiver. The GNSS position data can be, for example, GPS position data, Beidou position data, etc.

[0207] In an embodiment of the present disclosure, the camera initial pose estimation data at time t0 refers to the camera pose estimation data at the time corresponding to the initial frame image determined by the cloud visual positioning service. In this embodiment, the camera pose estimation data at time t0 determined by the cloud visual positioning service can be used as the camera initial pose estimation data at time t0 corresponding to the initial frame image, and can also be used as the initial state data of the visual positioning time sequence model for subsequent prediction of the camera pose.

[0208] In an embodiment of the present disclosure, the pose estimation data can include position data such as latitude and longitude, and attitude data such as attitude yaw angle.

[0209] In an embodiment of the present disclosure, the IMU sensor data refers to the data detected by an IMU (Inertial Measurement Unit) sensor. For example, for a 6-axis IMU sensor composed of 3 accelerometers and 3 gyroscopes, the detected data can include 3-direction acceleration and 3-direction angular velocity.

[0210] In an embodiment of the present disclosure, the positioning time sequence model refers to a model pre-trained to be able to predict the visual positioning data at the next time. For example, when N = 1, after obtaining the camera initial pose estimation data at time t0, the camera pose estimation data at time t1 corresponding to the 1st frame image can be predicted according to the camera initial pose estimation data at time t0 and the acquired IMU sensor data at time t1. Then, when N = 2, the camera pose estimation data at time t2 corresponding to the 2nd frame image can be predicted according to the camera pose estimation data at time t1 and the acquired IMU sensor data at time t2, and so on, so that the camera pose estimation data at the time corresponding to all images can be obtained. That is, for the Nth frame image collected by the camera after the initial frame image, the IMU sensor data at time tN-1 is acquired, the IMU sensor data at time tN-1 and the camera pose estimation data at time tN-1 are used to predict the camera pose estimation data at time tN corresponding to the Nth frame image, and so on, until the camera pose estimation data at the time corresponding to the last frame image is obtained. N N N-1 ​​The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, where N is an integer greater than or equal to 1.

[0211] In one embodiment of the present disclosure, the visual positioning timing model can be implemented based on a Kalman filter. In this embodiment, the visual positioning timing model includes a prediction part based on a motion model and a correction part based on observation results. Figure 2 is a schematic diagram of a motion model according to an embodiment of the present disclosure, as shown in FIG. Figure 2 As shown, in this motion model, the three state quantities of target position P, target posture R and target velocity V can be estimated. In this embodiment, the motion model can satisfy the following update relationship:

[0212] R N+1 =R N exp((ω N dt) × )

[0213] V N+1 =V N +(R N a N +g)dt

[0214] P N+1 =P N +V N dt

[0215] Among them, ω N and a N They are respectively represented as angular velocity and acceleration in IMU sensing data.

[0216] Based on the above visual positioning timing model, we can N The state at time t is used to predict N+1 The state of the moment.

[0217] In this embodiment, for the images collected after the initial frame image, the camera pose estimation data at the corresponding moment predicted by the visual positioning time series model can be directly used as the positioning result of the image at the corresponding moment. However, if it is determined that the set visual positioning correction condition is met when the Nth frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, and then based on the t N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtainN the positioning result at the time t0. N The position data at the time t0 refers to the position data in the camera pose estimation data at the time t0. N The position data at the time t0 refers to the position data in the camera pose estimation data at the time t0.

[0218] The above embodiment can obtain accurate camera pose estimation data based on the spatiotemporal continuity assumption and in combination with historical time series information by means of the visual positioning time series model, thereby better improving the problem of unstable single-frame prediction in visual positioning and effectively improving the accuracy of visual positioning. Meanwhile, since the Kalman filter algorithm is very efficient, the calculation cost of the terminal can be basically ignored, that is, the calculation consumption can be greatly saved.

[0219] In an embodiment of the present disclosure, the first request module 601 can be configured to:

[0220] obtain the initial frame image collected by the camera and the GNSS position data output by the GNSS receiver, and send a positioning request to the visual positioning service, wherein the positioning request carries the initial frame image and the GNSS position data, so that the visual positioning service calculates and returns the camera initial pose estimation data at the time t0 based on the initial frame image and the GNSS position data.

[0221] Since the historical time series information at the time t0 is difficult to predict, in this embodiment, the camera initial pose estimation data at the time t0 is obtained by means of the visual positioning service of the cloud. Specifically, the terminal obtains the GNSS position data at the time t0 and the initial frame image, and sends the GNSS position data at the time t0 and the initial frame image to the visual positioning service of the cloud in the positioning request, so as to obtain the camera pose estimation data at the time t0 calculated by the visual positioning service of the cloud based on the GNSS position data at the time t0 and the initial frame image. The camera pose estimation data at the time t0 obtained by the visual positioning service of the cloud can be used as the camera initial pose estimation data at the time t0.

[0222] Similarly, the second request module 603 can be configured to:

[0223] If the set visual positioning correction condition is met when the Nth frame image is collected, the Nth frame image collected by the camera and the position data at the time tN are obtained. N

[0224] The Nth frame image and the position data at the time tN are sent to the visual positioning service in a positioning request, so that the visual positioning service calculates and returns the visual positioning pose at the time tN based on the Nth frame image and the position data at the time tN. N N N the visual positioning pose at the time tN.​​​

[0225] Considering that the visual positioning timing model has a certain estimation bias, in order to improve the accuracy of the camera pose estimation data, the camera pose estimation data generated by the visual positioning timing model needs to be corrected by referring to the cloud camera pose estimation data. However, considering the control of cloud computing cost, not every frame of image needs to refer to the cloud camera pose estimation data, but first determines whether the visual positioning correction condition set when collecting the Nth frame of image is met, if yes, the Nth frame of image collected by the camera and the position data at t N time are obtained, and a positioning request is sent to the visual positioning service, wherein the positioning request carries the Nth frame of image and the position data at t N time, so that the visual positioning service calculates and returns the visual positioning pose at t N time based on the Nth frame of image and the position data at t N time. Subsequently, the visual positioning pose at t N time can be corrected for the camera pose estimation data at t N time output by the visual positioning timing model, to obtain the final positioning result at t N time. The visual positioning correction condition can include: correction is performed once every M frames of image, or the scene of collecting the image is a scene set for visual positioning correction, such as GNSS positioning signal being blocked to cause inaccurate GNSS position, encountering a fork or turning to cause certain error of IMU data, uncertainty of the currently used visual positioning timing model exceeding a preset threshold (i.e., the error of the visual positioning timing model is large), or communication network interruption or unstable data transmission.

[0226] In another embodiment of the present disclosure, when the Nth frame of image is collected, the position data in the camera pose estimation data at t N time can not be directly used as the position data at t N time, but similar to the process of obtaining the initial camera pose estimation data, the GNSS position data output by the GNSS receiver at t N time is used as the position data at t N time, and the visual positioning service is requested to return the visual positioning pose at t N time based on the GNSS position data at t N time and the Nth frame of image. At this time, the second request module 603 can be configured to:

[0227] if the visual positioning correction condition set when the Nth frame of image is collected is met, the Nth frame of image collected by the camera and the position data at tN GNSS position data output by the GNSS receiver at the time t N-1 ;

[0228] sending a positioning request to the visual positioning service, wherein the positioning request carries the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ; N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ; N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ; N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ;

[0229] However, in actual applications, the GNSS position data corresponding to some image frames may drift unpredictably, causing the cloud to be unable to calculate or obtain accurate camera pose estimation data, at which time, the GNSS position data needs to be corrected, and the corrected GNSS position data is sent to the cloud, so that the cloud calculates camera pose estimation data based on the corrected GNSS position data, thereby avoiding visual positioning errors caused by GNSS positioning errors. Considering that the visual positioning timing model implemented based on the Kalman filter can predict the pose data at the current time from the pose data at the previous time through a motion model, when the GNSS position data needs to be corrected, the position data in the pose data at the current time predicted by using the visual positioning timing model is used as the correction data of the GNSS position data that has drifted, and the position data in the camera pose estimation data at the time t N is used as the position data at the time t N. N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ; N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ;

[0230] That is, in an embodiment of the present disclosure, the apparatus further comprises:

[0231] The first sending module is configured to, in response to receiving the retransmission position data request sent by the visual positioning service, send the position data in the camera pose estimation data at the time t N to the visual positioning service, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the position data. N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ; N GNSS position data output by the GNSS receiver at the time t N-1, so that the visual positioning service calculates and returns the visual positioning pose at the time t N based on the Nth image and the GNSS position data output by the GNSS receiver at the time t N-1 ;

[0232] The retransmission position data request refers to a request sent by the cloud visual positioning service to the terminal to make the terminal retransmit position data, because the GNSS position data previously sent to the cloud visual positioning service cannot be used or cannot be used to obtain accurate camera pose estimation data.

[0233] In an embodiment of the present disclosure, the collection time of the Nth image and the collection time of the GNSS position data can be the same or different, as long as the GNSS position data is the GNSS position data closest to the position of the Nth image in the position data output by the GNSS receiver.

[0234] In an embodiment of the present disclosure, the correction module 604 can be configured to:

[0235] For the visual positioning pose at the t N time, asynchronous correction is performed to obtain the visual positioning pose at the t N time after asynchronous correction.

[0236] Based on the visual positioning pose at the t N time after asynchronous correction, the camera pose estimation data output by the visual positioning time sequence model at the t N time is corrected to obtain the positioning result at the t N time.

[0237] In actual applications, due to network transmission, cloud computing and other factors, there is a certain time overhead for the terminal to call the cloud camera pose estimation data calculation service, which causes an asynchronous phenomenon between the cloud and the terminal, that is, when the terminal sends a positioning request and receives the camera pose estimation data sent by the cloud after a certain time, the camera pose estimation data received at this time is not the cloud camera pose estimation data corresponding to the time when the terminal sends the positioning request. Therefore, in this embodiment, the camera pose estimation data at the t N time corresponding to the Nth image needs to be corrected asynchronously, and then the camera pose estimation data at the t N time after asynchronous correction is used to correct the camera pose estimation data output by the visual positioning time sequence model at the t N time to obtain the positioning result at the t N time.

[0238] In an embodiment of the present disclosure, the asynchronous correction of the visual positioning pose at the t N time to obtain the visual positioning pose at the t N time after asynchronous correction can be configured as:

[0239] After sending a positioning request to the visual positioning service, the positioning request sending time t T is recorded, and the state quantity of the visual positioning time sequence model at the t T time is backed up.

[0240] The t Tthe speed value and the IMU sensor data of each time after the time t T+P the time t N the visual positioning pose at the time t

[0241] the visual positioning pose at the time t N the visual positioning pose at the time t T correct the state quantity of the visual positioning time series model at the time t

[0242] the visual positioning pose at the time t T the visual positioning pose at the time t N the visual positioning pose at the time t

[0243] In this embodiment, the correction of the visual positioning pose at the time t N is realized by correcting the state quantity of the visual positioning time series model at the time when the positioning request is sent, and predicting by using the corrected state quantity of the visual positioning time series model. Specifically, as shown in Figure 3 , wherein, Figure 3 the serial number in represents the execution order, ① the terminal sends a visual positioning request to the cloud visual positioning service, records the time when the positioning request is sent, denoted as t T , and backs up the state quantity of the visual positioning time series model at this time; ② from the time t T , the speed value and the IMU sensor data of each time after the time t T are recorded, until ③ the time t T+P , the visual positioning pose at the time t N corresponding to the Nth frame of image sent by the cloud is received, that is, the request result stops; ④ the visual positioning pose at the time t N corresponding to the Nth frame of image is used to correct the state quantity of the visual positioning time series model at the time t T ; ⑤ the state quantity of the visual positioning time series model at the time t T after correction, the speed value and the IMU sensor data of each time recorded before are input into the visual positioning time series model, and the camera pose estimation data of the last time obtained by prediction and deduction is taken as the camera pose estimation data of the Nth frame of image corresponding to the time t N after correction, and then ⑥ the speed value and the IMU sensor data can be combined to continue to predict and deduce the camera pose estimation data of the subsequent time.

[0244] In an embodiment of the present disclosure, the visual positioning pose at the time t N is used to correct the state quantity of the visual positioning time series model at the time t TThe part that corrects the state quantity of the visual positioning time sequence model at the time t N can be configured as:

[0245] The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. N The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. T The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained.

[0246] As mentioned above, the visual positioning time sequence model based on the Kalman filter includes a prediction part based on a motion model and a correction part based on an observation result, so in this embodiment, the visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. T The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. T The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained.

[0247] Similarly, in an embodiment of the present disclosure, the visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. N The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. N The part that corrects the camera pose estimation data at the time t N output by the visual positioning time sequence model can be configured as:

[0248] The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. N The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained. N The visual positioning pose at the time t N is input as an observation quantity into the correction part of the visual positioning time sequence model to perform a correction operation, so that the corrected state quantity of the visual positioning time sequence model at the time t N is obtained.

[0249] Figure 7 A structural block diagram of a cloud server according to an embodiment of the present disclosure is shown, which can be realized by software, hardware or a combination of the two to become part or all of an electronic device. As shown in Figure 7 The positioning apparatus includes:

[0250] The acquisition module 701 is configured to acquire position data and images carried by a positioning request sent by a terminal in response to receiving the positioning request.

[0251] The calculation module 702 is configured to calculate camera pose estimation data based on the position data and images.

[0252] The second sending module 703 is configured to send the camera pose estimation data to the terminal as an input of a pre-trained visual positioning time sequence model running on the terminal or for correcting a result output by the visual positioning time sequence model.

[0253] With the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter level and below) positioning results of a positioned object (such as a smart phone, a car, a robot, etc.). In order to obtain high-precision positioning results, the prior art provides a positioning scheme based on a high-precision GNSS+INS integrated navigation system, which obtains the positioning results (latitude and longitude position and attitude) of the positioned object at each moment by means of the high-precision GNSS+INS integrated navigation system. The inventors found in the research of the scheme that, although the scheme can obtain high-precision positioning results, the sensors and chips used by the high-precision GNSS+INS integrated navigation system are expensive, and the high cost leads to poor applicability of the scheme. Therefore, it is urgent to provide a more universal positioning scheme that can guarantee positioning accuracy and is low in cost.

[0254] In view of the above problems, in this embodiment, a cloud server is proposed, which obtains a visual positioning pose based on only IMU sensor data obtained by a civilian IMU and a visual positioning time sequence model trained in advance by a terminal and sends the visual positioning pose to the terminal, so that the terminal performs positioning. Compared with the prior art, the cloud server does not need expensive sensors and chips, and further, the technical scheme combines images collected by a camera in the positioning process, guarantees the accuracy of the final positioning result, and is a low-cost positioning scheme with stronger universality.

[0255] In an embodiment of the present disclosure, the computing module 702 can be configured to:

[0256] determine a candidate reference image with a distance to the position data within a preset range from the position data from a reference image database, wherein the reference image database stores images with pre-computed pose data;

[0257] calculate an image similarity between the image and the candidate reference image, and determine the candidate reference image with an image similarity satisfying a preset condition as the reference image corresponding to the image;

[0258] calculate relative pose data between the image and the reference image;

[0259] obtain pose data of the reference object, and calculate pose data of the image based on the pose data of the reference object and the relative pose data, and take the pose data of the image as camera pose estimation data.

[0260] In this embodiment, the camera pose estimation data corresponding to the time when the terminal sends the image is determined based on the comparison between the reference image with pre-calculated pose data and the terminal-sent image. Specifically, first, one or more candidate reference images with distances to the location data sent by the terminal within a preset range are determined from the reference image database, i.e., one or more images with closer distances to the location data are taken as the candidate reference images, wherein the reference image database refers to a database storing images with pre-calculated accurate pose data that can be used as a location reference in the future, and the preset range can be set according to the actual application needs; then, the image similarity between the terminal-sent image and the candidate reference images is calculated, and the candidate reference image with an image similarity satisfying a preset condition, such as being greater than a certain preset similarity threshold, is determined as the reference image corresponding to the image; then, the relative pose data between the image and the reference image is calculated; finally, the pre-calculated pose data of the reference object is obtained, and based on the pose data of the reference object and the previously calculated relative pose data, such as adding the pose data of the reference object and the relative pose data, the pose data of the image can be obtained, which can be taken as the camera pose estimation data corresponding to the time when the image is sent, such as, if the image is an initial frame image, the obtained camera pose estimation data is the camera initial pose estimation data, which can be taken as the initial input of the pre-trained visual positioning time sequence model running on the terminal; if the image is the Nth frame image, the obtained camera pose estimation data is the camera pose estimation data at time t N , which can be used to correct the output result of the visual positioning time sequence model at the corresponding time.

[0261] In an embodiment of the present disclosure, the part of calculating the relative pose data between the image and the reference image can be configured as:

[0262] The image and the reference image are input into a pre-trained relative pose data calculation model to calculate the relative pose data between the image and the reference image.

[0263] In this embodiment, the calculation of the relative pose data is obtained by means of a pre-trained deep neural network model or other relative pose model, wherein the pre-training process of the relative pose model belongs to the technology that should be mastered by those skilled in the art, and the present disclosure does not repeat it here.

[0264] In an embodiment of the present disclosure, when the image is an initial frame image, the location data is GNSS location data, i.e., the GNSS location data is used to calculate the camera initial pose estimation data; when the image is the Nth frame image collected after the initial frame image, the location data isN the position data in the camera pose estimation data at time t N the position data in the camera pose estimation data at time t N the camera pose estimation data at time t is predicted by the visual positioning time series model according to t N the IMU sensor data at time t and the camera pose estimation data at time t N-1 the camera pose estimation data at time t is predicted by the visual positioning time series model according to t

[0265] In another embodiment of the present disclosure, when the image is the Nth frame image acquired after the initial frame image, the position data can also not directly use the position data in the camera pose estimation data at time t N but similar to the calculation process of the camera initial pose estimation data, the position data in the camera pose estimation data at time t is used first to calculate the camera pose estimation data at time t N GNSS position data at time t. At this time, the device can further comprise:

[0266] a third sending module configured to send a retransmission position data request to the terminal when a candidate reference image with a distance from the GNSS position data within a preset range cannot be determined from the reference image database according to the GNSS position data, so that the terminal sends the position data in the camera pose estimation data at time t N predicted by the visual positioning time series model to the visual positioning service;

[0267] determine a candidate reference image with a distance from the position data within a preset range from the reference image database according to the received position data.

[0268] As mentioned above, in actual application, the GNSS position data corresponding to some image frames can have unpredictable drift, causing the cloud to be unable to calculate or obtain accurate camera pose estimation data. At this time, it is necessary to correct the GNSS position data, so that the cloud can calculate the camera pose estimation data according to the corrected position data, thereby avoiding the visual positioning error caused by the GNSS position positioning error. Therefore, further, when it is found that a candidate reference image with a distance from the GNSS position data within a preset range cannot be determined from the reference image database according to the GNSS position data sent by the terminal, a retransmission position data request is sent to the terminal, so that the terminal sends the position data in the camera pose estimation data at time t N predicted by the visual positioning time series model to the cloud; after receiving the new position data, the cloud determines a candidate reference image with a distance from the position data within a preset range from the reference image database according to the new position data again.

[0269] Figure 7 The technical terms and technical features involved in the embodiments shown and related embodiments are the same or similar to Figure 6 The technical terms and technical features mentioned in the embodiments shown and related embodiments are the same or similar to Figure 7 The explanations and descriptions of the technical terms and technical features involved in the embodiments shown and related embodiments can refer to the above explanations and descriptions of the technical terms and technical features involved in the embodiments shown and related embodiments. Figure 6 The explanations and descriptions of the technical terms and technical features involved in the embodiments shown and related embodiments can refer to the above explanations and descriptions of the technical terms and technical features involved in the embodiments shown and related embodiments.

[0270] Figure 8 A structural block diagram of a positioning device according to another embodiment of the present disclosure is shown, which can be realized by software, hardware or a combination of both as part of or all of an electronic device. As shown in the figure, Figure 8 The positioning device comprises:

[0271] The terminal 801 is configured to request a visual positioning service to return camera initial pose estimation data at time t0 of the camera based on initial frame images captured by a camera and GNSS position data output by a GNSS receiver, obtain IMU sensor data at time tN of an Nth frame image captured by the camera after the initial frame image, input the IMU sensor data at time tN and the camera pose estimation data at time t0 into a pre-trained visual positioning timing model, and predict camera pose estimation data at time tN. N N N-1 N N is an integer greater than or equal to 1, and if a set visual positioning correction condition is met when the Nth frame image is captured, the terminal 801 is configured to request the visual positioning service to return a visual positioning pose at time tN based on the position data at time t0 and the Nth frame image, correct the camera pose estimation data at time tN output by the visual positioning timing model based on the visual positioning pose at time tN, and obtain a positioning result at time tN. N N N N N

[0272] The cloud 802 is configured to calculate camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame images and send it to the terminal in response to receiving a positioning request sent by the terminal.

[0273] ​​​​​​​​As mentioned above, with the development of technology, more and more scenarios (such as unmanned driving, AR navigation, etc.) need to obtain high-precision (meter-level and below) positioning results of the located objects (such as smart phones, cars, robots, etc.). In order to obtain high-precision positioning results, the existing technology provides a positioning solution based on a high-precision GNSS+INS combined navigation system. This solution uses the high-precision GNSS+INS combined navigation system to obtain the positioning results (latitude and longitude positions and postures) of the located objects at every moment. In the process of studying this solution, the inventors found that although this solution can obtain high-precision positioning results, the sensors and chips used in the high-precision GNSS+INS combined navigation system are expensive. The excessively high cost makes the applicability of this solution not strong. Therefore, there is an urgent need to provide a more universal positioning solution that can guarantee positioning accuracy and is low-cost.

[0274] Taking the above problems into consideration, in this embodiment, a positioning device is proposed. The device performs positioning only based on the IMU sensor data obtained by the civilian-grade IMU, the visual positioning timing model pre-trained by the terminal, and the visual positioning posture provided by the visual positioning service from the cloud. Compared with the existing technology, it does not require expensive sensors and chips. Furthermore, this technical solution combines the images collected by the camera and the visual positioning posture provided by the visual positioning service during the positioning process, ensuring the accuracy of the final positioning result. It is a low-cost and more universal positioning solution.

[0275] In one embodiment of the present disclosure, the positioning device may be implemented as a positioning system including a terminal and a cloud.

[0276] In one embodiment of the present disclosure, the portion of the cloud that calculates the camera's initial pose estimation data at time t0 based on the GNSS position data and the initial frame image may be configured as follows:

[0277] Determining, from a reference image database based on the GNSS position data, candidate reference images whose distances to the GNSS position data are within a preset range, wherein the reference image database stores images having pre-calculated pose data;

[0278] Calculating the image similarity between the initial frame image and the candidate reference image, and determining the candidate reference image whose image similarity meets a preset condition as the reference image corresponding to the initial frame image;

[0279] Calculating relative pose data between the initial frame image and the reference image;

[0280] The pose data of the reference object is obtained, and the pose data of the initial frame image is calculated based on the pose data of the reference object and the relative pose data, and is used as the initial pose estimation data of the camera at time t0.

[0281] In one embodiment of the present disclosure, the portion of calculating the relative pose data between the initial frame image and the reference image may be configured as follows:

[0282] The initial frame image and the reference image are input into a pre-trained relative pose data calculation model to calculate the relative pose data between the initial frame image and the reference image.

[0283] In one embodiment of the present disclosure, the t N The position data at the moment t N The position data in the camera pose estimation data at the moment.

[0284] In one embodiment of the present disclosure,

[0285] The cloud is further configured to, when a candidate reference image whose distance to the GNSS position data is within a preset range cannot be determined from a reference image database based on the GNSS position data, send a request to resend the position data to the terminal, and determine, based on the received position data, a candidate reference image whose distance to the position data is within a preset range from the reference image database;

[0286] The terminal is further configured to, in response to receiving the request to resend the location data, N The position data in the camera pose estimation data at the moment is sent to the cloud.

[0287] In one embodiment of the present disclosure,

[0288] The terminal is further configured to N The visual positioning pose at the moment is asynchronously corrected to obtain the asynchronously corrected pose t N Visual positioning pose at the moment.

[0289] In one embodiment of the present disclosure, the terminal is N The visual positioning pose at the moment is asynchronously corrected to obtain the asynchronously corrected pose t N The visual positioning pose at the moment is also configured as:

[0290] After sending a positioning request to the visual positioning service, record the time t when the positioning request is sent T , and back up t T The state quantity of the visual positioning timing model at the moment;

[0291] Record T The speed value and IMU sensor data at each moment after time t T+P Received at time t NVisual positioning posture at the moment;

[0292] Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment;

[0293] The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning pose at the moment.

[0294] Figure 8 The technical terms and technical features involved in the embodiments shown and related Figures 1-7 The technical terms and technical features mentioned in the embodiments shown and related are the same or similar. Figure 8 The explanation and description of the technical terms and technical features involved in the embodiments shown and related can refer to the above Figures 1-7 The explanations of the illustrated and related embodiments will not be repeated here.

[0295] The present disclosure also discloses a navigation service, wherein, based on the positioning method described above, a positioning result of a navigated object is obtained, and based on the positioning result, a navigation guidance service for a corresponding scenario is provided to the navigated object. The corresponding scenario is one or a combination of AR navigation, elevated navigation, or main and auxiliary road navigation.

[0296] The present disclosure also discloses an electronic device, which includes a memory and a processor; wherein:

[0297] The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement any of the above method steps.

[0298] Figure 9 It is a structural diagram of a computer system suitable for implementing the positioning method according to an embodiment of the present disclosure.

[0299] like Figure 9 As shown, the computer system 900 includes a processing unit 901, which can execute various processes in the above-mentioned embodiments according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the computer system 900 are also stored in the RAM 903. The processing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0300] The following components are connected to the I / O interface 905: an input part 906 including a keyboard, a mouse, and the like; an output part 907 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage part 908 including a hard disk, and the like; and a communication part 909 including a network interface card such as a LAN card, a modem, and the like. The communication part 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as necessary. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 910 as necessary, so that a computer program read out therefrom is installed in the storage part 908 as necessary. Among them, the processing unit 901 can be implemented as a CPU, a GPU, a TPU, an FPGA, an NPU, and the like.

[0301] In particular, according to embodiments of the present disclosure, the method described above can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a non-transitory computer readable medium, the computer program containing program code for executing the positioning method. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 909, and / or installed from the removable medium 911.

[0302] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0303] The units or modules described in the embodiments of the present disclosure can be implemented by software, or can be implemented by hardware. The described units or modules can also be provided in a processor, and the names of the units or modules do not constitute a limitation on the units or modules themselves in some cases.

[0304] As another aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which can be the computer readable storage medium included in the apparatus in the above-mentioned embodiments; or can exist separately and not be assembled into the apparatus. The computer readable storage medium stores one or more programs for being executed by one or more processors to perform the method described in the embodiments of the present disclosure.

[0305] The above description is merely the preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the inventive scope of the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) can also be covered.

Claims

1. A positioning method, comprising: Based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, request the visual positioning service in the cloud to return the camera's initial pose estimation data at time t0, wherein the GNSS position data is the GNSS position data of the position data output by the GNSS receiver that is closest to the position where the initial frame image was captured; For the Nth frame image captured by the camera after the initial frame image, obtain t N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, where N is an integer greater than or equal to 1; If the set visual positioning correction condition is met when the Nth frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, where the t N The position data at the moment t N The position data in the camera pose estimation data at time t N GNSS position data output by the GNSS receiver at all times; Based on the N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

2. The method according to claim 1, wherein requesting the visual positioning service to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver comprises: Obtain the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, and send a positioning request to the visual positioning service, wherein the positioning request carries the initial frame image and GNSS position data, so that it calculates and returns the camera's initial pose estimation data at time t0 based on the initial frame image and GNSS position data.

3. The method according to claim 1 or 2, wherein if the set visual positioning correction condition is satisfied when the N-th frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, including: If the set visual positioning correction condition is met when the Nth frame image is collected, the Nth frame image collected by the camera and t N Location data at the moment; Send a positioning request to the visual positioning service, wherein the positioning request carries the Nth frame image and t N The position data at the moment, so that the visual positioning service is based on the Nth frame image and t N Calculate and return the position data at time t N Visual positioning pose at the moment.

4. The method according to claim 3, wherein the visual positioning correction condition comprises: Correction is performed every M frames of image, or the scene of the captured image is the set scene for visual positioning correction.

5. The method according to any one of claims 1-2 and 4, wherein the N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning results at the moment, including: After sending a positioning request to the visual positioning service, record the time t when the positioning request was sent. T , and back up t T The state quantity of the visual positioning timing model at the moment; Record T The speed value and IMU sensor data at each moment after time t T+P Received at time t N Visual positioning posture at the moment; Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment; The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning posture at the moment; Based on the asynchronously corrected t N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

6. A positioning method, comprising: In response to receiving the positioning request sent by the terminal, the cloud obtains the position data and image carried in the positioning request, and calculates the camera pose estimation data based on the position data and image. When the image is the initial frame image, the position data is the GNSS position data output by the GNSS receiver that is closest to the position where the initial frame image was taken; when the image is the Nth frame image collected after the initial frame image, the position data is t N The position data in the camera pose estimation data at time t N The GNSS position data output by the GNSS receiver at the moment t N The camera pose estimation data at the moment is obtained by the pre-trained visual positioning time series model running on the terminal according to t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is predicted; The cloud sends the camera pose estimation data to the terminal as input to the visual positioning timing model or for correcting the result output by the visual positioning timing model when a set visual positioning correction condition is met.

7. The method according to claim 6, wherein the step of calculating camera pose estimation data based on the position data and the image comprises: Determining, from a reference image database according to the position data, candidate reference images whose distances from the position data are within a preset range, wherein the reference image database stores images having pre-calculated pose data; Calculating the image similarity between the image and the candidate reference image, and determining the candidate reference image whose image similarity meets a preset condition as the reference image corresponding to the image; Calculating relative pose data between the image and the reference image; The pose data of the reference object is obtained, and the pose data of the image is calculated based on the pose data of the reference object and the relative pose data, and the pose data is used as the camera pose estimation data.

8. The method according to claim 7, wherein the calculating the relative pose data between the image and the reference image is implemented as follows: The image and the reference image are input into a pre-trained relative pose data calculation model to calculate the relative pose data between the image and the reference image.

9. A positioning method, comprising: The terminal requests the visual positioning service in the cloud to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, wherein the GNSS position data is the GNSS position data of the position data output by the GNSS receiver that is closest to the position at which the initial frame image was captured; In response to receiving the positioning request sent by the terminal, the cloud calculates the camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame image, and sends the estimated data to the terminal; The terminal obtains t for the Nth frame image captured by the camera after the initial frame image. N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, wherein N is an integer greater than or equal to 1, if the set visual positioning correction condition is met when the N-th frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, and based on the t N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N The positioning result at the moment, wherein the t N The position data at the moment t N The position data in the camera pose estimation data at time t N GNSS position data output by the GNSS receiver at this moment.

10. The method according to claim 9, wherein the cloud calculates the camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame image, comprising: Determining, from a reference image database based on the GNSS position data, candidate reference images whose distances to the GNSS position data are within a preset range, wherein the reference image database stores images having pre-calculated pose data; Calculating the image similarity between the initial frame image and the candidate reference image, and determining the candidate reference image whose image similarity meets a preset condition as the reference image corresponding to the initial frame image; Calculating relative pose data between the initial frame image and the reference image; The pose data of the reference object is obtained, and the pose data of the initial frame image is calculated based on the pose data of the reference object and the relative pose data, and is used as the initial pose estimation data of the camera at time t0.

11. The method according to claim 10, wherein the calculating the relative pose data between the initial frame image and the reference image is implemented as follows: The initial frame image and the reference image are input into a pre-trained relative pose data calculation model to calculate the relative pose data between the initial frame image and the reference image.

12. The method according to any one of claims 9 to 11, wherein the t N The position data at the moment t N The position data in the camera pose estimation data at the moment.

13. The method according to any one of claims 9 to 11, further comprising: After sending a positioning request to the visual positioning service, record the time t when the positioning request was sent. T , and back up t T The state quantity of the visual positioning timing model at the moment; Record T The speed value and IMU sensor data at each moment after time t T+P Received at time t N Visual positioning posture at the moment; Using the t N The visual positioning pose at time t T Correcting the state quantity of the visual positioning timing model at the moment; The corrected t T The state quantity of the visual positioning timing model at time t, the recorded speed value and the IMU sensor data are input into the visual positioning timing model to obtain the asynchronously corrected t N Visual positioning pose at the moment.

14. A positioning device comprising: A first request module is configured to request a visual positioning service located in the cloud to return initial camera pose estimation data at time t0 of the camera based on an initial frame image captured by the camera and GNSS position data output by the GNSS receiver, wherein the GNSS position data is GNSS position data of the position data output by the GNSS receiver that is closest to the position at which the initial frame image was captured; The prediction module is configured to obtain t N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, where N is an integer greater than or equal to 1; The second request module is configured to, if the set visual positioning correction condition is met when the Nth frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, where the t N The position data at the moment t N The position data in the camera pose estimation data at time t N GNSS position data output by the GNSS receiver at all times; The correction module is configured to N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N Positioning result at the moment.

15. A cloud server comprising: The acquisition module is configured to, in response to receiving a positioning request sent by the terminal, acquire the position data and image carried in the positioning request, and calculate the camera pose estimation data based on the position data and image, wherein, when the image is an initial frame image, the position data is the GNSS position data output by the GNSS receiver that is closest to the position where the initial frame image was taken; when the image is the Nth frame image collected after the initial frame image, the position data is t N The position data in the camera pose estimation data at time t N The GNSS position data output by the GNSS receiver at the moment t N The camera pose estimation data at the moment is obtained by the pre-trained visual positioning time series model running on the terminal according to t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is predicted; A sending module is configured to send the camera pose estimation data to the terminal as an input of the visual positioning timing model or to correct the result output by the visual positioning timing model when a set visual positioning correction condition is met.

16. A positioning device comprising: The terminal is configured to request the visual positioning service in the cloud to return the camera's initial pose estimation data at time t0 based on the initial frame image captured by the camera and the GNSS position data output by the GNSS receiver, wherein the GNSS position data is the GNSS position data of the position closest to the position where the initial frame image was captured among the position data output by the GNSS receiver, and obtain t for the Nth frame image captured by the camera after the initial frame image. N The IMU sensor data at time t N IMU sensor data at time t N-1 The camera pose estimation data at the moment is input into the pre-trained visual positioning time series model to predict the t N The camera pose estimation data at the moment, wherein N is an integer greater than or equal to 1, if the set visual positioning correction condition is met when the N-th frame image is collected, then based on the t N The position data at the moment and the Nth frame image are requested to return the t N The visual positioning pose at the moment, and based on the t N The visual positioning pose at the moment is the output of the visual positioning timing model t N The camera pose estimation data at time t is corrected to obtain N The positioning result at the moment, wherein the t N The position data at the moment t N The position data in the camera pose estimation data at time t N GNSS position data output by the GNSS receiver at all times; The cloud is configured to, in response to receiving a positioning request sent by the terminal, calculate the camera initial pose estimation data at time t0 based on the GNSS position data and the initial frame image, and send the estimated data to the terminal.

17. A computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by a processor, the method steps described in any one of claims 1 to 13 are implemented.

18. A navigation service, wherein: Based on the method according to any one of claims 1 to 13, a positioning result of a navigated object is obtained, and based on the positioning result, a navigation guidance service of a corresponding scene is provided for the navigated object.

19. The navigation service according to claim 18, wherein: The corresponding scenario is a combination of one or more of AR navigation, elevated navigation, and main and auxiliary road navigation.

Citation Information

Patent Citations

  • Aircraft navigation method, aircraft navigation device and aircraft navigation system

    CN110231028A