Visual positioning method and electronic equipment

Through the visual positioning method, using the camera to acquire images and combine the image feature processing of the server, the problem of low positioning accuracy of existing electronic devices is solved, high-precision positioning and attitude measurement are achieved, and user experience is improved.

CN113672756BActive Publication Date: 2025-08-12HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010580807.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-14
Filing Date
2020-06-23
Publication Date
2025-08-12
Estimated Expiration
2040-06-23

AI Technical Summary

Technical Problem

Existing electronic equipment positioning technology relies on electromagnetic wave signals and is easily disturbed by buildings and atmospheric ionosphere, resulting in low positioning accuracy and inability to meet the needs of high-precision positioning, especially in augmented reality applications.

Method used

The visual positioning method is adopted to collect images through the camera and send visual positioning requests to the server. The server locates based on image features, combining the number of types of contour lines in the image and equipment posture to improve positioning accuracy.

Benefits of technology

It improves the positioning accuracy and attitude measurement accuracy of electronic devices, meets the needs of high-precision positioning, and enhances the user experience, especially in the combination of virtual and real and autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113672756B_ABST
    Figure CN113672756B_ABST
Patent Text Reader

Abstract

A visual positioning method and electronic device relate to the field of visual positioning technology and can be applied to electronic devices with cameras. The method specifically includes: detecting a first event for triggering a visual positioning process, determining whether the number of types of contour lines in a first image captured by the camera is greater than or equal to a first threshold; if so, sending a first visual positioning request to a server, the first visual positioning request including a first image and a first geographic location, the first geographic location being the geographic location of the electronic device itself measured when the first image was captured; and finally receiving a first visual positioning result sent from the server in response to the first visual positioning request, the first visual positioning result including a second geographic location. Compared with positioning based on electromagnetic wave signals, this technical solution helps to improve positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application No. 202010405196.8, filed with the Patent Office of China on May 14, 2020, entitled “A visual positioning method based on image semantic information in a large scene”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of computer vision, and in particular to a visual positioning method and electronic equipment. Background Art

[0003] At present, electronic devices such as mobile phones and tablets can provide users with services such as maps, navigation, and virtual-reality integration, and these services are implemented by relying on positioning technology. In the prior art, electronic devices are usually positioned based on electromagnetic wave signals (such as satellite signals or base station signals). Take base station signals as an example. The principle of positioning of electronic devices based on base station signals is that the electronic device determines the distance between itself and the base station based on the time it takes for the electromagnetic wave signal to be transmitted from the base station to itself, and then calculates its own position (i.e. the location of the electronic device) based on the distance between itself and the base station and the location of the base station. However, the transmission of electromagnetic wave signals is easily interfered with by buildings, the atmospheric ionosphere, etc., resulting in a large deviation between the distance determined based on the time it takes for the electromagnetic wave signal to be transmitted from the base station to the electronic device and the actual distance between the base station and the electronic device, which seriously affects the positioning accuracy of the electronic device. Summary of the Invention

[0004] The present application provides a visual positioning method and electronic device, which help to improve the positioning accuracy of electronic devices.

[0005] In a first aspect, a visual positioning method is provided in an embodiment of the present application, which is applied to an electronic device, wherein the electronic device includes a camera. The method specifically includes: detecting a first event, wherein the first event is used to trigger a visual positioning process; then, determining whether the number of types of contour lines in a first image captured by the camera is greater than or equal to a first threshold; if the number of types of contour lines in the first image is greater than or equal to the first threshold, sending a first visual positioning request to a server, wherein the first visual positioning request includes the first image and a first geographic location, wherein the first geographic location is the geographic location of the electronic device measured when the camera captures the first image; finally, receiving a first visual positioning result sent from the server in response to the first visual positioning request, wherein the first visual positioning result includes a second geographic location, wherein the second geographic location is the geographic location of the electronic device when the camera captures the first image, and the accuracy of the second geographic location is higher than the accuracy of the first geographic location.

[0006] In the embodiments of the present application, since the electronic device can send a first visual positioning request to the server, and the first visual positioning request includes a first image and a first geographic location, this helps the server to visually position the electronic device based on the first image, which helps improve positioning accuracy compared to positioning based on electromagnetic wave signals. Furthermore, since the electronic device sends the first visual positioning request to the server when the number of contour line types in the first image is greater than or equal to a first threshold, the server is more likely to successfully locate the device based on the first image.

[0007] In one possible design, the first visual positioning request further includes first camera parameters and / or a first device posture. The first camera parameters are camera parameters used by the camera when capturing the first image, and the first device posture is used to indicate at least one of the altitude angle, pitch angle, and roll angle of the electronic device itself measured when the camera captured the first image. This helps improve the accuracy of the server's positioning based on the first image.

[0008] In one possible design, the electronic device, after determining that the pitch angle of the electronic device is within a first angle range and the roll angle of the electronic device is within a second angle range, further determines whether the number of types of contour lines in the first image captured by the camera is greater than or equal to a first threshold. This helps increase the likelihood that the number of types of contour lines in the first image captured by the camera is greater than or equal to the first threshold.

[0009] In one possible design, the first visual positioning result further includes a second device posture, which indicates at least one of the altitude angle, pitch angle, and roll angle of the electronic device when the camera captures the first image, and the accuracy of the second device posture is higher than that of the first device posture. This helps the electronic device obtain a more accurate device posture.

[0010] In one possible design, upon receiving the second visual positioning result sent from the server in response to the first visual positioning request, the electronic device sends a second visual positioning request to the server when the proportion of repeated content with the first image in the second image captured by the camera is less than or equal to a second threshold, and the number of types of contour lines in the second image is greater than or equal to the first threshold, after the second visual positioning result is used to indicate that positioning based on the first image has failed. The second visual positioning request includes the second image and a third geographic location, and the third geographic location is the geographic location of the electronic device itself measured when the camera captured the second image. This helps to increase the possibility of successful server positioning based on the second image.

[0011] In one possible design, if the number of contour line types in the first image is less than the first threshold, the user is prompted to adjust the camera's shooting angle. This facilitates interaction between the user and the electronic device, allowing the user to be informed that the first image currently captured by the camera does not meet visual positioning requirements.

[0012] In a second aspect, a visual positioning method is provided for an embodiment of the present application, specifically including: a server receiving a first visual positioning request from an electronic device, the first visual positioning request including a first image and a first geographic location; then, the server extracts image features of the first image; and according to the first geographic location, selects Q candidate geographic locations from M candidate geographic locations of 360 panoramic images collected in a panoramic map, the distance between each of the Q candidate geographic locations and the first geographic location is less than or equal to a first threshold, Q is less than or equal to M, and M and Q are positive integers; the server determines a second geographic location from the Q candidate geographic locations, the second geographic location having the highest similarity between the image features of the 360 panoramic image and the first image among the Q candidate geographic locations, and returns a first visual positioning result to the electronic device, the first visual positioning result including the second geographic location.

[0013] In the embodiment of the present application, since the server can perform positioning based on the first image and the first geographical location from the electronic device, it helps to improve positioning accuracy compared to positioning based on electromagnetic wave signals.

[0014] In one possible design, the image features of the first image include contour line indications of N feature points in the first image, as well as the orientation angles and altitude angles of the N feature points in a first coordinate system, where N is a positive integer. The first coordinate system is the reference coordinate system of the 360-degree panoramic image captured at the candidate geographic location in the panoramic map. The contour line indications are used to indicate the type of contour line where the feature points are located. This helps simplify implementation and improve positioning accuracy.

[0015] In one possible design, the first visual positioning request further includes a first device posture;

[0016] The server may extract the image features of the first image based on the following manner:

[0017] The server performs semantic segmentation on the first image to obtain a semantic map of the first image, and obtains contour line indications of N feature points in the first image, as well as orientation angles and altitude angles of the N feature points in a second coordinate system based on the semantic map of the first image; the second coordinate system is a reference coordinate system of the first image; then, the server obtains the orientation angles and altitude angles of the N feature points in the first coordinate system based on the first device posture and the orientation angles and altitude angles of the N feature points in the second coordinate system.

[0018] By unifying the first image and the 360-degree panoramic image captured at the candidate geographical location in the panoramic map into the same reference coordinate system, the reliability of positioning is improved.

[0019] In one possible design, the first visual positioning request further includes first camera parameters, where the first camera parameters are camera parameters used to capture the first image;

[0020] The server may perform semantic segmentation on the first image to obtain a semantic graph of the first image based on the following method:

[0021] The server performs image processing on the first image to obtain an intermediate image, and performs semantic segmentation on the intermediate image to obtain a semantic graph of the first image; the camera parameters of the intermediate image are second camera parameters; or,

[0022] The server performs semantic segmentation on the first image to obtain a semantic map of an intermediate image, and performs image processing on the semantic map of the intermediate image to obtain a semantic map of the first image, wherein camera parameters of the semantic map of the first image are the second camera parameters;

[0023] The second camera parameters are camera parameters used when capturing a 360-degree panoramic image at a candidate geographical location in the panoramic map.

[0024] The above technical solution helps to further improve the reliability of positioning.

[0025] In one possible design, the similarity between the image features of the 360-degree panoramic image collected at the second geographical location and the image features of the first image satisfies the following expression:

[0026]

[0027] Wherein, (x, y, h) is the second geographic location, offset is the offset of one of the orientation angles in the orientation angle offset set, Loss(x, y, h, offset) is used to characterize the similarity between the image features of the 360-degree panoramic image collected at the second geographic location and the image features of the first image, Wi is the weight value of the i-th contour line in the first image, Y(i) is the set of orientation angles of all feature points on the i-th contour line in the first image, j is the orientation angle of a feature point on the i-th contour line in the first image, and P I (i, j) is the height angle of the feature point on the i-th contour line in the first image when the orientation angle is j, r is the total number of contour line types in the first image, P M(x ,y,h)(i,j+offset) is the altitude angle of the feature point with the orientation angle j+offset on the i-th contour line in the 360-degree panoramic image collected at the second geographical location. This helps to simplify the implementation.

[0028] In one possible design, after the server determines that the number of contour line types in the first image is greater than or equal to a first threshold, it selects Q candidate geographic locations from M candidate geographic locations collected from the 360-degree panoramic image in the panoramic map. This helps to increase the probability of successful positioning based on the first image.

[0029] In one possible design, the highest similarity between the image features of the 360-degree panoramic image collected at the second geographical location and the image features of the first image is within the image feature similarity range required for visual positioning accuracy, thereby helping to improve positioning accuracy.

[0030] In one possible design, when the highest similarity between the image features of the 360 panoramic image collected at the second geographic location and the image features of the first image is not within the image feature similarity range required by the visual positioning accuracy, the server returns a second visual positioning result to the electronic device, and the second visual positioning result is used to indicate that positioning based on the first image has failed.

[0031] In a third aspect, an electronic device is provided in an embodiment of the present application, comprising a camera, one or more processors, a memory, and one or more computer programs; wherein the camera is used to capture images; the computer program is stored in the memory, and the computer program is called when the processor is running, so that the electronic device executes the first aspect and any possible design method of the first aspect.

[0032] The fourth aspect is a server of an embodiment of the present application, which includes one or more processors, memories, and computer programs; wherein the computer program is stored in the memory; when the processor calls the computer program during operation, the server executes the second aspect and any possible design method of the second aspect.

[0033] The fifth aspect is a chip provided in an embodiment of the present application, which is coupled to a memory in an electronic device so that the chip calls a computer program stored in the memory during operation to implement the above-mentioned aspects of the embodiment of the present application and any possible design method provided in each aspect.

[0034] In the sixth aspect, a computer storage medium is provided in an embodiment of the present application, which stores a computer program. When the computer program runs on an electronic device, the electronic device executes the above-mentioned aspects and any possible design method of each aspect.

[0035] In the seventh aspect, a computer program product is provided in an embodiment of the present application. When the computer program product is run on an electronic device, the electronic device executes the above aspects and any possible design method of each aspect.

[0036] In an eighth aspect, an embodiment of the present application provides a communication system comprising an electronic device and a server, wherein the electronic device is configured to execute the method of the first aspect and any possible design of the first aspect; and the server is configured to execute the method of the second aspect and any possible design of the second aspect.

[0037] In addition, the technical effects brought about by any possible design method in the third to eighth aspects can be found in the technical effects brought about by different design methods in the method part, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0039] Figure 2 A system architecture diagram of an embodiment of the present application;

[0040] Figure 3 A schematic diagram of an image according to an embodiment of the present application;

[0041] Figure 4 A schematic diagram of the pitch angle, roll angle, and heading angle of a mobile phone according to an embodiment of the present application;

[0042] Figure 5 A schematic diagram of an image according to an embodiment of the present application;

[0043] Figure 6 Schematic diagram of the altitude angle and orientation angle of the feature point in an embodiment of the present application;

[0044] Figure 7 A flowchart of a method for obtaining image features of a 360-degree panoramic image at a candidate geographical location in a panoramic map according to an embodiment of the present application;

[0045] Figure 8 A schematic diagram of a large-scale 3D model according to an embodiment of the present application;

[0046] Figure 9 A flowchart of a visual positioning method according to an embodiment of the present application is shown;

[0047] Figure 10 A schematic flow chart of another visual positioning method according to an embodiment of the present application;

[0048] Figure 11 This is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0049] Figure 12 A result diagram of a server according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0050] Electronic devices can provide users with services such as maps, navigation, and virtual-reality integration, and these services are implemented by relying on positioning technology. Among them, the more accurate the positioning of the electronic device, the more reliable the services that the electronic device provides to users that rely on positioning technology, and the better the user experience of using services that rely on positioning technology on electronic devices. However, in the prior art, electronic devices are usually positioned based on electromagnetic wave signals such as satellite signals (such as GPS signals), base station signals, Wi-Fi signals, or Bluetooth signals. This method is easily affected by the environment (such as buildings, atmospheric ionosphere, etc.) and has low positioning accuracy. Moreover, this method based on electromagnetic wave signal positioning can only obtain the device's geographic location with low accuracy, and cannot obtain the device's posture, and cannot be applied to specific applications (such as augmented reality (AR) applications. Therefore, the existing positioning method cannot meet the positioning needs of electronic devices.

[0051] In view of this, an embodiment of the present application provides a visual positioning method that can realize the positioning of electronic devices in combination with images, which not only helps to improve the positioning accuracy, but also can obtain a more precise device posture, thereby meeting the positioning requirements of electronic devices and improving user experience.

[0052] It should be understood that in this application, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" is just a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in this application, "multiple" means two or more than two. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, or a, b and c.

[0053] Throughout this application, the terms "exemplary," "in some embodiments," and "in other embodiments" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" should not be construed as preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner.

[0054] It should be pointed out that the words "first", "second", etc. involved in this application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0055] The embodiments of the present application can be applied to scenarios that combine the virtual and the real, such as adding virtual elements to a real environment to achieve a surreal sensory experience. Furthermore, the embodiments of the present application can also be applied to scenarios such as autonomous driving and in-vehicle navigation, without limitation. For example, the embodiments of the present application can be applied to other application scenarios that rely on positioning.

[0056] It should be understood that the electronic devices of the embodiments of the present application may be mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0057] For example, Figure 1 As shown in FIG, it is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 1As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0058] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0059] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices, or two or more different processing units may be integrated into a single device.

[0060] The controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of reading and executing instructions.

[0061] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0062] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0063] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the electronic device.

[0064] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.

[0065] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0066] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface, enabling the function of playing music through Bluetooth headphones.

[0067] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the electronic device's camera function. The processor 110 and the display 194 communicate via the DSI interface to implement the electronic device's display function.

[0068] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0069] USB interface 130 is an interface that complies with USB standards and specifications, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 130 can be used to connect a charger to charge the electronic device, or to transfer data between the electronic device and peripheral devices. USB interface 130 can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.

[0070] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0071] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the electronic device's wireless charging coil. While charging the battery 142, the charging management module 140 can also power the electronic device through the power management module 141.

[0072] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193 and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0073] The wireless communication function of the electronic device can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor.

[0074] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0075] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied in electronic devices. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0076] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0077] The wireless communication module 160 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate and amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0078] In some embodiments, the antenna 1 of the electronic device is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS) and / or satellite based augmentation system (SBAS).

[0079] The electronic device implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0080] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLED, or a quantum dot light-emitting diode (QLED). In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than one.

[0081] The electronic device can realize the shooting function through the ISP, camera 193, video codec, GPU, display 194 and application processor.

[0082] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0083] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0084] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device selects a frequency, the DSP performs a Fourier transform on the frequency energy.

[0085] Video codecs are used to compress or decompress digital video. Electronic devices may support one or more video codecs. This allows them to play or record videos in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0086] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic devices, such as image recognition, face recognition, speech recognition, and text comprehension.

[0087] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0088] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0089] The electronic device can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0090] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0091] The speaker 170A, also called a "speaker," is used to convert audio electrical signals into sound signals. The electronic device can listen to music or make hands-free calls through the speaker 170A.

[0092] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device receives a call or voice message, the voice can be heard by placing the receiver 170B close to the human ear.

[0093] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device can be provided with at least one microphone 170C. In other embodiments, the electronic device can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.

[0094] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0095] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be located on display screen 194. There are many types of pressure sensors 180A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. The electronic device determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 194, the electronic device detects the intensity of the touch operation based on pressure sensor 180A. The electronic device can also calculate the location of the touch based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.

[0096] The gyroscope sensor 180B can be simply referred to as a gyroscope, and can be used to determine the motion posture of an electronic device. In some embodiments, the angular velocity of the electronic device around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for shooting anti-shake. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the electronic device's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shake of the electronic device through reverse motion to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenes.

[0097] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device calculates altitude, assists positioning and navigation through the air pressure value measured by the air pressure sensor 180C.

[0098] The magnetic sensor 180D, also known as a magnetometer, includes a Hall effect sensor. The electronic device can use the magnetic sensor 180D to detect the opening and closing of a flip case. In some embodiments, when the electronic device is a flip phone, the electronic device can detect the opening and closing of the flip cover based on the magnetic sensor 180D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.

[0099] The acceleration sensor 180E, also known as an accelerometer, detects the magnitude of acceleration of an electronic device in all directions (generally three axes). When the electronic device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0100] Distance sensor 180F is used to measure distance. The electronic device can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device can use distance sensor 180F to measure distance to achieve fast focus.

[0101] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device emits infrared light outward through the light emitting diode. The electronic device uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device. When insufficient reflected light is detected, the electronic device can determine that there is no object near the electronic device. The electronic device can use the proximity light sensor 180G to detect when the user holds the electronic device close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.

[0102] The ambient light sensor 180L senses ambient light brightness. The electronic device can adaptively adjust the brightness of the display screen 194 based on the perceived ambient light. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device is in a pocket to prevent accidental touches.

[0103] Fingerprint sensor 180H is used to collect fingerprints. Electronic devices can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.

[0104] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device uses the temperature detected by the temperature sensor 180J to implement a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device heats the battery 142 to prevent the electronic device from shutting down abnormally due to low temperature. In other embodiments, when the temperature is lower than another threshold, the electronic device boosts the output voltage of the battery 142 to prevent abnormal shutdown due to low temperature.

[0105] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied on or near the touch sensor. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device, in a location different from that of the display screen 194.

[0106] The bone conduction sensor 180M can obtain vibration signals. In some embodiments, the bone conduction sensor 180M can obtain vibration signals from the vibrating bones of the human body. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulse signals. In some embodiments, the bone conduction sensor 180M can also be set in headphones to form bone conduction headphones. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bones of the human body obtained by the bone conduction sensor 180M to implement voice functions. The application processor can parse heart rate information based on the blood pressure pulse signals obtained by the bone conduction sensor 180M to implement heart rate detection functions.

[0107] Keys 190 include a power button, a volume button, and the like. Keys 190 may be mechanical keys or touch-sensitive keys. The electronic device may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device.

[0108] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0109] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, and the like.

[0110] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and separated from the electronic device by inserting it into or removing it from the SIM card interface 195. The electronic device can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. Electronic devices interact with the network through SIM cards to implement functions such as calls and data communications. In some embodiments, the electronic device uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device and cannot be separated from the electronic device.

[0111] The following embodiments will be based on Figure 1 Taking the mobile phone with the structure shown as an example, the visual positioning method of the embodiment of the present application is introduced in detail.

[0112] Figure 2 A system architecture diagram of an embodiment of the present application is shown. Figure 2 As shown, the system architecture of the embodiment of the present application includes a mobile phone and a server. It should be noted that in the embodiment of the present application, the server can be a cloud server or a local server, etc., and this is not limited.

[0113] Specifically, the mobile phone triggers the visual positioning process and sends a visual positioning request to the server. The visual positioning request includes the i-th frame image captured by the camera, camera parameters, the device's low-precision posture, and the device's low-precision geographic location. i is a positive integer. Upon receiving the visual positioning request from the mobile phone, the server executes the visual positioning method and returns the visual positioning result to the mobile phone.

[0114] The i-th frame image captured by the camera can be a frame image captured by the camera after the mobile phone triggers the visual positioning process. It should be understood that in the embodiment of the present application, the i-th frame image captured by the camera can also be referred to as the i-th frame picture captured by the camera. Considering the data transmission cost, the image format can be jpeg. Of course, in the embodiment of the present application, the format of the i-th frame image captured by the camera can also be tif, bmp, etc., without limitation.

[0115] In an embodiment of the present application, a mobile phone can trigger a visual positioning process in a scenario where positioning is required. For example, the mobile phone can trigger the visual positioning process when a first event is detected. For example, the mobile phone can trigger the visual positioning process in response to an operation of opening a first application. The first application can be an application that supports the visual positioning function, such as Cyberverse, a camera, etc. For example, the operation of opening the first application can be an operation of clicking the icon of the first application, a voice command operation, a quick gesture operation, or other operations, etc., which is not limited in the embodiment of the present application. And / or, after the first application is started, the mobile phone triggers the visual positioning process periodically and / or through an event.

[0116] For example, when the mobile phone displays the interface of the first application on the display screen, when it detects that the distance between the current geographical location of the mobile phone and the geographical location of the most recent visual positioning meets the visual positioning requirements, the visual positioning process is triggered. In this case, the current geographical location of the mobile phone can be determined by the mobile phone based on electromagnetic wave signals (such as GPS signals or base station signals, etc.), or it can be determined based on location-based services (LBS), and there is no limitation on this. For example, the mobile phone can determine that the distance between the current geographical location of the mobile phone and the geographical location of the most recent visual positioning meets the visual positioning requirements when the distance between the current geographical location of the mobile phone and the geographical location of the most recent visual positioning reaches a certain threshold.

[0117] For another example, the mobile phone can also periodically trigger the visual positioning process when running the first application. It should be noted that the period for triggering the visual positioning process can be pre-set before the mobile phone leaves the factory, or can be set by the user according to his own needs, and there is no limitation on this.

[0118] For another example, when a mobile phone supports a local positioning system (such as a simultaneous localization and mapping (SLAM) system), the visual positioning process is triggered after the local positioning system is initialized. For example, the mobile phone triggers the initialization of the local positioning system in response to the operation of opening an application that supports the local positioning system, or the mobile phone is unable to obtain device positioning information (for example, the device's geographic location and / or device posture are lost due to interference) when the application that supports the local positioning system is in the startup state, and the local positioning system initialization is retriggered.

[0119] Of course, the above is only an example of triggering the visual positioning process and does not constitute a limitation on triggering the visual positioning process. The embodiment of the present application does not limit the way in which the mobile phone triggers the visual positioning process.

[0120] Furthermore, in some embodiments, the i-th frame of image captured by the camera can be an image captured by the camera after the mobile phone triggers the visual positioning process, with content richness that meets the visual positioning requirements. This helps improve the server's probability of successful visual positioning, reduces the number of visual positioning requests sent by the mobile phone to the server, and thus reduces the server's computing pressure.

[0121] For example, the image currently captured by the camera is the i-th frame image captured by the camera. The mobile phone can use a binary classification network model to determine whether the content richness of the i-th frame image captured by the camera meets the visual positioning requirements. The binary classification network model can be obtained by performing image training on multiple frames of images that are known to meet the visual positioning requirements and multiple frames of images that are known to not meet the visual positioning requirements. Of course, in the embodiment of the present application, the mobile phone can also use other methods to determine whether the content richness of the i-th frame image captured by the camera meets the visual positioning requirements, and there is no limitation on this.

[0122] It should be noted that images with content richness that meets the requirements of visual positioning refer to images with richer semantic types or images with a larger number of contour types. Semantic types may include buildings, mountains, roads, skies, rivers, etc. Contours refer to the boundaries between different semantic types in an image. Figure 3 Taking the image shown as an example, in this case, the semantic types involved in the i-th frame image captured by the camera include buildings and sky. Figure 3 The contour lines in the image shown include the boundary between the sky and the building.

[0123] Specifically, the mobile phone uses the i-th frame image captured by the camera as the input of the binary classification network model, and determines whether the i-th frame image captured by the camera meets the visual positioning requirements based on the output of the binary classification network model. In other embodiments, when the mobile phone determines that the content richness of the i-th frame image captured by the camera does not meet the visual positioning requirements based on the output of the binary classification network model, it can prompt the user to adjust the camera's shooting angle so that the camera can capture an image with content richness that meets the visual positioning requirements. For example, the mobile phone can play a voice prompt message to the user through a speaker and / or display a prompt message to the user through a display screen to prompt the user to adjust the camera's shooting angle.

[0124] Furthermore, in some other embodiments, the mobile phone can first determine whether its current posture meets the image acquisition requirements. If its current posture meets the image acquisition requirements, the mobile phone then determines whether the content richness of the image currently captured by the camera meets the visual positioning requirements. If the current posture of the mobile phone meets the image acquisition requirements, this helps increase the probability that the content richness of the image captured by the camera meets the visual positioning requirements. Moreover, compared with determining whether the image content richness meets the visual positioning requirements, determining whether the posture meets the image acquisition requirements consumes less computing resources and is easier to implement, thereby helping to reduce the processing requirements of the mobile phone.

[0125] For example, the mobile phone can determine whether its current posture meets the image acquisition requirements based on the following methods:

[0126] The mobile phone obtains its current pitch angle and roll angle based on information from a posture sensor (e.g., one or more of a gyroscope, a magnetic sensor, an accelerometer, and / or a gravity sensor). The mobile phone then determines whether its current pitch angle is within a first angle range and whether its current roll angle is within a second angle range. If the mobile phone's current pitch angle is within the first angle range and its current roll angle is within the second angle range, the mobile phone determines that its current posture meets the image requirements. This helps ensure that images captured by the mobile phone through the camera contain rich semantic types (e.g., buildings, sky, ground, etc.), improves the probability and reliability of successful visual positioning, reduces the number of invalid visual positioning requests sent by the mobile phone to the server, and reduces server pressure. In other embodiments, if the mobile phone's current pitch angle is not within the first angle range and / or its current roll angle is not within the second angle range, the mobile phone determines that its current posture does not meet the image requirements and prompts the user to adjust the phone's posture. For example, the mobile phone may play a voice prompt through a speaker and / or display a prompt on a display screen to prompt the user to adjust the phone's posture so that the phone's posture meets the image acquisition requirements.

[0127] It should be noted that in the embodiment of the present application, the first angle range and the second angle range can be pre-set by the mobile phone before leaving the factory, or can be set by the user according to his or her own needs, etc. The embodiment of the present application does not limit the setting method of the first angle range and the second angle range. For example, when the mobile phone is in portrait mode, the first angle range can be -20° to 40°, and the second angle range can be 75° to 105°. For another example, when the mobile phone is in landscape mode, the first angle range can be -20° to 40°, and the second angle range can be -15° to 15°.

[0128] Furthermore, in some embodiments, when the visual positioning result returned by the server to the mobile phone indicates that the visual positioning has failed, the mobile phone may prompt the user to adjust the shooting angle of the camera and / or the posture of the mobile phone, and after the camera captures the j-th frame image that meets the visual positioning requirements, send a visual positioning request to the server again, and the visual positioning request includes the j-th frame image captured by the camera, camera parameters, low-precision posture of the device and low-precision geographic location of the device.

[0129] In this case, after the user adjusts the camera's shooting angle and / or the phone's posture, the phone can first determine whether its current posture meets the image acquisition requirements. If so, the phone then determines whether the image currently captured by the camera meets the visual positioning requirements. If so, the phone sends another visual positioning request to the server.

[0130] Taking the case where the image currently captured by the camera is the j-th frame image captured by the camera, and the visual positioning request sent by the mobile phone to the server includes the i-th frame image captured by the camera, as an example, the mobile phone receives a visual positioning result returned from the server indicating that positioning based on the i-th frame image failed, for example, the mobile phone can determine whether the j-th frame image captured by the camera meets the visual positioning requirements in the following ways:

[0131] The phone determines whether the proportion of content in the jth frame captured by the camera that overlaps with the ith frame, and the content richness of the jth frame, meet the visual positioning requirements. If both the proportion of content in the jth frame captured by the camera that overlaps with the ith frame, and the content richness of the jth frame, meet the visual positioning requirements, the phone determines that the jth frame meets the visual positioning requirements. This helps prevent the phone from sending too many invalid visual positioning requests to the server, reducing server pressure.

[0132] If the proportion of repeated content in the j-th frame image captured by the camera and the i-th frame image captured by the camera does not meet the visual positioning requirements, and / or the content richness of the j-th frame image captured by the camera does not meet the visual positioning requirements, the mobile phone determines that the j-th frame image captured by the camera does not meet the visual positioning requirements.

[0133] For example, the mobile phone can determine whether the proportion of the content repeated in the j-th frame image captured by the camera and the i-th frame image captured by the camera meets the visual positioning requirements based on the following method:

[0134] The mobile phone determines whether the change in the phone's posture is within the range required by visual positioning based on the low-precision posture of the device when the camera captured the j-th image frame and the low-precision posture of the device when the camera captured the i-th image frame, and / or determines whether the change in the phone's position is within the range required by visual positioning based on the low-precision geographic location of the device when the camera captured the j-th image frame and the low-precision geographic location of the device when the camera captured the i-th image frame. If the change in the phone's posture and / or the change in the phone's position are both within the range required by visual positioning, the mobile phone determines that the proportion of content in the j-th image frame captured by the camera that is repeated with the i-th image frame captured by the camera meets the visual positioning requirements.

[0135] Furthermore, in some embodiments, when the change in the posture of the mobile phone is not within the range of the visual positioning requirements, and / or the change in the position of the mobile phone is not within the range of the visual positioning requirements, the mobile phone determines that the proportion of the content repeated in the j-th frame image captured by the camera and the i-th frame image captured by the camera does not meet the visual positioning requirements.

[0136] For example, the range of visual positioning requirements for changes in the phone's posture is no less than 40°, and the range of visual positioning requirements for changes in the phone's distance is no less than 10 meters. When the phone's pitch angle changes by 20° and the phone's distance changes by 8 meters, the phone determines that the proportion of the repeated content in the j-th frame image captured by the camera and the i-th frame image captured by the camera does not meet the visual positioning requirements.

[0137] For another example, the mobile phone may also perform image content analysis on the i-th and j-th frames captured by the camera to determine the proportion of content in the j-th frame captured by the camera that overlaps with the i-th frame. The mobile phone then determines whether the proportion of content in the j-th frame captured by the camera that overlaps with the i-th frame meets the visual positioning requirements.

[0138] The above is only an example of a specific implementation method for a mobile phone to determine whether the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera meets the visual positioning requirements, and does not constitute a limitation on the embodiments of the present application. In the embodiments of the present application, other methods can be used to determine whether the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera meets the visual positioning requirements.

[0139] Among them, the mobile phone can first determine whether the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera meets the visual positioning requirements. When the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera meets the visual positioning requirements, it can then determine whether the content richness of the j-th frame image captured by the camera meets the visual positioning requirements. Alternatively, the mobile phone can also first determine whether the content richness of the j-th frame image captured by the camera meets the visual positioning requirements. When the content richness of the j-th frame image meets the visual positioning requirements, it can then determine whether the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera meets the visual positioning requirements. Alternatively, the mobile phone can simultaneously determine whether the proportion of the content repeated with the i-th frame image captured by the camera in the j-th frame image captured by the camera and the content richness of the j-th frame image meet the visual positioning requirements. The embodiments of the present application are not limited to this.

[0140] The following describes the camera parameters, the low-precision posture of the device, and the low-precision geographic location of the device, taking a visual positioning request including the i-th frame image captured by the camera, camera parameters, the low-precision posture of the device, and the low-precision geographic location of the device as an example.

[0141] The camera parameters are the camera parameters used when the camera captures the i-th frame image, and may include in-camera parameters, exposure parameters, etc. In the embodiment of the present application, the in-camera parameters can be understood as parameters related to the camera's own characteristics, such as focal length, pixels, etc. In particular, when the camera lens is a special lens with large distortion, such as a fisheye lens, the in-camera parameters also include distortion correction parameters. For lenses with smaller distortion, the in-camera parameters may not include distortion correction parameters. It should be noted that the camera parameters can be set by the user according to his or her own needs, or the camera parameters are set by the mobile phone before leaving the factory, or the camera parameters can also be automatically adjusted by the mobile phone in combination with different function settings (such as anti-shake function or automatic zoom function). Alternatively, some of the camera parameters are set by the user according to his or her own needs, and another part of the parameters may be set by the mobile phone before leaving the factory.

[0142] The low-precision posture of the device is the posture of the mobile phone measured when the camera captures the i-th frame image. Specifically, the low-precision posture of the device is used to indicate at least one of the pitch angle, roll angle and heading angle of the mobile phone when the camera captures the i-th frame image. For example, the pitch angle, roll angle and heading angle of the mobile phone can be referenced to the world coordinate system. For example, Figure 4As shown, the pitch angle of the phone is the angle of rotation of the phone around the X-axis in the world coordinate system, which is used to indicate the up and down orientation of the device when shooting, that is, the camera is facing up or down. The roll angle of the phone is the angle of rotation of the phone around the Z-axis in the world coordinate system, which is used to indicate the degree of left and right tilt of the device when shooting. The heading angle of the phone is the angle of rotation of the phone around the Y-axis in the world coordinate system, which is used to indicate the direction of shooting.

[0143] For example, the low-precision posture of the device can be obtained by the mobile phone by measuring the information of its own posture sensor (such as one or more of a gyroscope sensor, a magnetic sensor, an acceleration sensor and / or a gravity sensor). For example, by measuring the information of the gravity sensor and the magnetic sensor when the camera captures the i-th frame of image, the mobile phone's orientation angle is the orientation angle of the mobile phone when the camera captures the i-th frame of image. Usually, the error of the magnetic sensor is relatively large, within 30 degrees. For another example, by measuring the information of the gravity sensor when the camera captures the i-th frame of image, the mobile phone's pitch angle and roll angle are obtained, which are the pitch angle and roll angle of the mobile phone when the camera captures the i-th frame of image, and the error is usually within 2 degrees.

[0144] For another example, the low-precision device posture can also be obtained based on the posture information of the local map on the mobile phone's simultaneous localization and mapping (SLAM) system. It should be noted that the above is only an example of how the mobile phone can obtain the low-precision device posture. In the embodiments of this application, the mobile phone can also obtain the low-precision device posture through other methods.

[0145] The low-precision geographic location of the device is the geographic location of the mobile phone measured when the camera captures the i-th frame of image, which may include longitude, latitude and altitude, and its accuracy range is usually within 40m. For example, the low-precision geographic location of the device can be determined by the mobile phone based on electromagnetic wave signals (such as GPS signals or base station signals). For another example, the low-precision geographic location of the device can also be determined by the mobile phone based on location-based services (LBS). It should be noted that the above is only an example of how the mobile phone obtains the low-precision geographic location of the device. In the embodiment of the present application, the mobile phone can also obtain the low-precision geographic location of the device through other means.

[0146] The following describes the server's visual positioning method in detail, taking the example of a visual positioning request including the i-th frame image captured by the camera, camera parameters, low-precision device posture, and low-precision device geographic location.

[0147] First, M candidate geographic locations in a panoramic map and image features of 360-degree panoramic images captured at the M candidate geographic locations with the first camera parameters are pre-stored in the server. In an embodiment of the present application, the panoramic map can also be referred to as a large scene map. M is a positive integer. Taking a 360-degree panoramic image captured at one of the M candidate geographic locations as an example, the image features of the 360-degree panoramic image may include parameter information of multiple feature points in the 360-degree panoramic image. The parameter information of each feature point may include the orientation angle of the feature point, the altitude angle of the feature point, and the type of contour line where the feature point is located. Among them, the contour line can be used to represent the boundary line between different semantic types of content in the 360-degree panoramic image, which can be continuous or discontinuous. For example, Figure 5 Taking the image shown as an example, the semantic type of the content in the dark gray area is building, the semantic type of the content in the light gray area above the dark gray area is sky, and the semantic type of the content in the black area below the dark gray area is ground. Figure 5 The image shown includes two types of contour lines: contour lines between buildings and the sky, and contour lines between buildings and the ground.

[0148] It should be noted that the reference coordinate system of the 360-degree panoramic image can be the world coordinate system. In this case, the orientation angle and altitude angle of the feature points of the 360-degree panoramic image are based on the world coordinate system. Figure 6 As shown, the origin O in the world coordinate system is the location of the observation point (photographing device). The Z axis in the world coordinate system can be in the north-south direction, pointing to the north, the X axis can be in the east-west direction, pointing to the east, and the Y axis is the height direction. Point P' is the projection of point P on the horizontal ground. Among them, the plane formed by the X axis, the origin O and the Y axis is the horizontal ground, the OP direction is the shooting direction, and the angle between OP and OP' is is the altitude angle of point P, and the angle δ between OP′ and the Z axis is the orientation angle of point P.

[0149] For example, the orientation angle and altitude angle of the feature point can be expressed using integers at a certain angular resolution (for example, 10 / degree), which helps to reduce the amount of data of the parameter information of the feature point. In the embodiment of the present application, the size of the parameter information of a feature point can be only 1kb.

[0150] For example, in the embodiment of the present application, the candidate geographic location and the image features of the 360-degree panoramic image collected at the candidate geographic location may be stored in the format of Table 1:

[0151] Table 1

[0152]

[0153] Of course, the above is merely an example of the storage format of the candidate geographic location and the image features of the 360-degree panoramic image collected at the candidate geographic location, and the embodiments of the present application do not limit this.

[0154] For example, Figure 7 FIG. 1 is a flow chart of a method for obtaining image features of a candidate geographic location in a panoramic map and a 360-degree panoramic image captured at the candidate geographic location in an embodiment of the present application, which specifically includes the following steps:

[0155] 701. Construct a large scene 3D model. For example, the level of detail of the large scene 3D model can be lod2 (basic model) or lod1 (block model).

[0156] In some embodiments, a large-scale 3D model is constructed based on a target image. The target image is an image of one or more regions on Earth captured from a high altitude by a satellite or drone, and may be obtained from a third party. The specific implementation of constructing a large-scale 3D model based on the target image can be found in the prior art and will not be further described here.

[0157] 702. Place virtual photographing devices (e.g., virtual cameras) in user-accessible areas of the large-scale 3D model, and record the geographic location of each virtual photographing device and the first camera parameters used by each virtual photographing device to capture 360-degree panoramic images. The interval between two adjacent virtual photographing devices is a fixed value (e.g., 1 meter). It should be noted that the interval between two adjacent virtual photographing devices and the first camera parameters can be set by the R&D personnel based on actual needs. For ease of calculation, the virtual photographing devices are placed with a pitch angle of 0 and a roll angle of 0.

[0158] For example, the area accessible to the user may be a road, a beach, a mountain, a river, or other area accessible to the user.

[0159] For example, in Figure 8 In the large scene 3D model shown, each point represents a virtual camera device, and the interval between two adjacent points is 1 meter.

[0160] At step 703, image features of the 360-degree panoramic image captured by each virtual camera device using the first camera parameters are extracted. For example, the image features of the 360-degree panoramic image include parameter information of multiple feature points in the 360-degree panoramic image. The parameter information of each feature point may include the feature point's orientation angle, the feature point's altitude angle, and the type of contour line on which the feature point resides. It should be noted that the orientation angle and altitude angle of the feature points involved in step 703 may be referenced to the world coordinate system to facilitate subsequent visual positioning.

[0161] The orientation angle of the feature point in the 360-degree panoramic image is within the range of 0° to 360°, and the altitude angle is within the range of -40° to 70°.

[0162] In some embodiments, image features of the 360-degree panoramic image captured by each virtual shooting device using the first camera parameters may be extracted based on the following method:

[0163] First, obtain the semantic map of the 360-degree panoramic image captured by each virtual camera device with the first camera parameter. The resolution of the semantic map of the 360-degree panoramic image is set by the developer according to actual needs. For example, the resolution of the semantic map of the 360-degree panoramic image is 0.1 degree / pix. For example, Figure 5 The image shown is a semantic map of a 360-degree panoramic image in an embodiment of the present application. Then, based on the semantic map, image features of the 360-degree panoramic image are obtained. It should be noted that the resolution of the semantic map of the 360-degree panoramic image is the resolution used to extract the image features of the 360-degree panoramic image.

[0164] The above is only an example of a method for obtaining image features of a candidate geographic location in a panoramic map and a 360-degree panoramic image collected at the candidate geographic location. The embodiment of the present application does not limit the method for obtaining image features of a candidate geographic location in a panoramic map and a 360-degree panoramic image collected at the candidate geographic location. For example, the image features of a candidate geographic location in a panoramic map and a 360-degree panoramic image collected at the candidate geographic location in the embodiment of the present application can also be obtained through manual collection.

[0165] It should be noted that Figure 7 The method for obtaining the image features of the candidate geographic location in the panoramic map and the 360-degree panoramic image collected at the candidate geographic location can be executed on one or more computing devices (such as computers or servers). Figure 7 The computing device of the method shown may be a server for executing the visual positioning method of the embodiment of the present application, or may not be a server for executing the visual positioning method of the embodiment of the present application. Figure 7 In the case where the computing device of the method shown is not a server for executing the visual positioning method of the embodiment of the present application, after obtaining the candidate geographic location in the panoramic map and the image features of the 360 panoramic image collected at the candidate geographic location, the computing device also needs to upload the candidate geographic location in the panoramic map, the image features of the 360 panoramic image collected at the candidate geographic location, and the relevant parameters used to obtain the 360 panoramic image and image features (such as the first camera parameters, the resolution of the semantic map of the 360 panoramic image, etc.) to the server for executing the visual positioning method of the embodiment of the present application, so that the server can perform visual positioning.

[0166] For example, Figure 9 As shown in FIG, a visual positioning method according to an embodiment of the present application specifically includes the following steps:

[0167] 901. The server receives a visual positioning request from a mobile phone. The visual positioning request includes a first image, second camera parameters, a low-precision device pose, and a low-precision device geographic location. The first image is a frame captured by the mobile phone's camera, the second camera parameters are the parameters used by the mobile phone's camera when capturing the first image, and the low-precision device pose and low-precision device geographic location are measured by the mobile phone's camera when capturing the first image.

[0168] Specifically, for the second camera parameters, the low-precision posture of the device and the low-precision geographic location of the device, please refer to the above introduction on the camera parameters, low-precision posture of the device and the low-precision geographic location of the device on the mobile phone side, which will not be repeated here.

[0169] 902. The server obtains a semantic map of the first image based on the first camera parameter and the second camera parameter. For example, the resolution of the semantic map of the first image is a first resolution. The first resolution is pre-configured in the server and is used to extract image features of the 360-degree panoramic image collected at the candidate geographical location in the panoramic map. For example, when using Figure 7 When the method shown extracts image features of a 360-degree panoramic image collected at a candidate geographic location in a panoramic map, the resolution used to extract image features of the 360-degree panoramic image collected at a candidate geographic location in the panoramic map is 0.1 degree / pix, and the first resolution is 0.1 degree / pix.

[0170] It should be noted that the first camera parameters are the camera parameters used for capturing the 360-degree panoramic image at the candidate geographic location in the panoramic map. In some embodiments, if the second camera parameters are different from the first camera parameters, the server may obtain the semantic graph of the first image based on the following method:

[0171] First, the server performs image processing on the first image according to the first camera parameters and the second camera parameters, and converts the first image into a second image, and the camera parameters of the second image are the second camera parameters. Then, the server performs semantic segmentation processing on the second image to obtain a semantic map of the first image. This helps to unify the camera parameters used for the 360 panoramic image collected at the candidate geographical location in the panoramic map, and improve the reliability of visual positioning. For example, the algorithm used for the semantic segmentation processing of the second image can be a semantic segmentation algorithm of the deeplab series, or other algorithms (such as RefineNet, PSPNet, CASENET), etc., which is not limited in the embodiments of the present application. Alternatively, the server first performs semantic segmentation processing on the first image to obtain a semantic map of an image, and then performs image processing on the semantic map of the image obtained by performing speech segmentation processing on the first image according to the first camera parameters and the second camera parameters to obtain a semantic map of the first image, thereby achieving unification with the camera parameters used for the 360 panoramic image collected at the candidate geographical location in the panoramic map.

[0172] In addition, it can be understood that when the first camera parameters are the same as the second camera parameters, the server can perform semantic segmentation processing on the first image to obtain a semantic map of the first image without performing image conversion.

[0173] 903. The server extracts image features of the first image based on the semantic graph of the first image. For example, the image features of the first image include parameter information of multiple feature points. The parameter information of each feature point includes the orientation angle, altitude angle, and type of contour line on which the feature point is located. For example, the type of contour line on which the feature point is located can be represented by a contour line indicator. The contour line indicator can be a numerical value, a character, or the like, without limitation.

[0174] It should be noted that the number of feature points belonging to different contour line types in the first image may be the same or different.

[0175] For example, the reference coordinate system of the first image is the mobile phone coordinate system. Therefore, the orientation angle and altitude angle of the feature point of the first image are based on the mobile phone coordinate system, that is, the orientation angle and altitude angle of the feature point extracted from the first image are the orientation angle and altitude angle of the feature point in the mobile phone coordinate system. The mobile phone coordinate system here can be understood as the coordinate system of the local positioning system (such as the SLAM system) in the mobile phone, or the mobile phone coordinate system can also be understood as a coordinate system with a certain position on the mobile phone (such as the center of mass, the location of the camera) as the origin, the long side of the mobile phone display screen as the X-axis (or Y-axis), the short side as the Y-axis (or X-axis), and the axis perpendicular to the plane of the display screen as the Z-axis.

[0176] Furthermore, in some embodiments, the server may first determine whether the number of types of contour lines in the first image is greater than or equal to R based on the semantic map of the first image. If the number of types of contour lines in the first image is greater than or equal to R, the server then obtains the image features of the first image based on the semantic map of the first image, thereby helping to increase the probability of successful visual positioning. In other embodiments, when the number of types of contour lines in the first image is less than R, the server returns a visual positioning result to the mobile phone, and the visual positioning result is used to indicate a failure of image-based positioning. It should be noted that the value of R may be pre-configured in the server. For example, the server may adjust the value of R according to a certain strategy or algorithm so that the value of R can better meet the needs of visual positioning.

[0177] 904. The server converts the low-precision posture of the device and the orientation angle and altitude angle of the feature point of the first image in the mobile phone coordinate system into the orientation angle and altitude angle of the feature point of the first image in the world coordinate system.

[0178] 905. The server selects Q candidate geographic locations from the M candidate geographic locations in the panoramic map based on the device's low-precision geographic location. The distance between each of the Q candidate geographic locations and the device's low-precision geographic location is less than or equal to a first threshold. The value of the first threshold can be an empirical value pre-configured in the server, or can be determined based on the accuracy of the device's low-precision geographic location obtained by the mobile phone, etc., and is not limited to this. For example, if the accuracy of the device's low-precision geographic location obtained by the mobile phone is 30 meters, the value of the first threshold can be greater than or equal to 30 meters.

[0179] It should be noted that step 905 is not necessarily sequentially executed with steps 902, 903, and 904. However, step 905 is executed after step 901 and before step 906. Steps 902 to 904 are also executed after step 901 and before step 906. For example, step 905 is executed before step 902. For another example, step 905 and step 902 are executed simultaneously.

[0180] 906. The server determines a first geographic location from the Q candidate geographic locations based on the image features of the 360-degree panoramic image and the image features of the first image collected at the Q candidate geographic locations, wherein the first geographic location has the highest similarity between the image features of the 360-degree panoramic image and the first image among the Q candidate geographic locations.

[0181] The similarity of image features is described by taking the similarity between the image features of a 360-degree panoramic image collected at a candidate geographic location k among the Q candidate geographic locations and the image features of the first image as an example.

[0182] For example, the similarity of image features satisfies the following expression (1):

[0183]

[0184] Among them, Loss (x, y, h, offset) is used to indicate the similarity between the image features of the 360 panoramic image collected at the candidate geographic location k and the image features of the first image. (x, y, h) is the candidate geographic location k, which can be longitude, latitude and altitude respectively. Offset is an offset of the orientation angle, which can take all values in the orientation angle offset set. The orientation angle offset set is pre-configured in the server and can also be adjusted in real time based on the current calculation results. Wi is the weight value of the i-th contour line in the first image, which is pre-configured in the server. Y(i) is the orientation angle set of all feature points on the i-th contour line in the first image, and j is the orientation angle of a feature point on the i-th contour line in the first image. P I (i, j) is the height angle of the feature point on the i-th contour line in the first image when the orientation angle is j, r is the total number of contour line types in the first image, P M(x ,y,h)(i,j+offset) is the altitude angle of the feature point on the i-th contour line in the 360 panoramic image collected at the candidate geographic location k when the orientation angle is j+offset.

[0185] It should be noted that the above is an example of a method for calculating the similarity of image features of different images by evaluating or indicating the similarity of image features through Loss, and does not constitute a limitation on the method for calculating the similarity of image features. It should be understood that in actual implementation, there can be many detailed adjustments and optimizations for the expression of the similarity of image features. The expression (1) is only an example to illustrate the general idea and does not constitute a limitation on the method for calculating the similarity of image features in the embodiment of the present application. It should also be noted that when evaluating or indicating the similarity of image features through Loss, the smaller the value of Loss, the higher the similarity of image features, and conversely, the lower the similarity of image features.

[0186] In addition, in the embodiment of the present application, in addition to evaluating the similarity of image features by Loss, the similarity of image features can also be evaluated or indicated by the intersection over union (IOU) of the images. For example, the server can calculate the IOU of the image reported by the mobile phone and the 360-degree panoramic image collected at the candidate geographical location in the panoramic map based on the image reported by the mobile phone, the low-precision location of the device, the low-precision posture of the device, and the 360-degree panoramic image collected at the candidate geographical location in the panoramic map.

[0187] For example, the server can traverse all values in the orientation angle offset set for the image features of the 360 panoramic images collected at Q candidate geographic locations, calculate the similarity with the image features of the first image, and then determine the first geographic location from the Q candidate geographic locations based on the similarity with the image features of the first image obtained by the above calculation.

[0188] Alternatively, the server may first select Y candidate geographic locations from Q candidate geographic locations. Adjacent geographic locations within the Y candidate geographic locations are separated by a second threshold. The server then iterates through all values in the heading angle offset set for the image features of the 360-degree panoramic images captured at the Y candidate geographic locations and calculates their similarity with the image features of the first image. The server then determines a second geographic location from the Y candidate geographic locations based on the similarity calculated with the image features of the first image. The second geographic location has the highest similarity between the 360-degree panoramic image and the image features of the first image among the Y candidate geographic locations. Based on the second geographic location, the server then selects Z candidate geographic locations from the Q candidate geographic locations, where the distance between each of the Z candidate geographic locations and the second geographic location is less than or equal to a third threshold. Specifically, the second threshold is greater than the third threshold, and the second and third thresholds are preconfigured in the server. For example, the second threshold may be 10 meters, and the third threshold may be 5 meters. The server iterates through all values in the heading angle offset set for the image features of the 360-degree panoramic images captured at these Z candidate geographic locations and calculates their similarity with the image features of the first image. Finally, based on the similarity between the image features of the panoramic map collected at the Z candidate locations and the first image, the server determines a first location from the Z candidate locations. This location has the highest similarity between the image features of the 360-degree panoramic image and the first image among the Z candidate locations. It should be noted that the first location also has the highest similarity between the image features of the 360-degree panoramic image and the first image among the Q candidate locations. This first location is the first location mentioned in step 906. This helps reduce the amount of data required for the server's visual positioning calculations.

[0189] 907. The server returns a visual positioning result to the mobile phone, where the visual positioning result includes the first geographic location determined in step 906.

[0190] Furthermore, in some embodiments, the server first determines whether the highest similarity between the image features of the 360-degree panoramic image collected at the Q candidate geographic locations and the first image is within a first range. If the highest similarity of the image features is within the first range, the server returns the visual positioning result to the mobile phone, and the visual positioning result includes the first geographic location determined in step 906. In other embodiments, if the highest similarity of the image features is not within the first range, the server returns the visual positioning result to the mobile phone, and the visual positioning result is used to indicate that the visual positioning failed. This helps to improve the accuracy of visual positioning. It should be noted that the first range can be pre-configured in the server and can be the image feature similarity range required for visual positioning accuracy.

[0191] In other embodiments of the present application, the server may also determine the high-precision device posture of the mobile phone based on the image features of the 360-degree panoramic image captured at the candidate geographic location in the panoramic map and the low-precision device posture reported by the mobile phone. The high-precision device posture can be used to indicate at least one of the heading angle, roll angle, and pitch angle of the mobile phone. For example, the high-precision device posture may include at least one of the heading angle, roll angle, and pitch angle of the mobile phone, or may be a rotation matrix Rx, etc. The rotation matrix Rx can be referred to in the relevant description below.

[0192] For example, the server may also determine the orientation angle of the mobile phone based on the offset used to determine the first geographic location and the low-precision device posture. For example, if the low-precision device posture includes the orientation angle of the mobile phone as α, and the offset used to determine the first geographic location is δ, the server may determine the orientation angle of the mobile phone as α + δ based on the low-precision device posture and the offset used to determine the first geographic location.

[0193] For example, the server can determine the roll angle and pitch angle of the mobile phone based on the SVD decomposition algorithm according to the image features of the 360 panoramic image collected at the candidate geographic location in the panoramic map and the low-precision posture of the device reported by the mobile phone, as well as the offset used to determine the first geographic location.

[0194] For example, the pitch angle and roll angle of a mobile phone can be determined based on the following method:

[0195] First, the server determines the adjusted orientation angle of each feature point in the first image according to the offset used when determining the first geographic location and the orientation angle of each feature point in the first image.

[0196] Secondly, the server searches for the altitude angles of the feature points of the 360-degree panoramic image when the adjusted orientation angles of the feature points in the first image are the same from the image features of the 360-degree panoramic image collected at the M candidate geographical locations in the panoramic map.

[0197] Then, the server normalizes the adjusted orientation angle and altitude angle of each feature point in the first image to obtain a first matrix, wherein each column parameter in the first matrix indicates the normalized coordinates of a feature point in the first image; and normalizes the orientation angle and altitude angle of the feature point found in the image features of the 360-degree panoramic image collected at M candidate geographic locations to obtain a second matrix, wherein each column parameter in the second matrix indicates the normalized coordinates of a feature point in the 360-degree panoramic image; wherein the orientation angles of the feature points indicated by the parameters at the same position in the first matrix and the second matrix are the same. For example, in an embodiment of the present application, the feature point can be mapped onto a unit sphere based on the orientation angle and altitude angle of the feature point, and the 3D coordinates of the projection point of the feature point on the unit sphere in the world coordinate system are the normalized coordinates of the feature point. Of course, in an embodiment of the present application, the normalized coordinates of the feature point can also be obtained by other means, which is not limited to this.

[0198] Finally, the server uses the SVD decomposition algorithm to obtain the rotation matrix Rx based on the first matrix and the second matrix. The matrix obtained by multiplying the rotation matrix Rx by the first matrix on the left is closest to the second matrix.

[0199] Take the first matrix as P I , the second matrix is P M For example, P I and P M are all matrices with 3 rows and n columns, where n is the total number of feature points in the first image. First, based on the SVD decomposition algorithm, we obtain the matrix that satisfies ||P M -R X1 *P I ||The value of Rx1 when it is minimum. Then, RX1*P I Convert it into the altitude angle and orientation angle of each feature point in the first image. And find the image features of the 360 panoramic image collected at the M candidate geographical locations that match the image features of RX1*P I The orientation angles and elevation angles of the feature points with the same orientation angles as the feature points in the first image are normalized by processing the orientation angles and elevation angles of the feature points found in the image features of the 360-degree panoramic image collected at the M candidate geographical locations to obtain the matrix P. M1 . Continue based on the SVD decomposition algorithm and get the result satisfying ||P M1 -R X2 *R X1 *P I The value of Rx2 when || is minimum, and so on, based on the SVD decomposition algorithm, we can get the value that satisfies ||P Mi -R Xi *......*R X2 *R X1 *PI ||Minimum Rx i The value of . Among them, P Mi See P M1 The method of obtaining is not described here. When the value of i is equal to the fourth threshold, or Rx i When it is approximately the unit matrix, Rx is R Xi *......*R X2 *R X1 It should be understood that in Rx i When R is approximately the unit matrix, Xi *......*R X2 *R X1 *P I With P Mi Basically overlap. For example, in an embodiment of the present application, when the values of the matrix other than the diagonal are approximately 0, the matrix is determined to be approximately the identity matrix. Wherein, when the absolute value of the values of the matrix other than the diagonal is less than or equal to a fifth threshold (e.g., 0.0001), the values of the matrix other than the diagonal are determined to be approximately 0.

[0200] Furthermore, the server can determine the high-precision posture of the device based on the rotation matrix Rx and the low-precision posture of the device, that is, the pitch angle, roll angle, and heading angle of the adjusted mobile phone. For example, the method of adjusting the heading angle of the mobile phone can refer to the above introduction of adjusting the heading angle of the mobile phone based on the offset. For example, for the pitch angle and roll angle of the mobile phone, the server can convert the rotation matrix Rx into Euler angles, and adjust the pitch angle and roll angle of the mobile phone in the low-precision posture of the device based on the Euler angles converted from the rotation matrix Rx.

[0201] It should be noted that Figure 9 This is merely an example and does not constitute a limitation on the visual positioning method of the embodiment of the present application.

[0202] Of course, it is understandable that in the embodiment of the present application, when the mobile phone is pre-configured with the image features of the 360-degree panoramic image collected at M candidate geographical locations in the panoramic map, the mobile phone can execute the visual positioning process after triggering the visual positioning process. Figure 9 In steps 902 to 906 of the visual positioning method shown, the mobile phone uses the first geographical location determined in step 906 as the geographical location of itself when the camera captures the first image.

[0203] The above-mentioned visual positioning method can help mobile phones obtain a high-precision geographic location. A large number of tests using the above-mentioned method to obtain visual positioning results have shown that in effective scenarios, 99% of the geographic location positioning errors are less than 5 meters, 90% of the geographic location positioning errors are less than 3 meters, and 75% of the geographic location positioning errors are less than 2 meters. Compared with the existing technology, the positioning accuracy of the geographic location is greatly improved. Moreover, the visual positioning method of the embodiment of the present application can also obtain a high-precision device posture. A large number of tests using the above-mentioned method to obtain visual positioning results have shown that the pitch angle and roll angle errors of the mobile phone are within 1°, 99% of the heading angle errors are within 3 degrees, and 90% of the heading angle errors are within 1°. Compared with the existing technology, the accuracy of obtaining the device posture is greatly improved.

[0204] It should be noted that, in some other embodiments, the server can also perform visual positioning in combination with multiple frames of images. For example, after triggering the visual positioning process, the mobile phone reports the jth frame of image to the server before reporting the i-th frame of image to the server, but the server fails to locate based on the j-th frame of image. Then, after receiving the i-th frame of image reported by the mobile phone, the server can use the low-precision geographic location and low-precision posture of the device when collecting the i-th frame of image, as well as the low-precision geographic location and low-precision posture of the device when collecting the j-th frame of image, to determine that the position change difference of the content in the i-th frame of image and the j-frame of image is small, and then splice the i-th frame of image and the j-frame of image into one frame of image, and perform visual positioning based on the spliced image. This helps to increase the amount of information in the image and improve the accuracy and success rate of visual positioning. For specific methods, please refer to Figure 9 The visual positioning method shown will not be described in detail here. When the server determines that the position change difference between the content in the i-th frame image and the j-th frame image is large, visual positioning is performed based on the i-th frame image. Of course, the above method can be applied to the visual positioning of three or more frames of images. When three or more frames of images are used for visual positioning, one or more frames with more contour line types (i.e., richer semantic types) in the multiple frames of images can be selected (if the position change of the content of the image is small in the selected multiple frames of images), and the selected multiple frames of images can be spliced to obtain one frame of image, and then visual positioning is performed based on the spliced image.

[0205] Furthermore, when visual positioning is performed in combination with multiple frames, the server selects one or more frames from the multiple frames to obtain a first image and extracts the image features of the first image. Then, based on the similarity between the image features of the first image and the 360-degree panoramic images collected at the Q candidate geographic locations, the server selects the candidate geographic locations from the Q candidate geographic locations and arranges them in the top N positions in descending order based on the similarity of the image features. The first image is a frame selected from the multiple frames or an image obtained by splicing the selected multiple frames. For information about the Q candidate geographic locations, please refer to Figure 9 The relevant introduction in [1] will not be repeated here. Based on the image features of images other than the first image in the multiple frames, each of the N candidate locations is scored, and the location with the highest N candidate location information score is used as the first location included in the visual positioning result fed back to the mobile phone.

[0206] The following example uses a first image and a second image, with N being 2. The server determines geographic locations 1 and 2 from the Q candidate geographic locations based on the similarity between the image features of the first image and 360-degree panoramic images collected at the Q candidate geographic locations. Geographic location 1 has the highest similarity between its 360-degree panoramic image and the image features of the first image among the Q candidate geographic locations, and Geographic location 2 has the second highest similarity between its 360-degree panoramic image and the image features of the first image among the Q candidate geographic locations. The server determines the similarity between the image features of the second image and a 360-degree panoramic image collected at Geographic location 3 among the Q candidate geographic locations, and also determines the similarity between the image features of the second image and a 360-degree panoramic image collected at Geographic location 4 among the Q candidate geographic locations. Geographic location 3 is determined based on the relative geographic location relationship between the first and second images and Geographic location 1, and Geographic location 4 is determined based on the relative geographic location relationship between the first and second images and Geographic location 2. The server's score for geographic location 1 is F1 = K11 * L11 + K21 * L21, where L11 indicates the highest similarity between the first image and the 360-degree panoramic image captured at geographic location 1, and L21 indicates the highest similarity between the second image and the 360-degree panoramic image captured at geographic location 3. K11 and K12 can be weight coefficients, which can be related to the similarity of the image features or predefined. For example, when the similarity of the image features is within range 1, the corresponding weight coefficient is K11, and when the similarity is within range 2, the corresponding weight coefficient is K21. Similarly, the server's score for geographic location 2 is F2 = K12 * L12 + K22 * L22, where L12 indicates the highest similarity between the first image and the 360-degree panoramic image captured at geographic location 2, and L22 indicates the highest similarity between the second image and the 360-degree panoramic image captured at geographic location 4. K12 and K22 can be weight coefficients. When F1 is less than F2, the first location information included in the visual positioning result returned by the server to the mobile phone is geographic location 2, which helps to further improve the accuracy of visual positioning.

[0207] Of course, the above is only an example of geographic location scoring and does not constitute a limitation on the geographic location scoring method in the embodiment of the present application. In the embodiment of the present application, geographic location scoring can also be performed in other ways.

[0208] It should also be noted that in Figure 9In the visual positioning method shown, when the camera parameters used by the camera in the mobile phone to capture images are the same as the camera parameters used to capture 360-degree panoramic images at the candidate geographic location in the panoramic map, the camera parameters may not be included in the visual positioning request, and the server does not need to perform image processing based on the camera parameters. In addition, when the reference coordinate system of the first image included in the visual positioning request is the same as the reference coordinate system used to capture the 360-degree panoramic image at the candidate geographic location in the panoramic map, the low-precision posture of the device may not be included in the visual positioning request, and the server does not need to perform coordinate system conversion based on the altitude angle and orientation angle of the feature point in the first image.

[0209] Based on the above embodiments, the present invention provides a visual positioning method. Figure 10 As shown, the specific steps include:

[0210] 1001. An electronic device detects a first event, which is used to trigger a visual positioning process.

[0211] 1002. The electronic device determines that the content richness of the i-th frame image captured by the camera meets the visual positioning requirement, and sends a first visual positioning request to the server. The first visual positioning request includes the i-th frame image and a first geographic location measured by the electronic device when the camera captures the i-th frame image.

[0212] 1003. After receiving the first visual positioning request from the electronic device, the server extracts image features of the i-th image frame and selects Q candidate geographic locations from M candidate geographic locations in the panoramic map based on the first geographic location. A distance between each of the Q candidate geographic locations and the first geographic location is less than or equal to a first threshold, where Q is less than or equal to M, and M and Q are positive integers.

[0213] 1004. The server determines a second geographic location from the Q candidate geographic locations. The second geographic location has the highest similarity between the image features of the 360-degree panoramic image and the i-th frame image among the Q candidate geographic locations.

[0214] 1005. The server returns a first visual positioning result to the electronic device. The first visual positioning result includes the second geographic location.

[0215] In some embodiments, the first visual positioning request may further include camera parameters and / or a low-precision device posture when the camera captures the i-th frame of image. In the case where the first visual positioning request includes the low-precision device posture when the camera captures the i-th frame of image, the first visual positioning result may further include a high-precision device posture, where the high-precision device posture is used to indicate at least one of the altitude angle, pitch angle, and heading angle of the electronic device when the camera captures the i-th frame of image.

[0216] In other embodiments, if the server fails to select a candidate geographic location from the M candidate geographic locations of the panoramic map based on the first geographic location, or if the highest similarity between the image features of the 360 panoramic image and the first image among the Q candidate geographic locations is also within the image feature similarity range required for visual positioning accuracy, then the server returns a second visual positioning result to the electronic device, and the second visual positioning result is used to indicate that positioning based on the i-th frame image has failed.

[0217] about Figure 10 The specific implementation of the visual positioning method shown can be found in the relevant introduction of the above embodiments, which will not be repeated here.

[0218] The above embodiments can be used alone or in combination with each other to achieve different technical effects.

[0219] In the embodiments provided in the present application above, the methods provided in the embodiments of the present application are introduced from the perspective of electronic devices and servers as execution entities. In order to implement the various functions in the methods provided in the embodiments of the present application above, the electronic device or server may include a hardware structure and / or a software module to implement the above functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether a function of the above functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.

[0220] like Figure 11 As shown, the embodiment of the present application discloses an electronic device 1100, which may include: a camera 1101, one or more processors 1102, a memory 1103, and one or more computer programs. For example, the above-mentioned components may be connected via one or more communication buses. The one or more computer programs are stored in the above-mentioned memory 1103 and configured to be executed by the one or more processors 1102 to implement the embodiment of the present application. Figure 10 The functions implemented by the electronic device side of the visual positioning method shown.

[0221] In some embodiments, the electronic device 1100 may also include a display screen 1104 and / or a microphone 1105, where the display screen 1104 is used to display prompt information for adjusting the shooting angle of the camera 1101 and / or adjusting the device posture, and the microphone 1105 is used to display prompt information for adjusting the shooting angle of the camera 1101 and / or adjusting the device posture.

[0222] like Figure 12As shown, a server 1200 disclosed in an embodiment of the present application includes: one or more processors 1201, a memory 1202, and one or more computer programs. The one or more computer programs are stored in the memory 1202 and configured to be executed by the one or more processors 1201 to implement the embodiment of the present application. Figure 9 or Figure 10 The functions implemented on the server side of the visual positioning method shown.

[0223] In addition, an embodiment of the present application further discloses a communication system, including an electronic device 1100 and a server 1200 .

[0224] The processors involved in the above embodiments may be general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application may be directly implemented as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in a memory, and the processor reads instructions from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0225] Those skilled in the art will appreciate that the units and algorithm steps described in the various examples in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application.

[0226] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0227] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0228] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0229] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0230] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0231] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A visual positioning method, characterized in that: Applied to an electronic device, the electronic device includes a camera, and the method includes: A first event is detected, where the first event is used to trigger a visual positioning process; Determining whether the number of types of contour lines in the first image captured by the camera is greater than or equal to a first threshold; If the number of types of contour lines in the first image is greater than or equal to the first threshold, sending a first visual positioning request to the server, where the first visual positioning request includes the first image and a first geographic location, where the first geographic location is the geographic location of the electronic device measured when the camera captures the first image; receiving a first visual positioning result sent by the server in response to the first visual positioning request, the first visual positioning result including a second geographic location, the second geographic location being the geographic location of the electronic device when the camera captures the first image, the second geographic location having a higher accuracy than the first geographic location, the second geographic location being determined by the server extracting image features of the first image and selecting Q candidate geographic locations from M candidate geographic locations for capturing a 360-degree panoramic image on a panoramic map based on the first geographic location; In which, the image features of the first image include contour line indications of N feature points in the first image, and orientation angles and altitude angles of the N feature points in a first coordinate system, where N is a positive integer, the first coordinate system is a reference coordinate system of a 360-degree panoramic image collected at a candidate geographic location in the panoramic map, and the contour line indication is used to indicate the type of contour line where the feature point is located; the distance between each of the Q candidate geographic locations and the first geographic location is less than or equal to a first threshold, Q is less than or equal to M, and M and Q are positive integers; and the second geographic location has the highest similarity in image features between the 360-degree panoramic image and the first image among the Q candidate geographic locations.

2. The method according to claim 1, wherein The first visual positioning request also includes first camera parameters and / or first device posture, the first camera parameters are camera parameters used by the camera when capturing the first image, and the first device posture is used to indicate at least one of the altitude angle, pitch angle and roll angle measured by the electronic device when the camera captures the first image.

3. The method according to claim 1, wherein Before determining whether the number of types of contour lines in the first image captured by the camera is greater than or equal to a first threshold, the method further includes: It is determined that the pitch angle of the electronic device is within a first angle range and the roll angle of the electronic device is within a second angle range.

4. The method according to any one of claims 1 to 3, characterized in that The first visual positioning result also includes: a second device posture, which is used to indicate at least one of the altitude angle, pitch angle and roll angle of the electronic device when the camera captures the first image, and the accuracy of the second device posture is higher than the accuracy of the first device posture; the first device posture is used to indicate at least one of the altitude angle, pitch angle and roll angle of the electronic device measured by the camera when the first image is captured.

5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: receiving a second visual positioning result sent from the server in response to the first visual positioning request, where the second visual positioning result indicates a failure of positioning based on the first image; When the proportion of repeated content in the second image captured by the camera and the first image is less than or equal to a second threshold, and the number of types of contour lines in the second image is greater than or equal to the first threshold, a second visual positioning request is sent to the server, where the second visual positioning request includes the second image and a third geographic location, where the third geographic location is the geographic location of the electronic device measured when the camera captures the second image.

6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the number of types of contour lines in the first image is less than the first threshold, the user is prompted to adjust the shooting angle of the camera.

7. A visual positioning method, characterized in that: The method comprises: The server receives a first visual positioning request from an electronic device, where the first visual positioning request includes a first image and a first geographic location; The server extracts image features of the first image, where the image features of the first image include contour line indications of N feature points in the first image, and orientation angles and altitude angles of the N feature points in a first coordinate system, where N is a positive integer, the first coordinate system is a reference coordinate system of a 360-degree panoramic image captured at a candidate geographical location in a panoramic map, and the contour line indications are used to indicate types of contour lines where the feature points are located; The server selects, based on the first geographic location, Q candidate geographic locations from M candidate geographic locations from which 360-degree panoramic images are collected in the panoramic map, where a distance between each of the Q candidate geographic locations and the first geographic location is less than or equal to a first threshold, and Q is less than or equal to M, where M and Q are positive integers; The server determines a second geographical location from the Q candidate geographical locations, where the second geographical location has the highest similarity between the 360-degree panoramic image and the image features of the first image among the Q candidate geographical locations; The server returns a first visual positioning result to the electronic device, where the first visual positioning result includes the second geographic location.

8. The method according to claim 7, wherein The first visual positioning request also includes a first device posture; The server extracting image features of the first image includes: The server performs semantic segmentation on the first image to obtain a semantic map of the first image, and obtains, based on the semantic map of the first image, contour line indications of N feature points in the first image, and orientation angles and altitude angles of the N feature points in a second coordinate system; the second coordinate system is a reference coordinate system for the first image; The server obtains the orientation angles and altitude angles of the N feature points in the first coordinate system according to the first device posture and the orientation angles and altitude angles of the N feature points in the second coordinate system.

9. The method according to claim 8, wherein The first visual positioning request further includes first camera parameters, where the first camera parameters are camera parameters used to capture the first image; The server performs semantic segmentation on the first image to obtain a semantic graph of the first image, including: The server performs image processing on the first image to obtain an intermediate image, and performs semantic segmentation on the intermediate image to obtain a semantic graph of the first image; the camera parameters of the intermediate image are second camera parameters; or, The server performs semantic segmentation on the first image to obtain a semantic map of an intermediate image, and performs image processing on the semantic map of the intermediate image to obtain a semantic map of the first image, wherein camera parameters of the semantic map of the first image are the second camera parameters; The second camera parameters are camera parameters used when capturing a 360-degree panoramic image at a candidate geographical location in the panoramic map.

10. The method according to any one of claims 7 to 9, characterized in that: The similarity between the image features of the 360-degree panoramic image collected at the second geographical location and the image features of the first image satisfies the following expression: Wherein, (x, y, h) is the second geographic location, offset is the offset of one of the orientation angles in the orientation angle offset set, Loss(x, y, h, offset) is used to characterize the similarity between the image features of the 360-degree panoramic image collected at the second geographic location and the image features of the first image, Wi is the weight value of the i-th contour line in the first image, Y(i) is the set of orientation angles of all feature points on the i-th contour line in the first image, j is the orientation angle of a feature point on the i-th contour line in the first image, and P Ι (i, j) is the height angle of the feature point on the i-th contour line in the first image when the orientation angle is j, r is the total number of contour line types in the first image, P M(x,y,h) (i, j+offset) is the altitude angle of the feature point with the orientation angle j+offset on the i-th contour line in the 360-degree panoramic image collected at the second geographical location.

11. The method according to any one of claims 7 to 9, characterized in that: Before the server selects Q candidate geographical locations from the M candidate geographical locations for collecting 360-degree panoramic images in the panoramic map, the method further includes: The server determines that the number of types of contour lines in the first image is greater than or equal to a first threshold.

12. The method according to any one of claims 7 to 9, characterized in that: The highest similarity between the image features of the 360-degree panoramic image collected at the second geographical location and the image features of the first image is within the image feature similarity range required by the visual positioning accuracy.

13. The method according to claim 12, wherein: The method further comprises: When the highest similarity between the image features of the 360 panoramic image collected at the second geographic location and the image features of the first image is not within the image feature similarity range required by the visual positioning accuracy, the server returns a second visual positioning result to the electronic device, and the second visual positioning result is used to indicate that positioning based on the first image has failed.

14. An electronic device, characterized in that: The electronic device includes a camera, one or more processors, a memory, and one or more computer programs; Wherein, the camera is used to collect images; The computer program is stored in the memory, and when the processor is executed, the computer program is called, so that the electronic device executes the method according to any one of claims 1 to 6.

15. A server, characterized in that: The server includes one or more processors, memory, and a computer program; The computer program is stored in the memory; and when the processor is running, the computer program is called, so that the server executes the method according to any one of claims 7 to 13.

16. A chip, characterized in that: The chip is coupled to a memory in an electronic device so that the chip calls a computer program stored in the memory during operation to implement the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 13.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program, and when the computer program runs on the electronic device, the electronic device executes the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 13.

18. A communication system, characterized in that: The method comprises an electronic device and a server, wherein the electronic device is used to execute the method according to any one of claims 1 to 6; and the server is used to execute the method according to any one of claims 7 to 13.

Citation Information

Patent Citations

  • Panoramic map database acquisition system and vision-based positioning and navigating method

    CN103398717A

  • Method and system for positioning interaction with robot by using camera device

    CN110919644A