Visual positioning evaluation method and electronic device
The pose estimation confidence value is obtained in visual positioning evaluation through the target algorithm model, which solves the problems of difficulty in judging the accuracy of visual positioning and high computing resource consumption in the prior art, and achieves fast and accurate visual positioning evaluation, which is suitable for a wide range of application scenarios.
Patent Information
- Application Number
- CN201911137563.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2039-11-19
AI Technical Summary
In the vision-based spatial positioning evaluation, the prior art lacks an effective accuracy judgment mechanism, which leads to the system deviating from the cruise path without knowing it, and consumes a lot of computing resources and time.
Through the target algorithm model, the feature information of the image is obtained and the pose estimation confidence value is calculated, avoiding analyzing a large number of images, simplifying the calculation process, and saving computing resources and time.
It realizes the rapid and accurate evaluation of visual positioning in scenarios such as lighting and seasons, and has a wider range of applications, low cost and low system complexity, and is suitable for applications such as autonomous driving and machine navigation.
Smart Images

Figure CN112907658B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of electronic technologies, and in particular, to a visual positioning evaluation method and an electronic device. Background Art
[0002] Vision-based spatial positioning has important applications in various scenarios such as robot navigation, autonomous driving, and augmented reality. When applying vision-based spatial positioning, it is necessary to judge the accuracy of vision-based spatial positioning in real time. If there is no mechanism for judging the accuracy of visual positioning, the system will deviate from the cruise path without knowing it.
[0003] Currently, the method for evaluating the confidence value of visual positioning only using visual information is mainly to calculate the difference between the feature descriptors of the positioning image and the matching image, mainly to find the similarity between the matching image and the positioning image. This requires a high overlapping area between the matching image and the positioning image. In order to maintain the reliability of the confidence value results in a large range, it is necessary to densely sample at various angles and positions in space, analyze a large number of images, thereby generating a large amount of data, which will consume a large amount of computing resources and computing time. Summary of the Invention
[0004] Embodiments of the present application provide a visual positioning evaluation method and an electronic device. The electronic device obtains a pose estimation confidence value corresponding to the feature information of the image through a target algorithm model, without analyzing a large number of images, and can save computing resources and computing time.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0006] On the one hand, the technical solution of the present application provides a visual positioning evaluation method, which includes: First, the electronic device obtains a first image. Then, the electronic device determines a first estimated pose corresponding to the first image according to the first image. After that, the electronic device calculates first feature information of the first image according to the first image and the first estimated pose. Finally, the electronic device obtains a first pose estimation confidence value corresponding to the first estimated pose according to the first feature information by using a target algorithm model. The target algorithm model is used to represent the corresponding relationship between the feature information of the image and the pose estimation confidence value, and the pose estimation confidence value is used to represent the credibility of the estimated pose.
[0007] In this solution, the electronic device can obtain a pose estimation confidence value corresponding to the feature information of the image by using a target algorithm model, without analyzing a large number of images, without generating a large amount of data, the calculation process is simple, and it can better save computing resources and computing time.
[0008] In a possible implementation, the ratio of the maximum value to the minimum value of the pose estimation confidence value obtained by using the target algorithm model is less than a first preset value.
[0009] That is to say, the variance of the pose estimation confidence value obtained by using the target algorithm model is small. The smaller the variance, the smaller the fluctuation of the data. Therefore, the pose estimation confidence value obtained by using the target algorithm model is relatively stable, which is convenient for subsequent operations such as pose correction.
[0010] In another possible implementation, the pose estimation confidence value obtained by using the target algorithm model is within a preset interval, and the preset interval includes a threshold value; if the first pose estimation confidence value is greater than or equal to the threshold value, the first estimated pose is highly credible; or, if the first pose estimation confidence value is less than the threshold value, the first estimated pose is less credible; the method further includes: the electronic device re-determines the second estimated pose corresponding to the first image according to the first image.
[0011] That is to say, the electronic device divides whether the estimated pose is credible through the threshold value, and makes timely adjustments to the pose estimation according to the relationship between the first pose estimation confidence value and the threshold value, which can assist the autonomous positioning and navigation of the electronic device.
[0012] In another possible implementation, before the electronic device acquires the first image, the method further includes: the electronic device constructs a reference data set, the reference data set includes multiple groups of data, and each group of data in the multiple groups of data includes matching second image data and third image data, the second image data is the data of the acquired second image, and the third image data is the data of the image in the database; the reference data set includes a first data set. Then, the electronic device calculates the feature information of the second image data in the first data set according to the second image data and the third image data in the first data set. After that, the electronic device obtains a first pose error according to the pose corresponding to the second image in the first data set and the third estimated pose; the third estimated pose is the estimated pose obtained by the electronic device according to the third image data in the first data set. Finally, the electronic device determines the target algorithm model according to the feature information of the second image in the first data set and the first pose error.
[0013] That is to say, before the electronic device acquires the first image, it can train the algorithm model so as to determine the pose estimation confidence value corresponding to the estimated pose based on the trained target algorithm model.
[0014] In another possible implementation, the electronic device determines a target algorithm model according to the feature information of the second image in the first dataset and the first pose error, including: The electronic device determines a reference algorithm model according to the feature information of the second image data in the first dataset. Then, the electronic device obtains an estimated pose error based on the reference algorithm model and the feature information of the second image data in the first dataset, and the estimated pose error is used to represent the accuracy of the third estimated pose. After that, the electronic device calculates the difference value between the estimated pose error and the first pose error. Finally, the electronic device determines the target algorithm model according to the reference algorithm model and the difference value.
[0015] That is to say, the electronic device can use the difference value as the supervision value for model training, and determine the target algorithm model according to the difference value and the reference algorithm model.
[0016] In another possible implementation, there are multiple reference algorithm models determined by the electronic device according to the feature information of the second image data in the first dataset, and different reference algorithm models correspond to different difference values. The electronic device determines the target algorithm model according to the reference algorithm model and the difference value, including: The electronic device determines the reference algorithm model corresponding to the minimum value among the different difference values as the target algorithm model.
[0017] That is to say, the electronic device determines the reference algorithm model with the smallest difference degree between the output estimated pose error and the true pose error as the target algorithm model.
[0018] In another possible implementation, the feature information of the second image data in the first dataset may include at least one of the number of inliers in the second image data in the first dataset, the number of matching pairs between the 2D points of the second image data and the 3D point cloud of the third image data, the covariance matrix of the inliers in the horizontal and vertical directions of the second image data, the reprojection error of the inliers in the second image data, or the number of feature points in the second image data. Among them, the inliers in the second image data are the points in the second image data where the pixel error between the fourth image data and the second image data is less than the second preset value, and the fourth image data is the 2D image data projected according to the 3D point cloud of the third image data and the third estimated pose.
[0019] That is to say, the electronic device can determine the target algorithm model according to the feature information of the second image in the first dataset including the above information and the first pose error.
[0020] In another possible implementation, the reference data set includes a second data set. The second data set includes matching second image data and third image data. The second image data is the data of the second image collected, and the third image is the data of the image in the database. After the electronic device determines the target algorithm model according to the feature information of the second image data in the first data set and the first pose error, the method further includes: the electronic device uses the second data set to evaluate the accuracy of the target algorithm model.
[0021] That is to say, the electronic device can use the second data set to test the already constructed target algorithm model to evaluate the accuracy of the target algorithm model.
[0022] In another possible implementation, the reference data set includes a third data set, and the third data set is used for: evaluating the accuracy of the reference algorithm model during the process in which the electronic device determines the target algorithm model according to the feature information of the second image in the first data set and the first pose error.
[0023] That is to say, the electronic device can test the reliability of the model during the construction of the algorithm model through the third data set.
[0024] In another possible implementation, the algorithm model is an extreme gradient boosting tree model.
[0025] In this solution, the extreme gradient boosting tree model occupies less resources and can be pre-loaded into the memory, which can further reduce the calculation time of the pose estimation confidence value and facilitate the application of this method to small devices with limited computing power to achieve visual space positioning accuracy estimation.
[0026] On the other hand, the present technical solution provides a visual positioning evaluation device, which is included in the electronic device and has the function of implementing the behavior of the electronic device in the above aspects and the possible implementations of the above aspects. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, an acquisition module or unit, a calculation module or unit, etc.
[0027] On the other hand, the present technical solution provides an electronic device, including: one or more processors; a memory; a plurality of applications; and one or more computer programs; one or more computer programs are stored in the memory, and one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device is caused to execute the visual positioning evaluation method in any one of the above aspects and any possible implementation.
[0028] On the other hand, the present technical solution provides a computer-readable storage medium, including computer instructions, which, when running on an electronic device, enable the electronic device to execute the visual positioning evaluation method in any one of the possible implementations in any of the above aspects.
[0029] On the other hand, the present technical solution provides a computer program product, which, when running on an electronic device, enables the electronic device to execute the visual positioning evaluation method in any one of the possible designs in any of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0031] Figure 2 FIG. is a flowchart of a visual positioning evaluation method provided by an embodiment of the present application;
[0032] Figure 3 FIG. is a flowchart of another visual positioning evaluation method provided by an embodiment of the present application;
[0033] Figure 4 FIG. is a flowchart of another visual positioning evaluation method provided by an embodiment of the present application;
[0034] Figure 5 FIG. is a schematic diagram of program code for regression analysis provided by an embodiment of the present application;
[0035] Figure 6 FIG. is a flowchart of another visual positioning evaluation method provided by an embodiment of the present application;
[0036] Figure 7 FIG. is a schematic structural diagram of another electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.
[0038] Evaluating the accuracy of vision-based spatial positioning is crucial for various scenarios such as robot navigation, autonomous driving, and augmented reality. In the existing technologies for evaluating the confidence value of vision-based positioning using only visual information, the evaluation results are only valid in a small number of scenarios. When environmental variables such as lighting and seasons change significantly, the positioning image may have a large difference from the matching image stored in the database, and the confidence value given by this method will have a large error. Therefore, this method is easily affected by the environment and has a relatively limited scope of use.
[0039] In another method for evaluating the accuracy of vision-based spatial positioning, it is also necessary to rely on other information besides the visual content itself, such as data from non-visual sensors in autonomous driving and information on the robot's pose in robot navigation. Evaluating the confidence value of vision-based positioning using the above information requires additional sensors, which results in higher costs and higher system complexity.
[0040] The embodiment of the present application provides a vision-based positioning evaluation method, which can be applied to an electronic device. The electronic device acquires a first image; then, based on the first image, determines a first estimated pose corresponding to the first image. After that, the electronic device calculates first feature information of the first image according to the first image and the first estimated pose. Finally, the electronic device obtains a first pose estimation confidence value corresponding to the first estimated pose according to the first feature information by using a target algorithm model. Among them, the target algorithm model is used to represent the correspondence between the feature information of the image and the pose estimation confidence value, and the pose estimation confidence value is used to represent the credibility of the estimated pose. Thus, the electronic device can obtain the pose estimation confidence value corresponding to the feature information of the image by using the target algorithm model, without analyzing a large number of images, without generating a large amount of data, the calculation process is simple, and it can better save computing resources and computing time.
[0041] The vision-based positioning evaluation method provided by the embodiment of the present application can be applied to scenarios where environmental variables such as lighting and seasons change significantly, and has a wider scope of application. Moreover, the method provided by the embodiment of the present application does not require additional sensors, has lower costs, and lower system complexity, which is conducive to the combination of vision-based positioning with various applications such as autonomous driving and machine navigation.
[0042] Among them, the pose involved in the embodiment of the present application is used to indicate the position and orientation corresponding to the image. Exemplarily, the position can be represented by a rectangular coordinate system (x, y, z) space, and the orientation can be represented by a (yaw, roll, pitch) space. Among them, yaw represents the parameter of rotation around the z-axis; roll represents the parameter of rotation around the x-axis; pitch represents the parameter of rotation around the y-axis.
[0043] The message reminder method provided by the embodiments of this application can be applied to electronic devices with visual positioning functions, such as autonomous driving vehicles, virtual-real fusion devices, machine navigation devices, mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of this application do not impose any restrictions on the specific types of electronic devices.
[0044] Exemplarily, Figure 1 The structural schematic diagram of the electronic device 100 is shown. The electronic device 100 may include a processor 110, a memory 120, an antenna 1, an antenna 2, a mobile communication module 130, a wireless communication module 140, a sensor module 150, a camera 160, a display screen 170, a key 180, an indicator 190, etc. Among them, the sensor module 150 may include a gyroscope sensor 150A, a barometric pressure sensor 150B, a magnetic sensor 150C, an acceleration sensor 150D, etc.
[0045] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0046] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0047] The processor 110 may include a spatial computing (SC) module 111 and a simultaneous localization and mapping (SLAM) module 112. Among them, the SC module 111 may be used to perform visual spatial positioning estimation for determining the pose corresponding to an image based on the image. The SLAM module 112 may be used to perform real-time tracking of the pose of the electronic device based on the pose estimation result of the SC module 111, so as to achieve autonomous positioning and navigation of the electronic device. In some embodiments, the SC module 111 and the SLAM module 112 may be integrated in the same processing unit of the processor 110.
[0048] The wireless communication function of the electronic device 100 may be implemented by the antenna 1, the antenna 2, the mobile communication module 130, the wireless communication module 140, the modulation and demodulation processor, and the baseband processor, etc.
[0049] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 may be used to cover a single or multiple communication frequency bands. Different antennas may also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 may be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.
[0050] The memory 120 is used to store the application program code for executing the solution of this application and is controlled by the processor 110 to execute. The memory 120 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor through a bus. The memory may also be integrated with the processor.
[0051] The memory 120 may store a pose estimation module 121. The pose estimation module 121 is a program code module including the visual positioning evaluation method provided in the embodiments of the present application, and is used to perform a confidence evaluation on the estimated pose calculated by the SC module 111 to judge the reliability of the pose estimation result. Thereby, the SLAM module 112 can perform real-time tracking of the pose of the electronic device based on the result of the confidence evaluation of the pose estimation module 121, so as to realize the autonomous positioning and navigation of the electronic device. That is to say, the visual positioning evaluation method provided in the embodiments of the present application can serve as a communication bridge between the SC module 111 and the SLAM module 112 to complete the functions of autonomous positioning and navigation of the electronic device.
[0052] In some embodiments, the SC module 111, the SLAM module 112, and the pose estimation module 121 may be integrated in the same processing unit of the processor 110.
[0053] The mobile communication module 130 may provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 130 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 130 may receive electromagnetic waves by the antenna 1, filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 130 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 and radiate it out. In some embodiments, at least some functional modules of the mobile communication module 130 may be disposed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be disposed in the same device.
[0054] The wireless communication module 140 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 140 may be one or more devices integrating at least one communication processing module. The wireless communication module 140 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 140 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0055] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 130, and antenna 2 is coupled to wireless communication module 140, such that electronic device 100 can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0056] Camera 160 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element may be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, electronic device 100 may include one or N camera 160s, where N is a positive integer greater than 1.
[0057] The display screen 170 is used to display images, videos, etc. The display screen 170 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 170, where N is a positive integer greater than 1.
[0058] The button 180 includes a power-on button, a volume button, etc. The button 180 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100.
[0059] The indicator 190 can be an indicator light and can be used to indicate the charging state, the change in battery level, and can also be used to indicate messages, missed calls, notifications, etc.
[0060] The gyroscope sensor 150A can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 150A. The gyroscope sensor 150A can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 150A detects the angle of jitter of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and makes the lens offset the jitter of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 150A can also be used in navigation and somatosensory game scenarios.
[0061] The barometric pressure sensor 150B is used to measure barometric pressure. In some embodiments, the electronic device 100 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 150B to assist in positioning and navigation.
[0062] The magnetic sensor 150C can be used to determine the motion posture of the electronic device 100 according to the earth's magnetic field.
[0063] The acceleration sensor 150D can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). For example, the magnitude of the acceleration of the electronic device 100 in the (yaw, roll, pitch) three directions. When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0064] In the embodiment of the present application, the camera 160 can be used to capture an image. Then, the electronic device 100 can perform pose estimation on the captured image through the SC module 111 in the processor 110. The processor 110 can use the pose estimation module 121 stored in the memory 120 to evaluate the estimated pose and obtain a pose estimation confidence value. After that, the SLAM module 112 in the processor 110 can perform real-time tracking of the pose of the electronic device based on the result of the confidence evaluation of the pose estimation module 1202, so as to achieve autonomous positioning and navigation of the electronic device. Thus, the electronic device 100 does not need to analyze a large number of images to evaluate the pose reliability, and can save computing resources and computing time.
[0065] For ease of understanding, the following embodiments of the present application will take an electronic device with Figure 1 the structure shown as an example, and specifically describe the visual positioning evaluation method applied to the electronic device in combination with the accompanying drawings.
[0066] See Figure 2 , the visual positioning evaluation method may include:
[0067] 201. The electronic device acquires a first image.
[0068] Among them, the electronic device can be a mobile phone, an autonomous driving vehicle, a virtual-reality fusion device, a machine navigation device, etc. The first image can be an image in an actual application scenario, and the first image can be used to obtain pose data during visual positioning.
[0069] In some embodiments, for the electronic device to acquire a first image, it may include: the electronic device acquires an image captured by the electronic device. For example, when the electronic device is a mobile phone, in the shooting scenario, the mobile phone can acquire the image captured by the camera. For another example, when the electronic device is an autonomous driving vehicle, in the autonomous driving scenario, the autonomous driving vehicle can acquire the image collected by the camera.
[0070] In some other embodiments, the electronic device obtaining the first image may include: the electronic device obtaining an image received by the electronic device from other devices. The first image may be an image received by the electronic device in real time or an image received by the electronic device in advance. For example, when the electronic device is an autonomous driving vehicle, the first image may be an image received by the autonomous driving vehicle in real time from other devices or networks such as another autonomous driving vehicle, a mobile phone, a tablet computer, a communication network, etc.; or an image received by the autonomous driving vehicle in advance from other devices or networks such as another autonomous driving vehicle, a mobile phone, a tablet computer, a communication network, etc.
[0071] Among them, in scenarios such as autonomous driving and robot navigation that require real-time visual positioning, the first image is usually an image received in real time.
[0072] In some other embodiments, the electronic device obtaining the first image may include: the electronic device obtaining an image stored in the electronic device. For example, the image stored in the electronic device may be an image captured and stored by the electronic device, or an image downloaded and stored by the electronic device, or an image stored in the electronic device by other means. The embodiments of the present application do not limit the manner in which the image is stored in the electronic device.
[0073] It can be understood that the manner in which the electronic device obtains the first image is not limited to the above manner, and the embodiments of the present application do not limit the manner in which the electronic device obtains the first image.
[0074] That is to say, the electronic device can obtain an image as the first image through various channels, so as to subsequently determine the first estimated pose and its corresponding first pose estimation confidence value based on the first image.
[0075] 202. The electronic device determines a first estimated pose corresponding to the first image according to the first image.
[0076] The first estimated pose refers to the estimated pose corresponding to the first image obtained by the electronic device according to the first image, and is an estimated value of the actual pose corresponding to the first image.
[0077] In some embodiments, the electronic device may utilize Figure 1 the SC module 111 shown to determine the first estimated pose corresponding to the first image. It can be understood that the electronic device may also utilize other pose estimation modules or program codes for pose estimation to determine the first estimated pose corresponding to the first image. The embodiments of the present application do not limit the manner in which the electronic device determines the first estimated pose corresponding to the first image.
[0078] 203. The electronic device calculates first feature information of the first image according to the first image and the first estimated pose.
[0079] Among them, the first feature information of the first image is used to describe the features of the pixel points of the first image. The first feature information may include various information. For example, it may include at least one of the number of inliers in the first image, the number of matching pairs between the 2D points of the first image and the 3D point cloud of the matching image of the first image, the covariance matrix of the inliers in the horizontal and vertical directions of the first image, the reprojection error of the inliers in the first image, or the number of feature points in the first image, etc.
[0080] The inliers of the first image are the points in the first image where the pixel error between the first image and the 2D image projected according to the 3D point cloud of the matching image of the first image and the first estimated pose is less than a second preset value. The 2D image projected according to the 3D point cloud of the matching image of the first image and the first estimated pose can be the 2D image obtained by projecting the 3D point cloud of the matching image of the first image along the first estimated pose, and can be simply referred to as the first projected image.
[0081] The second preset value is a predefined value. Exemplarily, the second preset value can be the ratio of the pixels of the first image to the first projected image; for example, the second preset value can be 1.1, 1.5, 1.8, 2, 3, or 5, etc. The embodiments of the present application do not limit the magnitude of this ratio. Again exemplarily, the second preset value can be the pixel difference between the first image and the first projected image; for example, when the pixel value is in the range of [0, 255], the second preset value can be 10, 20, 35, 50, 70, or 100, etc. The embodiments of the present application do not limit the magnitude of this pixel difference. Again exemplarily, the second preset value can be the percentage of the pixel difference between the first image and the first projected image and the pixel value of the first image; for example, the second preset value can be 5%, 10%, 16%, 20%, 30%, or 40%, etc. The embodiments of the present application do not limit the magnitude of this percentage.
[0082] The number of matching pairs between the 2D points of the first image and the 3D point cloud of the matching image of the first image refers to the number of point pairs where the 2D points of the first image and the corresponding 3D point cloud of the matching image of the first image correspond to the same object. For example, when the first image and its matching image include multiple objects such as people, buildings, and roads, if the 2D points of the first image and the corresponding 3D point cloud of the matching image of the first image both correspond to the object of the building, it can be considered that the 2D points of the first image and the corresponding 3D point cloud of the matching image of the first image are matching point pairs.
[0083] The covariance matrix of the inliers in the horizontal and vertical directions of the first image can be the covariance matrix of the inliers in the x and y directions of the first image.
[0084] The reprojection error of the inliers in the first image can be the pixel error between the inliers in the first image and the points in the corresponding first projected image.
[0085] The number of feature points of the first image is the number of feature points of the first image extracted based on a feature extraction algorithm. The feature extraction algorithm can be a convolutional neural network (CNN), scale-invariant feature transform (SIFT), oriented FAST and rotated BRIEF (ORB), decision tree, support vector machine, etc. The embodiments of the present application do not limit the type of the feature extraction algorithm.
[0086] In some implementation manners, when extracting the feature points of the first image based on the feature extraction algorithm, the electronic device can add various environmental variable information such as texture and illumination in the actual usage scenario, so that the extracted feature points of the first image can be feature points covering various environmental variable information, making the first feature information of the first image have better robustness to changes in environmental variables and a wider applicable range. Therefore, the electronic device can subsequently use the target algorithm model according to the first feature information to obtain a first pose estimation confidence value that has better robustness to changes in environmental variables.
[0087] 204. The electronic device uses the target algorithm model according to the first feature information to obtain a first pose estimation confidence value corresponding to the first estimated pose.
[0088] Among them, the target algorithm model is used to represent the correspondence between the feature information of the image and the pose estimation confidence value, and the pose estimation confidence value is used to represent the credibility of the estimated pose. The electronic device can use the target algorithm model according to the first feature information to obtain a first pose estimation confidence value corresponding to the first estimated pose, so as to represent the credibility of the first estimated pose. The first pose estimation confidence value can subsequently be used by the electronic device as a reference for pose correction and the like.
[0089] The pose estimation confidence value obtained by using the target algorithm model can be within a preset interval. The preset interval is the predefined value range of the pose estimation confidence value. For example, the preset interval can be [0, 1], [0, 2], [1, 3], or [5, 6], etc. The embodiments of the present application do not limit the preset interval.
[0090] In some embodiments, the ratio of the maximum value to the minimum value of the pose estimation confidence value obtained by using the target algorithm model is less than a first preset value. The first preset value is a preset value. For example, the first preset value can be 30, 50, or 60, etc. The embodiments of the present application do not limit the magnitude of the first preset value. At the same time, the minimum value of the pose estimation confidence value may be 0; for example, when the preset interval is [0, 1], the minimum value of the pose estimation confidence value can be 0. Since 0 cannot be used as a divisor, when the minimum value of the pose estimation confidence value is 0, the ratio of the maximum value to the minimum value of the pose estimation confidence value is meaningless. Therefore, the fact that the ratio of the maximum value to the minimum value of the pose estimation confidence value is less than the first preset value actually means that the ratio of the maximum value of the pose estimation confidence value to the non-zero minimum value is less than the first preset value.
[0091] The pose estimation confidence value can be subsequently used by the electronic device as a reference for pose correction and the like. In the embodiments of the present application, the ratio of the maximum value to the minimum value of the pose estimation confidence value obtained by using the target algorithm model is less than the first preset value, that is to say, the variance of the pose estimation confidence value is small. The smaller the variance, the smaller the fluctuation of the data, so the pose estimation confidence value is more stable, which is convenient for subsequent operations such as pose correction.
[0092] In some embodiments, the preset interval may include a threshold. If the first pose estimation confidence value is greater than or equal to the threshold, the credibility of the first estimated pose is high; if the first pose estimation confidence value is less than the threshold, the credibility of the first estimated pose is low. Therefore, this threshold can also be called a confidence threshold.
[0093] In some embodiments, if the first pose estimation confidence value is greater than or equal to the threshold, the first estimated pose is credible; the first estimated pose can be used for subsequent operations such as pose correction. If the first pose estimation confidence value is less than the threshold, the first estimated pose is not credible. Further, if the first pose estimation confidence value is less than the threshold and the first estimated pose is not credible, the electronic device can also re-determine the second estimated pose corresponding to the first image according to the first image.
[0094] The threshold is a value preset and included in the preset interval. For example, when the preset interval is [0, 1], the threshold can be 0.3, 0.4, 0.5, 0.6, or 0.7, etc. The embodiments of the present application do not limit the magnitude of the threshold.
[0095] That is to say, the electronic device divides whether the estimated pose is credible through the threshold, and makes timely adjustments to the pose estimation according to the relationship between the first pose estimation confidence value and the threshold, which can assist the autonomous positioning and navigation of the electronic device.
[0096] In the solution described in steps 201 - 204, first, the electronic device acquires a first image; then, based on the first image, the electronic device determines a first estimated pose corresponding to the first image. After that, the electronic device calculates first feature information of the first image according to the first image and the first estimated pose. Finally, the electronic device uses the target algorithm model to obtain a first pose estimation confidence value corresponding to the first estimated pose according to the first feature information. Among them, the target algorithm model is used to represent the correspondence between the feature information of the image and the pose estimation confidence value, and the pose estimation confidence value is used to represent the credibility of the estimated pose. Thus, the electronic device can use the target algorithm model to obtain the pose estimation confidence value corresponding to the feature information of the image, without analyzing a large number of images, without generating a large amount of data, the calculation process is simple, and it can better save computing resources and computing time.
[0097] In some embodiments, before the electronic device acquires the first image, the algorithm model can be learned and trained so as to determine the first pose estimation confidence value corresponding to the first estimated pose based on the trained target algorithm model.
[0098] Exemplarily, the algorithm model can include a tree model. For example, a decision tree model, a boosting tree model, an extreme gradient boosting tree model, etc. The type of the algorithm model is not limited in the embodiments of the present application.
[0099] Among them, the tree model occupies less resources and can be pre-loaded into the memory in advance, which can reduce the calculation time of the pose estimation confidence value, so that this method can be applied to small devices with limited computing power to achieve visual spatial positioning accuracy estimation. For example, before step 204 above, the electronic device can pre-load the target algorithm model into the memory. Some tests have found that the pre-loaded tree model can be less than 200KB, and in this case, the calculation time of the tree model can be less than 0.5ms.
[0100] By using the tree model as the algorithm model to calculate the pose estimation confidence value, the electronic device can save computing resources and computing time, and is convenient for small devices with limited computing power to use.
[0101] Before step 201, referring to Figure 3 , the method may further include:
[0102] 301. The electronic device constructs a reference data set.
[0103] Among them, the reference data set includes multiple groups of data, and each group of data in the multiple groups of data includes matching second image data and third image data. The second image data is the data of the acquired second image, and the third image data is the data of the image in the database; the reference data set includes a first data set.
[0104] Exemplarily, the second image data may include 2D image data; the third image data may include 2D image data and 3D point cloud data.
[0105] 302. The electronic device calculates the feature information of the second image data in the first data set according to the second image data and the third image data in the first data set.
[0106] Among them, the first data set may also be referred to as a training set, which is used to train an algorithm model. Optionally, the training set can cover as large an actual usage scenario as possible, so that when the electronic device uses the trained target algorithm model for pose confidence evaluation, the first images in various scenarios can obtain reliable pose estimation confidence values through the target algorithm model. However, the pose evaluation results in the prior art are only valid in some scenarios, and the scope of use is relatively limited.
[0107] Among them, the feature information of the second image data in the first data set includes at least one of the following: the number of inliers in the second image data in the first data set, the number of matching pairs between the 2D points of the second image data and the 3D point cloud of the third image data, the covariance matrix of the inliers in the horizontal and vertical directions of the second image data, the reprojection error of the inliers in the second image data, or the number of feature points in the second image data; among them, the inliers in the second image data are the points in the second image data where the pixel error between the fourth image data and the second image data is less than the second preset value, and the fourth image data is the 2D image data projected according to the 3D point cloud of the third image data and the third estimated pose.
[0108] Among them, the feature information of the second image data in the first data set is similar to the first feature information of the first image involved in the foregoing step 203, and will not be elaborated here.
[0109] 303. The electronic device obtains a first pose error according to the pose corresponding to the second image in the first data set and the third estimated pose.
[0110] Among them, the pose corresponding to the second image in the first data set may be the pose corresponding to the second image collected while collecting the second image in the first data set. That is, the pose corresponding to the second image is also collected while collecting the second image, and this process can be implemented by an image acquisition device with a pose acquisition function. The specific type of this device is not limited in the embodiments of the present application. Since the pose corresponding to the second image in the first data set is collected simultaneously when the second image is collected, the pose corresponding to the second image in the first data set can represent the actual pose corresponding to the second image, and therefore, it can also be referred to as the true pose of the second image.
[0111] Among them, the third estimated pose is the estimated pose obtained by the electronic device based on the third image data in the first dataset. Since the collected second image data matches the third image data in the database, the third estimated pose can also be referred to as the estimated pose corresponding to the second image.
[0112] In some embodiments, the electronic device can use Figure 1 the SC module 111 shown in the figure to obtain the estimated pose according to the third image data in the first dataset. It can be understood that the electronic device can also use other pose estimation modules or program codes for pose estimation to obtain the estimated pose according to the third image data in the first dataset. The embodiments of the present application do not limit the manner in which the electronic device obtains the estimated pose according to the third image data in the first dataset.
[0113] Among them, the first pose error is used to represent the pose difference between the pose corresponding to the second image in the first dataset and the third estimated pose. Since the second image in the first dataset can be referred to as the true pose of the second image, and the third estimated pose can be referred to as the estimated pose corresponding to the second image, the first pose error reflects the difference between the true pose and the corresponding estimated pose of the second image in the first dataset. Therefore, the first pose error can also be referred to as the first true pose error.
[0114] Exemplarily, the first pose error can be the Euclidean distance between the pose corresponding to the second image in the first dataset and the third estimated pose in the (x, y, z) space and the Euclidean distance in the (yaw, roll, pitch) space. Among them, the Euclidean distance in the (x, y, z) space is used to represent the position error, and the Euclidean distance in the (yaw, roll, pitch) space is used to represent the attitude error, so as to reflect the pose difference between the true pose and the corresponding estimated pose of the second image in the first dataset through the Euclidean distances in these two spaces.
[0115] It can be understood that there are other representation forms of the first pose error, and the embodiments of the present application do not limit the representation form of the first pose error.
[0116] 304. The electronic device determines the target algorithm model according to the feature information of the second image in the first dataset and the first pose error.
[0117] That is to say, before the electronic device acquires the first image, it can train the algorithm model so as to determine the pose estimation confidence value corresponding to the estimated pose based on the trained target algorithm model.
[0118] Among them, referring to Figure 4 , step 304 may specifically include:
[0119] 401. The electronic device determines a reference algorithm model according to the feature information of the second image data in the first dataset.
[0120] Among them, the electronic device can first determine the initial range of the parameters in the algorithm model according to the feature information of the second image data in the first dataset, and then take random numbers within the initial range to generate corresponding multiple reference algorithm models.
[0121] In some implementation manners, when the electronic device takes random numbers within the initial range to generate corresponding multiple reference algorithm models, it may include: the electronic device takes random numbers within the initial range and generates corresponding multiple reference algorithm models through regression analysis. Among them, regression analysis refers to a statistical analysis method for determining the quantitative relationship of interdependence between two or more variables. Exemplarily, the program code for regression analysis can be as Figure 5 shown.
[0122] In some embodiments, the electronic device can determine a reference algorithm model according to the feature information of the second image data in the first dataset and the database form.
[0123] Among them, the database form can be used to represent the situation of the images involved in the actual usage scenario. Exemplarily, the database form includes the richness of the image texture in the actual usage scenario and the image point source situation in the actual usage scenario, etc. For example, the image point source situation in the actual usage scenario can include the density of the image point sources in the actual usage scenario. Image texture represents the repeated local patterns in the image and their arrangement rules. The embodiments of the present application do not limit the type of the database form.
[0124] In some other embodiments, in addition to the feature information of the second image data in the first dataset and the database form, the electronic device can also combine other relevant information to determine the reference algorithm model. The embodiments of the present application do not limit the information relied on by the electronic device when determining the reference algorithm model.
[0125] 402. The electronic device obtains an estimated pose error according to the reference algorithm model and the feature information of the second image data in the first dataset.
[0126] Among them, the estimated pose error is used to represent the accuracy of the third estimated pose.
[0127] The input of the above reference algorithm model is the feature information of the image, and the output is the estimated pose error. When the electronic device obtains the estimated pose error according to the reference algorithm model and the feature information of the second image data in the first dataset, it takes the feature information of the second image data in the first dataset as the input of the reference model and obtains the output of the estimated pose error.
[0128] 403. The electronic device calculates the difference value of the estimated pose error compared to the first pose error.
[0129] As can be seen from the foregoing, the first pose error reflects the difference between the true pose and the corresponding estimated pose of the second image in the first dataset, and the first pose error can be referred to as the first true pose error. The difference value of the estimated pose error compared to the first pose error reflects the degree of difference between the estimated pose error and the true pose error.
[0130] 404. The electronic device determines a target algorithm model based on the reference algorithm model and the difference value.
[0131] Among them, in the embodiments of the present application, the difference value is used as the supervision value for model training to determine the target algorithm model.
[0132] That is to say, the electronic device can use the difference value as the supervision value for model training, and determine the target algorithm model according to the difference value and the reference algorithm model.
[0133] In some implementation manners, as can be seen from the foregoing, the reference algorithm models determined by the electronic device according to the feature information of the second image data in the first dataset may include multiple ones, and different reference algorithm models correspond to different difference values. Refer to Figure 6 , step 404 may include:
[0134] 601. The electronic device determines the reference algorithm model corresponding to the minimum value among different difference values as the target algorithm model.
[0135] It can be understood that it is expected that the difference between the estimated pose error obtained according to the trained target algorithm model and the true pose error is very small, so that the target algorithm model can more accurately evaluate the confidence value of visual positioning. Therefore, the electronic device determines the reference algorithm model corresponding to the minimum value among different difference values as the target algorithm model.
[0136] That is to say, the electronic device determines the reference algorithm model with the smallest difference degree between the output estimated pose error and the true pose error as the target algorithm model.
[0137] In some embodiments, the reference dataset includes a second dataset, the second dataset includes matched second image data and third image data, the second image data is the data of the collected second image, and the third image is the data of the image in the database.
[0138] The above-mentioned second dataset can be used to evaluate the accuracy of the target algorithm model subsequently, and the second dataset can also be referred to as the test dataset. Therefore, refer to Figure 6 , after step 601, the method may further include:
[0139] 602. The electronic device evaluates the accuracy of the target algorithm model using the second data set.
[0140] For example, the reference data set includes a total of 2,000 groups of data. Among them, 1,200 groups can be used as the first data set for model training, and 600 groups of data can be used as the second data set for evaluating the trained target algorithm model.
[0141] Exemplarily, the electronic device can evaluate the accuracy of the target algorithm model through the false positive rate and the false negative rate. When the false positive rate is less than or equal to the third preset value and the false negative rate is less than or equal to the fourth preset value, it can be considered that the accuracy of the target algorithm model is relatively high; otherwise, the accuracy of the target algorithm is relatively low. Among them, the false positive rate refers to the probability that the pose estimation confidence value output by the target algorithm model is less than the threshold when the actual pose error is relatively small. That is, the probability of misjudging the actually accurate pose estimation as an inaccurate pose estimation. The false negative rate refers to the probability that the pose estimation confidence value output by the target algorithm model is greater than or equal to the threshold when the actual pose error is relatively large. That is, the probability of misjudging the actually inaccurate pose estimation as a relatively accurate pose estimation. The third preset value and the fourth preset value are both pre-set values. And the third preset value and the fourth preset value are associated with the defined division of the actual pose error size.
[0142] Exemplarily, assume that the scenario where the position error is greater than or equal to 5 units or the attitude error is greater than or equal to 10 units is a scenario with a relatively large actual pose error. Some tests found that in the test set composed of 600 groups of data in different scenarios with a large flow of people indoors and outdoors, the situation where the scenario with a position error greater than or equal to 5 units or an attitude error greater than or equal to 10 units is below the confidence threshold can reach 92%, while the situation where the scenario with a position error less than 5 units or an attitude error less than 10 units is above the confidence threshold can reach 99%. Thus, the situation where the scenario with a position error greater than or equal to 5 units or an attitude error greater than or equal to 10 units is above the confidence threshold is 8%, and the situation where the scenario with a position error less than 5 units or an attitude error less than 10 units is below the confidence threshold is 1%. That is, the false positive rate is 1% and the false negative rate is 8%. At this time, if the third preset value is 2% and the fourth preset value is 10%, then the accuracy of the target algorithm model is relatively high.
[0143] Exemplarily, assume that a scenario where the position error is greater than or equal to 3 units or the attitude error is greater than or equal to 5 units is a scenario with a relatively large actual pose error. Some tests found that in a test set composed of 600 groups of data in different scenarios with a large number of people indoors and outdoors, the situation where the scenario with a position error greater than or equal to 3 units or an attitude error greater than or equal to 5 units is below the confidence threshold reached 79%, while the situation where the scenario with a position error less than 3 units or an attitude error less than 5 units is above the confidence threshold reached 93%. Thus, the situation where the scenario with a position error greater than or equal to 3 units or an attitude error greater than or equal to 5 units is above the confidence threshold is 21%, and the situation where the scenario with a position error less than 3 units or an attitude error less than 5 units is below the confidence threshold is 7%. That is, the false positive rate is 7% and the false negative rate is 21%. At this time, if the third preset value is 7% and the fourth preset value is 30%, the accuracy of the target algorithm model is relatively high.
[0144] It can be understood that the manner in which the electronic device evaluates the accuracy of the target algorithm model using the second data set is not limited to this, and the embodiments of the present application do not limit the manner in which the electronic device evaluates the accuracy of the target algorithm model using the second data set.
[0145] That is to say, the electronic device can use the second data set to test the already constructed target algorithm model to evaluate the accuracy of the target algorithm model.
[0146] In some implementation manners, when the accuracy of the target algorithm model is relatively low, the electronic device can re - execute steps 401 to 404 to re - perform model training. The electronic device can loop through this process until the accuracy of the target algorithm model is relatively high and the target algorithm model is successfully modeled.
[0147] In other implementation manners, when the accuracy of the target algorithm model is relatively low, the electronic device can increase the sample size of the training set and re - execute steps 401 to 404 to re - perform model training.
[0148] In some embodiments, the reference data set may include a third data set, and the third data set is used to: evaluate the accuracy of the reference algorithm model during the process in which the electronic device determines the target algorithm model based on the feature information of the second image in the first data set and the first pose error.
[0149] Among them, the third data set can be referred to as a validation set, which is used to validate the algorithm model during the training process of the algorithm model. For example, the reference data set includes a total of 2,000 groups of data. Among them, 1,200 groups can be used as the first data set for model training, 600 groups of data can be used as the second data set for evaluating the trained target algorithm model; 200 groups of data can be used for the third data set to validate the model during the algorithm model training.
[0150] Exemplarily, the electronic device can perform model validation by a process similar to steps 302 - 304, that is, replacing the first data set in steps 302 - 304 with the third data set to test the reliability of the determined target algorithm model. It can be understood that the manner of using the third data set for model validation is not limited to this, and the embodiments of the present application do not limit the manner of using the third data set for model validation.
[0151] That is to say, the electronic device can test the reliability of the model through the third data set during the process of algorithm model construction.
[0152] It can be understood that in order to implement the above functions, the electronic device includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0153] This embodiment can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0154] In the case of dividing each functional module corresponding to each function, Figure 7 shows a possible composition schematic diagram of the electronic device 700 involved in the above embodiment, as Figure 7 shown, the electronic device 700 may include: a construction unit 701, a calculation unit 702, an acquisition unit 703, and a determination unit 704.
[0155] Among them, the building block 701 can be used to support the electronic device 700 to execute the above-mentioned step 301, etc., and / or for other processes of the technologies described herein.
[0156] The computing unit 702 can be used to support the electronic device 700 to execute the above-mentioned step 203, step 204, step 302, step 303, step 402, step 403, step 602, etc., and / or for other processes of the technologies described herein.
[0157] The obtaining unit 703 can be used to support the electronic device 700 to execute the above-mentioned step 201, etc., and / or for other processes of the technologies described herein.
[0158] The determining unit 704 can be used to support the electronic device 700 to execute the above-mentioned step 202, step 304, step 401, step 404, step 601, etc., and / or for other processes of the technologies described herein.
[0159] It should be noted that all relevant contents of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.
[0160] The electronic device 700 provided in this embodiment is used to execute the above-mentioned antenna gain adjustment method, so it can achieve the same effect as the above implementation method.
[0161] In the case of adopting an integrated unit, the electronic device 700 may include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the actions of the electronic device 700. For example, it can be used to support the electronic device 700 to execute the steps performed by the above-mentioned building block 701, computing unit 702, obtaining unit 703, and determining unit 704. The storage module can be used to support the electronic device 700 to store program codes, data, etc. The communication module can be used to support the communication between the electronic device 700 and other devices, such as the communication with a wireless access device.
[0162] Among them, the processing module can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device for interacting with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.
[0163] Embodiments of the present application also provide a computer-readable storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the above-related method steps to implement the visual positioning evaluation method in the above embodiments.
[0164] Embodiments of the present application also provide a computer program product, which, when running on a computer, causes the computer to execute the above-related steps to implement the visual positioning evaluation method executed by the electronic device in the above embodiments.
[0165] In addition, embodiments of the present application also provide a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the visual positioning evaluation method executed by the electronic device in the above method embodiments.
[0166] Among them, the electronic device, the computer-readable storage medium, the computer program product or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, which will not be elaborated here.
[0167] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0168] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0169] The unit described as a separation component may or may not be physically separated. The component shown as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0170] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0172] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A visual positioning evaluation method, characterized in that, The method includes: The electronic device acquires a first image; The electronic device determines a first estimated pose corresponding to the first image according to the first image; The electronic device calculates first feature information of the first image according to the first image and the first estimated pose; The electronic device obtains a first pose estimation confidence value corresponding to the first estimated pose by using a target algorithm model according to the first feature information. The target algorithm model is used to represent the correspondence between the feature information of the image and the pose estimation confidence value. The pose estimation confidence value is used to represent the credibility of the estimated pose. The target algorithm model is determined according to the feature information of the second image in the first dataset and the first pose error. The first pose error is used to represent the pose difference between the pose corresponding to the second image in the first dataset and the third estimated pose.
2. The method according to claim 1, characterized in that, The ratio of the maximum value to the minimum value of the pose estimation confidence value obtained by using the target algorithm model is less than a first preset value.
3. The method according to claim 1 or 2, characterized in that, The pose estimation confidence value obtained by using the target algorithm model is within a preset interval, and the preset interval includes a threshold value; If the first pose estimation confidence value is greater than or equal to the threshold value, the first estimated pose has a high credibility; Or If the first pose estimation confidence value is less than the threshold value, the first estimated pose has a low credibility; The method further includes: The electronic device re-determines a second estimated pose corresponding to the first image according to the first image.
4. The method according to claim 1 or 2, characterized in that, Before the electronic device acquires the first image, the method further includes: The electronic device constructs a reference dataset, which includes multiple groups of data. Each group of data in the multiple groups of data includes matched second image data and third image data. The second image data is the data of the acquired second image, and the third image data is the data of the image in the database. The reference dataset includes the first dataset; The electronic device calculates the feature information of the second image data in the first dataset according to the second image data and the third image data in the first dataset; The electronic device obtains a first pose error according to the pose corresponding to the second image in the first dataset and the third estimated pose. The third estimated pose is the estimated pose obtained by the electronic device according to the third image data in the first dataset; The electronic device determines the target algorithm model according to the feature information of the second image in the first dataset and the first pose error.
5. The method according to claim 4, characterized in that, The electronic device determines the target algorithm model according to the feature information of the second image in the first dataset and the first pose error, including: The electronic device determines a reference algorithm model according to the feature information of the second image data in the first dataset; The electronic device obtains an estimated pose error according to the reference algorithm model and the feature information of the second image data in the first dataset. The estimated pose error is used to represent the accuracy of the third estimated pose; The electronic device calculates the difference value of the estimated pose error compared with the first pose error; The electronic device determines the target algorithm model according to the reference algorithm model and the difference value.
6. The method according to claim 5, characterized in that, There are multiple reference algorithm models determined by the electronic device according to the feature information of the second image data in the first data set, and different reference algorithm models correspond to different difference values; The electronic device determines the target algorithm model according to the reference algorithm model and the difference value, including: The electronic device determines the reference algorithm model corresponding to the minimum value among different difference values as the target algorithm model.
7. The method according to claim 4, characterized in that, The feature information of the second image data in the first data set includes at least one of the number of inliers of the second image data in the first data set, the number of matching pairs between the 2D points of the second image data and the 3D point cloud of the third image data, the covariance matrix of the inliers in the horizontal and vertical directions of the second image data, the reprojection error of the inliers of the second image data, or the number of feature points of the second image data; Wherein, the inliers of the second image data are the points in the second image data whose pixel error between the fourth image data and the second image data is less than a second preset value, and the fourth image data is 2D image data projected according to the 3D point cloud of the third image data and the third estimated pose.
8. The method according to claim 4, characterized in that, The reference data set includes a second data set, and the second data set includes matching second image data and third image data. The second image data is the data of the collected second image, and the third image is the data of the image in the database; After the electronic device determines the target algorithm model according to the feature information of the second image data in the first data set and the first pose error, the method further includes: The electronic device evaluates the accuracy of the target algorithm model by using the second data set.
9. The method according to claim 4, characterized in that, The reference data set includes a third data set, and the third data set is used for: evaluating the accuracy of the reference algorithm model during the process of the electronic device determining the target algorithm model according to the feature information of the second image in the first data set and the first pose error.
10. The method according to claim 1 or 2, characterized in that, The algorithm model is an extreme gradient boosting tree model.
11. An electronic device, characterized in that, Including: One or more processors; A memory; Multiple applications; And one or more computer programs; wherein, the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the following steps: Obtain a first image; Determine a first estimated pose corresponding to the first image according to the first image; Calculate first feature information of the first image according to the first image and the first estimated pose; According to the first feature information, a first pose estimation confidence value corresponding to the first estimated pose is obtained by using a target algorithm model, where the target algorithm model is used to represent the correspondence between the feature information of an image and the pose estimation confidence value, and the pose estimation confidence value is used to represent the credibility of the estimated pose; the target algorithm model is determined according to the feature information of the second image in the first dataset and the first pose error; the first pose error is used to represent the pose difference between the pose corresponding to the second image in the first dataset and the third estimated pose.
12. The electronic device according to claim 11, characterized in that, The ratio of the maximum value to the minimum value of the pose estimation confidence value obtained by using the target algorithm model is less than a first preset value.
13. The electronic device according to claim 11 or 12, characterized in that, The pose estimation confidence value obtained by using the target algorithm model is within a preset interval, and the preset interval includes a threshold value. If the first pose estimation confidence value is greater than or equal to the threshold value, the first estimated pose is highly credible. Or If the first pose estimation confidence value is less than the threshold value, the first estimated pose is less credible; the electronic device further performs the following steps: according to the first image, re-determine the second estimated pose corresponding to the first image.
14. The electronic device according to claim 11 or 12, characterized in that, Before acquiring the first image, the electronic device further performs the following steps: Construct a reference dataset, where the reference dataset includes multiple groups of data, and each group of data in the multiple groups of data includes matching second image data and third image data, the second image data is the data of the acquired second image, and the third image data is the data of the image in the database; the reference dataset includes the first dataset. According to the second image data and the third image data in the first dataset, calculate the feature information of the second image data in the first dataset. According to the pose corresponding to the second image in the first dataset and the third estimated pose, obtain a first pose error; the third estimated pose is the estimated pose obtained by the electronic device according to the third image data in the first dataset. According to the feature information of the second image in the first dataset and the first pose error, determine the target algorithm model.
15. The electronic device according to claim 14, characterized in that, The determining the target algorithm model according to the feature information of the second image in the first dataset and the first pose error includes: Determine a reference algorithm model according to the feature information of the second image data in the first dataset. Obtain an estimated pose error according to the reference algorithm model and the feature information of the second image data in the first dataset, where the estimated pose error is used to represent the accuracy of the third estimated pose. Calculate the difference value of the estimated pose error compared to the first pose error. Determine the target algorithm model according to the reference algorithm model and the difference value.
16. The electronic device according to claim 15, characterized in that, There are multiple reference algorithm models determined according to the feature information of the second image data in the first dataset, and different reference algorithm models correspond to different difference values. The determining the target algorithm model according to the reference algorithm model and the difference value includes: Determine the reference algorithm model corresponding to the minimum value among different said difference values as the target algorithm model.
17. The electronic device according to claim 14, characterized in that, The feature information of the second image data in the first data set includes at least one of the number of inliers of the second image data in the first data set, the number of matching pairs between the 2D points of the second image data and the 3D point cloud of the third image data, the covariance matrix of the inliers in the horizontal and vertical directions of the second image data, the reprojection error of the inliers of the second image data, or the number of feature points of the second image data. Among them, the inliers of the second image data are the points in the second image data where the pixel error between the fourth image data and the second image data is less than a second preset value, and the fourth image data is 2D image data projected according to the 3D point cloud of the third image data and the third estimated pose.
18. The electronic device according to claim 14, characterized in that, The reference data set includes a second data set, and the second data set includes matching second image data and third image data. The second image data is the data of the collected second image, and the third image is the data of the image in the database. After determining the target algorithm model according to the feature information of the second image data in the first data set and the first pose error, the electronic device further performs the following steps: Use the second data set to evaluate the accuracy of the target algorithm model.
19. The electronic device according to claim 14, wherein The reference data set includes a third data set, and the third data set is used to: evaluate the accuracy of the reference algorithm model during the process of determining the target algorithm model according to the feature information of the second image in the first data set and the first pose error.
20. A computer-readable storage medium, wherein Includes computer instructions, which when running on an electronic device, cause the electronic device to execute the visual positioning evaluation method according to any one of claims 1-10.
Citation Information
Patent Citations
3d Human Models Applied To Pedestrian Pose Classification
CN103886315A
Visual odometer, positioning method thereof, robot and storage medium
CN108955718A