Visual positioning method and electronic equipment

By matching two sets of relevant data for the starting floor and the target floor in visual positioning and using the data with higher scores for positioning, combined with a confidence threshold, the error problem in cross-floor positioning is solved, achieving higher accuracy and faster positioning.

CN121788596APending Publication Date: 2026-04-03HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In visual positioning technology, when positioning across floors, the positioning results are prone to discrepancies with the actual situation, especially when the user has not reached the new floor but the system indicates that they have.

Method used

By acquiring the query image, two sets of relevant data are matched, namely the departure floor and the target arrival floor. The data with the higher score is used for positioning, and the location information is determined by combining the confidence threshold to ensure accuracy and speed.

Benefits of technology

It improves visual positioning accuracy, reduces positioning errors, shortens positioning latency, and ensures accurate output of location information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788596A_ABST
    Figure CN121788596A_ABST
Patent Text Reader

Abstract

The invention provides a visual positioning method and electronic equipment, the method is applied to first electronic equipment, the method comprises the steps that a query image of a first position point where second electronic equipment is located is acquired, and the query image is an image acquired at the first position point by the second electronic equipment in response to confirmation operation of arriving at a second floor; the second floor is a target arrival floor of the second electronic equipment; according to the query image, first data and second data corresponding to the first position point are obtained, the first data are data corresponding to the first position point in a data set of a first floor, the second data are data corresponding to the first position point in a data set of a second floor, and the first floor is a departure floor of the second electronic equipment; and determining target position information of the first position point according to the data with the higher score in the first data and the second data. According to the scheme, positioning is carried out by mainly utilizing the data with relatively high scores in the two groups of related data of the query image, so that the visual positioning precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual positioning technology, and more particularly to a visual positioning method and electronic device. Background Technology

[0002] In the field of visual positioning technology, a user only needs to capture an image at a certain location, and the user's location, i.e., the location of the electronic device, can be determined based on that image using visual positioning. However, in practical applications, it has been found that when crossing floors, the positioning results are prone to discrepancies with the actual situation. For example, the user may not have reached the new floor, but the positioning result indicates that the user has already arrived at the new floor.

[0003] Therefore, improving the accuracy of visual positioning is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a visual positioning method and an electronic device that can improve visual positioning accuracy.

[0005] In a first aspect, a visual positioning method is provided, applied to a first electronic device. The method includes: acquiring a query image of a first location point of a second electronic device, wherein the query image is an image captured by the second electronic device at the first location point in response to a confirmation operation indicating that it has reached a second floor; the second floor is the target floor reached by the second electronic device; based on the query image, acquiring first data and second data corresponding to the first location point, wherein the first data is the data corresponding to the first location point in the data set of the first floor, and the second data is the data corresponding to the first location point in the data set of the second floor, wherein the first floor is the departure floor of the second electronic device; and determining the target location information of the first location point based on the data with the higher score among the first data and the second data.

[0006] In this technical solution, two sets of relevant data for the departure floor and the target arrival floor are matched separately using the query image. The data with the higher score from the two sets of matched data is used for positioning. Since positioning is based on data with a higher degree of matching with the query image and of better quality, the accuracy of visual positioning can be effectively improved. In addition, since only one set of relevant data is used for positioning each time, rather than using all the data of a certain floor, let alone all the data of two floors, the positioning speed is relatively fast, effectively reducing latency.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, when determining the location information of the first location point based on the data with the higher score among the first data and the second data, it may include: if the score of the first data is higher, determining the first predicted location information of the first location point based on the first data; if the confidence level of the first predicted location information is greater than or equal to the first confidence level threshold, determining the first predicted location information as the target location information; or, if the confidence level of the first predicted location information is less than the first confidence level threshold, determining the target location information based on the second data.

[0008] In this implementation, if the first data score is high, it means that the second electronic device is likely still on the first floor or very close to the first floor. Under this premise, a confidence threshold is set for it. Only the predicted location information that exceeds the first confidence threshold will be output as the target location information. Otherwise, the second data with a lower score will be used to determine the target location information, which can ensure the accuracy of the target location information.

[0009] In one example, when the confidence level of the first predicted location information is less than the first confidence threshold, determining the target location information based on the second data can include: determining the second predicted location information of the first location point based on the second data; determining the second predicted location information as the target location information when the confidence level of the second predicted location information is greater than or equal to the second confidence threshold, where the second confidence threshold is greater than the first confidence threshold; or, prompting to re-acquire the query image when the confidence level of the second predicted location information is less than the second confidence threshold. In this example, when the target location information is determined using the second data with a lower score, a higher confidence threshold is used to constrain the output of the location information. If the confidence level is less than the second confidence threshold, prompting to re-acquire the query image avoids the situation of "choosing a slightly better result from two very bad results," where both sets of data actually have low scores, and neither set of data will yield accurate location information. Using the lower-scoring data will, in principle, only yield location information with even lower accuracy. If, under such circumstances, the output target location information still exceeds the higher second confidence threshold, it means that the output target location information is indeed accurate enough. This can be understood as a penalty for data with poor ratings, setting a higher confidence threshold for them. In particular, if the second data should have a higher rating than the first data, but instead has a lower rating, it means that the user may have mistakenly confirmed that they have reached the second floor.

[0010] In one example, when determining the location information of a first location point based on the data with the higher score between the first and second data, the process may include: if the score of the second data is higher, determining the third predicted location information of the first location point based on the second data; if the confidence level of the third predicted location information is greater than or equal to a third confidence threshold, determining the third predicted location information as the target location information, where the third confidence threshold is less than a first confidence threshold; or, if the confidence level of the third predicted location information is less than the third confidence threshold, determining the target location information based on the first data. In this example, the higher score of the second data point indicates a higher probability that the first location point is on the second floor. The first location point is the position where the user confirms they have arrived on the second floor, and in principle, the first location point should be a location point on the second floor. Therefore, the higher score of the second data point better aligns with this principle. Thus, the third confidence threshold is set lower than the first confidence threshold. Combining this with the fact that the second confidence threshold is greater than the first confidence threshold, that is, the second confidence threshold (used when calculating location information using the second data point with a lower score) > the first confidence threshold (used when calculating location information using the first data point with a higher score) > the third confidence threshold (used when calculating location information using the second data point with a higher score).

[0011] In one example, when the confidence level of the third predicted location information is less than the third confidence threshold, determining the target location information based on the first data may include: determining the third predicted location information of the first location point based on the first data; when the confidence level of the third predicted location information is greater than or equal to the fourth confidence threshold, determining the third predicted location information as the target location information, where the fourth confidence threshold is greater than the third confidence threshold; or, when the confidence level of the third predicted location information is less than the fourth confidence threshold, prompting a re-acquisition of the query image. In this example, when calculating location information using the second data with a higher score, if the confidence level is too low, the location information is calculated using the first data with a lower score, and a higher fourth confidence threshold is used. This is similar to the penalty for poorly scored data mentioned above, to avoid the situation of "choosing a slightly better result from two very poor results," and will not be elaborated further.

[0012] In conjunction with the first aspect, in some implementations of the first aspect, when obtaining the first data and the second data corresponding to the first location point based on the query image, it may include: determining the data in the data set of the first floor that is within a first preset distance range from the first location point as the first data; and determining the data in the data set of the second floor that is within a first preset distance range from the first location point as the second data.

[0013] In this implementation, taking only data within a certain range can effectively reduce the amount of data needed to determine location information, thereby reducing the amount of computation and shortening the latency.

[0014] In one example, if the first data and / or second data determined based on a first preset distance range are empty, the method further includes: determining the data in the data set of the first floor that is within a second preset distance range from the first location point as the first data, where the second preset distance range is greater than the first preset distance range; and determining the data in the data set of the second floor that is within a second preset distance range from the first location point as the second data. In this example, expanding the data range when no data can be obtained based on the first preset distance range, and re-obtaining the first and second data based on the second preset distance range, can further ensure the amount of data obtained. This ensures that the computational load is reduced and the latency is shortened, while also avoiding the inability to match suitable data due to an excessively small query range.

[0015] In another example, if the first and second data determined based on the second preset distance range are empty, the above method also includes prompting the user to re-acquire the query image. In this example, if matching is performed based on a large range (the second preset distance range) and no suitable data is found, especially if there are no matching data for either floor, it indicates that the collected query image is likely faulty. Therefore, prompting the user to re-acquire the query image is appropriate. It should be understood that in this example, the main consideration is that if the data for the two floors differs significantly, it is easy to only match data for one floor, resulting in only the first or second data being non-empty. In such cases, it is not necessary to re-acquire the query image; only the non-empty data set needs to be used for subsequent calculations. Therefore, this example only prompts the user to re-acquire the query image when both sets of data are empty.

[0016] In conjunction with the first aspect, in some implementations of the first aspect, the above method also includes:

[0017] Determine a first similarity array and a second similarity array, wherein the first similarity array includes the similarity between each frame of at least one frame of images in the first data and the query image, and the second similarity array includes the similarity between each frame of images in at least one frame of images in the second data and the query image;

[0018] Based on the first similarity array and the second similarity array, determine the scores of the first data and the second data, respectively.

[0019] A first similarity array and a second similarity array are determined. The first similarity array includes the similarity between each frame of at least one frame in the first data and the query image. The second similarity array includes the similarity between each frame of at least one frame in the second data and the query image. Based on the first and second similarity arrays, the scores for the first data and the second data are determined respectively. In this implementation, the two sets of similarity arrays corresponding to the first and second data are obtained first, and then the scores are determined based on the two sets of similarity arrays respectively.

[0020] In one example, when determining the score of the first data and the score of the second data based on the first similarity array and the second similarity array, it may include: determining the score of the first data and the score of the second data by taking the maximum value in the first similarity array and the maximum value in the second similarity array, respectively; or, determining the score of the first data and the score of the second data by taking the average value in the first similarity array and the average value in the second similarity array, respectively. In this example, the score is determined by removing the maximum value.

[0021] In another example, at least one frame in the first dataset has a similarity greater than a first similarity threshold with the query image; at least one frame in the second dataset also has a similarity greater than the first similarity threshold with the query image. This example provides a filtering (selection) method (basis): only images with similarity thresholds within a certain range are included in either the first or second dataset. This eliminates the need to use the entire dataset from both the first and second datasets for subsequent calculations, effectively reducing computational load and improving processing efficiency and accuracy. However, it should be understood that this selection method may result in an empty set, meaning the first and / or second datasets may be empty, indicating that no data with a similarity less than the first similarity threshold can be found in either the first or second dataset. In such cases, the adaptive expansion of the search scope described above can be adopted, using a second similarity threshold that is less than the first similarity threshold. The above steps can be repeated, or a prompt to retrieve the query image again can be provided.

[0022] In another example, the method further includes prompting the user to retrieve the query image if the first layer of the dataset does not contain images with a similarity greater than a first similarity threshold, or if the second layer of the dataset does not contain images with a similarity greater than the first similarity threshold. In this example, the user is prompted to retrieve the query image whenever the similarity threshold is not met.

[0023] In another example, when the first data and / or second data determined based on the first similarity threshold are empty, the first data and second data are determined based on the second similarity threshold, such that the similarity between each frame in at least one frame of the first data and the query image is greater than the second similarity threshold; and the similarity between each frame in at least one frame of the second data and the query image is greater than the second similarity threshold, where the second similarity threshold is less than the first similarity threshold. It should be understood that this solution is equivalent to reducing the value of the first similarity threshold.

[0024] In one implementation, the first data includes the N images from the first floor's dataset that have the highest similarity to the query image, where N is a positive integer; the second data includes the N images from the second floor's dataset that have the highest similarity to the query image. This implementation retrieves the N most similar images from the entire floor's dataset, meaning N is almost never empty, as it would only be empty if there were no images in the entire floor's dataset, a situation that is highly improbable. However, this approach also results in the N images having very low actual similarity scores, only ranking high. In such cases, it's more appropriate to suggest retrieving the search images.

[0025] In a second aspect, a visual positioning device is provided, comprising a unit consisting of software and / or hardware for performing any of the methods of the first aspect.

[0026] Thirdly, an electronic device is provided, comprising: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to enable the electronic device to implement any of the methods of the first aspect.

[0027] Fourthly, a chip system is provided, the chip system being applied to an electronic device, the chip system including one or more processors, the one or more processors being configured to invoke computer instructions to enable the electronic device to implement any of the methods of the first aspect.

[0028] Optionally, the chip system also includes a memory electrically connected to the processor.

[0029] Optionally, the chip system may also include a communication interface.

[0030] Fifthly, a computer-readable storage medium is provided, the computer-readable storage medium including instructions that, when executed on an electronic device, enable the electronic device to implement any of the methods of the first aspect.

[0031] In a sixth aspect, a computer program product is provided, comprising a computer program that, when executed by an electronic device, can implement any of the methods of the first aspect. Attached Figure Description

[0032] Figure 1 This is a schematic diagram illustrating the implementation principle of visual positioning applicable to an embodiment of this application.

[0033] Figure 2 This is a schematic diagram of a cross-floor movement scenario according to an embodiment of this application.

[0034] Figure 3 This is a schematic diagram of an applicable interactive scenario according to an embodiment of this application.

[0035] Figure 4 This is a schematic flowchart of a visual positioning method according to an embodiment of this application.

[0036] Figure 5 This is a schematic flowchart illustrating another visual positioning method according to an embodiment of this application.

[0037] Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0038] The embodiments of this application will now be described in conjunction with the accompanying drawings.

[0039] Figure 1This is a schematic diagram illustrating the implementation principle of visual positioning applicable to an embodiment of this application. Visual positioning is a technology that utilizes images for localization, such as the widely used visual positioning system / service (VPS) technology. In visual positioning, a feature map needs to be constructed for some spatial areas using technologies such as simultaneous localization and mapping (SLAM), and the feature map is stored in a visual positioning device in the cloud. The feature map contains 3D point cloud data (hereinafter referred to as point cloud data) and sample images corresponding to different locations within the spatial area. Based on this, the user can use the device to be positioned to take a photo of their current location and send this photo as a query image to the visual positioning device. After receiving the query image, the visual positioning device performs image retrieval on the query image in the sample images of the feature map and matches the most similar sample image. Since the location of the sample image within the spatial area is known (recorded in the feature map), the actual location information of the query image within the spatial area can be determined based on the matched sample image. This completes the determination of location information based on visual positioning, and allows the output of the confidence level of the location information.

[0040] Based on the determined location information of the device to be located, visual positioning can further determine the pose information of the device according to the needs of the actual application. The visual positioning device continues to extract 2D feature points from the query image. Simultaneously, based on the correlation between the 2D feature points of the sample image in the cloud feature map and the 3D feature points in the 3D point cloud, the 3D feature point data associated with the most similar sample image is determined. Then, the 2D feature points of the query image are matched with the determined 3D feature point data, and the actual pose information of the device to be located within the spatial region is calculated. Additionally, confidence data for the pose information can be determined as needed.

[0041] Because the location information of the sample images is known and highly accurate when constructing feature maps of a spatial region, visual positioning, which uses image matching, can theoretically achieve high-precision positioning. This allows visual positioning to be widely applied in various indoor and outdoor positioning scenarios. Furthermore, visual positioning can also acquire the pose information of the device's location, making it well-suited for applications in virtual reality, augmented reality, and the metaverse. Through visual positioning technology, users can obtain an excellent immersive operating experience in virtual reality, augmented reality, and the metaverse.

[0042] Traditional maps only display a geographical location, while visual positioning pinpoints the exact location down to each floor within a building, offering higher accuracy. This often leads to situations involving multiple floors in visual positioning. However, in practical applications, it has been found that positioning errors can easily occur in these scenarios, such as indicating that the user has reached a new floor before actually arriving at it. Analysis revealed that user error can cause this problem. If a user confirms arrival at a new floor before actually reaching it, the visual positioning device will only be able to determine the current location based on the data from that new floor, resulting in a positioning error.

[0043] Combination Figure 2 For example, Figure 2 This is a schematic diagram of a cross-floor movement scenario according to an embodiment of this application. Figure 2 As shown in (a), suppose a user wearing a smart wearable device (an example of a device to be located) moves from floor A (the starting floor) to floor B (the destination floor). When the user reaches the stairwell, the smart wearable device will prompt the user to start crossing floors. Assuming the user... Figure 2 As shown in (a), during the journey, confirming arrival at floor B via clicking or voice control at location A or B will cause a positioning error, displaying an incorrect navigation icon on the smart wearable device. In other words, the user is prompted to turn via the navigation icon even though they haven't actually reached floor B. Figure 2 As shown in (b). Assume the user is... Figure 2 If location C (as shown in (a)) confirms that floor B has been reached, no incorrect navigation icon will appear.

[0044] To address the aforementioned issues, this application provides a novel visual positioning method. This method utilizes a query image to match two sets of relevant data for each floor: the departure floor and the target arrival floor. The method then uses the data with the higher score from the two sets of matched data for positioning. Since positioning is based on data with a higher degree of matching and superior quality to the query image, the accuracy of visual positioning can be effectively improved. Furthermore, because positioning is performed using only one set of relevant data at a time, rather than using all data from a single floor, or even all data from both floors, the positioning speed is relatively fast, effectively reducing latency.

[0045] Figure 3 This is a schematic diagram illustrating an applicable interactive scenario according to an embodiment of this application. For example... Figure 3As shown, the AR glasses are an example of a device to be located, image #1 is an example of a query image, and the cloud device is an example of a visual positioning device. When a user walks while wearing the AR glasses, and the user clicks to confirm reaching a new floor, the AR glasses capture image #1 and send it to the cloud device. The cloud device then executes... Figure 4 or Figure 5 The steps shown obtain the target location information and send it to the AR glasses. If the cloud device cannot obtain the target location information, it will send a command to the AR glasses to re-acquire the image, prompting the AR glasses to prompt the user to re-acquire and query the image.

[0046] It should be understood that Figure 3 This explanation uses the interaction between two devices as an example. However, with technological advancements, powerful computing devices are likely to emerge, allowing the device to be located and the visual positioning device to be the same device. In other words, the first electronic device and the second electronic device in this application would be integrated into a single electronic device. In this case, the difference in the interaction scheme between the two electronic devices lies only in that the steps performed by the second electronic device are also performed by the first electronic device. Therefore, this application's solution is applicable not only to current visual positioning systems that include both a device to be located and a visual positioning device, but also to future electronic devices capable of performing steps from both the device to be located and the visual positioning device.

[0047] Figure 4 This is a schematic flowchart illustrating a visual positioning method according to an embodiment of this application. The following is a description of... Figure 4 Each step is described below. This method is applied to a first electronic device, which can be, for example, a visual positioning device. The visual positioning device can be a server, host computer, cloud device, or other electronic device. The second electronic device, i.e., the device to be located, can be a smart wearable device, in-vehicle terminal, mobile phone, laptop computer, or other terminal device that can be carried, worn, or moved by the user. Smart wearable devices can be smart bracelets, virtual reality (VR) terminals, augmented reality (AR) terminals, or mixed reality (MR) terminals, etc.

[0048] S401. Obtain the query image of the first location point where the second electronic device is located.

[0049] The query image is an image captured at a first location point by the second electronic device in response to a confirmation operation that it has reached the second floor; the second floor is the target floor reached by the second electronic device.

[0050] In one implementation, the second electronic device captures an image at a first location and sends it to the first electronic device. When the first electronic device executes step S401, it receives the captured image, which is the aforementioned query image.

[0051] In another implementation, the second electronic device stores the real-time captured image for querying into a preset storage unit, and the first electronic device reads the stored image from the preset storage unit to obtain the aforementioned query image.

[0052] Other possible ways to obtain the query image will not be listed one by one.

[0053] Step S401 can be executed in response to the user confirming that they have arrived at the second floor through clicking or other operations on the second electronic device, or in response to the user confirming that they have arrived at the second floor through voice control commands to the second electronic device, or it can be executed under the trigger of other user operations; there is no limitation.

[0054] S402. Based on the queried image, obtain the first data and the second data corresponding to the first location point.

[0055] The first data is the data corresponding to the first location point in the data set on the first floor, and the second data is the data corresponding to the first location point in the data set on the second floor. The first floor is the departure floor of the second electronic device, and the second floor is the target arrival floor of the second electronic device.

[0056] In one implementation, step S402 may include: determining data within a first preset distance range from the first location point in the data set of the first floor as first data; and determining data within a first preset distance range from the first location point in the data set of the second floor as second data. Selecting only data within a certain range can effectively reduce the amount of data needed to determine location information, thereby reducing computational load and shortening latency.

[0057] In one example, if the first data and / or second data determined based on a first preset distance range are empty, the method further includes: determining the data in the data set of the first floor that is within a second preset distance range from the first location point as the first data, where the second preset distance range is greater than the first preset distance range; and determining the data in the data set of the second floor that is within a second preset distance range from the first location point as the second data. In this example, expanding the data range when no data can be obtained based on the first preset distance range, and re-obtaining the first and second data based on the second preset distance range, can further ensure the amount of data obtained. This ensures that the amount of computation is reduced and the latency is shortened, while also avoiding the inability to match suitable data due to an excessively small query range. This method of using a larger distance range to match and determine the first and second data again when no data can be matched within the first distance range can be called an adaptive expansion of the search range.

[0058] In another example, if the first and second data determined based on the second preset distance range are empty, the above method also includes prompting the user to re-acquire the query image. In this example, if matching is performed based on a large range (the second preset distance range) and no suitable data is found, especially if there are no matching data for either floor, it indicates that the collected query image is likely faulty. Therefore, prompting the user to re-acquire the query image is appropriate. It should be understood that in this example, the main consideration is that if the data for the two floors differs significantly, it is easy to only match data for one floor, resulting in only the first or second data being non-empty. In such cases, it is not necessary to re-acquire the query image; only the non-empty data set needs to be used for subsequent calculations. Therefore, this example only prompts the user to re-acquire the query image when both sets of data are empty.

[0059] It should also be understood that in one example of the above implementation, two queries are performed progressively based on the first and second preset distance ranges. The first query is performed if any set of data is empty (either the first or second set of data is empty), then a second query is performed based on the second preset distance range. The second query only prompts for retrieving the image if both sets of data are empty. This execution process ensures that an appropriate amount of data is obtained while preventing the need to execute subsequent steps when data is empty.

[0060] In one implementation, when the first data and / or the second data determined based on the first preset distance range are empty, the above method further includes: determining, as the first data, the data within a second preset distance range from the first position point in the data set of the first floor, where the third preset distance range is s1 times the first preset distance range, and s1 is a real number greater than 1; determining, as the second data, the data within a fourth preset distance range from the first position point in the data set of the second floor, where the fourth preset distance range is s2 times the first preset distance range, and s2 is a real number greater than 1. In this implementation, by setting expansion coefficients, the distance range for re - retrieval is increased by a multiple, and the two floors respectively correspond to expansion coefficients s1 and s2.

[0061] In an example, s1 < s2. In this example, since the target floor the user is going to is the second floor, when setting the coefficients, it can be biased towards setting a relatively larger coefficient for the second floor, which can further improve the effect of the solution of this application.

[0062] S403. Determine the target position information of the first position point according to the data with a higher score in the first data and the second data.

[0063] In one implementation, step S403 may include: when the score of the first data is higher, determining the first predicted position information of the first position point according to the first data; when the confidence of the first predicted position information is greater than or equal to the first confidence threshold, determining the first predicted position information as the target position information; or when the confidence of the first predicted position information is less than the first confidence threshold, determining the target position information according to the second data. In this implementation, when the score of the first data is higher, it indicates that the second electronic device is very likely still on the first floor or very close to the first floor. Under this premise, a confidence threshold is set, and only the predicted position information exceeding the first confidence threshold will be output as the target position information. Otherwise, the second data with a lower score will be used to determine the target position information, which can ensure the accuracy of the target position information.

[0064] It should be understood that in the solution of this application, there is no limitation on the calculation method of the confidence. For example, a decision tree, a regression tree, or a deep neural network, etc. can be used, as long as the confidence of the position information can be obtained, and there are no other limitations.

[0065] In one example, when the confidence level of the first predicted location information is less than the first confidence threshold, determining the target location information based on the second data can include: determining the second predicted location information of the first location point based on the second data; determining the second predicted location information as the target location information when the confidence level of the second predicted location information is greater than or equal to the second confidence threshold, where the second confidence threshold is greater than the first confidence threshold; or, prompting to re-acquire the query image when the confidence level of the second predicted location information is less than the second confidence threshold. In this example, when the target location information is determined using the second data with a lower score, a higher confidence threshold is used to constrain the output of the location information. If the confidence level is less than the second confidence threshold, prompting to re-acquire the query image avoids the situation of "choosing a slightly better result from two very bad results," where both sets of data actually have low scores, and neither set of data will yield accurate location information. Using the lower-scoring data will, in principle, only yield location information with even lower accuracy. If, under such circumstances, the output target location information still exceeds the higher second confidence threshold, it means that the output target location information is indeed accurate enough. This can be understood as a penalty for data with poor ratings, setting a higher confidence threshold for them. In particular, if the second data should have a higher rating than the first data, but instead has a lower rating, it means that the user may have mistakenly confirmed that they have reached the second floor.

[0066] In one example, when determining the location information of a first location point based on the data with the higher score between the first and second data, the process may include: if the score of the second data is higher, determining the third predicted location information of the first location point based on the second data; if the confidence level of the third predicted location information is greater than or equal to a third confidence threshold, determining the third predicted location information as the target location information, where the third confidence threshold is less than a first confidence threshold; or, if the confidence level of the third predicted location information is less than the third confidence threshold, determining the target location information based on the first data. In this example, the higher score of the second data point indicates a higher probability that the first location point is on the second floor. The first location point is the position where the user confirms they have arrived on the second floor, and in principle, the first location point should be a location point on the second floor. Therefore, the higher score of the second data point better aligns with this principle. Thus, the third confidence threshold is set lower than the first confidence threshold. Combining this with the fact that the second confidence threshold is greater than the first confidence threshold, that is, the second confidence threshold (used when calculating location information using the second data point with a lower score) > the first confidence threshold (used when calculating location information using the first data point with a higher score) > the third confidence threshold (used when calculating location information using the second data point with a higher score).

[0067] In one example, when the confidence level of the third predicted location information is less than the third confidence threshold, determining the target location information based on the first data may include: determining the third predicted location information of the first location point based on the first data; when the confidence level of the third predicted location information is greater than or equal to the fourth confidence threshold, determining the third predicted location information as the target location information, where the fourth confidence threshold is greater than the third confidence threshold; or, when the confidence level of the third predicted location information is less than the fourth confidence threshold, prompting a re-acquisition of the query image. In this example, when calculating location information using the second data with a higher score, if the confidence level is too low, the location information is calculated using the first data with a lower score, and a higher fourth confidence threshold is used. This is similar to the penalty for poorly scored data mentioned above, to avoid the situation of "choosing a slightly better result from two very poor results," and will not be elaborated further.

[0068] The above implementation mainly involves setting different confidence thresholds for different rating scenarios. A higher confidence threshold is set for floors with lower ratings. This avoids the situation of "choosing a slightly better result from two very bad results." In other words, if both sets of data have low ratings, using either set will not yield good results. Using the lower-rated data should, in principle, only yield worse results. If, under these circumstances, a higher confidence threshold is still achieved, the system's judgment error can be corrected, resulting in a relatively accurate result. For example, suppose the total score is 100 points, with the first data point at 16 and the second at 10. Objectively speaking, using either set of data to determine location information will not yield accurate results, as the confidence level is relatively low. In this case, setting a higher confidence threshold for the lower-rated second data point can prevent the output of location information with such low confidence that it does not meet the requirements. It should be understood that the specific values ​​mentioned above are only for understanding the scheme and are not intended to be limiting. In the above implementation, for data with higher scores, different confidence thresholds are set based on whether the data is on the first or second floor. When the first data has a higher score (meaning that the first location point is closer to the starting floor, or may even still be on the starting floor), it means that the user is likely to have mistakenly confirmed "having reached the new floor" or that the user has given up on crossing floors. This is because the first location point is the position the user is in when confirming that they have reached the new floor. If they have truly reached the new floor, the second data should have a higher score, not the first data. Therefore, under this premise, the first confidence threshold is set relatively high, but not higher than the confidence threshold corresponding to the data with lower scores. When the second data has a higher score (meaning that the first location point is very likely already on the target floor), it means that the user is very likely to have correctly confirmed "having reached the new floor". Therefore, under this premise, the third confidence threshold is set to be the lowest.

[0069] For scoring, the first and second data can be evaluated based on the same indicators to obtain evaluation scores, and then the scores of the first and second data can be determined based on a certain strategy.

[0070] The evaluation metrics that can be used include image similarity and Euclidean distance. For example, assuming we use similarity as the evaluation metric, and the first dataset includes N1 images and the second dataset includes N2 images, after matching (retrieval) the query image, we retrieve N1 images with high similarity to the query image from the first dataset and N2 images with high similarity to the query image from the second dataset. Then, we can determine the score for the first dataset based on the similarity scores of these N1 images, and the score for the second dataset based on the similarity scores of the N2 images. For example, the highest score among N1 and N2 scores can be taken separately; or, the average of N1 and N2 scores can be taken; or, the N1 images can be classified first, and the average similarity score under each category can be taken, followed by a weighted sum of the averages of several categories to obtain the score for the first data; similarly, the N2 images can be classified, and the average similarity score under each category can be taken, followed by a weighted sum of the averages of several categories to obtain the score for the second data. Other calculable scoring methods are not listed here, such as using support vector machines. In other words, this application does not limit the method used to calculate the scores for the first and second data, as long as a suitable method is used and the same calculation method is applied to both the first and second data. N1 and N2 are non-negative integers that are not entirely zero.

[0071] To illustrate the Euclidean distance metric, suppose the first set of data includes N1 images, and the second set includes N2 images. After matching (retrieval) the query images, N1 images from the first set of data whose Euclidean distance to the query images falls within a preset range are retrieved, and N2 images from the second set of data whose Euclidean distance to the query images falls within the same preset range are retrieved. Then, the score for the first set of data can be determined based on the Euclidean distance scores of these N1 images, and the score for the second set of data can be determined based on the Euclidean distance scores of the N2 images. For example, the highest score among N1 scores and the highest score among N2 scores can be taken separately; or, the average of N1 scores and the average of N2 scores can be taken; or, the N1 images can be classified first, and then the maximum Euclidean distance score under each category can be taken, followed by a weighted sum of the maximum values ​​from several categories to obtain the score for the first data; similarly, the N2 images can be classified, and then the maximum Euclidean distance score under each category can be taken, followed by a weighted sum of the maximum values ​​from several categories to obtain the score for the second data, etc. Other computable scoring methods are not listed here, such as using support vector machines, etc. In other words, this application does not limit the method used to calculate the scores for the first and second data, as long as a suitable method is used and the same calculation method is applied to both the first and second data. N1 and N2 are non-negative integers that are not entirely zero.

[0072] In one implementation, the method further includes: determining a first similarity array and a second similarity array, wherein the first similarity array includes the similarity between each frame of at least one frame in the first data and the query image, and the second similarity array includes the similarity between each frame of at least one frame in the second data and the query image; and determining the score of the first data and the score of the second data based on the first similarity array and the second similarity array, respectively. In this implementation, the main approach is to first obtain two sets of similarity arrays corresponding to the first data and the second data, and then determine the score based on each of the two sets of similarity arrays.

[0073] In one example, when determining the score of the first data and the score of the second data based on the first similarity array and the second similarity array, it may include: determining the score of the first data and the score of the second data by taking the maximum value in the first similarity array and the maximum value in the second similarity array, respectively; or, determining the score of the first data and the score of the second data by taking the average value in the first similarity array and the average value in the second similarity array, respectively. In this example, the score is determined by removing the maximum value.

[0074] In another example, at least one frame in the first dataset has a similarity greater than a first similarity threshold with the query image; at least one frame in the second dataset also has a similarity greater than the first similarity threshold with the query image. This example provides a filtering (selection) method (basis): only images with similarity thresholds within a certain range are included in either the first or second dataset. This eliminates the need to use the entire dataset from both the first and second datasets for subsequent calculations, effectively reducing computational load and improving processing efficiency and accuracy. However, it should be understood that this selection method may result in an empty set, meaning the first and / or second datasets may be empty, indicating that no data with a similarity less than the first similarity threshold can be found in either the first or second dataset. In such cases, the adaptive expansion of the search scope described above can be adopted, using a second similarity threshold that is less than the first similarity threshold. The above steps can be repeated, or a prompt to retrieve the query image again can be provided.

[0075] In another example, the method further includes prompting the user to retrieve the query image if the first layer of the dataset does not contain images with a similarity greater than a first similarity threshold, or if the second layer of the dataset does not contain images with a similarity greater than the first similarity threshold. In this example, the user is prompted to retrieve the query image whenever the similarity threshold is not met.

[0076] In another example, when the first data and / or second data determined based on the first similarity threshold are empty, the first data and second data are determined based on the second similarity threshold, such that the similarity between each frame in at least one frame of the first data and the query image is greater than the second similarity threshold; and the similarity between each frame in at least one frame of the second data and the query image is greater than the second similarity threshold, where the second similarity threshold is less than the first similarity threshold. It should be understood that this solution is equivalent to reducing the value of the first similarity threshold.

[0077] In one implementation, the first data includes the N images from the first floor's dataset that have the highest similarity to the query image, where N is a positive integer; the second data includes the N images from the second floor's dataset that have the highest similarity to the query image. This implementation retrieves the N most similar images from the entire floor's dataset, meaning N is almost never empty, as it would only be empty if there were no images in the entire floor's dataset, a situation that is highly improbable. However, this approach also results in the N images having very low actual similarity scores, only ranking high. In such cases, it's more appropriate to suggest retrieving the search images.

[0078] In one example, the similarity threshold-based method and the similarity ranking-based method can be combined to select images that simultaneously meet both similarity criteria and similarity ranking. This approach combines the advantages of both methods to select more accurate images.

[0079] It should also be understood that the range of the first and second data points is determined here based on similarity. The text above also provides methods for determining the range of the first and second data points based on a first preset distance range and a second preset distance range. These are different schemes, and they can be combined. Distance-based filtering is simpler and more convenient, suitable for situations with a large amount of data in the entire floor image. Similarity-based filtering is more accurate but computationally more demanding, suitable for situations with a smaller amount of data in the entire floor image. Combining the two allows for initial distance-based filtering followed by further similarity-based filtering.

[0080] Figure 4 The method described primarily utilizes the query image to match two sets of relevant data for each floor: the departure floor and the target arrival floor. The method then uses the data with the higher score from the two sets of matched data for localization. Because localization is based on data with a higher degree of matching and better quality than the query image, it effectively improves visual localization accuracy. Furthermore, since only one set of relevant data is used for localization at a time, rather than using all data from a single floor, or even all data from both floors, the localization speed is relatively fast, effectively reducing latency.

[0081] Figure 5 This is a schematic flowchart illustrating another visual positioning method according to an embodiment of this application. Figure 5 It can be Figure 4 An example of the method shown.

[0082] S501, The visual positioning device determines that the device to be positioned (worn or carried by the user) has arrived at the stairwell.

[0083] A visual positioning device is one example of a first electronic device, and the device to be positioned is one example of a second electronic device. When the device to be positioned is a vehicle-mounted terminal, the user moves synchronously with the device.

[0084] S502, The visual positioning device determines whether the user wants to move from floor A to floor B.

[0085] Floor A and floor B are examples of the departure floor and the destination floor, respectively, which are examples of the first floor and the second floor.

[0086] It should be understood that determining whether a user wants to cross floors actually means determining whether a second electronic device wants to cross floors.

[0087] S503, The visual positioning device obtains the location of the stairwell where the user begins to cross floors from the device to be positioned.

[0088] The location of a stairwell can be represented, for example, by a point of interest (POI). It should be understood that in this application, visual positioning is a more refined form of positioning, and latitude and longitude information is insufficient to represent different locations. For example, the latitude and longitude information of the same stairwell on different floors is the same; therefore, it is necessary to use POIs to represent the location information of different locations.

[0089] S504. In response to the user's confirmation of arrival at floor B, the device to be located acquires an image of the current location and sends the image as a query image to the visual positioning device.

[0090] This step is an example of S401.

[0091] S505. The visual positioning device obtains data within a certain range in floor A and data within the same range in floor B based on the query image.

[0092] The certain range in this step is an example of the first preset distance range and the second preset distance range mentioned above. The data obtained are the first data and the second data. The relevant content can be referred to above and will not be repeated here.

[0093] This step is an example of S402.

[0094] S506. The visual positioning device compares the retrieval results of the first data and the second data, selects the data with relatively better retrieval results, and saves the data with relatively poor retrieval results.

[0095] The search results here are the ratings mentioned above. Related content will not be repeated here, as described above.

[0096] S507. The visual positioning device uses the data with relatively good search results to perform data matching and pose calculation to obtain location information and pose information.

[0097] This application primarily addresses location information.

[0098] S508. When the first data is the better search result, the visual positioning device selects a confidence threshold C1, and when the second data is the better search result, it selects a confidence threshold C2, where C1 > C2.

[0099] C1 could be an example of the first confidence threshold, and C2 could be an example of the third confidence threshold.

[0100] S509. When the confidence level of the location information calculated in step S507 is less than the confidence level threshold set in step S508, the visual positioning device switches to step S511; otherwise, it executes step S510.

[0101] S510, the visual positioning device outputs location information.

[0102] S511. The visual positioning device uses the data with poor search results to perform data matching and pose calculation to obtain location information and pose information.

[0103] S512. When the confidence level of the location information calculated in step S511 is less than the confidence level threshold C3, the visual positioning device proceeds to step S513; otherwise, it proceeds to step S510.

[0104] C3>C1.

[0105] C3 could be an example of a second confidence threshold or a fourth confidence threshold.

[0106] S513, The visual positioning device sends a first instruction, which instructs the device to be positioned to reacquire the query image.

[0107] S514. The device to be located receives the first instruction and prompts the user to re-acquire images through screen display and / or voice prompts.

[0108] Steps S506-S511 are an example of step S403.

[0109] The methods of the embodiments of this application have been described above with reference to the accompanying drawings. It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially, these steps are not necessarily executed in the order shown in the figures. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the steps or stages of other steps. The apparatus of the embodiments of this application will now be described with reference to the accompanying drawings.

[0110] This application also provides a visual positioning device, which includes an acquisition unit and a processing unit. This device can be integrated into a visual positioning device capable of performing complex visual positioning calculations, such as a host, server, or cloud server. This device can be used to execute any of the visual positioning methods described above. For example, the acquisition unit can be used to execute step S401, and the processing unit can be used to execute steps S402-S403. This device can also be used to execute... Figure 5 Each step. In one implementation, the device may further include a storage unit for storing relevant data. This storage unit may be integrated into any of the aforementioned units, or it may be a unit independent of all the aforementioned units.

[0111] Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. It should be understood that in this application, the visual positioning device mainly acquires images from the device to be positioned and then performs subsequent steps. Therefore, the device to be positioned is actually the one acquiring the images, while the visual positioning device only needs to receive the images and does not necessarily need to have a camera function. Figure 6 As shown, the electronic device 1000 includes: at least one processor 1001 ( Figure 6 (Only one is shown) a processor, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable on the at least one processor 1001, wherein the processor 1001 executes the computer program 1003 to implement the steps of any of the methods described above. That is, the electronic device 1000 can be either the training device or the inference device described above.

[0112] Those skilled in the art will understand that Figure 6 This is merely an example of an electronic device and does not constitute a limitation on electronic devices. In practice, electronic devices may include more or fewer components than those shown in the illustration, or combinations of certain components, or different components. For example, they may also include input / output modules, network access modules, etc.

[0113] Processor 1001 may include one or more processing units, such as: a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network process unit (NPU), other general-purpose processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. Different processing units may be independent devices or integrated into one or more processors.

[0114] The controller can be the nerve center and command center of the electronic device 1000. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0115] In some embodiments, memory 1002 may be an internal storage unit of electronic device 1000, such as a hard disk or memory of electronic device 1000. In other embodiments, memory 1002 may be an external storage device of electronic device 1000, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on electronic device 1000. Optionally, memory 1002 may include both internal and external storage units of electronic device 1000. Memory 1002 is used to store operating system, application programs, bootloaders, data, and other programs, such as program code of computer programs. Memory 1002 may also be used to temporarily store data that has been output or will be output.

[0116] The processor 1001 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 1001 is a cache memory. This memory can store instructions or data that the processor 1001 has just used or that are used repeatedly. If the processor 1001 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 1001, and thus improves the efficiency of the system.

[0117] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0118] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0119] This application also provides an electronic device, which includes: one or more processors and a memory; the memory is coupled to one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to enable the electronic device to perform the steps in any of the above methods.

[0120] This application also provides a chip system applied to an electronic device. The chip system includes one or more processors, which invoke computer instructions to cause the electronic device to perform the steps in any of the methods described above. Optionally, the chip system further includes a memory electrically connected to the processor. Optionally, the chip system may also include a communication interface.

[0121] This application also provides a computer-readable storage medium storing instructions that, when executed by an electronic device, can implement any of the methods described above. This computer-readable medium may include at least: any entity or device capable of carrying computer program code (instructions) to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0122] This application also provides a computer program product, which includes a computer program that, when executed by an electronic device, can implement any of the above-described methods. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form.

[0123] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0124] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0125] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0128] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0129] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0130] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0131] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A visual positioning method, applied to a first electronic device, characterized in that, include: Acquire a query image of the first location point where the second electronic device is located. The query image is an image captured by the second electronic device at the first location point in response to a confirmation operation that it has reached the second floor. The second floor is the target floor reached by the second electronic device. Based on the query image, first data and second data corresponding to the first location point are obtained. The first data is the data corresponding to the first location point in the data set of the first floor, and the second data is the data corresponding to the first location point in the data set of the second floor. The first floor is the departure floor of the second electronic device. Based on the data with higher scores in the first data and the second data, the target location information of the first location point is determined.

2. The method according to claim 1, characterized in that, The step of determining the location information of the first location point based on the data with higher scores in the first data and the second data includes: If the score of the first data is high, the first predicted location information of the first location point is determined based on the first data. If the confidence level of the first predicted location information is greater than or equal to the first confidence threshold, the first predicted location information is determined as the target location information; or... If the confidence level of the first predicted location information is less than the first confidence threshold, the target location information is determined based on the second data.

3. The method according to claim 2, characterized in that, When the confidence level of the first predicted location information is less than the first confidence threshold, determining the target location information based on the second data includes: Based on the second data, determine the second predicted location information of the first location point; When the confidence level of the second predicted location information is greater than or equal to the second confidence threshold, the second predicted location information is determined as the target location information, where the second confidence threshold is greater than the first confidence threshold; or... If the confidence level of the second predicted location information is less than the second confidence threshold, a prompt will be made to re-acquire the query image.

4. The method according to claim 2, characterized in that, The step of determining the location information of the first location point based on the data with higher scores in the first data and the second data includes: If the score of the second data is high, the third predicted location information of the first location point is determined based on the second data; If the confidence level of the third predicted location information is greater than or equal to a third confidence threshold, the third predicted location information is determined as the target location information, where the third confidence threshold is less than the first confidence threshold; or... If the confidence level of the third predicted location information is less than the third confidence level threshold, the target location information is determined based on the first data.

5. The method according to claim 4, characterized in that, When the confidence level of the third predicted location information is less than the third confidence threshold, determining the target location information based on the first data includes: Based on the first data, determine the third predicted location information of the first location point; When the confidence level of the third predicted location information is greater than or equal to the fourth confidence threshold, the third predicted location information is determined as the target location information, wherein the fourth confidence threshold is greater than the third confidence threshold; or, When the confidence level of the third predicted location information is less than the fourth confidence threshold, a prompt is made to re-acquire the query image.

6. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the first data and second data corresponding to the first location point based on the query image includes: The data within a first preset distance from the first location point in the data set of the first floor are determined as the first data; The data within the first preset distance range from the first location point in the data set of the second floor are determined as the second data.

7. The method according to claim 6, characterized in that, If the first data and / or the second data determined based on the first preset distance range are empty, the method further includes: The data within a second preset distance range from the first location point in the data set of the first floor are determined as the first data, where the second preset distance range is greater than the first preset distance range; The data within the second preset distance range from the first location point in the data set of the second floor are determined as the second data.

8. The method according to claim 7, characterized in that, If the first data and the second data determined based on the second preset distance range are empty, the method further includes: The system prompts you to retrieve the query image again.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: A first similarity array and a second similarity array are determined. The first similarity array includes the similarity between each frame of at least one frame of the first data and the query image. The second similarity array includes the similarity between each frame of at least one frame of the second data and the query image. Based on the first similarity array and the second similarity array, the scores of the first data and the second data are determined respectively.

10. The method according to claim 9, characterized in that, The step of determining the score of the first data and the score of the second data based on the first similarity array and the second similarity array respectively includes: The maximum value in the first similarity array and the maximum value in the second similarity array are respectively determined as the score of the first data and the score of the second data; or, The average value of the first similarity array and the average value of the second similarity array are respectively determined as the score of the first data and the score of the second data.

11. The method according to claim 9 or 10, characterized in that, In the first data, the similarity between each frame of at least one image and the query image is greater than a first similarity threshold; in the second data, the similarity between each frame of at least one image and the query image is greater than the first similarity threshold.

12. The method according to claim 11, characterized in that, The method further includes: If the data set of the first floor does not contain images with a similarity greater than the first similarity threshold, or if the data set of the second floor does not contain images with a similarity greater than the first similarity threshold, a prompt will be made to retrieve the query image again.

13. The method according to any one of claims 1 to 9, characterized in that, The first data includes the N frames of images in the data set of the first floor that have the highest similarity to the query image, where N is a positive integer; the second data includes the N frames of images in the data set of the second floor that have the highest similarity to the query image.

14. An electronic device, characterized in that, The electronic device includes: one or more processors, and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 13.

15. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 13.