Method and apparatus for device positioning, electronic device, and storage medium
By combining LiDAR data and image data, and utilizing feature semantic maps and multimodal large models, the problem of low device positioning success rate was solved, achieving more efficient device positioning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies have low device positioning success rates, especially when autonomous moving devices lose their location in indoor or outdoor environments, making it impossible to accurately pinpoint their own location and plan subsequent movement routes.
By obtaining query information about the range of the target device, and combining it with a pre-built feature semantic map and a multimodal large model, the target device is located using LiDAR data and image data, thus determining the first range information and the second location information of the target device.
It improves the success rate of device location after location loss, especially when combining image and text information, it can more accurately determine the device's location.
Smart Images

Figure CN119273759B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of positioning technology, and in particular to a method, apparatus, electronic device, and storage medium for device positioning. Background Technology
[0002] With the development of electronic information technology, various devices capable of autonomously moving in specific environments (indoor or outdoor) to perform specific tasks have emerged in daily life, such as robotic vacuum cleaners and food delivery robots. When these devices move in specific environments, they sometimes experience location loss. When location loss occurs, the device cannot determine its position in the environment, making it unable to plan its subsequent movement route based on its current location.
[0003] In related technologies, after a location loss occurs, a device can use its own camera to capture images of its surrounding environment and match these images with a pre-built map of the environment to re-determine its location. This method relies solely on image matching; however, the images used to build the map have a single perspective, and the data used for matching is limited to this single modality, resulting in a low matching success rate and sometimes failing to determine the device's location. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus, electronic device, and storage medium for device positioning to solve the problem of low positioning success rate in the prior art.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] The first aspect of this application provides a method for device positioning, comprising:
[0007] Obtain query information related to the location of the target device;
[0008] Based on the matching results of the query information and the pre-constructed feature semantic map, the first range information of the target device is determined, wherein the first range information represents the range in which the target device is located;
[0009] The target sensor data of the target device is processed according to the map data within the first range information to obtain the second location information of the target device. The second location information represents the location of the target device within the range indicated by the first range information. The target sensor data includes at least one of lidar data and image data.
[0010] Optionally, the feature semantic map includes multiple coordinate data and visual features corresponding to the coordinate data. The coordinate data represents the location of an object in the environment where the target device is located, and the visual features corresponding to the coordinate data are the visual features of the corresponding object.
[0011] Optionally, the method for constructing the feature semantic map includes:
[0012] Obtain multiple coordinate data of the environment in which the target device is located, as well as multiple key images of the environment captured by the device.
[0013] The key image is processed according to a multimodal large model to obtain the pixel features corresponding to each pixel in the key image;
[0014] For each coordinate data, the pixel features of each pixel corresponding to the coordinate data are fused according to the mapping relationship between the pixels contained in the key image and the coordinate data to obtain the visual features corresponding to the coordinate data.
[0015] Optionally, the step of processing the key image according to the multimodal large model to obtain the pixel features corresponding to each pixel in the key image includes:
[0016] The global and local features of the key images are extracted based on the multimodal large model.
[0017] The global features of the image and the local features of the pixels in the key image are fused to obtain the pixel features corresponding to the pixels in the key image.
[0018] Optionally, fusing the pixel features of each pixel corresponding to the coordinate data to obtain the visual features corresponding to the coordinate data includes:
[0019] From the multiple pixel features, select the key region pixel features corresponding to the pixel points located in the key regions of the image;
[0020] Filter out pixel features related to the occluded object from the pixel features of the key region to obtain the filtered pixel features;
[0021] The filtered pixel features corresponding to the coordinate data are fused to obtain the visual features corresponding to the coordinate data.
[0022] Optionally, the query information includes at least one of image information and text information, wherein the image information is obtained by the target device capturing objects within its range, and the text information is obtained by the target device's human-computer interaction module.
[0023] Optionally, determining the first range information of the target device based on the matching results of the query information and the pre-built feature semantic map includes:
[0024] When the query information includes the image information, global image features and key object features of the image information are extracted based on a multimodal large model;
[0025] The global image features and the key object features are fused to obtain the query image features;
[0026] Based on the matching results of the query image features and the feature semantic map, the first range information of the target device is determined.
[0027] A second aspect of this application provides a device for positioning equipment, comprising:
[0028] The acquisition unit is used to obtain query information related to the range of the target device;
[0029] The determining unit is configured to determine the first range information of the target device based on the matching results of the query information and the pre-constructed feature semantic map, wherein the first range information represents the range in which the target device is located;
[0030] The processing unit is configured to process the target sensor data of the target device based on the map data within the first range information to obtain the second location information of the target device. The second location information represents the location of the target device within the range indicated by the first range information. The target sensor data includes at least one of lidar data and image data.
[0031] A third aspect of this application provides an electronic device, comprising:
[0032] One or more processors;
[0033] A storage device on which one or more programs are stored;
[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement the device positioning method provided by any of the first aspects of this application.
[0035] The electronic device in the third aspect can be a target device. The target device can obtain query information and report the query information to the server used to run the multimodal large model. The server obtains the first range information based on the query information and sends the first range information to the target device. After obtaining the first range information, the target device determines the second location information based on the map data and target sensor data within the first range information.
[0036] A fourth aspect of this application provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the device positioning method provided in any one of the first aspects of this application.
[0037] The beneficial effects of this plan are as follows:
[0038] First, the first range information of the target device is determined based on query information and a feature semantic map related to the target device's location. Then, based on the first range information, precise positioning is performed using map data within the corresponding range, as well as the target device's LiDAR data and image data, to obtain the target device's second location information within the corresponding range. Through this method, this solution can improve the success rate of positioning by combining multimodal data, namely the target device's LiDAR data and image data, after the target device's positioning has been lost. Furthermore, the query information can include image information and text information; in this case, the solution can further improve the success rate of repositioning by combining multimodal query information. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0040] Figure 1 This is a flowchart of a device positioning method disclosed in an embodiment of this application;
[0041] Figure 2 This is a schematic diagram of a first range information and a second position information disclosed in an embodiment of this application;
[0042] Figure 3 This is a flowchart of a method for constructing a feature semantic map disclosed in an embodiment of this application;
[0043] Figure 4 This is a schematic diagram of the structure of a device for positioning disclosed in an embodiment of this application;
[0044] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0046] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0047] Furthermore, in this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0048] This application provides a method for device positioning. Please refer to [link to relevant documentation]. Figure 1 Here is a flowchart of the method, which may include the following steps.
[0049] S101, obtain query information related to the range of the target device.
[0050] S102, based on the matching results of the query information and the pre-built feature semantic map, determine the first range information of the target device, which represents the range of the target device.
[0051] S103, the target sensor data of the target device is processed according to the map data within the first range information to obtain the second location information of the target device. The second location information represents the location of the target device within the range indicated by the first range information. The target sensor data includes at least one of lidar data and image data.
[0052] The execution subject of the above-described device positioning method can be a target device that has lost its location. In this case, after determining that it has lost its location, the target device can execute step S101 to obtain query information. Then, the target device executes S102 to send the query information to the server connected to the target device. After receiving the query information, the server matches the query information with the feature semantic map, determines the first range information based on the matching result, and sends the first range information to the target device. After obtaining the first range information, the target device can directly execute S103 based on the first range information, that is, process the target sensor data locally on the target device to obtain the second location information. Alternatively, after obtaining the first range information, the target device can send the target sensor data to the server to trigger the server to process the target sensor data in the manner of step S103, and then receive the second location information fed back by the server.
[0053] The execution entity of the above-described device positioning method can be a server. In this case, in step S101, the server can receive query information reported by the target device. This query information can be reported to the server after the target device determines that it has lost its location. After receiving the query information, the server executes step S102 based on the query information, matching the query information with a feature semantic map to obtain first range information. Then, the server executes step S103 based on the first range information to determine the second location information of the target device and sends the second location information to the target device. The target sensor data can be sent to the server by the target device along with the query information in S101, or it can be sent by the target device to the server after the server determines the first range information.
[0054] The target device in this embodiment can be any device that can autonomously move along a specific route and perform corresponding tasks within a certain range of environments (such as indoor environments like living rooms and classrooms, and outdoor environments like playgrounds and ball courts). For example, the target device can be a robot vacuum cleaner.
[0055] The target device can determine whether it has lost its location using any existing method; this embodiment does not limit this.
[0056] The first range information can be any information that can characterize the range of the target device. In some embodiments, the first range information may include the location information of an object in the environment as the center and the distance information as the radius. In this case, the range indicated by the first range information can be a circular range with the location of the object as the center and the distance information as the radius. Generally, the accuracy of the first range information can be within 1 meter.
[0057] The second location information can include any information that can accurately pinpoint the location of the target device. For example, the second location information can include the orientation and distance information of the target device relative to an object in the environment. Generally, the accuracy of the second location information can reach the centimeter level or higher.
[0058] As an example, see Figure 2 The first range information may include: a radius of 1 meter, with the center of the circle being the location (Xa, Ya) of object A. Based on this first range information, it can be determined that the target device is currently located within 1 meter of object A. The second location information may include: the target device is located due east of object A, at a distance of 0.3 meters from object A. Based on this second location information, the target device can be accurately located. Figure 2 The location shown is P.
[0059] The aforementioned object A can be any object in the environment where the target device is located. For example, object A can be a desk in a classroom, a roadblock placed on a playground, or a shelf in a supermarket.
[0060] The beneficial effects of this embodiment are as follows:
[0061] First, the first range information of the target device is determined based on query information and a feature semantic map related to the target device's location. Then, based on the first range information, precise positioning is performed using map data within the corresponding range, as well as the target device's LiDAR data and image data, to obtain the target device's second location information within the corresponding range. Through this method, this solution can improve the success rate of positioning by combining multimodal data, namely the target device's LiDAR data and image data, after the target device's positioning has been lost. Furthermore, the query information can include image information and text information; in this case, the solution can further improve the success rate of repositioning by combining multimodal query information.
[0062] The method provided in this embodiment can be executed by a target device. After determining that a location loss has occurred, the target device can execute step S101 to obtain query information (text or image). Then, the target device sends the query information to the server to trigger the server to process the obtained query information using a multimodal large model to obtain first range information. Then, the target device receives the first range information representing the approximate location fed back by the server (equivalent to step S102). Next, the target device obtains more accurate second location information based on the first range information and target sensor data.
[0063] Before applying the method of this embodiment to locate the device, the server can... Figure 3 The method shown constructs a feature semantic map of the environment in which the target device is located.
[0064] A feature semantic map can include point cloud data of the environment in which the target device is located (hereinafter referred to as the environment) and the corresponding visual features. Specifically, the point cloud data can include multiple coordinate data, which are three-dimensional coordinate data, specifically the three-dimensional spatial coordinates of a point on the surface of any object in the environment. Therefore, these coordinate data can characterize the location of objects in the environment in which the target device is located. For example, the point cloud data can include multiple coordinate data, which correspond to multiple points on the surface of a sofa in a living room environment. By combining the coordinate data of multiple points on the sofa surface, the location of the sofa in the living room environment can be determined.
[0065] Each coordinate data point corresponds to a visual feature, which characterizes the appearance of the object at that point. In other words, the visual feature corresponding to the coordinate data is specifically the visual feature of the corresponding object. Referring to the previous example, for multiple coordinate data points on the sofa surface, the visual feature of each coordinate data point reflects the appearance of the corresponding point on the sofa surface.
[0066] Please see Figure 3 Methods for constructing feature semantic maps can include:
[0067] A1, obtains multiple coordinate data of the environment in which the target device is located, as well as multiple key images of the environment captured by the camera;
[0068] A2, process key images based on multimodal large model to obtain pixel features corresponding to each pixel in the key image;
[0069] A3. For each coordinate data point, based on the mapping relationship between the pixels in the key image and the coordinate data, the pixel features of each pixel corresponding to the coordinate data are fused to obtain the visual features corresponding to the coordinate data.
[0070] On the one hand, the server needs to obtain point cloud data of the environment, that is, multiple coordinate data of the environment. The server can obtain point cloud data of the environment through radar scanning or image reconstruction, or it can use both methods at the same time.
[0071] For radar scanning solutions, the server can connect to a scanning device in the environment that is equipped with LiDAR, receive scanning data fed back by the scanning device, and determine the point cloud data of the environment based on the scanning data and the pose of the LiDAR itself.
[0072] For image reconstruction schemes, the server can receive multiple frames of images captured by a scanning device in the environment using a depth camera, as well as the pose of the depth camera when the images were captured. It can then use any existing image reconstruction method, such as the Neural Radiance Field in 3D Vision (Nerf) method or the Simultaneous Localization and Mapping (SLAM) method, to process the pose of the depth camera and the images to obtain point cloud data of the environment.
[0073] The methods for determining point cloud data described above can be found in relevant existing technologies and will not be elaborated further.
[0074] On the other hand, the server needs to obtain multiple key images of the environment, as well as the camera pose corresponding to each key image. The camera pose refers to the position and orientation of the depth camera when capturing the corresponding image frame. If the server obtains point cloud data through LiDAR, it also needs to call the depth camera of the scanning device to capture multiple key images at certain shooting intervals, and then receive the key images and corresponding camera poses reported by the scanning device. If the server obtains point cloud data through image reconstruction, it can select several images as key images from the multiple images provided by the scanning device according to a specific algorithm. The methods for capturing and selecting key images can be found in relevant existing technologies and will not be elaborated further.
[0075] In step A2, after obtaining multiple key images, the server can use a pre-built multimodal large model with image processing and feature extraction capabilities to process each key image to obtain the pixel features corresponding to each pixel in each key image, and then obtain the visual features corresponding to the coordinate data based on the pixel features.
[0076] Optionally, to make the obtained visual features more accurately represent the appearance of the object at the point corresponding to the coordinate data, for each key image frame, the server can obtain the pixel features of that key image frame in the following manner:
[0077] Extract global and local image features of key images based on a multimodal large model;
[0078] By fusing global features of the image with local features of the pixels in the key image, the pixel features corresponding to the pixels in the key image are obtained.
[0079] After a key frame image is processed by a multimodal large model, a global image feature and multiple local image features can be obtained. The global image feature can characterize the number, type, and location of objects displayed in the overall key frame image, while the local image features reflect the appearance attributes of objects displayed in different local areas of the key frame image, such as the color, shape, and material of the objects.
[0080] For any pixel in any frame of key image, the multimodal large model can determine which local region the pixel belongs to in the key image and where it is located in that local region. Then, the multimodal large model can fuse the position of the pixel in the local region, the local image features of the local region where the pixel is located, and the global image features of the key image to obtain the pixel features corresponding to the pixel.
[0081] In step A3, the server also needs to determine the mapping relationship between pixels in the key image and coordinate data in the point cloud data. A mapping relationship between a pixel and coordinate data can be understood as the pixel being obtained by projecting the point corresponding to that coordinate data on the object surface onto the camera's imaging plane. In other words, if a pixel P1 and coordinate data are mapped, it means that P1 is the projection of point P2 on the object surface corresponding to that coordinate data onto the plane.
[0082] In step A3, the server can project the coordinate data in the point cloud data onto the pixels of each key image frame based on the camera pose corresponding to each frame obtained in advance. This determines the mapping relationship between the coordinate data and the pixels, that is, it determines which pixels are projected from which coordinate data. The method for determining the above mapping relationship can be found in relevant prior art, and will not be described in detail here.
[0083] Once the mapping relationship between pixels and coordinate data is established, for each coordinate data point, the server can directly fuse the pixel features of multiple pixels corresponding to that coordinate data, and the result becomes the visual feature corresponding to that coordinate data. The fusion method can be to add multiple pixel features according to certain weights.
[0084] Alternatively, to improve the accuracy of the obtained visual features, the server can first post-process the pixel features, and then fuse the post-processed pixel features of the pixels corresponding to the coordinate data to obtain the visual features corresponding to the coordinate data:
[0085] B1, selects the key region pixel features corresponding to the pixel points located in the key region of the image from multiple pixel features;
[0086] B2 filters out pixel features related to the occluded object from the pixel features of the key region to obtain the filtered pixel features;
[0087] B3 fuses the filtered pixel features corresponding to the coordinate data to obtain the visual features corresponding to the coordinate data.
[0088] The key image region in step B1 can be a rectangular or circular area within a certain distance from the center of the key image. As an example, assuming the key image is a 1080*720 resolution image, the key image region could be a rectangular area of 540*360 pixels located in the center of the image. Filtering out the key region pixel features corresponding to pixels located within the key image region can be done by deleting the pixel features of pixels not located within the key image region, retaining only the pixel features of pixels located within the key image region, such as the region located in the center of the key image mentioned above. These retained pixel features are called key region pixel features.
[0089] The purpose of filtering by pressing B1 is that pixels at the edges of images captured by depth cameras may have severe distortion and cannot accurately reflect the appearance of the corresponding object surface. Deleting the pixel features of these pixels can improve the accuracy of the visual features obtained by subsequent pixel feature fusion.
[0090] In step B2, the server can use relevant existing technologies to identify which objects in the key image are occluded, and determine the occluded objects as occluded objects.
[0091] For example, object A can be identified in key image 1, but in another key image 2, object A is occluded by object B, so the server can determine that object A belongs to the occluded object.
[0092] The pixel features related to the occluded object refer to the pixel features of pixels belonging to the occluded object in the key image. Referring to the previous example, after determining that object A is an occluded object, the pixel features of pixels belonging to object A in key image 1 are the pixel features related to the occluded object A.
[0093] After identifying the occluded object, the server can delete the pixel features of the pixels belonging to the occluded object from all the retained key area pixel features, and use the remaining undeleted pixel features as the filtered pixel features in step B2.
[0094] Based on the previous example, after determining that object A is an occluded object, the server can delete the pixel features of the pixels belonging to object A in key image 1.
[0095] The purpose of filtering out pixel features related to the occluded object is to:
[0096] When visual features reflect the appearance of a distant, occluded object, matching query information with visual features may result in a successful match with a distant object, leading to inaccurate initial range information. For example, the target device may actually be located near object B, but the query information matches the visual features of a distant object A that is occluded by object B. This situation may cause the server to identify the target device as being near object A, resulting in inaccurate initial range information. Deleting pixel features related to the occluded object ensures that the visual features obtained by pixel feature fusion only reflect the appearance of unoccluded objects closer to the target device, thus avoiding the above situation to some extent.
[0097] After post-processing in steps B1 and B2 to obtain the filtered pixel features, the server can fuse the filtered pixel features of the corresponding pixels for each coordinate data to obtain the visual features corresponding to the coordinate data.
[0098] The pixel corresponding to the coordinate data refers to the pixel that has a mapping relationship with the coordinate data. As an example, the server projects each coordinate data in the point cloud data onto each key image frame based on the camera pose corresponding to each key image frame obtained in advance. It then determines that there is a mapping relationship between coordinate data 1 and pixel 1 of key image 1, pixel 2 of key image 2, and pixel 3 of key image 3. After the above post-processing, the pixel features corresponding to pixels 1 to 3 are all retained as filtered pixel features. Therefore, the server can fuse pixel feature 1 corresponding to pixel 1, pixel feature 2 corresponding to pixel 2, and pixel feature 3 corresponding to pixel 3 to obtain visual feature 1 corresponding to coordinate data 1.
[0099] In this embodiment, the query information may include at least one of image information and text information. The image information is obtained by the target device capturing objects within its range, and the text information is obtained by the target device's human-computer interaction module.
[0100] When the query information includes image information, and the target device discovers that it has lost its location and needs to obtain query information, it can stop in place, control its own depth camera to rotate 100 degrees, take multiple frames of images of the surroundings, and upload these images to the server as query information.
[0101] When the query information includes text, and the target device discovers that its location has been lost and needs to obtain query information, it can output a prompt message through its human-computer interaction module. This prompt message prompts the user to input information that represents the target device's current location. Then, the human-computer interaction module collects the user's feedback after outputting the prompt message and obtains the text information as the query information based on the user's feedback. For example, the user's feedback can be voice, such as "Currently located near the second table on the left in the first row." The target device can perform speech recognition on the above voice and obtain the content as the query information to report to the server. Alternatively, the user's feedback can also include text directly entered by the user via keyboard or touchpad. In this case, the target device can directly report the text as the query information to the server.
[0102] Human-computer interaction modules can include various modules that enable human-computer interaction, such as displays, speakers, microphones, buttons, and touchpads.
[0103] After the query information is reported to the server, if the query information includes image information, the server can match the query information with the feature semantic map in the following way to obtain the first range information:
[0104] C1, when the query information includes image information, extract global image features and key object features of the image information based on the multimodal large model;
[0105] C2, which fuses global image features and key object features to obtain query image features;
[0106] C3 determines the first range information of the target device based on the matching results of the query image features and the feature semantic map.
[0107] In step C1, the server can utilize a multimodal large model to extract global image features of the image information. If the image information includes multiple frames, such as the multiple frames captured by the target device rotating its depth camera in the aforementioned example, the server can extract global image features from each frame separately, and then fuse the global image features of the multiple frames to obtain the global image features of the image information.
[0108] On the other hand, the server can use a multimodal large model to identify key objects displayed in the image information, and then extract the image features of these key objects as key object features. Which objects are specifically identified as key objects can be determined by the multimodal large model; generally, the multimodal large model can identify objects that are larger in size and closer to the target device as target objects.
[0109] In step C2, the server can perform a weighted summation of the global image features and the key object features of the image information according to certain weights, and use the result as the fused query image features.
[0110] In step C3, the server can calculate the cosine similarity between the visual features corresponding to each coordinate data in the query image features and the feature semantic map, respectively. The coordinate data whose cosine similarity satisfies the target filtering criteria are determined as target coordinate data. Then, the first range information is determined based on these target coordinate data. The target coordinate data may include one or more coordinate data.
[0111] For example, after calculation, it is found that the cosine similarity between visual feature 1 corresponding to coordinate data 1 and the query image feature satisfies the target filtering condition, and the cosine similarity between visual feature 2 corresponding to coordinate data 2 and the query image feature satisfies the target filtering condition. Therefore, it is determined that the target coordinate data includes coordinate data 1 and coordinate data 2.
[0112] The target filtering criteria can be set as needed. For example, it can be set to the corresponding cosine similarity being greater than or equal to a specific similarity threshold, or it can be set to the corresponding cosine similarity being ranked in the top N positions (sorted from largest to smallest) among all coordinate data, where N is a preset positive integer.
[0113] When multiple target coordinate data are determined, if these target coordinate data all correspond to the same object, for example, coordinate data 1 and coordinate data 2 in the above example both correspond to points on the sofa surface, then the location of the corresponding object can be determined as the center of the first range information, and the range within a preset distance around the center can be determined as the range indicated by the first range information (i.e., Figure 2 (The area within the dashed circle).
[0114] When multiple target coordinate data are determined, if these target coordinate data correspond to the same object respectively, for example, coordinate data 1 corresponds to a point on the surface of the sofa and coordinate data 2 corresponds to a point on the surface of the coffee table in the above example, the midpoint of the multiple target coordinate data can be calculated, and the location of the midpoint can be determined as the center of the first range information. The range within a preset distance around the center can be determined as the range indicated by the first range information.
[0115] If the query information includes text information, the server can obtain the first range of information as follows:
[0116] The server uses a multimodal large model to process text information, obtains query text features, and then calculates the cosine similarity between the visual features corresponding to each coordinate data in the feature semantic map and the query text features. The coordinate data whose cosine similarity meets the target selection criteria are determined as target coordinate data, and the first range information is determined based on the target coordinate data.
[0117] The methods for determining the target coordinate data and the methods for determining the first range information based on the target coordinate data can be found in the aforementioned embodiments and will not be repeated here.
[0118] The advantage of obtaining the first range information in the above manner is that, in related technologies, manual semantic annotation of various objects in the map is required before relocalization using only images, which is costly. In this embodiment, the first range information corresponding to the text information can be directly determined through a multimodal large model, without the need for prior manual semantic annotation, thus reducing the cost of applying the method of this embodiment.
[0119] The multimodal large model used in this embodiment can be an open-source and pre-trained multimodal large model in the relevant technical field, such as various large language models; or it can be a large model obtained by fine-tuning the aforementioned model using relevant image data and text data. Compared with directly training the entire model, this fine-tuning method can effectively reduce the amount of computation and save the consumed computing resources.
[0120] If the query information includes image information and text information, the server can process the image information and text information in the two ways described above respectively, obtain the range determined based on the image information and the range determined based on the text information, and then determine the intersection or union of the two ranges as the first range information.
[0121] In step S103, the second position information can be determined as follows:
[0122] The map data may include sensor data pre-scanned at multiple locations using other scanning devices. For example, within the range corresponding to the first range information, the map data may include sensor data 1 pre-scanned at location 1, sensor data 2 pre-scanned at location 2, and sensor data 3 pre-scanned at location 3. The sensor data scanned at each location may include LiDAR data scanned at the corresponding location using LiDAR, and image data captured at the corresponding location using a depth camera.
[0123] After obtaining the target sensor data, the target sensor data can be compared with the sensor data at each location within the range indicated by the first range information. The similarity between the target sensor data and the sensor data at each location within the range indicated by the first range information is calculated. The locations with similarity greater than a certain threshold, or the locations ranked in the first K positions after sorting the similarity from largest to smallest, are determined as candidate locations, where K is a preset positive integer.
[0124] If only one candidate location is determined, it can be determined that the candidate location is the location of the target device, and the location information representing the candidate location is the second location information of the target device;
[0125] If multiple candidate locations are identified, the midpoint between these locations can be calculated. The location of this midpoint is then determined as the location of the target device. The location information representing the midpoint is the second location information of the target device.
[0126] Based on the previous example, assuming that the similarity between sensor data 1 and target sensor data is greater than a certain threshold, then it can be determined that position 1 of the scanned sensor data 1 is the location of the target device, and the location information corresponding to position 1 is the second location information;
[0127] Assuming that the similarity between sensor data 1 and target sensor data is greater than a certain threshold, and the similarity between sensor data 2 and target sensor data is greater than a certain threshold, then position 1 and position 2 can be determined as candidate positions. The midpoint between position 1 and position 2 is determined as the location of the target device, and the location information representing the location of this midpoint is the second location information of the target device.
[0128] This application also provides a device for positioning equipment; please refer to [link to relevant documentation]. Figure 4 The device may include the following units.
[0129] The obtaining unit 401 is used to obtain query information related to the range of the target device.
[0130] The determining unit 402 is used to determine the first range information of the target device based on the matching results of the query information and the pre-built feature semantic map. The first range information represents the range in which the target device is located.
[0131] The processing unit 403 is used to process the target sensor data of the target device according to the map data within the first range information to obtain the second location information of the target device. The second location information represents the location of the target device within the range indicated by the first range information. The target sensor data includes at least one of lidar data and image data.
[0132] Optionally, the feature semantic map includes multiple coordinate data and corresponding visual features. The coordinate data represents the location of the object in the environment where the target device is located, and the visual features corresponding to the coordinate data are the visual features of the corresponding object.
[0133] Optionally, the device may further include a construction unit 404 for constructing a feature semantic map in the following manner:
[0134] Obtain multiple coordinate data of the environment in which the target device is located, as well as multiple key images of the environment captured by the camera;
[0135] The key image is processed using a multimodal large model to obtain the pixel features corresponding to each pixel in the key image.
[0136] For each coordinate data, the pixel features of each pixel corresponding to the coordinate data are fused according to the mapping relationship between the pixels contained in the key image and the coordinate data to obtain the visual features corresponding to the coordinate data.
[0137] Optionally, when building unit 404 processes the key image based on the multimodal large model to obtain the pixel features corresponding to each pixel in the key image, it can be used for:
[0138] Extract global and local image features of key images based on a multimodal large model;
[0139] By fusing global features of the image with local features of the pixels in the key image, the pixel features corresponding to the pixels in the key image are obtained.
[0140] Optionally, when the building unit 404 fuses the pixel features of each pixel corresponding to the coordinate data to obtain the visual features corresponding to the coordinate data, it can be used for:
[0141] Select the key region pixel features corresponding to the pixel points located in the key regions of the image from multiple pixel features;
[0142] Filter out pixel features related to the occluded object from the pixel features of the key region to obtain the filtered pixel features;
[0143] After filtering the pixel features corresponding to the coordinate data, the visual features corresponding to the coordinate data are fused to obtain the visual features.
[0144] Optionally, the query information includes at least one of image information and text information. The image information is obtained by the target device from the objects within its range, and the text information is obtained by the target device's human-computer interaction module.
[0145] Optionally, when determining the first range information of the target device based on the matching results of the query information and the pre-built feature semantic map, the determining unit 402 can be used for:
[0146] When the query information includes image information, global image features and key object features of the image information are extracted based on a multimodal large model;
[0147] The query image features are obtained by fusing global image features and key object features;
[0148] Based on the matching results of the query image features and the feature semantic map, the first range information of the target device is determined.
[0149] The working principle of the device positioning apparatus provided in this embodiment can be found in the relevant steps of the device positioning method provided in any embodiment of this application, and will not be repeated here.
[0150] Another embodiment of the application also provides an electronic device, such as Figure 5 As shown, it specifically includes:
[0151] One or more processors 501.
[0152] Storage device 502, on which one or more programs are stored.
[0153] When one or more programs are executed by one or more processors 501, the one or more processors 501 implement the device positioning method as described in any of the above embodiments.
[0154] The electronic device in this embodiment can be a target device. The target device can obtain query information and report the query information to a server used to run a multimodal large model. The server obtains first range information based on the query information and sends the first range information to the target device. After obtaining the first range information, the target device determines the second location information based on the map data and target sensor data within the first range information.
[0155] Another embodiment of this application also provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a device positioning method as described in any of the above embodiments.
[0156] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for system or system embodiments, since they are fundamentally similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. Components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0157] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0158] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of device positioning, characterized by, The method comprises: obtaining query information related to the range in which the target device is located; the query information comprises at least one of image information obtained by the target device photographing an object in the range in which the target device is located and text information obtained by a human-computer interaction module of the target device from user input information representing the range in which the target device is located; determining first range information of the target device according to a matching result of the query information and a pre-constructed feature semantic map, the first range information being a circular range with a certain object in the map as a center and distance information as a radius, and being used to represent the range in which the target device is located; the feature semantic map comprises point cloud data of an environment in which the target device is located and visual features corresponding to the point cloud data; wherein the point cloud data comprises a plurality of coordinate data, each coordinate data corresponding to a visual feature; the visual feature represents an appearance of a corresponding object at the coordinate data; processing target sensor data of the target device according to map data in the first range information to obtain second position information of the target device, the second position information representing a position of the target device in the range indicated by the first range information, the target sensor data comprising at least one of lidar data and image data; the map data comprises sensor data pre-scanned at a plurality of positions; the accuracy of the second position information is higher than that of the first range information.
2. The method of claim 1, wherein, The feature semantic map comprises a plurality of coordinate data and visual features corresponding to the coordinate data, the coordinate data representing positions of objects in an environment in which the target device is located, and the visual features corresponding to the coordinate data being visual features of corresponding objects.
3. The method of claim 2, wherein, The method of constructing the feature semantic map comprises: obtaining a plurality of coordinate data of an environment in which the target device is located and a plurality of key images obtained by photographing the environment; processing the key images according to a multi-modal large model to obtain pixel features corresponding to each pixel point in the key images; for each coordinate data, fusing pixel features of each pixel point corresponding to the coordinate data according to a mapping relationship between the pixel points in the key images and the coordinate data to obtain a visual feature corresponding to the coordinate data.
4. The method of claim 3, wherein, The processing of the key images according to the multi-modal large model to obtain pixel features corresponding to each pixel point in the key images comprises: extracting image global features and image local features of the key images according to the multi-modal large model; fusing the image global features and the image local features of the local in which the pixel points are located in the key images to obtain the pixel features corresponding to the pixel points in the key images.
5. The method of claim 3, wherein, The fusing of the pixel features of each pixel point corresponding to the coordinate data to obtain the visual feature corresponding to the coordinate data comprises: selecting key region pixel features in which corresponding pixel points are located in key regions of images from a plurality of the pixel features; filtering out pixel features related to occluded objects from the key region pixel features to obtain filtered pixel features; Fuse the filtered pixel features of the pixel points corresponding to the coordinate data to obtain visual features corresponding to the coordinate data.
6. The method of claim 1, wherein, The first range information of the target device is determined according to a matching result of the query information and a pre-constructed feature semantic map. In a case where the query information includes the image information, global image features and key object features of the image information are extracted according to a multi-modal large model; Fuse the global image features and the key object features to obtain query image features; The first range information of the target device is determined according to a matching result of the query image features and the feature semantic map.
7. An apparatus for positioning a device, characterized by Comprise: An obtaining unit is configured to obtain query information related to a range in which a target device is located; The query information includes at least one of image information and text information, the image information is obtained by the target device shooting an object in the range in which the target device is located, and the text information is obtained by a human-computer interaction module of the target device from user input information representing the range in which the target device is located; A determining unit is configured to determine first range information of the target device according to a matching result of the query information and a pre-constructed feature semantic map, the first range information being a circular range with a certain object in the map as a center and distance information as a radius, and representing the range in which the target device is located; the feature semantic map includes point cloud data of an environment in which the target device is located and visual features corresponding to the point cloud data; the point cloud data includes a plurality of coordinate data, each coordinate data corresponding to a visual feature; and the visual feature represents an appearance of a corresponding object at the coordinate data. A processing unit is configured to process target sensor data of the target device according to map data in the first range information to obtain second position information of the target device, the second position information representing a position of the target device in the range indicated by the first range information, the target sensor data including at least one of lidar data and image data; the map data includes sensor data scanned in advance at a plurality of positions; and the second position information has a higher accuracy than the first range information.
8. An electronic device, comprising: Comprise: One or more processors; Storage having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the device positioning method according to any one of claims 1 to 6.
9. A computer storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the device positioning method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Cooperative localization method based on multi-modal map
CN113932814A
Positioning method and device, robot and computer readable storage medium
CN116295338A