A lightweight feature map construction and storage method for unknown complex scenes
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
然而,此类未知复杂场景具有陆标形态多变、纹理稀疏、结构弱表达、环境干扰强的特点,给高效、鲁棒的环境感知与地图构建带来了严峻挑战
[0010]This invention provides a lightweight feature map construction and storage method for unknown and complex scenes. The method constructs a defined configuration based on preset landmark types and feature attribute sets, filters and extracts feature vectors of real-world landmarks from multimodal perception data, and dynamically calculates feature confidence based on distance. This confidence drives the matching and updating mechanism of the feature map library, ultimately constructing a three-layer lightweight scene representation consisting of "feature vectors - structured feature maps - semantic text". Thus, this invention effectively overcomes the problems of low mapping efficiency, cumbersome maps, and poor adaptability caused by limited storage and computing resources in traditional dense mapping methods when facing large-scale unknown and complex scenes. Through lightweight feature extraction, confidence-guided active learning, and structured semantic representation, it significantly improves the mapping efficiency, storage economy, robustness, and efficiency of scene revisit localization for unknown and complex scenes.
Smart Images

Figure CN122551087A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent autonomous technology, and in particular to a lightweight feature map construction and storage method for unknown and complex scenarios. Background Technology
[0002] As deep space exploration and autonomous field exploration missions expand into more complex and unknown environments, higher demands are placed on the scene understanding and autonomous navigation capabilities of probes. In extreme environments such as Mars and the Moon, where artificial landmarks are scarce, probes must rely on the identification and memorization of natural terrain features (such as impact craters, rocks, and dunes) and mission traces (such as tire tracks) to achieve large-scale scene mapping, path planning, and revisit positioning. However, such unknown and complex scenes are characterized by varied landmark shapes, sparse textures, weak structural representation, and strong environmental interference, posing a severe challenge to efficient and robust environmental perception and map construction.
[0003] Existing technologies face a prominent contradiction when dealing with unknown and complex scenarios: on the one hand, to achieve reliable identification and matching of diverse and weakly characterized landmarks, maps need to contain rich and distinguishable feature information; on the other hand, the stringent resource constraints (computing power, storage space, energy consumption) of the detection platform require maps to be extremely lightweight. Existing methods either prioritize the former, resulting in excessively high resource requirements, or prioritize the latter, sacrificing representational ability and practicality. Consequently, it is difficult to construct lightweight maps with high distinguishability and strong matchability for large-scale unknown and complex scenarios with limited resources.
[0004] Therefore, there is an urgent need to provide a lightweight feature map construction and storage method for unknown and complex scenarios. Summary of the Invention
[0005] This invention provides a lightweight feature map construction and storage method for unknown and complex scenes. This method achieves lightweight map construction with high discriminative power and strong matching ability for large-scale unknown and complex scenes. The technical solution is as follows: On the one hand, this invention provides a lightweight feature map construction and storage method for unknown and complex scenarios, including: Based on the preset land landmark types and their corresponding feature attribute sets, construct the land landmark definition configuration; Based on the landmark definition configuration, real-world landmark regions belonging to the landmark type are selected from the multimodal perception data of the current scene. Feature attributes corresponding to the landmark definition configuration are extracted within the region to generate feature vectors of real-world landmarks. The confidence coefficient of the feature vectors of real-world landmarks is calculated by combining the feature attributes of the real-world landmarks. The feature map database is updated based on the feature vectors and the confidence coefficients. Based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, a structured feature map representing the distribution relationship between landmarks is constructed, and semantic text describing the current scene is generated; using the three-layer representation composed of the feature vectors, the structured feature map, and the semantic text, the feature map of the current scene is constructed and stored.
[0006] On the other hand, a lightweight feature map construction and storage device for unknown and complex scenarios is provided, the device comprising: The construction module is used to build the landmark definition configuration based on the preset landmark types and corresponding feature attribute sets; The calculation module is used to filter out real-scene landmark regions belonging to the landmark type from the multimodal perception data of the current scene based on the landmark definition configuration, extract feature attributes corresponding to the landmark definition configuration in the region, generate feature vectors of real-scene landmarks, and calculate the confidence coefficient of the feature vectors of real-scene landmarks by combining the feature attributes of the real-scene landmarks. The update module is used to update the feature map library based on the feature vector and the confidence coefficient; The generation module is used to construct a structured feature map representing the distribution relationship between landmarks based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, and generate semantic text describing the current scene; using the three-layer representation composed of the feature vectors, the structured feature map and the semantic text, the feature map of the current scene is constructed and stored.
[0007] On the other hand, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in any one of the present specification.
[0008] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in any one of the present specification.
[0009] On the other hand, the present invention provides a computer program product, characterized in that it includes a computer program, which, when executed by a processor, implements the steps of any of the methods described in this specification.
[0010] This invention provides a lightweight feature map construction and storage method for unknown and complex scenes. The method constructs a defined configuration based on preset landmark types and feature attribute sets, filters and extracts feature vectors of real-world landmarks from multimodal perception data, and dynamically calculates feature confidence based on distance. This confidence drives the matching and updating mechanism of the feature map library, ultimately constructing a three-layer lightweight scene representation consisting of "feature vectors - structured feature maps - semantic text". Thus, this invention effectively overcomes the problems of low mapping efficiency, cumbersome maps, and poor adaptability caused by limited storage and computing resources in traditional dense mapping methods when facing large-scale unknown and complex scenes. Through lightweight feature extraction, confidence-guided active learning, and structured semantic representation, it significantly improves the mapping efficiency, storage economy, robustness, and efficiency of scene revisit localization for unknown and complex scenes. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a lightweight feature map construction and storage method for unknown and complex scenarios provided by an embodiment of the present invention; Figure 2 This is a hardware architecture diagram of a computer device provided in an embodiment of the present invention; Figure 3 This is a structural diagram of a lightweight feature map construction and storage device for unknown and complex scenarios provided by an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0014] The specific implementation of this method is described below.
[0015] Please refer to Figure 1 This invention provides a lightweight feature map construction and storage method for unknown and complex scenarios, the method comprising: Step 100: Construct the landmark definition configuration based on the preset landmark types and corresponding feature attribute sets; Step 102: Based on the landmark definition configuration, select real-world landmark areas belonging to the landmark type from the multimodal perception data of the current scene, extract feature attributes corresponding to the landmark definition configuration within the area, generate feature vectors of real-world landmarks, and calculate the confidence coefficient of the feature vectors of real-world landmarks by combining the feature attributes of real-world landmarks. Step 104: Update the feature map database based on the feature vectors and confidence coefficients; Step 106: Based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, construct a structured feature map representing the distribution relationship between landmarks, and generate semantic text describing the current scene; use the three-layer representation composed of feature vectors, structured feature map and semantic text to complete the construction and storage of the feature map of the current scene.
[0016] In this invention, a defined configuration is first constructed based on a preset set of landmark types and feature attributes. Feature vectors of real-world landmarks are then selected and extracted from multimodal perception data. Simultaneously, feature confidence is dynamically calculated based on distance, driving the matching and updating mechanism of the feature map library. Ultimately, a three-layer lightweight scene representation of "feature vectors - structured feature maps - semantic text" is constructed. Thus, this invention effectively overcomes the problems of low mapping efficiency, cumbersome maps, and poor adaptability caused by limited storage and computing resources in traditional dense mapping methods when facing large-scale unknown and complex scenes. Through lightweight feature extraction, confidence-guided active learning, and structured semantic representation, it significantly improves the mapping efficiency, storage economy, robustness, and efficiency of scene revisit localization for unknown and complex scenes.
[0017] The following description Figure 1 The execution method for each step is shown.
[0018] First, regarding step 100: In some implementations, the feature attributes include landform type, spatial coordinates, geometric shape, color features, texture attributes, scale parameters, and first observation time.
[0019] In this invention, the feature attributes contained in the feature vector are carefully designed to achieve efficient and lightweight representation of the multi-dimensional attributes of land landmarks.
[0020] Specifically, the feature attributes include, but are not limited to, the following: spatial coordinates, used to determine the position of the landmark in a global or local reference frame; geometric morphology, shape descriptors such as contour, curvature, and volume extracted from point cloud data; color features, color histograms or dominant color information statistically obtained from image regions; texture attributes, texture descriptors reflecting surface roughness and pattern regularity; and scale parameters, such as physical size measurements like height, area, and volume. These elements together constitute a compact digital profile of the landmark, greatly reducing the storage capacity required by traditional maps while ensuring sufficient distinguishability, and providing a robust and computable data foundation for subsequent feature matching, scene recognition, and revisit localization.
[0021] Regarding step 102: In some implementations, real-world landmark areas belonging to the landmark type are selected from the multimodal perception data of the current scene, including: Semantic segmentation is performed on images in multimodal perception data to obtain semantic labels and region masks. Based on the region masks, the real-world landmark regions corresponding to the landmark types are determined.
[0022] In this invention, the system performs semantic segmentation on the acquired image data. This process, based on a preset landmark definition configuration, outputs a semantic label and region mask for each pixel. The region mask precisely delineates continuous pixel regions in the image belonging to the preset landmark type (such as rocks or impact craters), thereby separating these regions from the complex background and defining them as real-world landmark regions. This step not only achieves preliminary identification and spatial localization of the target landmark, but more importantly, it provides precise spatial guidance for subsequent multimodal feature extraction: the system will extract features directionally from the corresponding spatial locations of other modal data such as point clouds and spectra based on the region mask, ensuring the targeting and accuracy of feature attribute extraction and effectively suppressing interference from irrelevant background information. This is a key prerequisite for achieving lightweight, highly discriminative feature construction.
[0023] In some implementations, the confidence coefficient is determined in the following ways: When the distance between the real-world landmark and the sensing device is less than or equal to the minimum confidence distance threshold, the confidence coefficient saturates to 1, indicating that the corresponding feature vector is valid. When the distance between the real-world landmark and the sensing device is greater than or equal to the maximum effective distance threshold, the confidence coefficient is truncated to 0, indicating that the corresponding feature vector is invalid. When the distance between the real-world landmark and the sensing device is between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient decreases as the distance increases.
[0024] In this invention, when the distance between the real-world landmark and the sensing device is less than or equal to the minimum confidence distance threshold, the confidence coefficient saturates to 1, indicating that the sensor data quality is high and the features are stable and reliable at close range. When the distance is greater than or equal to the maximum effective distance threshold, the confidence coefficient is truncated to 0, indicating that the feature has failed due to insufficient resolution or noise interference, and is therefore directly discarded. When the distance is between the above two thresholds, the confidence coefficient decays exponentially. This mechanism allows the feature confidence to gradually decrease with increasing distance, so that in subsequent map updates and matching stages, the system can prioritize decisions based on high-confidence features, effectively suppressing errors introduced by low-quality data from long distances, and ensuring the accuracy, robustness, and efficiency of the feature map during continuous expansion.
[0025] Specifically, when the distance lies between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient is determined by the following formula: In the formula, CS(d) is the confidence coefficient, d is the distance between the real-world landmark and the sensing device, and d0 is the minimum confidence distance threshold. This is the attenuation coefficient.
[0026] Regarding step 104: In some implementations, the feature map database is updated based on feature vectors and confidence coefficients, including: The similarity between the feature vectors of real-world landmarks and the existing feature vectors of landmarks in the feature map library is calculated. If the similarity is lower than the preset threshold, the real-world landmark is determined to be a newly added landmark. If the confidence coefficient of the real-world landmark is greater than zero, its feature vector and associated semantic information are stored in the feature map library. If the similarity is not lower than the preset threshold, the real-world landmark is determined to be an existing landmark, and the feature vector of the corresponding landmark in the feature map library is updated based on the confidence coefficient.
[0027] In this invention, the system calculates the similarity between the real-time generated feature vectors of real-world landmarks and the feature vectors of existing landmarks of the same type in the feature map library. If the similarity is lower than a preset threshold, the current landmark is determined to be a new landmark, and when its confidence coefficient is greater than zero (i.e., the feature is valid), a learning process is triggered to store its feature vector and related semantic context in the map library, thereby expanding the system's cognitive scope. If the similarity is not lower than the threshold, it is determined to be an existing landmark, and based on its confidence coefficient, the feature representation of the corresponding landmark in the library is iteratively optimized by integrating the current observation data. For example, as the detector gradually approaches the landmark, high-resolution features at close range are added. This mechanism gives the system the dual ability to "actively identify new landmarks" and "continuously optimize known landmarks" in unknown scenarios, enabling the feature map library to dynamically evolve with the exploration process and continuously deepen the structured cognition of the environment.
[0028] Regarding step 106: In this invention, after extracting features from all landmarks and updating the map database within the current scene, the system enters the stage of structured scene representation. First, the system utilizes the feature vectors and spatial locations of each landmark to construct a structured feature map. This map, in the form of a graph structure or topological relationships, accurately records the relative spatial distribution and topological connections between landmark nodes, forming a machine-understandable digital description of the scene layout. Simultaneously, based on predefined domain rules, the system automatically generates a scene semantic text, summarizing the core content of the current scene in natural language (e.g., "The scene center is a large impact crater, with several rock blocks distributed to its east"). Finally, the feature vectors, the structured feature map, and the semantic text together constitute a unified, complementary, three-layer lightweight scene representation. This representation system compresses high-dimensional, redundant raw perceptual data into extremely simple structured knowledge, achieving an order-of-magnitude reduction in storage overhead while fully preserving the key information required for scene recognition and revisit localization, thus completing the construction of the feature map for the current scene.
[0029] In one specific implementation, the system constructs a feature map of the current scene based on a three-layer representation consisting of feature vectors, a structured feature map, and semantic text. First, it performs a coarse semantic matching: using a text embedding model, the semantic text of the current scene and the semantic text of historical scenes stored in the feature map library are converted into vector representations. Cosine similarity is then used for rapid semantic alignment, efficiently filtering out several candidate scenes with the most similar semantic descriptions from a large-scale map library. Subsequently, a fine structure-feature matching is performed on the candidate scenes: each candidate scene is compared with the structured feature map of the current scene, sequentially checking the consistency of landmark node types and the structural similarity of spatial distribution relationships between landmarks. For nodes that match in both type and structure, the apparent similarity of their local feature vectors is further compared. If the combined result of the above multi-layer matching meets preset conditions, the current scene is determined to be an existing scene, and localization is successfully achieved; otherwise, it is treated as a new scene, and the mapping process is initiated. This judgment step, through a layered strategy of "rapid semantic filtering and fine structural verification," significantly improves the recognition efficiency in a large-scale scene library while ensuring matching robustness, efficiency, and accuracy.
[0030] Furthermore, the following supplementary explanation is provided regarding the construction method of this invention: 1) Predefined: the types of landmarks that are expected to exist in the target scene, and their characteristic attributes and evaluation methods are defined; 2) Reality perception and mapping: a. Identify landmarks in the actual scene based on multi-source, multi-dimensional sensing data and recognition methods, and determine the type of the real-scene landmarks by referring to predefined landmark types, and identify their characteristic attributes. Characteristic attributes include the following features: morphological attributes (shape, size, height, roughness, slope, particle size), mechanical attributes (pressure bearing capacity, density, etc.), and orientation attributes (relative positional relationship between landmarks, relative positional relationship with the current sensing location, etc.); b. Set confidence coefficients for the identified feature elements based on the resolution and accuracy of the perceived data and the identification method used. Each feature element of each target is assigned a confidence coefficient, and the confidence coefficient range can be set to 0-1. c. Establish a real-scene landmark database for the identified real-scene landmarks and their feature attributes. Each real-scene landmark in the database corresponds to a feature vector. If a perceived real-scene landmark cannot be classified according to the predefined landmark classification, it is determined to be a new landmark, and its landmark type and feature elements are supplemented and updated and included in the real-scene landmark database. d. Construct a local feature map based on the identified real-world landmarks and their characteristic attributes. This map only retains the identified real-world landmarks and their characteristic attributes, without requiring detailed 3D information. The aim is to retain the effective map information required for planning, navigation, and revisiting, while significantly reducing storage requirements. e. After each perception result update, the real-scene landmark database (real-scene landmarks and their feature elements) is updated, and the local feature map is also updated. f. The feature maps constructed through multiple perception and recognition processes are fused and stitched together to expand and construct a global feature map, recording the effective scene features along the entire mobile detection trajectory to ensure the accuracy and richness of the feature map information.
[0031] like Figure 2 , Figure 3 As shown, embodiments of the present invention provide a lightweight feature map construction and storage device for unknown and complex scenarios. The device can be implemented through software, hardware, or a combination of both. From a hardware perspective, such as... Figure 2 The diagram shown is a hardware architecture diagram of a computing device for a lightweight feature map construction and storage device for unknown and complex scenarios provided in an embodiment of the present invention. (Except for...) Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown, the computing device in the embodiment may also include other hardware, such as a forwarding chip responsible for processing packets. Taking software implementation as an example, such as... Figure 3 As shown, as a logical device, it is formed by the CPU of its computing device reading the corresponding computer program from the non-volatile memory into memory and running it. The lightweight feature map construction and storage device for unknown and complex scenarios provided in this embodiment includes: Module 300 is used to construct the landmark definition configuration based on the preset landmark type and the corresponding feature attribute set; The calculation module 302 is used to filter out real-scene landmark areas belonging to the landmark type from the multimodal perception data of the current scene based on the landmark definition configuration, extract feature attributes corresponding to the landmark definition configuration within the area, generate feature vectors of real-scene landmarks, and calculate the confidence coefficient of the feature vectors of real-scene landmarks by combining the feature attributes of real-scene landmarks. Update module 304 is used to update the feature map library based on feature vectors and confidence coefficients; The generation module 306 is used to construct a structured feature map representing the distribution relationship between landmarks based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, and generate semantic text describing the current scene; using the three-layer representation composed of feature vectors, structured feature maps and semantic text, the feature map of the current scene is constructed and stored.
[0032] In some specific implementations, the construction module 300 can be used to perform the above step 100, the calculation module 302 can be used to perform the above step 102, the update module 304 can be used to perform the above step 104, and the generation module 306 can be used to perform the above step 106.
[0033] In some specific implementations, the construction module 300 is used to perform the following operations: The characteristic attributes include landform type, spatial coordinates, geometric shape, color features, texture attributes, scale parameters, and first observation time.
[0034] In some specific implementations, the calculation module 302 is used to perform the following operations: From the multimodal perception data of the current scene, select real-world landmark areas belonging to the landmark type, including: Semantic segmentation is performed on images in multimodal perception data to obtain semantic labels and region masks. Based on the region masks, the real-world landmark regions corresponding to the landmark types are determined.
[0035] In some specific implementations, the calculation module 302 is also used to perform the following operations: The confidence coefficient is determined in the following ways: When the distance between the real-world landmark and the sensing device is less than or equal to the minimum confidence distance threshold, the confidence coefficient saturates to 1, indicating that the corresponding feature vector is valid. When the distance between the real-world landmark and the sensing device is greater than or equal to the maximum effective distance threshold, the confidence coefficient is truncated to 0, indicating that the corresponding feature vector is invalid. When the distance between the real-world landmark and the sensing device is between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient decreases as the distance increases.
[0036] In some specific implementations, the calculation module 302 is also used to perform the following operations: When the distance falls between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient is determined by the following formula: In the formula, CS(d) is the confidence coefficient, d is the distance between the real-world landmark and the sensing device, and d0 is the minimum confidence distance threshold. This is the attenuation coefficient.
[0037] In some specific implementations, the update module 304 is used to perform the following operations: The feature map database is updated based on feature vectors and confidence coefficients, including: The similarity between the feature vectors of real-world landmarks and the existing feature vectors of landmarks in the feature map library is calculated. If the similarity is lower than the preset threshold, the real-world landmark is determined to be a newly added landmark. If the confidence coefficient of the real-world landmark is greater than zero, its feature vector and associated semantic information are stored in the feature map library. If the similarity is not lower than the preset threshold, the real-world landmark is determined to be an existing landmark, and the feature vector of the corresponding landmark in the feature map library is updated based on the confidence coefficient.
[0038] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the lightweight feature map construction and storage device for unknown and complex scenarios. In other embodiments of the present invention, the lightweight feature map construction and storage device for unknown and complex scenarios may include more or fewer components than illustrated, or combine some components, split some components, or arrange different components. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0039] The information interaction and execution process between the modules in the above-mentioned device are based on the same concept as the method embodiment of the present invention, and the specific details can be found in the description of the method embodiment of the present invention, and will not be repeated here.
[0040] This invention also provides a computing device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the lightweight feature map construction and storage method for unknown and complex scenarios according to any embodiment of this invention.
[0041] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to execute the lightweight feature map construction and storage method for unknown and complex scenarios according to any embodiment of this invention.
[0042] Embodiments of this application also provide a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium and executes the computer program, causing the computer device to perform any of the lightweight feature map construction and storage methods for unknown and complex scenarios described in the above embodiments.
[0043] Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the embodiments described above is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.
[0044] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0045] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0046] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0047] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion module connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion module execute some and all of the actual operations, thereby realizing the function of any of the above embodiments.
[0048] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0049] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lightweight feature map construction and storage method for unknown and complex scenarios, characterized in that, include: Based on the preset land landmark types and their corresponding feature attribute sets, construct the land landmark definition configuration; Based on the landmark definition configuration, real-world landmark regions belonging to the landmark type are selected from the multimodal perception data of the current scene. Feature attributes corresponding to the landmark definition configuration are extracted within the region to generate feature vectors of real-world landmarks. The confidence coefficient of the feature vectors of real-world landmarks is calculated by combining the feature attributes of the real-world landmarks. The feature map database is updated based on the feature vectors and the confidence coefficients. Based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, a structured feature map representing the distribution relationship between landmarks is constructed, and semantic text describing the current scene is generated; using the three-layer representation composed of the feature vectors, the structured feature map, and the semantic text, the feature map of the current scene is constructed and stored.
2. The method according to claim 1, characterized in that, The step of updating the feature map database based on the feature vector and the confidence coefficient includes: The similarity between the feature vectors of the real-world landmarks and the existing feature vectors of landmarks in the feature map library is calculated. If the similarity is lower than a preset threshold, the real-world landmark is determined to be a newly added landmark. If the confidence coefficient of the real-world landmark is greater than zero, its feature vector and associated semantic information are stored in the feature map library. If the similarity is not lower than the preset threshold, the real-world landmark is determined to be an existing landmark, and the feature vector of the corresponding landmark in the feature map library is updated based on the confidence coefficient.
3. The method according to claim 1 or 2, characterized in that, The confidence coefficient is determined in the following ways: When the distance between the real-world landmark and the sensing device is less than or equal to the minimum confidence distance threshold, the confidence coefficient saturates to 1, indicating that the corresponding feature vector is valid; When the distance between the real-world landmark and the sensing device is greater than or equal to the maximum effective distance threshold, the confidence coefficient is truncated to 0, indicating that the corresponding feature vector is invalid; When the distance between the real-world landmark and the sensing device is between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient decreases as the distance increases.
4. The method according to claim 3, characterized in that, When the distance is between the minimum confidence distance threshold and the maximum effective distance threshold, the confidence coefficient is determined by the following formula: In the formula, CS(d) is the confidence coefficient, d is the distance between the real-world landmark and the sensing device, and d0 is the minimum confidence distance threshold. This is the attenuation coefficient.
5. The method according to claim 1, characterized in that, The step of filtering out real-world landmark areas belonging to the landmark type from the multimodal perception data of the current scene includes: The images in the multimodal perception data are semantically segmented to obtain semantic labels and region masks. Based on the region masks, the real-world landmark regions corresponding to the landmark types are determined.
6. The method according to claim 1, characterized in that, The characteristic attributes include landform type, spatial coordinates, geometric shape, color features, texture attributes, scale parameters, and first observation time.
7. A lightweight feature map construction and storage device for unknown and complex scenarios, characterized in that, The device includes: The construction module is used to build the landmark definition configuration based on the preset landmark types and corresponding feature attribute sets; The calculation module is used to filter out real-scene landmark regions belonging to the landmark type from the multimodal perception data of the current scene based on the landmark definition configuration, extract feature attributes corresponding to the landmark definition configuration in the region, generate feature vectors of real-scene landmarks, and calculate the confidence coefficient of the feature vectors of real-scene landmarks by combining the feature attributes of the real-scene landmarks. The update module is used to update the feature map library based on the feature vector and the confidence coefficient; The generation module is used to construct a structured feature map representing the distribution relationship between landmarks based on the feature vectors and spatial locations of all landmarks in the current scene in the updated feature map library, and generate semantic text describing the current scene; using the three-layer representation composed of the feature vectors, the structured feature map and the semantic text, the feature map of the current scene is constructed and stored.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to implement the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1-6.