Semantic map construction method and device, semantic map navigation method and device, electronic equipment and storage medium
By collecting and fusing images from different perspectives, more rich and accurate semantic maps are generated, which solves the problems of lack of map content and low accuracy in the prior art, and realizes smarter and more accurate navigation services.
Patent Information
- Application Number
- CN202510450896.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the generation of semantic maps mainly relies on a single field-end perspective, resulting in relatively scarce map content and low accuracy, affecting the usage experience of navigation functions, and the application scenarios are limited to car parking scenarios and lack universality.
By using the field-end camera and the mobile camera to acquire different viewing angle images of the target field, identify and fuse the same objects in different acquired images, and generate a semantic map containing object position labels and identification labels.
It improves the richness and accuracy of the semantic map, supports the search tasks of open vocabulary, enhances the intelligence level of navigation services, and improves navigation accuracy.
Smart Images

Figure CN120141509A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of navigation technology, and in particular, to a method, apparatus, electronic device, and storage medium for semantic map construction and navigation based on a semantic map. Background Art
[0002] With the rapid development of deep learning technology, significant progress has also been made in semantic navigation technology. Semantic navigation refers to the integration of semantic understanding of user retrieval requests and environmental space information to provide users with accurate and efficient navigation services.
[0003] A semantic map is a map that contains description information such as object types, poses, and shapes. Through deep learning methods, a robot can use the semantic map to predict the target location to provide intelligent navigation services. For example, intelligent navigation services can be provided for large public places such as airports, hospitals, and shopping malls. In addition, by combining semantic navigation technology with vehicles, services such as driving navigation and positioning for parking can be provided for vehicles.
[0004] In the prior art, semantic maps can be mainly constructed through the following two schemes: one is to obtain visual image data of a robot in an environment and construct an open scene map representation based on the visual image data; the other is to obtain environmental information and reconstruct the real environment based on the environmental information to generate a corresponding semantic map.
[0005] The above-mentioned schemes both construct maps from a single field-end perspective, resulting in relatively scarce content and low accuracy of the generated maps, and affecting the usage experience of subsequent navigation functions. Moreover, the application scenarios of the above-mentioned technical solutions are mostly limited to the vehicle parking scenario and are not universal. Summary of the Invention
[0006] In view of this, various embodiments of the present disclosure provide a map construction and navigation method, apparatus, electronic device, and storage medium to at least partially solve the above problems.
[0007] According to a first aspect of the embodiments of the present disclosure, a method for constructing a semantic map is provided, including: collecting images of different perspectives of a target field using a field-end camera and a mobile-end camera to obtain a plurality of collected images of the target field, where the field-end camera is fixedly arranged in the target field and the mobile-end camera is movably arranged in the target field; identifying objects in each collected image to obtain a three-dimensional point cloud map and description information of each object corresponding to each collected image; and fusing the three-dimensional point cloud maps and description information of each object corresponding to each collected image according to the relative position relationship between the field-end camera and the mobile-end camera to generate a semantic map including position labels and identification labels of each object.
[0008] According to a second aspect of the embodiments of the present disclosure, there is provided a navigation method based on a semantic map, including: obtaining a search instruction for a target field; querying each object in the semantic map of the target field according to the search instruction, and determining a search target corresponding to the search instruction from each object, wherein the semantic map of the target field is generated by the semantic map construction method as described in the first aspect; determining a navigation starting point in the semantic map, and generating a navigation path from the navigation starting point to the search target.
[0009] According to a third aspect of the embodiments of the present disclosure, there is provided a semantic map construction device, including: an image acquisition module, configured to acquire different perspective images of a target field by using a field-side camera and a mobile-side camera, so as to obtain a plurality of acquired images of the target field, wherein the field-side camera is fixedly arranged in the target field, and the mobile-side camera is movably arranged in the target field; an image recognition module, configured to recognize objects in each acquired image, so as to obtain each three-dimensional point cloud map and each description information corresponding to each object in each acquired image; a map construction module, configured to fuse each three-dimensional point cloud map and description information corresponding to each object in each acquired image according to the three-dimensional point cloud maps and description information corresponding to each object in each acquired image, and the relative position relationship between the field-side camera and the mobile-side camera, and generate a semantic map including position labels and recognition labels of each object.
[0010] According to a fourth aspect of the embodiments of the present disclosure, there is provided a navigation device based on a semantic map, including: an obtaining module, configured to obtain a search instruction for a target field; a search module, configured to query each object in the semantic map of the target field according to the search instruction, and determine a search target corresponding to the search instruction from each object, wherein the semantic map of the target field is generated by the semantic map construction device as described in the third aspect; a navigation module, configured to determine a navigation starting point in the semantic map, and generate a navigation path from the navigation starting point to the search target.
[0011] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method described in the first aspect or the second aspect above.
[0012] According to a sixth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect or the second aspect above.
[0013] In summary, the semantic map construction and navigation scheme provided by various aspects of the present disclosure collect images from different field perspectives of the target field, identify and fuse the same objects in different collected images to construct a semantic map with richer and more accurate object description content, and can support the retrieval task of open vocabulary, which can improve the intelligent level of navigation services and navigation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is a processing flow chart of the semantic map construction method for an exemplary embodiment of the present disclosure.
[0016] Figure 2 It is a scene schematic diagram of collecting images of the target field for an exemplary embodiment of the present disclosure.
[0017] Figure 3 It is a processing flow chart of the navigation method based on the semantic map for an exemplary embodiment of the present disclosure.
[0018] Figure 4 It is a structural schematic diagram of the semantic map construction device for an exemplary embodiment of the present disclosure.
[0019] Figure 5 It is a structural schematic diagram of the navigation device based on the semantic map for an exemplary embodiment of the present disclosure.
[0020] Figure 6 It is a structural diagram of the electronic device for an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure, the following will clearly and detailedly describe the technical solutions in the embodiments of the present disclosure in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. Based on the embodiments in the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art should fall within the protection scope of the embodiments of the present disclosure.
[0022] The following will describe the specific implementation of the present disclosure in detail in conjunction with each drawing:
[0023] Semantic map construction method
[0024] Figure 1 The processing flow of the semantic map construction method according to an exemplary embodiment of the present disclosure mainly includes the following steps:
[0025] Step 102: Use the field camera and the mobile camera to collect images of different perspectives of the target field, and obtain multiple collected images of the target field.
[0026] The target field includes an indoor field or an outdoor field. Among them, the indoor field can be large indoor places such as airports, shopping malls, hospitals, libraries, indoor parking lots, etc., or small indoor places with a family as a unit; the outdoor field can be any type of outdoor place, including but not limited to: streets, outdoor parking lots, open commercial centers, etc. That is to say, the technical solution of the present disclosure is applicable to various scenarios.
[0027] In this embodiment, the field camera is fixedly arranged in the target field, and the mobile camera is movably arranged in the target field.
[0028] For example, Figure 2 In the shown example, when the target field 200 is a room (i.e., an indoor field), the field camera 210 can be a camera fixedly installed in the room, the mobile device 22 is a robot that can move in the room, and the mobile camera 220 is a camera installed on the robot; when the target field 200 is a street (i.e., an indoor field), the field camera 210 can be, for example, a camera fixedly installed on fixed facilities such as gantries and street lights, the mobile device 22 can be a vehicle traveling on the street, and the mobile camera 220 is a camera installed on the vehicle. The number of the field camera 210 and the mobile camera 220 set respectively is not limited to one shown in the drawings. Generally, since the mobile camera 220 can move relative to the target field 200, one mobile camera 220 is sufficient. The number of the field camera 210 can be adjusted arbitrarily according to actual needs. For example, two or more are set.
[0029] In some embodiments, one image can be collected from the target field 200 respectively through the field camera 210 and the mobile camera 220, and multiple collected images of the target field 200 are obtained. For example, when one field camera 210 and one mobile camera 220 are set in the target field 200, two collected images of the target field can be obtained.
[0030] In some embodiments, the collected images obtained through the field camera 210 and the mobile camera 220 are RGB images.
[0031] Step 104: Identify the objects in each collected image, and obtain the three-dimensional point cloud map and description information corresponding to each object in each collected image.
[0032] In some embodiments, the object categories and object positions in the captured images can be recognized to obtain the three-dimensional point cloud maps and description information of each object in the captured images.
[0033] Exemplarily, a trained vision-language multi-modal large model can be used to perform image encoding on each captured image, identify the category and position of each object in each captured image, and generate the three-dimensional point cloud maps and description information corresponding to each object for each captured image.
[0034] Step 106: According to the three-dimensional point cloud maps and description information corresponding to each object for each captured image, and the relative position relationship between the field-side camera and the mobile-side camera, fuse the three-dimensional point cloud maps and description information corresponding to each object for each captured image to generate a semantic map including the position labels and recognition labels of each object.
[0035] In some embodiments, the relative position relationship between the field-side camera and the mobile-side camera can be determined in the following manner:
[0036] Refer to Figure 2 , a mobile-side positioning label 222 can be set on the mobile device 22, and a field-side positioning label 212 can be set in the target field 200. The first pose of the field-side positioning label 212 in the coordinate system of the field-side camera 210 (refer to the dashed arrow 201), the second pose of the field-side positioning label 212 in the coordinate system of the mobile-side camera 220 (refer to the dashed arrow 203), and the third pose of the mobile-side positioning label 222 in the coordinate system of the field-side camera 210 (refer to the dashed arrow 202) can be solved respectively. According to the first pose, the second pose, and the third pose, the fourth pose of the mobile-side camera 220 in the coordinate system of the mobile-side positioning label 222 (refer to the dashed arrow 204) can be solved. According to the third pose and the fourth pose, the fifth pose of the mobile-side camera 220 in the coordinate system of the field-side camera 210 (refer to the dashed arrow 205) can be solved, and the relative position relationship between the field-side camera 210 and the mobile-side camera 220 can be determined according to the solution result of the fifth pose.
[0037] It should be noted that the solution method for the relative position relationship between the field-side camera 210 and the mobile-side camera 220 is not limited to the above example, and those skilled in the art can also adopt other known solutions to obtain it, and this embodiment does not limit it.
[0038] In some embodiments, the three-dimensional point cloud maps and description information corresponding to each object for each captured image can be recognized, the three-dimensional point cloud maps and description information corresponding to the same object in the target field for each captured image can be determined, and according to the relative position relationship between the field-side camera and the mobile-side camera and the coordinate information of each three-dimensional point cloud map, the three-dimensional point cloud maps corresponding to the same object for each captured image can be fused to obtain the fused point cloud maps and position labels corresponding to each object.
[0039] In some embodiments, geometric similarity calculations may be performed based on the 3D point cloud maps of each object corresponding to each captured image to determine the 3D point cloud maps of the same object in the target field corresponding to each captured image. Specifically, since the capture perspectives of different captured images are different, the 3D point cloud maps of the same object presented in different captured images are also different. The different 3D point cloud maps of the same object in different captured images can be determined by identifying the object contours of each object in each captured image.
[0040] In some embodiments, the description information of the same object corresponding to each captured image may be fused to obtain the recognition labels corresponding to each object. Specifically, semantic similarity calculations are performed based on the description information of each object in each captured image to determine the description information in each captured image that refers to the same object. For example, by performing semantic similarity calculations, description information such as "chair", "stool", and "seat" can be determined as recognition labels for representing the same object to support open-ended target search tasks.
[0041] In some embodiments, a 3D semantic map of the target field may be generated by combining the fused point cloud maps, position labels, and recognition labels corresponding to each object.
[0042] Specifically, the fused point cloud maps of each object may be superimposed according to the position labels of each object to construct a dynamic visual semantic map of the target field and determine the relative position relationships between the objects.
[0043] In summary, in this embodiment, by fusing multiple captured images from different perspectives to complement the environmental information in different captured images, the content of the generated semantic map can be made more rich and accurate, solving the problem of single information caused by generating a semantic map based on a single perspective.
[0044] Furthermore, when the objects in the target scene change (for example, new objects are added or removed in the target scene, or the positions of the objects change), only a new captured image needs to be obtained through the mobile camera, without re-obtaining the captured images of the field camera, to update the semantic map of the target field, improving the update efficiency of the semantic map and supporting the dynamic update of the semantic map.
[0045] In addition, in this embodiment, by fusing different description information referring to the same object to generate the recognition labels of each object in the semantic map, open-ended vocabulary target retrieval can be achieved, solving the problem that a fixed vocabulary map can only retrieve a limited number of targets.
[0046] Navigation method based on semantic map
[0047] Figure 3The processing flow of the navigation method based on the semantic map according to an exemplary embodiment of the present disclosure mainly includes the following steps:
[0048] Step 302, obtain a search instruction for the target field.
[0049] In some embodiments, the search instruction can be an instruction manually input by a human or an instruction automatically generated based on a triggered task (e.g., a vehicle parking task).
[0050] Step 304, query each object in the semantic map of the target field according to the search instruction, and determine a search target corresponding to the search instruction from each object.
[0051] In this embodiment, the semantic map of the target field is generated by the above-mentioned semantic map construction method.
[0052] Specifically, the search target corresponding to the search instruction can be determined from each object according to the position label and / or recognition label of each object marked in the semantic map of the target field.
[0053] In this embodiment, a large language model can be used to perform semantic parsing on the search instruction to obtain the search object description information and associated object description information of the search instruction. By parsing the search object corresponding to the search instruction and its associated objects, the accuracy of the search results can be effectively improved, and the search task of open vocabulary can be further supported, improving the intelligence level of the search task. For example, when the search instruction is "apple", the search object is the apple, and the associated object is the table for placing the apple.
[0054] In some embodiments, semantic similarity analysis can be performed according to the search object description information and associated object description information of the search instruction and the recognition labels of each object in the semantic map of the target field, and at least one candidate object and at least one associated object corresponding to the search instruction can be determined from each object. For example, when the search instruction is "a coffee shop near a gas station", each coffee shop near each gas station is each candidate object, and the associated object is the gas station.
[0055] In some embodiments, the relative position relationship between each candidate object and each associated object can be analyzed according to the position labels of each candidate object and each associated object marked in the semantic map, so as to determine the search target corresponding to the search instruction from each candidate object. For example, according to the relative position relationship between each coffee shop and the gas station, the coffee shop closest to the gas station can be determined as the search target.
[0056] Step 306, determine a navigation starting point in the semantic map, and generate a navigation path from the navigation starting point to the search target.
[0057] In some embodiments, the navigation starting point can be obtained by manual input by a person or by performing an automatic positioning operation through an automatic positioning system.
[0058] In some embodiments, the position tag of the search target can be determined as the navigation end point, and path planning is performed based on the semantic map of the target field, the navigation starting point, and the navigation end point to generate a navigation path.
[0059] In some embodiments, model predictive control (MPC) can be used to perform path following of the navigation path to achieve automatic movement control.
[0060] In summary, the navigation method based on the semantic map in this embodiment is implemented by using the semantic map generated by the above-described embodiment of the semantic map construction method. Since the content of the generated semantic map is richer and can support the search task of open vocabulary, not only can the accuracy of the navigation result be improved, but also the intelligent level of the navigation service can be improved.
[0061] Furthermore, by analyzing the search object description information and the associated object description information corresponding to the search instruction, and combining the relative position relationship between the search object and the associated object, the search target corresponding to the search instruction can be determined, which can further improve the accuracy of instruction parsing and the level of the navigation service.
[0062] Figure 4 FIG. is a schematic diagram of a semantic map construction device 400 according to an exemplary embodiment of the present disclosure, which mainly includes:
[0063] An image acquisition module 402, configured to acquire different perspective images of a target field by using a field camera and a mobile camera to obtain a plurality of acquired images of the target field, wherein the field camera is fixedly arranged in the target field, and the mobile camera is movably arranged in the target field;
[0064] An image recognition module 404, configured to recognize objects in each acquired image to obtain each three-dimensional point cloud map and each description information corresponding to each object in each acquired image;
[0065] A map construction module 406, configured to fuse the three-dimensional point cloud maps and description information corresponding to each object in each acquired image according to the three-dimensional point cloud maps and description information corresponding to each object in each acquired image and the relative position relationship between the field camera and the mobile camera, and generate a semantic map including the position tags and recognition tags of each object.
[0066] In some embodiments, the mobile camera is disposed on the mobile device. The image acquisition module 402 is further configured to: set a mobile positioning tag on the mobile device and set a field positioning tag in the target field; solve a first pose of the field positioning tag in the coordinate system of the field camera, a second pose of the field positioning tag in the coordinate system of the mobile camera, and a third pose of the mobile positioning tag in the coordinate system of the field camera; solve a fourth pose of the mobile camera in the coordinate system of the mobile positioning tag according to the first pose, the second pose, and the third pose; and solve a fifth pose of the mobile camera in the coordinate system of the field camera according to the third pose and the fourth pose to determine a relative position relationship between the field camera and the mobile camera.
[0067] In some embodiments, the image recognition module 404 is further configured to: recognize a three-dimensional point cloud map and a description information corresponding to each object in each acquired image, and determine the three-dimensional point cloud maps and the description information corresponding to the same object in the target field in each acquired image; fuse the three-dimensional point cloud maps corresponding to the same object in each acquired image according to the relative position relationship between the field camera and the mobile camera and the coordinate information of each three-dimensional point cloud map to obtain each fused point cloud map and each position tag corresponding to each object, and fuse the description information corresponding to the same object in each acquired image to obtain each recognition tag corresponding to each object; and generate a three-dimensional semantic map of the target field by combining each fused point cloud map, each position tag, and each recognition tag corresponding to each object.
[0068] In some embodiments, the image recognition module 404 is further configured to: perform geometric similarity calculation according to the three-dimensional point cloud maps corresponding to each object in each acquired image to determine the three-dimensional point cloud maps corresponding to the same object in the target field in each acquired image; and perform semantic similarity calculation according to the description information of each object in each acquired image to determine the description information pointing to the same object in each acquired image.
[0069] Navigation device based on semantic map
[0070] Figure 5 The following is a schematic diagram of a navigation device 500 based on a semantic map according to an exemplary embodiment of the present disclosure, which mainly includes:
[0071] An acquisition module 502, configured to acquire a search instruction for a target field.
[0072] A search module 504, configured to query each object in the semantic map of the target field according to the search instruction, and determine a search target corresponding to the search instruction from each object, where the semantic map of the target field is generated by the above-mentioned semantic map construction device 400.
[0073] A navigation module 506, configured to determine a navigation starting point in the semantic map and generate a navigation path from the navigation starting point to the search target.
[0074] The semantic map construction device 400 and the semantic map-based navigation device 500 according to the embodiments of the present disclosure can be used to implement the corresponding predicted map construction method and the semantic map-based navigation method in the foregoing method embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein. In addition, the function implementation of each module in the semantic map construction device 400 and the semantic map-based navigation device 500 according to the embodiments of the present disclosure can be referred to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated herein either.
[0075] The embodiments of the present disclosure provide a computer storage medium storing computer program code, which, when run by a processor, causes the processor to execute the map construction and navigation methods according to the embodiments of the present disclosure.
[0076] An exemplary embodiment of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and the computer program, when executed by the at least one processor, is used to cause the electronic device to execute the map construction and navigation methods according to the exemplary embodiments of the present disclosure.
[0077] Refer to Figure 6 , which shows a schematic structural diagram of an electronic device according to Embodiment 5 of the present application. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0078] As Figure 6 shown, the electronic device may include: a processor 602, a communication interface 604, a memory 606, and a communication bus 608.
[0079] Wherein:
[0080] The processor 602, the communication interface 604, and the memory 606 communicate with each other through the communication bus 608.
[0081] The communication interface 604 is configured to communicate with other electronic devices or servers.
[0082] The processor 602 is configured to execute the program 610, and specifically may execute the relevant steps in the foregoing embodiments of the semantic map construction and its navigation methods.
[0083] Specifically, the program 610 may include program code that includes computer operation instructions.
[0084] The processor 602 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0085] The memory 606 is used to store the program 610. The memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0086] The program 610 may include multiple computer instructions. Specifically, the program 610 may cause the processor 602 to execute the semantic map construction and its navigation method or corresponding operations described in any one of the foregoing multiple method embodiments through multiple computer instructions.
[0087] For the specific implementation of each step in the program 610, reference may be made to the corresponding descriptions in the corresponding steps and units in the foregoing method embodiments, and the corresponding beneficial effects are achieved, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0088] The embodiments of the present application also provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the foregoing multiple method embodiments. The computer storage medium includes but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, or magneto-optical disk, etc.
[0089] The embodiments of the present application also provide a computer program product, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the semantic map construction and its navigation method described in any one of the foregoing embodiments.
[0090] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data that have been authorized by the users or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0091] It should be pointed out that according to the needs of implementation, the various components / steps described in the embodiments of this application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.
[0092] The methods according to the embodiments of this application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium. Thus, the methods described herein can be stored as such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0093] Those of ordinary skill in the art will appreciate that the units and method steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0094] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The scope of patent protection of the embodiments of the present application shall be defined by the claims.
Claims
1. A method for constructing a semantic map, comprising: Using a field-side camera and a mobile-side camera to capture images of a target field from different perspectives, a plurality of captured images of the target field are obtained, wherein the field-side camera is fixedly disposed in the target field, and the mobile-side camera is movably disposed in the target field; Identify objects in each collected image, and obtain three-dimensional point cloud images and description information of each object corresponding to each collected image; According to the three-dimensional point cloud image and description information of each object corresponding to each acquired image, and the relative position relationship between the field-end camera and the mobile-end camera, the three-dimensional point cloud image and description information of each object corresponding to each acquired image are fused to generate a semantic map including the location label and identification label of each object.
2. The method for constructing a semantic map according to claim 1, wherein: The mobile terminal camera is arranged on the mobile device; The relative position relationship between the field-side camera and the mobile-side camera is determined by: Setting a mobile terminal positioning tag on the mobile device and setting a field terminal positioning tag in the target field; Solve the first pose of the field-side positioning tag in the coordinate system of the field-side camera, the second pose of the field-side positioning tag in the coordinate system of the mobile-side camera, and the third pose of the mobile-side positioning tag in the coordinate system of the field-side camera; According to the first posture, the second posture, and the third posture, solving a fourth posture of the mobile terminal camera in the coordinate system of the mobile terminal positioning tag; According to the third posture and the fourth posture, a fifth posture of the mobile-end camera in the coordinate system of the field-end camera is solved to determine the relative position relationship between the field-end camera and the mobile-end camera.
3. The method for constructing a semantic map according to claim 1, wherein: The method generates a semantic map including a location label and an identification label of each object by fusing the three-dimensional point cloud image and the description information corresponding to each collected image of each object according to the three-dimensional point cloud image and the description information corresponding to each collected image of each object and the relative position relationship between the field-side camera and the mobile-side camera, including: Identify the three-dimensional point cloud image and description information of each object corresponding to each collected image, and determine each three-dimensional point cloud image and description information of the same object in the target field corresponding to each collected image; According to the relative position relationship between the field-side camera and the mobile-side camera and the coordinate information of each three-dimensional point cloud image, the three-dimensional point cloud images corresponding to each collected image of the same object are fused to obtain each fused point cloud image and each position label corresponding to each object, and the description information corresponding to each collected image of the same object is fused to obtain each identification label corresponding to each object; The fused point cloud images, location tags, and identification tags corresponding to the objects are combined to generate a three-dimensional semantic map of the target field.
4. The method for constructing a semantic map according to claim 3, wherein: The identifying of the three-dimensional point cloud images and description information of each object corresponding to each collected image, and determining the three-dimensional point cloud images and description information of each object in the target field corresponding to each collected image, includes: Calculate geometric similarity based on the three-dimensional point cloud images of each object corresponding to each acquired image, and determine the three-dimensional point cloud images of the same object in the target field corresponding to each acquired image; According to each description information of each object in each collected image, semantic similarity calculation is performed to determine each description information pointing to the same object in each collected image.
5. A navigation method based on a semantic map, comprising: Obtain search instructions for the target area; Querying each object in the semantic map of the target field according to the search instruction, and determining the search target corresponding to the search instruction from each object, wherein the semantic map of the target field is generated by the semantic map construction method according to any one of claims 1 to 4; A navigation starting point in the semantic map is determined, and a navigation path from the navigation starting point to the search target is generated.
6. The navigation method based on semantic map according to claim 5, wherein: The step of querying each object in the semantic map of the target field according to the search instruction and determining the search target corresponding to the search instruction from each object includes: Performing semantic analysis on the search instruction to obtain search object description information and associated object description information of the search instruction; Performing semantic similarity analysis based on the search object description information and associated object description information of the search instruction and the identification labels of each object in the semantic map of the target field, and determining at least one candidate object and at least one associated object corresponding to the search instruction from the objects; According to the position labels of each candidate and each associated object identified in the semantic map, the relative position relationship between each candidate and each associated object is analyzed to determine the search target corresponding to the search instruction from each candidate.
7. A semantic map construction device, comprising: An image acquisition module, used to acquire images of a target field from different perspectives using a field-side camera and a mobile-side camera to obtain a plurality of acquired images of the target field, wherein the field-side camera is fixedly arranged in the target field, and the mobile-side camera is movably arranged in the target field; An image recognition module is used to recognize objects in each collected image and obtain three-dimensional point cloud images and description information of each object corresponding to each collected image; The map construction module is used to fuse the three-dimensional point cloud image and description information of each object corresponding to each collected image, and the relative position relationship between the field camera and the mobile camera, to generate a semantic map containing the location label and identification label of each object.
8. A navigation device based on a semantic map, comprising: An acquisition module, used to acquire a search instruction of a target field; a search module, configured to query each object in the semantic map of the target field according to the search instruction, and determine the search target corresponding to the search instruction from each object, wherein the semantic map of the target field is generated by the semantic map construction device according to claim 7; The navigation module is used to determine a navigation starting point in the semantic map and generate a navigation path from the navigation starting point to the search target.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to execute the semantic map construction method as described in any one of claims 1 to 4, or execute the semantic map-based navigation method as described in any one of claims 5 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program runs on a computer, the computer executes the semantic map construction method as described in any one of claims 1 to 4, or executes the semantic map-based navigation method as described in any one of claims 5 to 6.
Citation Information
Cited By
Model training method, electronic equipment and computer readable storage medium
CN120336859A
Robot object searching method, system and device and computer storage medium
CN121785330A