Method and device for constructing visual semantic map of body-aware robot

By calculating the angle between 3D points and the image normal and filtering the image center region, high-quality projection viewpoints are automatically selected. Semantic labels are determined by weighted voting of multiple images, which solves the problem of low efficiency of manual annotation in robot vision semantic map construction and realizes efficient and accurate semantic map construction.

CN121982237BActive Publication Date: 2026-07-03红象科技(北京)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
红象科技(北京)有限公司
Filing Date
2026-01-26
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In existing technologies, the construction of robot visual semantic maps relies on manual annotation, which is inefficient and costly.

Method used

By calculating the angle between 3D points and the image normal and filtering the image center region, high-quality projection viewpoints are automatically selected, and semantic labels are determined by multi-image weighted voting to construct a visual semantic map for embodied intelligent robots.

Benefits of technology

It improves the efficiency and accuracy of robot visual semantic map construction, and reduces the need and cost of manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982237B_ABST
    Figure CN121982237B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for constructing a visual semantic map for an embodied intelligent robot. The method includes acquiring a target 3D point P in a visual semantic map point cloud. f And based on P f Identify all visible images corresponding to the environmental images or video frames acquired by the robot; calculate P. f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold, resulting in a first set of filtered images; the width h of the i-th visible image in the first set is then used as the basis for further filtering. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; P is determined based on the second image set. f The corresponding tag categories were defined. This enabled the construction of a visual semantic map for the embodied intelligent robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and specifically to a method and apparatus for constructing a visual semantic map for an embodied intelligent robot. Background Technology

[0002] In the field of embodied intelligence, for robots to achieve autonomous localization and navigation, a semantic map needs to be constructed. First, the environment in which the robot is located is modeled, usually using point cloud methods. Then, the semantic categories of different areas of the model are labeled manually, such as coffee shops and express delivery stations. Finally, the robot completes localization and navigation between targets based on the semantic information of the map.

[0003] The above process is not efficient in terms of semantic information of manually labeled maps, and it also has the disadvantage of high cost due to the large amount of manual work involved. Summary of the Invention

[0004] The main objective of this invention is to provide a method and apparatus for constructing a visual semantic map for an embodied intelligent robot, so as to address the shortcomings of related technologies.

[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for constructing a visual semantic map for an embodied intelligent robot is provided, comprising acquiring a target 3D point P in a visual semantic map point cloud. f And based on the P f Identify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i ; Calculate the P f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold, resulting in a first set of filtered images; the width h of the i-th visible image in the first set is then used as the basis for further filtering. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; based on the second image set, the P is determined. f The corresponding tag category.

[0006] Optionally, based on the included angle α i The filtering of visible images based on the magnitude relationship between the included angle α and the preset angle threshold includes: if the included angle α i If the angle is greater than or equal to the preset angle threshold, then the visible image is filtered out, retaining the included angle α. i The visible images below the preset angle threshold are denoted as the first set of filtered images.

[0007] Optionally, based on the width h of the i-th visible image in the first image set i and high w i Determining the optimal projection area includes: setting the indentation ratio β, based on the width h of the i-th visible image. i and high w i Shrink the left and right sides of the image towards the center, respectively. i ×β, the top and bottom shrink towards the image center h respectively. i ×β forms the optimal projection area.

[0008] Optionally, based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set, including: if the projection point p i If a point does not fall within the optimal projection area of ​​the image, then the visible image is filtered out, and projection point p is selected from the resulting first filtered image set. i The images that fall into the corresponding optimal projection area are denoted as the second filtered image set.

[0009] Optionally, based on the projection points p of each image in the second filtered image set i The corresponding preset semantic labels are used to summarize the semantic scores of each label category, and the semantic label with the highest score is selected as P. f The labels include: performing semantic segmentation on each image in the second filtered image set, assigning a preset semantic label category set {c1,c2,...,cx} to each pixel in the image, and obtaining the P in each visible image. f The projection point p i The corresponding semantic labels; for each semantic label cy in the preset semantic label category set, traverse l images in the image set after the second filtering, and summarize all projection points p. i The contribution score corresponding to the image labeled cy is used to obtain the semantic score Scorecyf of cy.

[0010] Optionally, the projection point p i The contribution of the image labeled cy is divided into The semantic score of cy is .

[0011] Optionally, the preset angle threshold range is between [70°, 90°].

[0012] Optionally, the indentation ratio β is adaptively adjusted according to the resolution of the visible image, and the adjustment range is limited to [1%, 20%].

[0013] According to a second aspect of the present invention, an apparatus for constructing a visual semantic map for an embodied intelligent robot is provided, comprising a preprocessing unit for acquiring a target 3D point P in a visual semantic map point cloud. f And based on the P f Identify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i The first screening unit is used to calculate P. f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold to obtain a first set of filtered images; the second filtering unit is used to filter images based on the width h of the i-th visible image in the first set. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; the label determination unit is used to determine the P based on the second image set. f The corresponding tag category.

[0014] According to a third aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing the computer to perform the method described in any one of the first aspects.

[0015] According to a fourth aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the method described in any implementation of the first aspect.

[0016] This embodiment embodies a method and apparatus for constructing a visual semantic map for an intelligent robot, wherein the method includes acquiring a target 3D point P in the visual semantic map point cloud.f And based on the P f Identify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i ; Calculate the P f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold, resulting in a first set of filtered images; the width h of the i-th visible image in the first set is then used as the basis for further filtering. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; based on the second image set, the P is determined. f The corresponding tag categories. High-quality projection perspectives are automatically selected through angle detection and effective region detection, thus automating the construction of a visual semantic map for the embodied intelligent robot. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the method for constructing a visual semantic map for an embodied intelligent robot according to an embodiment of the present invention;

[0019] Figure 2 , Figure 3 This is a schematic diagram illustrating the application of the method for constructing a visual semantic map for an embodied intelligent robot according to an embodiment of the present invention;

[0020] Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of the invention described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] According to embodiments of the present invention, a method for constructing a visual semantic map for an embodied intelligent robot is provided, such as... Figure 1 As shown, steps 101 to 103 are included below:

[0025] Step 101: Obtain the target 3D point P in the visual semantic map point cloud. f And based on the P f Identify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i .

[0026] In this step, the robot performs point cloud modeling of its environment to obtain a visual semantic map point cloud, while simultaneously acquiring environmental images or video frames and selecting the target 3D point P from the point cloud. f Determine P f It can use the environmental image or video frame mapped by back projection as the P f The visible images, each with a preset width h. i and high w i(i is the index of the visible image, i is a positive integer), and the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i .

[0027] Step 102: Calculate P f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The size relationship between the image and the preset angle threshold is used to filter the visible images, resulting in the first set of filtered images.

[0028] As one optional implementation method in this embodiment, based on the included angle α i The filtering of visible images based on the magnitude relationship between the included angle α and the preset angle threshold includes: if the included angle α i If the angle is greater than or equal to the preset angle threshold, then the visible image is filtered out, retaining the included angle α. i Visible images that meet the preset angle threshold are denoted as the first set of filtered images.

[0029] In this step, see Figure 2 From all visible images, filter out low-quality images with excessively skewed projection views: by calculating 3D point P. f The angle α between the vector to the image center and the image normal ni i , make the included angle α i Images with an angle greater than or equal to 70°-90° are excluded, and only images with a smaller angle (a closer view to the front / frontal view) are retained for the next step.

[0030] The above steps are essentially the first round of quality control for the image's perspective. The smaller the angle, the better the image quality when the robot acquires the image. f The more orthogonal the shooting angle of the corresponding physical object (such as frontal or oblique frontal), the more P in the image. f Projection point p i The semantic features (such as object shape and texture) are clearer and more recognizable; conversely, an excessively large angle (close to 90°) means an extremely biased viewpoint (such as side view, oblique rear view, or edge side view). i The semantic features are easily distorted and ambiguous, making it difficult to accurately reflect the true semantics of an object.

[0031] This step removes images with distorted semantic features due to viewpoint bias, preventing such images from introducing errors in semantic score calculations. This ensures that subsequent images used for voting and scoring have a reliable viewpoint, reducing the risk of misjudgment of 3D point semantic labels from the source. Pre-filtering out worthless, low-quality images eliminates the need for subsequent effective region detection and semantic score statistics, reducing the robot's computational power consumption and improving the overall efficiency of semantic map construction.

[0032] Step 103: Based on the width h of the i-th visible image in the first image set i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set.

[0033] In this step, the images after angle detection and filtering undergo a second precise filtering process to eliminate inferior images with poor projection positions. The optimal projection area is then delineated at the center of the image, and P is... f The projection point p i Images not falling into this area are filtered out, retaining only p. i An image located in the central region of the image.

[0034] Image edge regions are susceptible to camera imaging distortion (such as lens distortion) and shooting interference (such as blurred edges and cluttered backgrounds), and semantic segmentation models typically have greater annotation errors for edge pixels than for the central region; while the central region of the image has clearer imaging and less distortion. i The pixel semantic annotation accuracy is higher. Based on reliable viewpoint, a reliable projection position selection condition is added, providing double assurance that all images entering the semantic score calculation are high-quality images with a positive viewpoint and centered position. i The pixel semantic labels are more accurate, providing high-quality data support for subsequent weighted voting.

[0035] Semantic segmentation models typically have lower annotation accuracy for edge pixels than for the center region, and filtering p i Images located at the edges can directly avoid the impact of edge pixel annotation errors on P. f The influence of semantic tags further improves the accuracy of 3D point semantic annotation.

[0036] As an optional implementation of this embodiment, based on the width h of the i-th visible image in the first image set... i and high w i Determining the optimal projection area includes: setting the indentation ratio β, based on the width h of the i-th visible image. i and high wi Shrink the left and right sides of the image towards the center, respectively. i ×β, the top and bottom shrink towards the image center h respectively. i ×β forms the optimal projection area.

[0037] See Figure 3 The indentation ratio β (1%-20%) can be flexibly adjusted according to the image resolution (e.g., β can be appropriately reduced for high-resolution images and appropriately increased for low-resolution images) to ensure that the delineation of the optimal projection area always matches the actual imaging quality of the image and avoid the failure of the screening criteria due to resolution differences.

[0038] As an optional implementation of this embodiment, the visible images in the first image set are filtered based on the judgment result to obtain the filtered second image set, including: if the projection point p i If a point does not fall within the optimal projection area of ​​the image, then the visible image is filtered out, and projection point p is selected from the resulting first filtered image set. i The images that fall into the corresponding optimal projection area are denoted as the second filtered image set.

[0039] Step 104: Determine P based on the second image set f The corresponding tag category.

[0040] As an optional implementation of this embodiment, each image in the second filtered image set undergoes semantic segmentation processing, assigning a preset semantic label category set {c1,c2,...,cx} to each pixel in the image, thereby obtaining the P in each visible image. f The projection point p i The corresponding semantic labels; for each semantic label cy in the preset semantic label category set, traverse l images in the image set after the second filtering, and summarize all projection points p. i The contribution score corresponding to the image labeled cy is used to obtain the semantic score Scorecyf of cy.

[0041] In this implementation, the p of a single image i Labels only provide qualitative information about whether a semantic category is supported (such as "supports table" or "supports chair"). The semantic score calculation based on the second image set transforms this qualitative information into quantitative data on the strength of support by summing the scores. The quantitative score clearly quantifies the competitive relationship between different semantic labels, providing a clear basis for comparison in the subsequent "selection of the highest-scoring label" and avoiding ambiguity in label determination.

[0042] As an optional implementation of this embodiment, the projection point p i The contribution of the image labeled cy is divided into The semantic score of cy is .

[0043] Semantic segmentation is performed on each of the visible images, assigning a preset semantic label category set {c1,c2,...,cx} (where x is a positive integer, and the semantic label category includes at least one of table, chair, and corridor) to each pixel in the image, thereby obtaining the P in each visible image. f The projection point p i (P) f The semantic label corresponding to the specific location projected onto the i-th visible image.

[0044] For each semantic label cy (y∈[1,x]) in the semantic label category set, traverse l images in the second filtered image set, summarize the contribution scores corresponding to all images with projection point pi labeled as cy, and obtain the semantic score of cy. Compare the semantic scores corresponding to all semantic labels cy, and select the semantic label corresponding to the maximum score as the semantic label of the target 3D point Pf;

[0045] Assign corresponding semantic labels to all 3D points in the visual semantic map point cloud to complete the construction of the visual semantic map for the embodied intelligent robot.

[0046] If a small number of high-quality images are mislabeled due to semantic segmentation errors (such as mislabeling "table" as "chair"), since the score is the weighted sum of all images, the contribution of a single incorrect label is low and unlikely to affect the final result. Compared to the logic of labeling a single image, the weighted summation of multiple images can dilute individual errors, making the label determination more resistant to interference and improving the adaptability in complex environments (such as slightly blurred images or a small number of misjudgments in semantic segmentation).

[0047] This method uses "angle detection" and "effective area detection" to select high-quality projection viewpoints, and then uses a voting and scoring mechanism to automatically construct a visual semantic map for embodied intelligent robots.

[0048] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0049] According to an embodiment of the present invention, an apparatus for constructing a visual semantic map for an embodied intelligent robot is also provided, including a preprocessing unit for acquiring a target 3D point P in a visual semantic map point cloud. f And based on the P fIdentify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i ;

[0050] The first screening unit is used to calculate P. f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold to obtain the first set of filtered images.

[0051] The second filtering unit is used to filter images based on the width h of the i-th visible image in the first image set. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set;

[0052] The label determination unit is used to determine the P based on the second image set. f The corresponding tag category.

[0053] According to embodiments of the present invention, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the methods described in any of the above embodiments.

[0054] According to embodiments of the present invention, the present invention also provides a readable storage medium storing computer instructions that enable a computer to perform the methods described in any of the above embodiments when executed.

[0055] According to embodiments of the present invention, the present invention also provides a computer program product that, when executed by a processor, can implement the methods described in any of the above embodiments.

[0056] Figure 4A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0057] like Figure 4 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0058] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0059] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the object matching method. For example, in some embodiments, the object matching method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the methods described above may be performed.

[0060] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0061] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0062] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

Claims

1. A method for constructing a visual semantic map for an embodied intelligent robot, characterized in that, Includes the following steps: Acquiring a target 3D point P in a visual semantic map point cloud f , and determining all corresponding visible images in an environment image or a video frame collected by a robot based on the P f Each visible image has a preset width h i and height w i , and a vector passing through the center point of each visible image and perpendicular to the direction of the image plane is defined as the image normal n i ​ Calculate the P f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold to obtain the first set of filtered images. Based on the width h of the i-th visible image in the first image set i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; Based on the second image set, determine the P f The corresponding tag category.

2. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 1, characterized in that, Based on the included angle α i The relationship between the magnitude of the image and a preset angle threshold is used to filter visible images, including: If the included angle α i If the angle is greater than or equal to the preset angle threshold, then the visible image is filtered out, retaining the included angle α. i The visible images below the preset angle threshold are denoted as the first set of filtered images.

3. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 1, characterized in that, Based on the width h of the i-th visible image in the first image set i and high w i Determining the optimal projection area includes: Set the indentation ratio β based on the width h of the i-th visible image. i and high w i Shrink the left and right sides of the image towards the center, respectively. i ×β, the top and bottom shrink towards the image center h respectively. i ×β forms the optimal projection area.

4. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 1, characterized in that, Based on the judgment results, the visible images in the first image set are filtered to obtain the filtered second image set, which includes: If the projection point p i If a point does not fall within the optimal projection area of ​​the image, then the visible image is filtered out, and projection point p is selected from the resulting first filtered image set. i The images that fall into the corresponding optimal projection area are denoted as the second filtered image set.

5. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 1, characterized in that, Based on the projection points p of each image in the image set after the second filtering, i The corresponding preset semantic labels are used to summarize the semantic scores of each label category, and the semantic label with the highest score is selected as P. f The tags include: Semantic segmentation is performed on each image in the second filtered image set, assigning a preset semantic label category set {c1,c2,...,cx} to each pixel in the image, thus obtaining the P in each visible image. f The projection point p i Corresponding semantic tags; For each semantic label cy in the preset semantic label category set, iterate through l images in the image set after the second filtering, and summarize all projection points p. i The contribution score corresponding to the image labeled cy is used to obtain the semantic score of cy. .

6. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 5, characterized in that, The projection point p i The contribution of the image labeled cy is divided into ; The semantic score of cy is .

7. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 1, characterized in that, The preset angle threshold range is between [70°, 90°].

8. The method for constructing a visual semantic map for an embodied intelligent robot according to claim 3, characterized in that, The indentation ratio β is adaptively adjusted according to the resolution of the visible image, and the adjustment range is limited to [1%, 20%].

9. A device for constructing a visual semantic map for an embodied intelligent robot, characterized in that, include The preprocessing unit is used to acquire the target 3D point P in the visual semantic map point cloud. f And based on the P f Identify all visible images corresponding to environmental images or video frames acquired by the robot, wherein each visible image has a preset width h. i and high w i And the vector passing through the center point of each visible image and perpendicular to the image plane is defined as the image normal n. i ; The first screening unit is used to calculate P. f The vector to the center point of the i-th visible image and the image normal n of that image i The included angle α between them i Based on the included angle α i The visible images are filtered based on their size relationship with a preset angle threshold to obtain the first set of filtered images. The second filtering unit is used to filter images based on the width h of the i-th visible image in the first image set. i and high w i Determine the optimal projection area; determine the target 3D point P. f Projecting the i-th visible image yields the projection point p. i Whether it falls within the optimal projection area of ​​the image; based on the judgment result, the visible images in the first image set are filtered to obtain the filtered second image set; The label determination unit is used to determine the P based on the second image set. f The corresponding tag category.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Positioning method and device based on point cloud semantic map and robot

    CN115597605A

  • Semantic map optimization method and system fusing semantic information SLAM mapping frame screening

    CN116363359A