Visual semantic raster map construction method based on multi-mode and pre-training model

Through the application of multimodal data fusion and pre-trained models, a visual semantic raster map is constructed, which solves the problems of insufficient semantic information and lack of multimodal data fusion in traditional methods, and improves the robot's understanding of the environment and the practicality of the map.

CN119963752AInactive Publication Date: 2025-05-09DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411831794.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional service robot map construction method relies on two-dimensional geometric features and three-dimensional point cloud features, resulting in insufficient semantic information and lack of multimodal data fusion, making it difficult to fully present the environment, and limiting the execution of robot tasks.

Method used

The visual semantic raster map construction method based on multimodal and pre-trained models is adopted to obtain depth images and color images through RGB-D cameras, and combined with pre-trained models such as LSeg and BLIP, semantic images are obtained, and depth, color and semantic information are integrated to build a raster map with both spatial layout, visual and semantic features.

Benefits of technology

It improves the robot's understanding of the environment, enhances the practicality and adaptability of the map in different scenarios, realizes a more natural interaction with humans, and improves the practicality and acceptability of map construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963752A_ABST
    Figure CN119963752A_ABST
Patent Text Reader

Abstract

The invention discloses a visual semantic raster map construction method based on multiple modes and a pre-training model. The method comprises the following steps: collecting data; obtaining a semantic image by using the pre-training model; global and local three-dimensional coordinates and corresponding color and semantic information are obtained; and constructing a grid map. According to the method, multi-modal data are fully fused, the depth image, the color image and the camera pose information of the RGB-D camera are utilized, semantic information is fused, the map has spatial layout, visual and semantic features, the understanding of the robot to the environment is improved, and the practicability and adaptability of the map in different scenes are enhanced. According to the method, each point in the map is endowed with semantic information, so that the robot can better understand the environment, more natural interaction with human beings is realized, and the practicability and acceptability of map construction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of service robots, and in particular to a method for constructing a visual semantic grid map based on multimodality and pre-training models. Background Art

[0002] Industrial robots operate in a known environment and only need to perform simple, repetitive tasks. Service robots, however, face largely unknown environments and require the construction of environmental maps to facilitate tasks such as path planning, obstacle avoidance, and object recognition. Building semantic maps, similar to human cognition, enables robots to better interact with users and efficiently complete service tasks. Human-robot interaction is a hot research topic in robotic mapping and navigation. Natural, concise, and user-friendly human-robot interfaces are crucial for the public acceptance of service robots. Building semantic environmental maps and interacting naturally with users can improve the quality of robot services in areas such as security inspections, elderly care, exhibition navigation, house cleaning, and search and rescue.

[0003] Currently, service robots build indoor maps by roaming and extracting information such as two-dimensional geometric features and three-dimensional point cloud features of the environment to construct raster maps, topological maps, and hybrid maps. There are two aspects that need improvement and exploration:

[0004] The first is a lack of semantic information. Traditional methods rely primarily on two-dimensional geometric features and three-dimensional point cloud features. This lack of semantic information when constructing maps makes it difficult for robots to understand their environment in the same way humans do and interact naturally. Consider assigning semantic information to each point in the map. This could enable robots to better understand their environment, enabling more natural interactions with humans and improving the practicality and acceptability of map construction.

[0005] Second, there is a lack of consideration for multimodal data fusion. Traditional map construction often relies on single-source data. For example, lidar only uses distance information, lacking color, texture, and semantic descriptions. This makes it difficult to fully represent the environment, limiting the robot's ability to perform tasks. Summary of the Invention

[0006] In order to solve the problems of traditional map construction methods that rely on two-dimensional geometric features and three-dimensional point cloud features, resulting in insufficient semantic information when constructing maps, and the lack of multimodal data fusion considerations in traditional map construction methods, the present invention provides a visual semantic raster map construction method based on multimodality and pre-trained models, which can enable maps to have both spatial layout, visual and semantic features, improve the robot's understanding of the environment, and enhance the practicality and adaptability of maps in different scenarios.

[0007] The basic idea of ​​the present invention is as follows: during the robot exploration phase, depth images and color images are acquired using an RGB-D camera, and the corresponding camera position and rotation state are acquired through an RGB-D real-time positioning and mapping system; semantic images are acquired through a language-driven semantic segmentation model LSeg, a visual language model BLIP, and color images; then, global and local three-dimensional coordinates are acquired using the depth image, camera intrinsic parameters, camera position, and posture information; coordinates, color, and semantic information are fused using camera intrinsic parameters, local three-dimensional coordinates, color images, and semantic images; finally, global three-dimensional coordinates are mapped to the corresponding three-dimensional grid using a standard conversion operation, and the corresponding color information and semantic information are filled in the information dictionary.

[0008] The technical solution adopted in the present invention is as follows:

[0009] A method for constructing a visual semantic grid map based on multimodality and pre-trained models includes the following steps:

[0010] A. Collect data

[0011] During the robot's exploration phase, an RGB-D camera with known camera intrinsic parameters is used to obtain depth images and color images, and the camera position and posture information is obtained through the RGB-D real-time positioning and mapping system.

[0012] B. Using pre-trained models to obtain semantic images

[0013] Semantic images are obtained with the help of the language-driven semantic segmentation model LSeg, the visual language model BLIP and color images.

[0014] C. Obtain global and local 3D coordinates and their corresponding color and semantic information

[0015] Depth images, camera intrinsic parameters, color images and semantic images are used to obtain global and local three-dimensional coordinates and their corresponding color and semantic information.

[0016] D. Build a raster map

[0017] Through coordinate conversion operations, the global three-dimensional coordinates are mapped to the three-dimensional grid, and the color and semantic information of the corresponding grid are filled in the information dictionary.

[0018] Furthermore, the method for collecting data in step A is as follows: when the robot enters an unknown environment for exploration, the robot obtains depth images and color images through the RGB-D camera with known camera intrinsic parameters carried by the robot, and obtains the world three-dimensional coordinates and quaternion rotation state of the camera position corresponding to all images through the RGB-D real-time positioning and mapping system.

[0019] The camera intrinsic parameters are parameters that describe the internal properties of the camera, including focal length and principal point coordinates.

[0020] Furthermore, the pre-trained model in step B includes a visual language model BLIP and a language-driven semantic segmentation model LSeg.

[0021] Furthermore, the method for obtaining a semantic image in step B is as follows: First, the visual language model BLIP is used to ask a template question, "What's in the room?" for each color image to obtain all object categories in the environment. The object categories are then combined with the language-driven semantic segmentation model LSeg to perform semantic segmentation on each color image. Each pixel is assigned an index corresponding to a list of object categories, linking the object category semantic information with the color image pixels to construct a semantic image corresponding to each color image.

[0022] Furthermore, the method for obtaining global and local three-dimensional coordinates and their corresponding color and semantic information in step C is as follows: the local three-dimensional coordinates of each pixel are obtained through the depth image and camera intrinsic parameters, and then the local three-dimensional coordinates are transformed into global three-dimensional coordinates using the camera position and rotation state; and then the camera intrinsic parameters and local three-dimensional coordinates are used to obtain the color and semantic information of the corresponding pixels in the corresponding color image and semantic image, thereby completing the fusion of semantic and color information and the depth image.

[0023] Furthermore, the method for obtaining local three-dimensional coordinates in step C includes the following steps:

[0024] C11. Merge into homogeneous coordinates

[0025] Assume that in the two-dimensional pixel coordinate system corresponding to the depth image p, the x-axis is the horizontal direction of the depth image, the right direction is positive, and the y-axis is the vertical direction of the depth image, the downward direction is positive; the width of the depth image is w, the height is h, and the total number of pixels is N; then the i-th pixel in the depth image is represented by the two-dimensional pixel coordinates [x i ,y i ] T Represents its position in the depth image plane, where x i Indicates the coordinate value of the pixel on the x-axis, y i Indicates the coordinate value of the pixel on the y-axis, where 0≤x i ≤w-1, 0≤x i ≤h-1, 1≤i≤N. According to the construction rules of homogeneous coordinates, the homogeneous coordinate matrix P of the depth image p with a dimension of 3×N is obtained 2D , the homogeneous coordinate matrix is ​​expressed as follows:

[0026]

[0027] Homogeneous coordinate matrix P 2DThe 1 in the third row indicates that an extra dimension is added to the coordinate representation of each pixel.

[0028] C12. Convert to three-dimensional coordinates

[0029] Assume that when shooting a depth image, the corresponding camera intrinsic parameter matrix is ​​K and the dimension is 3×3, then the inverse matrix of the camera intrinsic parameter matrix is ​​K -1 , multiply the inverse matrix of the camera intrinsic parameter matrix by the homogeneous coordinate matrix to convert the two-dimensional pixel coordinates into three-dimensional coordinates, and obtain the local three-dimensional coordinate matrix P of the depth image p with a dimension of 3×N C as follows:

[0030] P C =K -1 P 2D

[0031] The expanded matrix multiplication is expressed as follows:

[0032]

[0033] Among them, X i 、Y i and Z i k are the x-axis, y-axis, and z-axis coordinate values ​​corresponding to the i-th pixel after conversion in the local three-dimensional coordinate system of the depth image p. -1 ij is the inverse matrix K of the camera intrinsic parameter matrix -1 The element in the i-th row and j-th column, i = 1, 2, 3, j = 1, 2, 3.

[0034] X i 、Y i and Z i The calculation formula is:

[0035]

[0036] Furthermore, the standard form of the camera intrinsic parameter matrix K in step C12 is:

[0037]

[0038] where f x and f y are the focal lengths of the camera in the x and y directions, respectively, and c x and c y are the principal point coordinates in the x direction and the principal point coordinates in the y direction respectively.

[0039] C13. Incorporating depth information

[0040] Let the depth value corresponding to the i-th pixel in the depth image p be z i, then use a 1×N row vector Z = [z1, z2, …, z N to represent the depth information of all pixels. Multiply the local three-dimensional coordinate matrix P C of the depth image p by the depth information row vector Z to obtain the local three-dimensional coordinate matrix P world of the depth image p with a dimension of 3×N. The calculation formula is as follows:

[0041] P world = P C ⊙ Z

[0042] Here, ⊙ represents element-wise multiplication, that is:

[0043]

[0044] where X world,i , Y world,i and Z world,i are respectively the coordinate values of the x-axis, y-axis, and z-axis of the local three-dimensional coordinate system corresponding to the i-th pixel in the depth image p after integrating the depth information.

[0045] C14. Screening according to the effective depth range

[0046] Suppose the effective depth range of the depth image is [min_depth, max_depth]. Use this range to screen out the effective three-dimensional coordinates within the effective depth range. Then construct a 1×N mask vector M = [m1, m2, …, m N , and its elements are obtained through the following conditions:

[0047]

[0048] For the local three-dimensional coordinate matrix P world of the depth image p, screen it through this mask vector to obtain the local three-dimensional coordinate matrix P valid of the effective depth image p with a dimension of 3×N valid , N valid is the number of effective coordinates after screening, and N valid < N. The specific screening method is expressed as:

[0049]

[0050] That is, retain the local three-dimensional coordinates of the depth image p corresponding to the elements m i = 1 in the mask vector M, and discard the local three-dimensional coordinates of the depth image p where m i = 0.

[0051] C15. Reducing coordinate data according to the sampling rate

[0052] Let the sampling rate be s, where 0 < s ≤ 1. After sampling, the local three-dimensional coordinate matrix P of the final depth image p with dimension 3×M is obtained. final , where M is the number of remaining coordinates after sampling, and M < N. valid . Sampling is performed at equal intervals in the following manner:

[0053]

[0054] Here denotes slicing and selection of a two-dimensional matrix in the column direction with a step size of .

[0055] Through the above steps, the local three-dimensional coordinates corresponding to each pixel in each depth image are obtained.

[0056] Furthermore, the method for obtaining the global three-dimensional coordinates described in step C includes the following steps:

[0057] C21. Construct a homogeneous transformation matrix for each depth image

[0058] Let the world three-dimensional coordinates of the camera when shooting the i-th depth image u i be p i world = [x i world , y i world , z i world T , and the quaternion rotation state be q i = (a i , b i , c i , d i ). Among them, x i world , y i world , z i world are the coordinate values of the x-axis, y-axis, and z-axis in the world three-dimensional coordinate system of the camera calculated by the RGB-D simultaneous localization and mapping system when shooting the i-th depth image u i ; q i = (a i , b i , c i , d i ) is the rotation state in the world three-dimensional coordinate system calculated by the RGB-D simultaneous localization and mapping system when shooting the i-th depth image u i , where a i is the real part of the quaternion, and b i , c i , d​i is the quaternion q i The imaginary part of . Then construct the corresponding homogeneous transformation matrix T of the i-th depth image i The specific steps are as follows:

[0059] Using quaternion q i =(a i ,b i ,c i ,d i ) Construct the corresponding rotation matrix R i :

[0060]

[0061] Get the rotation matrix R i Then, construct the 4×4 homogeneous transformation matrix T according to the following formula i :

[0062]

[0063] The homogeneous transformation matrix T of the expanded representation i for:

[0064]

[0065] C22. Calculate the basic matrix for coordinate system transformation

[0066] The local 3D coordinate system corresponding to the first depth image u1 is selected as the global 3D coordinate system. The 3D coordinates of the camera corresponding to the first depth image u1 are expressed as p1 = [x1 world ,y1 world ,z1 world ] T , the quaternion rotation state is q1 = (a1, b1, c1, d1), and accordingly, the homogeneous transformation matrix constructed is T1. Further, the inverse matrix T1 of the homogeneous transformation matrix T1 is calculated -1 , and the inverse matrix will be used as the basic matrix for various subsequent coordinate transformation operations, providing an important support for the coordinate transformation and other related processing involved in the entire subsequent process.

[0067] C23. Calculate the global transformation matrix corresponding to each depth image

[0068] The i-th depth image u i The corresponding camera three-dimensional coordinates are p i =[x i world ,y i world ,z i world ] T, the quaternion rotation state is q i =(a i ,b i ,c i ,d i ). Accordingly, the homogeneous transformation matrix constructed is T i , and the initial transformation inverse matrix T1 -1 Multiply them together to get the global transformation matrix T gi =T1 -1 T i ;

[0069] C24. Coordinate transformation and obtaining global three-dimensional coordinates

[0070] Let the depth image u i The final local three-dimensional coordinate matrix is ​​U i,final , U i,final The jth coordinate in is p ui,j =[χ ui,j ,y ui,j ,z ui,j ] T , convert it into the following homogeneous coordinate form

[0071]

[0072] Then use the global coordinate transformation matrix T gi and Multiply to achieve the depth image u i The transformation from the local three-dimensional coordinate system to the global three-dimensional coordinate system gives the global three-dimensional coordinates in homogeneous coordinates as follows:

[0073]

[0074] Finally, extract the first three rows of data as the three-dimensional coordinates p in the global three-dimensional coordinate system global,i,j :

[0075]

[0076] where x global,i,j 、y global,i,j and z global,i,j They are the coordinate values ​​of the x-axis, y-axis, and z-axis in the global three-dimensional coordinate system.

[0077] Through the above steps, the local three-dimensional coordinates corresponding to each depth image are converted into global three-dimensional coordinates.

[0078] Furthermore, the method for obtaining color information and semantic information in step C includes the following steps:

[0079] C31. Coordinate transformation

[0080] Assume that U in step C24 i,final The jth coordinate in is p ui,j =[χ ui,j ,y ui,j ,z ui,j ] T .

[0081] At the same time, with the depth image u i Also captured is a color image u i,rgb , and for color image u i,rgb Correspondingly, a semantic image u with the same height, width and pixel arrangement is generated i,semantic So suppose that when taking color image u i,rgb When , the corresponding camera intrinsic parameter matrix is ​​K rgb .

[0082] Since in constructing the semantic image u i,semantic In the process, the pixel position is not changed, that is, the semantic image u i,semantic With color image u i,rgb The spatial layout at the pixel level is consistent, so the camera intrinsic parameter matrices corresponding to the two are the same, that is, the semantic image u i,semantic The corresponding camera intrinsic parameter matrix is ​​also K rgb Moreover, since the intrinsic parameter matrices are the same, the corresponding coordinate transformation results should also remain consistent when performing coordinate transformation operations.

[0083] Next, the depth image u is transformed into i The local three-dimensional coordinate p in the local three-dimensional coordinate system ui,j Convert to color image u i,rgb The camera coordinates in the camera coordinate system are expressed as follows:

[0084] P cam,ui,j rgb =K rgb P ui,j =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T

[0085] And because the semantic image u i,semantic With color image u i,rgb The coordinate transformation results are also consistent, so the semantic image u i,semantic The camera coordinates in the camera coordinate system are expressed as follows:

[0086] Pcam,ui,j semantic =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T .

[0087] C32. Extracting depth information

[0088] Extract color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic The coordinate value z on the z-axis cam,ui,j rgb , as depth information.

[0089] C33. Perspective Division and Homogeneous Coordinate Transformation

[0090] For color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic Perform perspective division and divide each coordinate component by z cam,ui,j rgb , get the color image u under homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic :

[0091]

[0092] C34. Calculate two-dimensional plane coordinates and extract color and semantic information

[0093] Using the result of step C33, extract the color image u under homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic The first two rows of data are as follows: i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system ui,j rgb and Pui,j semantic :

[0094]

[0095] where x ui,j rgb and y ui,j rgb are color images u i,rgb and semantic image u i,semantic The x-axis and y-axis coordinate values ​​in the two-dimensional pixel coordinate system.

[0096] Then use P ui,j rgb and P ui,j semantic , from the color image u of dimension H×W×3 i,rgb and a semantic image u of dimension H×W×1 i,semantic Extract the depth image u i The three-dimensional coordinate p in the local three-dimensional coordinate system ui,j Color information C RGB and semantic information S semantic , the formula is as follows:

[0097] C RGB =u i,rgb (x ui,j rgb ,y ui,j rgb ,:)

[0098] S semantic =u i,semantic (x ui,j rgb ,y ui,j rgb ,:)

[0099] Among them, H is the height of the color image and semantic image, W is the width of the color image and semantic image, and the semantic image u is the width of the color image and semantic image. i,semantic The semantic information in is the index of the item category list.

[0100] Furthermore, the method for constructing the grid map in step D is as follows:

[0101] The grids are divided into idle and occupied states. The grid size and two-dimensional scale are determined according to needs, as well as the information dictionary used to store the color information and semantic information of the grid in the occupied state. The three-dimensional coordinates are mapped to the corresponding three-dimensional grid by using flooring and coordinate offset. At the same time, the grid coordinates of the grid in the occupied state are used as the key to fill in the corresponding color information and semantic information in the information dictionary.

[0102] The mapping steps are as follows:

[0103] For the global three-dimensional coordinate p obtained in step C24 global,i,j =[x global,i,j ,y global,i,j ,z global,i,j ] T , and map it to the corresponding three-dimensional grid coordinates through the following steps:

[0104]

[0105] Where row, col and height are the coordinate values ​​of the row, column and height directions in the grid respectively. w and G h are the number of rows and columns of the grid, respectively, and cs is the side length of the grid.

[0106] The method of filling the information dictionary is as follows:

[0107] When a grid is determined to be occupied, its corresponding three-dimensional grid coordinates row, col and height are used as keys to convert the color information C corresponding to the grid into RGB and semantic information S semantic Store it in the information dictionary to complete the information filling operation during the raster map construction process.

[0108] Compared with the prior art, the present invention has the following beneficial effects:

[0109] 1. Traditional methods rely primarily on 2D geometric features and 3D point cloud features, but lack semantic information when constructing maps. This makes it difficult for robots to understand their environment in the same way humans do and interact naturally. This invention assigns semantic information to each point in the map, enabling robots to better understand their environment and interact more naturally with humans, improving the practicality and acceptability of map construction.

[0110] 2. Traditional map construction often relies on single-source data. For example, LiDAR uses only distance information, lacking color, texture, and semantic descriptions. This makes it difficult to fully represent the environment, limiting the robot's ability to perform tasks. This invention fully integrates multimodal data, utilizing depth images, color images, and camera pose information from RGB-D cameras, and incorporating semantic information. This allows the map to combine spatial layout, visual, and semantic features, improving the robot's understanding of the environment and enhancing the map's practicality and adaptability in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0112] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0113] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0114] In order to solve the problems of insufficient semantic information when constructing maps due to the reliance on two-dimensional geometric features and three-dimensional point cloud features in traditional map construction methods, as well as the lack of multimodal data fusion considerations in traditional map construction methods, the present invention provides a visual semantic grid map construction method based on multimodality and pre-trained models, which can construct a visual semantic grid topological map: it contains a three-dimensional grid for describing environmental geometric information, and a matching information dictionary for describing the color and semantic information of the grid. The information dictionary can be used to match the grid with color or semantic information.

[0115] like Figure 1 As shown in the figure, the process of constructing a visual semantic grid map based on multimodality and pre-trained models is as follows:

[0116] 1. Data Collection

[0117] When a robot enters an unknown environment, it can roam freely or be controlled by a human to explore the environment. During exploration, the robot acquires depth and color images using an RGB-D camera with known camera intrinsic parameters. The RGB-D real-time positioning and mapping system then acquires the 3D coordinates and quaternion rotation of the camera corresponding to the image.

[0118] Camera intrinsic parameters: Camera intrinsic parameters are parameters that describe the internal properties of the camera, mainly including focal length, principal point coordinates, etc. These parameters are usually fixed after the camera leaves the factory.

[0119] Focal length: Focal length is a very critical parameter in the camera optical system. It essentially determines the degree of zoom and viewing angle of the object when the camera is imaging. It is usually divided into horizontal focal length f x and vertical focal length f y , which correspond to the camera's zoom in and out of the object along the x-axis and y-axis, respectively. Focal length is generally expressed in millimeters (mm), but is often converted to pixels for calculations involving pixel coordinate conversion.

[0120] Principal point coordinates: The principal point coordinates are the principal point coordinates c in the x direction x and the y-direction principal point coordinate c y It determines the position of the center of the camera's imaging plane in the camera coordinate system. Simply put, light from an external object passes through the center of the camera lens and, ideally, is projected onto the principal point of the imaging plane. The coordinates of this principal point on the image plane are the principal point coordinates, usually measured in pixels.

[0121] RGB-D simultaneous localization and mapping system: A technology that uses color (RGB) and depth image information for simultaneous localization and mapping. The present invention does not restrict the use of any specific RGB-D simultaneous localization and mapping system, as long as the camera position and rotation state corresponding to the depth and color images can be obtained.

[0122] Quaternion is an effective mathematical tool for representing the rotational posture of an object in three-dimensional space. Compared with the traditional Euler angle representation, it avoids problems such as universal joint lock and can describe the rotation more smoothly and continuously. Quaternion q can generally be written as q = (a, b, c, d), where a is the real part and b, c, d are the imaginary parts, and it meets certain operation rules and constraints (a 2 +b 2 +c 2 +d 2 =1). Quaternions can be used to encode the rotation state of an object relative to a reference coordinate system. That is, the angles by which the object rotates around the x-axis, y-axis, and z-axis can be encoded in the quaternion.

[0123] 2. Obtaining Semantic Images Using Pre-trained Models

[0124] First, the visual language model BLIP is used to ask each color image a template question, "What's in the room?", to obtain all the object categories in the environment. The object categories are then combined with the language-driven semantic segmentation model LSeg to perform semantic segmentation on each color image. By assigning each pixel an index corresponding to a list of object categories, the semantic information of the object categories is linked to the image pixels, thus constructing a semantic image corresponding to each color image.

[0125] Pretrained models: Pretrained models are models that have been pre-trained on large datasets. These models typically perform well on certain general tasks and can serve as a starting point for subsequent specific tasks. The visual language model BLIP and the language-driven semantic segmentation model LSeg mentioned in this paper are both pretrained models.

[0126] Visual Language Inference (BLIP) Model: The BLIP model is a multimodal model that utilizes a hybrid encoder-decoder architecture and is jointly trained with three visual language objectives: image-text contrast learning to maximize the similarity between positive image-text pairs and minimize negative ones; image-text matching to learn joint representations; and image-conditional language modeling to generate text descriptions. The model also uses the CapFilt method to train on noisy data and then on clean data generated through generated captions and filtering. This joint training allows the model to effectively capture the relationship between images and text, enabling it to perform visual question answering, image-text retrieval, and image caption generation.

[0127] LSeg, a language-driven semantic segmentation model: LSeg is a multimodal learning-based model that focuses on semantic understanding and segmentation between images and text. By combining visual information and language descriptions, the model aims to accurately label and classify different regions in an image. The core idea of ​​LSeg is to use natural language descriptions to guide the segmentation of image content, enabling the model to understand the function and meaning of each part of the image. LSeg uses a deep learning architecture, including an image encoder and a language encoder. The former is responsible for extracting image features, and the latter converts text information into feature vectors. Through an interactive mechanism, LSeg can effectively integrate image and text information, thereby achieving more efficient semantic segmentation tasks.

[0128] 3. Obtaining global and local 3D coordinates and their corresponding color and semantic information

[0129] Through the depth image, the camera intrinsic parameters obtain the local three-dimensional coordinates of each pixel, and then use the camera position and rotation state to transform the local three-dimensional coordinates into global three-dimensional coordinates; then use the camera intrinsic parameters and local three-dimensional coordinates to obtain the color and semantic information of the corresponding pixels in the corresponding color image and semantic image, thereby completing the fusion of semantic and color information and depth image.

[0130] Get local 3D coordinates:

[0131] Merge into homogeneous coordinates; suppose that in the two-dimensional pixel coordinate system corresponding to the depth image p, the x-axis is the horizontal direction of the depth image, the right direction is positive, and the y-axis is the vertical direction of the depth image, the downward direction is positive; the width of the depth image is w, the height is h, and the total number of pixels is N; then the i-th pixel in the depth image is represented by the two-dimensional pixel coordinates [x i ,y i ] T Represents its position in the depth image plane, where x i Indicates the coordinate value of the pixel on the x-axis, y i Indicates the coordinate value of the pixel on the y-axis, where 0≤x i ≤w-1, 0≤x i≤h-1, 1≤i≤N. According to the construction rules of homogeneous coordinates, the homogeneous coordinate matrix P of the depth image p with a dimension of 3×N is obtained 2D , the homogeneous coordinate matrix is ​​expressed as follows:

[0132]

[0133] Homogeneous coordinate matrix P 2D The 1 in the third row indicates that an extra dimension is added to the coordinate representation of each pixel.

[0134] Homogeneous coordinates represent an N-dimensional vector using an N+1-dimensional vector, where the extra dimension is usually set to 1. Homogeneous coordinates allow for the unified representation of various geometric transformations, such as translation, rotation, scaling, and projection, using matrix multiplication.

[0135] A homogeneous coordinate matrix is ​​a matrix composed of homogeneous coordinates. Using a homogeneous coordinate matrix, multiple homogeneous coordinates can be operated and processed in a unified manner.

[0136] Convert to three-dimensional coordinates; when shooting a depth image, the corresponding camera intrinsic parameter matrix is ​​K and the dimension is 3×3, then the inverse matrix of the camera intrinsic parameter matrix is ​​K -1 , multiply the inverse matrix of the camera intrinsic parameter matrix by the homogeneous coordinate matrix to convert the two-dimensional pixel coordinates into three-dimensional coordinates, and obtain the local three-dimensional coordinate matrix P of the depth image p with a dimension of 3×N C ,

[0137] P C =K -1 P 2D

[0138] The expanded matrix multiplication is expressed as follows:

[0139]

[0140] Among them, X i 、Y i and Z i k are the x-axis, y-axis, and z-axis coordinate values ​​corresponding to the i-th pixel after conversion in the local three-dimensional coordinate system of the depth image p. -1 ij is the inverse matrix K of the camera intrinsic parameter matrix -1 The element in the i-th row and j-th column, i = 1, 2, 3, j = 1, 2, 3.

[0141] X i 、Y i and Z i The calculation formula is:

[0142]

[0143] In computer vision, the camera intrinsic parameter matrix is ​​a 3×3 matrix used to describe the inherent characteristics of a camera, usually denoted by K. It mathematically describes the internal geometric and optical properties of the camera as it maps points in three-dimensional space to the two-dimensional image plane. Its standard form is:

[0144]

[0145] where f x and f y are the camera's focal length in the x and y directions, c x and c y are the principal point coordinates in the x direction and the principal point coordinates in the y direction respectively.

[0146] A three-dimensional coordinate system is used to describe the position, orientation, and relationship of objects in three-dimensional space. It uses three mutually perpendicular axes—the x-axis, y-axis, and z-axis—to construct a reference framework for the entire three-dimensional space. This allows any point in space to be precisely located using a unique combination of coordinate values. The coordinate axis orientations are typically determined using the right-hand rule. Extend your right hand, with your thumb pointing toward the positive x-axis, your index finger pointing toward the positive y-axis, and your middle finger extended naturally and perpendicular to the plane defined by your thumb and index finger. The direction indicated by your middle finger is the positive z-axis. These three coordinate axes are mutually perpendicular and define the entire three-dimensional space. The x-axis typically represents the horizontal direction, or the left-right dimension. The y-axis typically represents the vertical direction, or the up-down dimension. The z-axis, following the right-hand rule, primarily represents depth, or the front-to-back dimension. It is often used to indicate the distance of an object from a reference plane (such as a camera imaging plane) or an observation point.

[0147] Therefore, a three-dimensional coordinate point p in the three-dimensional coordinate system can be represented by the three-dimensional coordinates [x p ,y p ,z p ] T Describes the position of the point in three-dimensional space, where x p Indicates the coordinate value of point p on the x-axis (horizontal direction, left and right dimensions), y p Indicates the coordinate value of point p on the y-axis (vertical direction, up and down dimension), z p Represents the coordinate value of point p on the z-axis (depth direction, front-to-back dimension).

[0148] A 3D coordinate matrix is ​​a matrix composed of 3D coordinates. It is used to describe the positions of multiple 3D coordinate points in 3D space.

[0149] Local three-dimensional coordinates are a type of three-dimensional coordinates, which are the coordinate representations in a three-dimensional coordinate system established for a specific local area or a specific object.

[0150] Integrate depth information; assume that the depth value corresponding to the i-th pixel in the depth image p is z i , then use a 1×N row vector Z = [z1, z2, …, z N to represent the depth information of all pixels, and multiply the local three-dimensional coordinate matrix P C of the depth image p by the depth information row vector Z to obtain the local three-dimensional coordinate matrix P world of the depth image p with a dimension of 3×N. The calculation formula is as follows:

[0151] P world = P C ⊙ Z

[0152] Here, ⊙ represents element-wise multiplication, that is:

[0153]

[0154] Among them, X world,i , Y world,i and Z world,i are respectively the coordinate values of the x-axis, y-axis and z-axis of the local three-dimensional coordinate system corresponding to the i-th pixel in the depth image p after integrating depth information.

[0155] Screen according to the effective depth range; assume that the effective depth range of the depth image is [min_depth, max_depth], and use this range to screen out the effective three-dimensional coordinates within the effective depth range. Then construct a 1×N mask vector M = [m1, m2, …, m N , and its elements are obtained by the following condition judgment:

[0156]

[0157] For the local three-dimensional coordinate matrix P world of the depth image p, screen through this mask vector to obtain the local three-dimensional coordinate matrix P valid of the effective depth image p with a dimension of 3×N valid , N valid is the number of effective coordinates after screening, and N valid < N. The specific screening method is expressed as:

[0158]

[0159] That is, retain the local three-dimensional coordinates of the depth image p corresponding to the elements m i = 1 in the mask vector M, and discard mi The local three-dimensional coordinates of the depth image p where

[0160] Reduce the coordinate data according to the sampling rate; let the sampling rate be s, 0 < s ≤ 1, and after sampling, obtain the local three-dimensional coordinate matrix P of the final depth image p with dimension 3×M final , where M is the number of remaining coordinates after sampling, and M < N valid . Sample at equal intervals in the following manner:

[0161]

[0162] Here denotes slicing and selection of a two-dimensional matrix in the column direction with a step size of .

[0163] Through the above steps, the local three-dimensional coordinates corresponding to each pixel in each depth image are obtained.

[0164] Obtain the global three-dimensional coordinates:

[0165] Let the world three-dimensional coordinates of the camera when shooting the i-th depth image u i be p i world = [x i world , y i world , z i world . [[ID=​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​The imaginary part of . Then construct the corresponding homogeneous transformation matrix T of the i-th depth image i The specific steps are as follows:

[0166] Using quaternion q i =(a i ,b i ,c i ,d i ) Construct the corresponding rotation matrix R i :

[0167]

[0168] Get the rotation matrix R i Then, construct the 4×4 homogeneous transformation matrix T according to the following formula i :

[0169]

[0170] So the specific homogeneous transformation matrix T i for:

[0171]

[0172] Calculate the basic matrix for coordinate system transformation; select the local 3D coordinate system corresponding to the first depth image as the global 3D coordinate system. The 3D coordinates of the camera corresponding to the first depth image are expressed as p1 = [x1 world ,y1 world ,z1 world ] T , the quaternion rotation state is q1 = (a1, b1, c1, d1), and accordingly, the homogeneous transformation matrix constructed is T1. Further, the inverse matrix T1 of the homogeneous transformation matrix T1 is calculated -1 , and the inverse matrix will be used as the basic matrix for various subsequent coordinate transformation operations, providing an important support for the coordinate transformation and other related processing involved in the entire subsequent process.

[0173] Calculate the global transformation matrix corresponding to each depth image; the camera three-dimensional coordinates corresponding to the i-th depth image u1 are p i =[x i world ,y i world ,z i world ] T , the quaternion rotation state is q i =(a i ,b i ,c i ,d i ). Accordingly, the homogeneous transformation matrix constructed is Ti , and the initial transformation inverse matrix T1 -1 Multiply them together to get the global transformation matrix T gi =T1 -1 T i ;

[0174] Coordinate transformation and acquisition of global three-dimensional coordinates; let the final depth image u i The local three-dimensional coordinate matrix is ​​U i,final , U i,final The jth coordinate in is p ui,j =[x ui,j ,y ui,j ,z ui,j ] T , convert it into the following homogeneous coordinate form

[0175]

[0176] Then use the global coordinate transformation matrix T gi and Multiply to achieve the depth image u i The transformation from the local three-dimensional coordinate system to the global three-dimensional coordinate system gives the global three-dimensional coordinates in homogeneous coordinates as follows:

[0177]

[0178] Finally, extract the first three rows of data as the three-dimensional coordinates p in the global three-dimensional coordinate system global,i,j :

[0179]

[0180] where x global,i,j 、y global,i,j and z global,i,j They are the coordinate values ​​of the x-axis, y-axis, and z-axis in the global three-dimensional coordinate system.

[0181] Through the above steps, the local three-dimensional coordinates corresponding to each depth image are converted into global three-dimensional coordinates.

[0182] Acquisition of color information and semantic information:

[0183] Coordinate transformation; for depth image u i , its final depth image u i The local three-dimensional coordinate matrix is ​​U i,final On this basis, let’s assume that U i,final The jth coordinate in is p ui,j =[x ui,j ,y ui,j ,z ui,j ]T .

[0184] At the same time, with the depth image u i Also captured is a color image u i,rgb , and for color image u i,rgb Correspondingly, a semantic image u with the same height, width and pixel arrangement is generated i,semantic So suppose that when taking color image u i,rgb When , the corresponding camera intrinsic parameter matrix is ​​K rgb .

[0185] Since in constructing the semantic image u i,semantic In the process, the pixel position is not changed, that is, the semantic image u i,semantic With color image u i,rgb The spatial layout at the pixel level is consistent, so the camera intrinsic parameter matrices corresponding to the two are the same, that is, the semantic image u i,semantic The corresponding camera intrinsic parameter matrix is ​​also K rgb Moreover, since the intrinsic parameter matrices are the same, the corresponding coordinate transformation results should also remain consistent when performing coordinate transformation operations.

[0186] Next, the depth image u is transformed into i The local three-dimensional coordinate p ui,j Convert to color image u i,rgb The camera coordinates in the camera coordinate system are expressed as follows:

[0187] P cam,ui,j rgb =K rgb P ui,j =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T

[0188] And because the semantic image u i,semantic With color image u i,rgb The coordinate transformation results are also consistent, so the semantic image u i,semantic The camera coordinates in the camera coordinate system are expressed as follows:

[0189] P cam,ui,j semantic =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T .

[0190] Extract depth information; extract color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic The coordinate value z on the z-axis cam,ui,j rgb , as depth information.

[0191] Perspective division and homogeneous coordinate transformation; for color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic Perform perspective division and divide each coordinate component by z cam,ui,j rgb , get the color image u under homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic :

[0192]

[0193] Calculate the two-dimensional plane coordinates and extract color and semantic information; extract the color image u under homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic The first two rows of data are as follows: i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system ui,j rgb and P ui,j semantic :

[0194]

[0195] where x ui,j rgb and y ui,j rgb are color images u i,rgb and semantic image u i,semantic The x-axis and y-axis coordinate values ​​in the two-dimensional pixel coordinate system.

[0196] Then use P ui,j rgb and P ui,j semantic , from the color image u of dimension H×W×3 i,rgb and a semantic image u of dimension H×W×1 i,semantic Extract the depth image u i The three-dimensional coordinate p in the local three-dimensional coordinate system ui,j Color information C RGB and semantic information S semantic , the formula is as follows:

[0197] C RGB =u i,rgb (x ui,j rgb ,y ui,j rgb ,:)

[0198] S semantic =u i,semantic (x ui,j rgb ,y ui,j rgb ,:)

[0199] Among them, H is the height of the color image and semantic image, W is the width of the color image and semantic image, and the semantic image u is the width of the color image and semantic image. i,semantic The semantic information in is the index of the item category list.

[0200] 4. Build a raster map

[0201] The grids are divided into idle and occupied states. The grid size and two-dimensional scale are determined according to needs, as well as the information dictionary used to store the color information and semantic information of the grid in the occupied state. The three-dimensional coordinates are mapped to the corresponding three-dimensional grid by using flooring and coordinate offset. At the same time, the grid coordinates of the grid in the occupied state are used as the key to fill in the corresponding color information and semantic information in the information dictionary.

[0202] Mapping steps:

[0203] For the global three-dimensional coordinate p obtained in the previous step global,i,j =[x global,i,j ,y global,i,j ,z global,i,j ] T , and map it to the corresponding three-dimensional grid coordinates through the following steps:

[0204]

[0205] Where row, col and height are the coordinate values ​​of the row, column and height directions in the grid respectively. w and G h are the number of rows and columns of the grid, respectively, and cs is the side length of the grid.

[0206] The method of filling the information dictionary is as follows:

[0207] When a grid is determined to be occupied, its corresponding three-dimensional grid coordinates row, col and height are used as keys to convert the color information C corresponding to the grid into RGB and semantic information S semantic Store it in the information dictionary to complete the information filling operation during the raster map construction process.

[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a visual semantic grid map based on multimodality and pre-training models, characterized by: The following steps are involved: A. Collecting Data In the robot exploration phase, the RGB-D camera with known camera intrinsic parameters is used to obtain depth images and color images, and the camera position and posture information is obtained through the RGB-D real-time positioning and mapping system; B. Using pre-trained models to obtain semantic images The semantic image is obtained by using the language-driven semantic segmentation model LSeg, the visual language model BLIP and the color image; C. Obtain global and local 3D coordinates and their corresponding color and semantic information Use depth images, camera intrinsic parameters, color images and semantic images to obtain global and local 3D coordinates and their corresponding color and semantic information; D. Build a raster map Through the coordinate conversion operation, the global three-dimensional coordinates are mapped to the three-dimensional grid, and the color and semantic information of the corresponding grid are filled in the information dictionary.

2. A method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for collecting data in step A is as follows: when the robot enters an unknown environment for exploration, the robot acquires depth images and color images through an RGB-D camera with known camera intrinsic parameters, and acquires the world three-dimensional coordinates and quaternion rotation states of the camera positions corresponding to all images through an RGB-D real-time positioning and mapping system; The camera intrinsic parameters are parameters that describe the internal properties of the camera, including focal length and principal point coordinates.

3. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The pre-trained model in step B includes a visual language model BLIP and a language-driven semantic segmentation model LSeg.

4. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for obtaining the semantic image described in step B is as follows: first, use the visual language model BLIP to ask a template question of "What's in the room?" to each color image, so as to obtain all the object categories of the environment; then combine the object category with the language-driven semantic segmentation model LSeg to perform semantic segmentation on each color image, so as to assign an index of a corresponding object category list to each pixel, and thus link the object category semantic information with the pixels of the color image, so as to construct a semantic image corresponding to each color image.

5. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for obtaining global and local three-dimensional coordinates and their corresponding color and semantic information described in step C is as follows: the local three-dimensional coordinates of each pixel are obtained through the depth image and the camera intrinsic parameters, and then the local three-dimensional coordinates are transformed into global three-dimensional coordinates using the camera position and rotation state; then the camera intrinsic parameters and the local three-dimensional coordinates are used to obtain the color and semantic information of the corresponding pixels in the corresponding color image and semantic image, thereby completing the fusion of the semantic and color information and the depth image.

6. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for obtaining local three-dimensional coordinates in step C comprises the following steps: C11. Merge into homogeneous coordinates Suppose that in the two-dimensional pixel coordinate system corresponding to the depth image p, the x-axis is the horizontal direction of the depth image, the right direction is positive, and the y-axis is the vertical direction of the depth image, the downward direction is positive; the width of the depth image is w, the height is h, and the total number of pixels is N; then the i-th pixel in the depth image is represented by the two-dimensional pixel coordinates [x i ,y i ] T represents its position in the depth image plane, where x i Indicates the coordinate value of the pixel on the x-axis, y i Indicates the coordinate value of the pixel on the y-axis, where 0≤x i ≤w-1, 0≤x i ≤h-1, 1≤i≤N; According to the construction rules of homogeneous coordinates, the homogeneous coordinate matrix P of the depth image p with a dimension of 3×N is obtained 2D , the homogeneous coordinate matrix is ​​expressed as follows: Homogeneous coordinate matrix P 2D The 1 in the third row of indicates that an extra dimension is added to the coordinate representation of each pixel; C12. Convert to three-dimensional coordinates Assume that when shooting a depth image, the corresponding camera intrinsic parameter matrix is ​​K and the dimension is 3×3, then the inverse matrix of the camera intrinsic parameter matrix is ​​K -1 , the inverse matrix of the camera intrinsic parameter matrix is ​​multiplied by the homogeneous coordinate matrix to convert the two-dimensional pixel coordinates into three-dimensional coordinates, and the local three-dimensional coordinate matrix P of the depth image p with a dimension of 3×N is obtained C as follows: P C =K -1 P 2D The expanded matrix multiplication is expressed as follows: Among them, X i , Y i and Z i are the x-axis, y-axis and z-axis coordinate values ​​corresponding to the i-th pixel after conversion in the local three-dimensional coordinate system of the depth image p; k -1 ij is the inverse matrix K of the camera intrinsic parameter matrix -1 The element in the i-th row and j-th column, i = 1, 2, 3, j = 1, 2, 3; X i , Y i and Z i The calculation formula is: C13. Incorporating depth information Let the depth value corresponding to the i-th pixel in the depth image p be z i , then use a 1×N row vector Z=[z1,z2,…,z N ] represents the depth information of all pixels, and the local three-dimensional coordinate matrix P of the depth image p C Multiply it with the depth information row vector Z to obtain the local three-dimensional coordinate matrix P of the depth image p with a dimension of 3×N world , the calculation formula is as follows: P world =P C ⊙Z Here ⊙ represents element-by-element multiplication, that is: Where X world,i , Y world,i and Z world,i are respectively the coordinate values ​​of the x-axis, y-axis and z-axis of the local three-dimensional coordinate system corresponding to the ith pixel in the depth image p after integrating the depth information; C14. Screening based on effective depth range Assume that the effective depth range of the depth image is [min_depth, max_depth], and use this range to filter out the valid 3D coordinates within the effective depth range, then construct a mask vector M = [m1, m2, ..., m N ], whose elements are determined by the following conditions: For the local three-dimensional coordinate matrix P of the depth image p world , filtering is performed through this mask vector to obtain a local three-dimensional coordinate matrix P of the valid depth image p with a dimension of 3×N valid , where N valid is the number of valid coordinates after filtering, and N valid <N; the specific filtering method is expressed as:​​ That is, keep the corresponding element m in the mask vector M i = 1, and discard m i = 0; C15. Reduce coordinate data according to sampling rate Let the sampling rate be s, where 0 < s ≤ 1. After sampling, the local three-dimensional coordinate matrix P of the final depth image p with dimension 3×M is obtained final , where M is the number of remaining coordinates after sampling, and M < N valid ; Sampling is performed at equal intervals in the following manner: here It means that the two-dimensional matrix is Perform slice selection; Through the above steps, the local three-dimensional coordinates corresponding to each pixel in each depth image are obtained.

7. A method for constructing a visual semantic grid map based on multimodality and pre-training models according to claim 6, characterized in that: The standard form of the camera intrinsic parameter matrix K in step C12 is: where f x and f y are the focal lengths of the camera in the x and y directions, respectively, and c x and c y are the principal point coordinates in the x direction and the principal point coordinates in the y direction respectively.

8. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The standard form of the camera intrinsic parameter matrix K in step C12 is: where f x and f y are the focal lengths of the camera in the x and y directions, respectively, and c x and c y are the principal point coordinates in the x direction and the principal point coordinates in the y direction respectively; The method for obtaining the global three-dimensional coordinates in step C comprises the following steps: C21. Construct a homogeneous transformation matrix for each depth image Suppose the i-th depth image u is taken i The world 3D coordinates of the camera at this time are p i world =[x i world ,y i world ,z i world ] T , the quaternion rotation state is q i =(a i ,b i ,c i ,d i ); where x i world ,y i world ,z i world To take the i-th depth image u i When , the coordinate values ​​of the x-axis, y-axis, and z-axis in the world three-dimensional coordinate system of the camera are calculated by the RGB-D real-time positioning and mapping system; q i =(a i ,b i ,c i ,d i ) is the i-th depth image u i The rotation state in the world three-dimensional coordinate system calculated by the RGB-D real-time positioning and mapping system when a i is the real part of the quaternion, b i ,c i ,d i is the quaternion q i The imaginary part of the i-th depth image is constructed by using the homogeneous transformation matrix T i The specific steps are as follows: Using quaternion q i =(a i ,b i ,c i ,d i ) Construct the corresponding rotation matrix R i : Get the rotation matrix R i Then, construct a 4×4 homogeneous transformation matrix T according to the following formula i : The homogeneous transformation matrix T of the unfolded representation i for: C22. Calculate the basic matrix for coordinate system transformation The local 3D coordinate system corresponding to the first depth image u1 is selected as the global 3D coordinate system; the 3D coordinate of the camera corresponding to the first depth image u1 is expressed as p1 = [x1 world ,y1 world ,z1 world ] T , the quaternion rotation state is q1 = (a1, b1, c1, d1), and accordingly, the homogeneous transformation matrix constructed is T1. Further, the inverse matrix T1 of the homogeneous transformation matrix T1 is calculated -1 , and the inverse matrix will be used as the basic matrix for various subsequent coordinate transformation operations, providing an important basis for the coordinate transformation and other related processing involved in the entire subsequent process; C23. Calculate the global transformation matrix corresponding to each depth image The i-th depth image u i The corresponding camera 3D coordinates are p i =[x i world ,y i world ,z i world ] T , the quaternion rotation state is q i =(a i ,b i ,c i ,d i ); accordingly, the homogeneous transformation matrix constructed is T i , and the initial transformation inverse matrix T1 -1 Multiply them together to get the global transformation matrix T gi =T1 -1 T i ; C24. Coordinate transformation and obtaining global three-dimensional coordinates Let the depth image u i The final local three-dimensional coordinate matrix is ​​U i,final , U i,final The jth coordinate in is p ui,j =[x ui,j ,y ui,j ,z ui,j ] T , convert it into the following homogeneous coordinate form Then, using the global coordinate transformation matrix T gi and Multiply to achieve the depth image u i The transformation from the local three-dimensional coordinate system to the global three-dimensional coordinate system, the global three-dimensional coordinates under homogeneous coordinates are as follows: Finally, the first three rows of data are extracted as the three-dimensional coordinates p in the global three-dimensional coordinate system. global,i,j : where x global,i,j ,y global,i,j and z global,i,j are the coordinate values ​​of the x-axis, y-axis, and z-axis in the global three-dimensional coordinate system respectively; Through the above steps, the local three-dimensional coordinates corresponding to each depth image are converted into global three-dimensional coordinates.

9. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for obtaining color information and semantic information in step C comprises the following steps: C31. Coordinate transformation Assume that U in step C24 i,final The jth coordinate in is At the same time, with the depth image u i Also captured is a color image u i,rgb , and for color image u i,rgb Correspondingly, semantic images u with the same height, width and pixel arrangement are generated i,semantic ; So suppose that when taking a color image u i,rgb When , the corresponding camera internal parameter matrix is ​​K rgb ; Since in constructing the semantic image u i,semantic In the process, the pixel position is not changed, that is, the semantic image u i,semantic With color image u i,rgb The spatial layout at the pixel level is consistent, so the camera intrinsic parameter matrices corresponding to the two are the same, that is, the semantic image u i,semantic The corresponding camera intrinsic parameter matrix is ​​also K rgb ; Moreover, since the internal parameter matrix is ​​the same, when performing coordinate transformation operations, the corresponding coordinate transformation results should also be consistent; Next, the depth image u is transformed by matrix multiplication. i The local three-dimensional coordinates p in the local three-dimensional coordinate system ui,j Convert to color image i,rgb The camera coordinates in the camera coordinate system are expressed as follows: P cam,ui,j rgb =K rgb P ui,j =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T And because the semantic image u i,semantic With color image u i,rgb The coordinate transformation results are also consistent, so the semantic image u i,semantic The camera coordinates in the camera coordinate system are expressed as follows: P cam,ui,j semantic =[x cam,ui,j rgb ,y cam,ui,j rgb ,z cam,ui,j rgb ] T ; C32. Extracting depth information Extract color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic The coordinate value z on the z-axis cam,ui,j rgb , as depth information; C33, perspective division and homogeneous coordinate transformation For color image u i,rgb and semantic image u i,semantic The camera coordinate P in the camera coordinate system cam,ui,j rgb and P cam,ui,j semantic Perform perspective division and divide each coordinate component by z cam,ui,j rgb , get the color image u in homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic : C34, calculate two-dimensional plane coordinates and extract color and semantic information Using the result of step C33, extract the color image u under homogeneous coordinates i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system img,ui,j rgb and P img,ui,j semantic The first two rows of data are used as the color image u i,rgb and semantic image u i,semantic The two-dimensional pixel coordinate P in the two-dimensional pixel coordinate system ui,j rgb and P ui,j semantic : where x ui,j rgb and ui,j rgb are color images u i,rgb and semantic image u i,semantic The coordinate values ​​of the x-axis and y-axis in the two-dimensional pixel coordinate system; Then use P ui,j rgb and P ui,j semantic , from a color image u of dimension H×W×3 i,rgb and a semantic image u of dimension H×W×1 i,semantic Extract the depth image u i The three-dimensional coordinates p in the local three-dimensional coordinate system ui,j Color information C RGB and semantic information S semantic , the formula is as follows: C RGB =u i,rgb (x ui,j rgb ,y ui,j rgb ,:) S semantic =u i,semantic (x ui,j rgb ,y ui,j rgb ,:) Among them, H is the height of the color image and the semantic image, W is the width of the color image and the semantic image, and the semantic image u i,semantic The semantic information in is the index of the item category list.

10. The method for constructing a visual semantic grid map based on multimodality and pre-training model according to claim 1, characterized in that: The method for constructing the grid map in step D is as follows: The grids are divided into an idle state and an occupied state, and the size and two-dimensional scale of the grids and an information dictionary for storing color information and semantic information of the grids in the occupied state are determined according to the requirements; the three-dimensional coordinates are mapped to the corresponding three-dimensional grids by rounding down and offsetting the coordinates, and the corresponding color information and semantic information are filled in the information dictionary with the grid coordinates of the grid in the occupied state as the key; The mapping steps are as follows: For the global three-dimensional coordinates obtained in step C24 Map it to the corresponding 3D grid coordinates by following the steps below: Where row, col and height are the coordinate values ​​of the row, column and height in the grid respectively. w and G h are the number of rows and columns of the grid, respectively, and cs is the length of the grid side; The method of filling the information dictionary is as follows: When a grid is determined to be occupied, its corresponding three-dimensional grid coordinates row, col and height are used as keys to convert the color information C corresponding to the grid into RGB and semantic information S semantic Stored in the information dictionary, complete the information filling operation during the raster map construction process.

Citation Information

Patent Citations

  • Topological map generation method based on visual fusion landmarks

    CN111210518A

  • Real-time map construction method and device fusing multi-modal asynchronous time data

    CN116592870A