Semantic segmentation method, apparatus, and system, as well as drilling machine
The integration of camera and radar data through multi-sensor calibration and semantic segmentation models addresses the challenge of environmental perception in excavators, enabling precise object detection and real-time control for unmanned operations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- JIANGSU XCMG STATE KEY LAB TECH CO LTD
- Filing Date
- 2024-11-06
- Publication Date
- 2026-04-15
AI Technical Summary
Accurately perceiving environmental information in the working area of an excavator is crucial for unmanned excavation, but existing systems struggle to integrate and process diverse sensory data effectively, leading to incomplete environmental awareness.
A semantic segmentation method that combines image information from cameras with radar point cloud data from lidar and four-dimensional millimeter-wave radar, using multi-sensor calibration and interactive semantic segmentation models to enhance environmental perception, enabling accurate detection of objects and environmental features.
Enables comprehensive and accurate environmental perception, providing real-time dynamic information for automated decision-making and control of excavators, enhancing their operational capabilities in various conditions.
Smart Images

Figure 2026512208000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application is based on Chinese Patent Application No. 202410220047.2 filed on February 27, 2024, and claims priority to this application. The disclosure of this application is hereby incorporated by reference in its entirety into this disclosure.
[0002] Technical Field This disclosure relates to the field of intelligentization and automation of construction machinery, and particularly to semantic segmentation methods, devices, and systems, as well as excavators.
Background Art
[0003] An excavator is a multi - functional construction machine and is widely applied to on - site construction in fields such as mining, water control engineering, transportation, and power engineering. An unmanned excavator can replace an excavator operator to carry out operations and construction at sites where landslides, toxic and harmful gases may occur, and at the same time to alleviate the problem of labor shortage caused by the aging of modern society. Therefore, the research and development of an excavator with an unmanned control function has great demand and practical significance in the industry. Cognizing the environmental information in the working area of an excavator is the basis for realizing unmanned excavation work. Accurately cognizing the environmental information in the working area of an excavator has become an urgent issue to be solved.
Summary of the Invention
Means for Solving the Problems
[0004] According to certain embodiments of the present disclosure, a semantic segmentation method is provided. This semantic segmentation method includes the steps of: acquiring image information and radar point cloud information of a work environment; determining a first feature representation of the image information; determining a second feature representation of the object's labeling information in the radar point cloud information based on the image information and the object's labeling information; and fusing the first feature representation of the image information with the second feature representation of the object's labeling information in order to extract the object from the image information.
[0005] In some embodiments, the step of acquiring image information and radar point cloud information of a work environment includes acquiring image information of the work environment using a camera device, acquiring first radar point cloud information of the work environment using a lidar, and / or acquiring second radar point cloud information of the work environment using a four-dimensional millimeter-wave radar.
[0006] In some embodiments, the semantic segmentation method further includes the step of determining object marker information, determining transformation matrices between the coordinate system of the excavator and the image coordinate system of the camera device, between the coordinate system of the excavator and the coordinate system of the lidar, and between the coordinate system of the excavator and the coordinate system of the four-dimensional millimeter-wave radar, respectively, depending on the installation position and angle of the camera device, lidar, and four-dimensional millimeter-wave radar on the excavator, determining the second marker information in the image coordinate system of the first marker information of the object in the first radar point cloud information, and / or determining the fourth marker information in the image coordinate system of the third marker information of the object in the second radar point cloud information, and determining the object marker information by fusing the original fifth marker information in the image coordinate system of the object in the image information with the second marker information in the image coordinate system and / or the fourth marker information in the image coordinate system.
[0007] In some embodiments, in order to determine the object's labeling information, merging the original fifth labeling information of the object in the image coordinate system with the second labeling information and / or fourth labeling information in the image coordinate system includes any of the following: performing a weighted average of the fifth labeling information and the second labeling information in the image coordinate system to obtain average labeling information as the object's labeling information; performing a weighted average of the fifth labeling information and the fourth labeling information in the image coordinate system to obtain average labeling information as the object's labeling information; and performing a weighted average of the fifth labeling information, the second labeling information and the fourth labeling information to obtain average labeling information as the object's labeling information, wherein the respective weights for weighting the fifth labeling information, the second labeling information and the fourth labeling information are configured to be the same, or are determined according to the mapping capabilities of the camera device, lidar and four-dimensional millimeter-wave radar under the current working conditions.
[0008] In some embodiments, in order to determine the object's labeling information, merging the original fifth labeling information of the object in the image coordinate system with the second labeling information and / or fourth labeling information in the image coordinate system includes selecting the labeling information corresponding to the best mapping quality as the object's labeling information, depending on the mapping quality of the camera device, lidar, and four-dimensional millimeter-wave radar under current working conditions, from the fifth and second labeling information in the image coordinate system, or from the fifth, second, and fourth labeling information in the image coordinate system.
[0009] In some embodiments, the step of determining a first feature representation of image information includes converting the image into a fixed-size image block, performing feature coding on this image block, and reducing the dimension of the feature coding data to a fixed-size first feature representation.
[0010] In some embodiments, the step of determining a second feature representation of the object's labeling information includes encoding each label point represented by the object's labeling information in order to obtain a second feature representation of the object's labeling information.
[0011] In some embodiments, the semantic segmentation method further includes the steps of determining mask information of an object based on image information, and determining a third feature representation of the mask information of an object based on the mask information of the object. The merging step includes merging the first feature representation of the image information with the second feature representation of the object's marking information and the third feature representation of the object's mask information in order to extract the object from the image information.
[0012] In some embodiments, the step of determining mask information of an object based on image information includes performing semantic segmentation based on image information to obtain an image region of the object, and generating mask information of the object based on the image region of the object.
[0013] In some embodiments, the step of determining a third feature representation of the object's mask information includes performing a convolution operation on the object's mask information and encoding it in order to obtain the third feature representation of the object's mask information.
[0014] According to some embodiments of the present disclosure, a semantic segmentation apparatus is provided. This semantic segmentation apparatus comprises a memory and a processor coupled to the memory, the processor being configured to perform a semantic segmentation method based on instructions stored in the memory.
[0015] According to certain embodiments of this disclosure, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program. When executed by a processor, this computer program embodies a semantic segmentation method.
[0016] According to certain embodiments of this disclosure, a computer program product is provided. This computer program product comprises a computer program, which, when executed by a processor, embodies a semantic segmentation method.
[0017] According to certain embodiments of the present disclosure, a semantic segmentation system is provided. This semantic segmentation system comprises a camera device for collecting image information of a work environment; a radar having a lidar for collecting first radar point cloud information of the work environment and / or a four-dimensional millimeter-wave radar for collecting second radar point cloud information of the work environment; and a semantic segmentation device configured to carry out a semantic segmentation method.
[0018] According to certain embodiments of this disclosure, an excavator is provided, which includes a semantic segmentation system.
[0019] The following briefly introduces the accompanying drawings necessary for describing the embodiments or related technologies. This disclosure can be better understood by the following detailed description with reference to the accompanying drawings.
[0020] The accompanying drawings in the following description are, needless to say, only some embodiments of the present disclosure. Those skilled in the art can obtain other accompanying drawings from these without inventive effort. [Brief explanation of the drawing]
[0021] [Figure 1] This diagram shows a schematic representation of the deployment of multi-sensors in an excavator and a schematic representation of the corresponding coordinate system according to some embodiments of this disclosure. [Figure 2] This is a schematic diagram of an imaging model of a camera device in the construction area of an excavator according to some embodiments of this disclosure. [Figure 3]Schematic flowchart of an interactive semantic segmentation method according to some embodiments of the present disclosure. [Figure 4] Schematic diagram of an interactive semantic segmentation model according to some embodiments of the present disclosure. [Figure 5] Schematic diagram of an interactive semantic segmentation method according to some embodiments of the present disclosure. [Figure 6] Schematic structural diagram of a semantic segmentation apparatus according to some embodiments of the present disclosure. [Figure 7] Schematic diagram of a semantic segmentation system according to some embodiments of the present disclosure. **Embodiments for Carrying Out the Invention**
[0022] It should be noted that the relative arrangements, mathematical formulas, and numerical values of the members and steps described in these embodiments do not limit the scope of the present disclosure and the appended claims unless otherwise specified.
[0023] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish multiple different steps, devices, or modules, and do not represent any specific technical meaning or an inevitable logical order between them.
[0024] It should also be understood that in the embodiments of the present disclosure, "a plurality of" can refer to two or more, and "at least one" can refer to one, two, or more.
[0025] It should also be understood that any component, data, or structure mentioned in the embodiments of the present disclosure is usually understood to be one or more unless explicitly defined or otherwise indicated in the context.
[0026] In addition, the term "and / or" in this disclosure is merely a correspondence to describe multiple associated objects. The term indicates that three relationships exist; for example, A and / or B indicates three situations: A exists alone, A and B exist together, and B exists alone. In addition, the letter " / " in this disclosure typically indicates an "or" relationship between multiple associated objects in a given context.
[0027] The descriptions of the various embodiments in this disclosure focus on highlighting the differences between the various embodiments, and it should be understood that their same or similar aspects may be referenced to one another, but are not described in detail one by one for the sake of simplification.
[0028] At the same time, for the sake of clarity, please understand that the dimensions of the various parts shown in the attached drawings are not depicted according to actual proportional relationships.
[0029] The following description of at least one exemplary embodiment is indeed illustrative and is not in any way limiting the present disclosure or its use or application.
[0030] Methods, techniques, and apparatus known to those skilled in the art may not be described in detail. However, these methods, techniques, and apparatus shall be considered as part of this specification where appropriate.
[0031] In the attached drawings, the same reference numerals and letters indicate the same items. Therefore, please note that once an item is defined in one attached drawing, further description of that item is not necessary in subsequent attached drawings.
[0032] In addition, to avoid obscuring this disclosure with unnecessary details, only processing steps and / or apparatus structures closely related to the solutions described herein are shown in the accompanying drawings, and other details less relevant to this disclosure are omitted. Similar reference numerals and letters in the accompanying drawings indicate similar items. Therefore, it should be noted that once an item is defined in one accompanying drawing, it is not necessary to elaborate on that item for subsequent accompanying drawings.
[0033] According to some embodiments of this disclosure, by comprehensively detecting multidimensional sensing information such as image information and radar information of the work environment, and by fusing feature representations of image information with feature representations of marked objects in the image information and radar information for semantic segmentation, objects can be detected more accurately from the image based on prompts for marked objects, so that the excavator can more comprehensively perceive the environment within the work area and provide a foundation for unmanned excavation work.
[0034] The semantic segmentation recognition system for smart drilling equipment in this disclosure comprises a multi-sensor hardware matching solution and a software algorithm architecture design solution, multi-sensor simultaneous calibration, and interactive semantic segmentation technology. The implementation methods of each part are detailed below. To achieve spatiotemporal synchronization of multiple types of sensors, a multi-sensor simultaneous calibration solution is designed using multiple sensors such as lidar, four-dimensional (4D) millimeter-wave radar, and cameras. Interactive semantic segmentation of laser point cloud data, 4D millimeter-wave radar point cloud data, and multiple images is achieved using a prompt-fused semantic segmentation model to realize interactive semantic segmentation of the surface shape and size of the environment as well as the surface texture of the environment. By fully utilizing the advantages of each of the lidar, 4D millimeter-wave radar, and camera devices, accurate semantic segmentation boundaries are obtained for the recognition system, providing effective support for comprehensive recognition of the environment in multiple different work scenes. While obtaining the three-dimensional coordinates of an object, it is also possible to obtain features such as the object's texture, appearance, and actual height, thus enabling the estimation of the object's type and providing a crucial foundation for decision-making planning of the excavator. Based on an interactive semantic segmentation method, it is possible to provide accurate, real-time dynamic information of the segmentation area. This information provides a sufficient cognitive foundation for the automated and smart planning and control of the excavator.
[0035] The hardware matching solution for multiple sensors is described below.
[0036] The first step in establishing the architecture of the perception system for smart drilling equipment is to select sensor types and implement hardware configuration solutions according to the working scenario of the unmanned drilling machine. The multiple sensors selected in the sensor hardware configuration solution for the unmanned drilling machine are mainly classified into three categories: lidar, camera devices, and 4D millimeter-wave radar.
[0037] Low-cost camera devices, such as cameras and video cameras, capable of identifying multiple different objects, offer advantages in areas such as accuracy in measuring object height and width, recognizing lane markings, and recognizing pedestrians. Therefore, they are essential sensors for functions such as obstacle recognition and road area detection. However, their operating range and distance measurement accuracy are inferior to 4D millimeter-wave radar, and they are easily affected by factors such as lighting and weather.
[0038] Lidar offers advantages such as high measurement accuracy and a wide scanning range, but its distance measurement is limited and susceptible to illumination. The azimuth-pitch resolution of a lidar can reach 0.1*0.1 degrees.
[0039] 4D millimeter-wave radar, an upgraded version of traditional millimeter-wave radar, offers better detection capabilities and higher resolution and accuracy. 4D refers to four dimensions: velocity, distance, horizontal angle, and vertical height. Compared to traditional three-dimensional (3D) millimeter-wave radar, 4D millimeter-wave radar adds height detection, integrating a fourth dimension into the traditional millimeter-wave radar. This gives 4D millimeter-wave radar the following advantages: It offers richer dimensions for acquiring information, including the ability to measure pitch angles; angular resolution can reach approximately 1 degree; longer detection range; maximum detection distance can exceed 300 meters; and a higher density of point clouds, enabling the formation of point cloud imaging-level outputs and data-driven image recognition. Furthermore, 4D millimeter-wave radar overcomes the effects of rain, haze, and dust, improving detection reliability. Currently, the azimuth-pitch resolution of 4D millimeter-wave radar can reach approximately 1*1 degree. 4D millimeter-wave radar has limitations due to the Doppler effect; that is, it still has shortcomings in recognizing scenes such as objects, vehicles, and pedestrians moving laterally, as well as vehicles at close range. Generally, 4D millimeter-wave radar can provide more reliable data support for the planning and control of smart drilling facilities.
[0040] Sensors such as camera systems, lidar, and millimeter-wave radar each have their own advantages and disadvantages. The solution of selecting multiple types of sensors in this disclosure can enable the complementary advantages of different types of sensors, allowing for the simultaneous perception of more comprehensive environmental information. Lidar and camera systems operate under fair weather conditions, while 4D millimeter-wave radar and camera systems operate under harsh environmental conditions such as rain, fog, and dust, so the unmanned excavator can operate normally under a variety of working conditions.
[0041] During specific implementation, a multi-sensor matching solution is designed based on an evaluation of the advantages and disadvantages of the multi-sensor configurations described above. Each sensor configuration solution has its own optimal application scenario. A detailed user manual is created to make it easier for customers to select the optimal sensor configuration solution according to their actual application scenario. This sensor configuration solution includes sensor type, number of sensors to be installed, sensor installation location, and similar elements to enable quick and accurate selection of multi-sensor types.
[0042] The multi-sensor simultaneous calibration solution is described below.
[0043] Figure 1 shows a schematic diagram of the deployment of multisensors in an excavator, as well as a schematic diagram of the corresponding coordinate system according to some embodiments of this disclosure.
[0044] Figure 1 shows a schematic diagram of the deployment of the camera 11, four-dimensional millimeter-wave radar 12, and lidar 13 on the excavator 14. The camera device 11, such as a camera or video camera, is installed directly in front of the excavator when viewed from below and rotates with the rotation of the excavator's cab. The four-dimensional millimeter-wave radar 12 is installed directly in front of the excavator when viewed from below and rotates with the rotation of the excavator's cab. The lidar 13 is installed directly above the excavator.
[0045] In Figure 1, the coordinate system of the excavator is represented by Ow, Xw, Yw, and Zw; the coordinate system of the camera device (e.g., camera or video camera) is represented by Oc, Xc, Yc, and Zc; the coordinate system of the lidar is represented by Ot, Xt, Yt, and Zt; and the coordinate system of the four-dimensional millimeter-wave radar is represented by Om, Xm, Ym, and Zm. All of these coordinate systems are three-dimensional.
[0046] In the coordinate system of the excavator, the center point of the excavator's base at the center of rotation of the horizontal ground surface is considered to be the origin Ow, the Y-axis Yw is aligned directly in front of the excavator's cab, the Z-axis Zw is perpendicular to the horizontal ground surface and upward, and the X-axis Xw is perpendicular to the plane of the Yw and Zw axes and extends to the right along the direct front of the cab.
[0047] The XYZ axis directions of the lidar coordinate system and the excavator coordinate system are the same, but their origins are different. The origin Ot of the lidar coordinate system is located at the center of the lidar.
[0048] In the coordinate system of a camera device (e.g., a camera or video camera), the center of the camera device is the origin Oc, the forward angle shooting direction of the camera device is the Y-axis Yc, the Z-axis Zc is perpendicular to Yc and upward, and the X-axis Xc is perpendicular to the plane of the Yc and Zc axes and extends to the right along the area immediately in front of the driver's cab.
[0049] In the coordinate system of a four-dimensional millimeter-wave radar, the center of the radar is the origin Om, the forward angle imaging direction of the radar is the Y-axis Ym, the Z-axis Zm is perpendicular to Ym and upward, and the X-axis Xm is perpendicular to the plane of the Ym and Zm axes and extends to the right along the front of the driver's cab. Therefore, if the forward angle imaging direction of the four-dimensional millimeter-wave radar and the camera are the same, the directions of the XYZ axes in the coordinate system of the four-dimensional millimeter-wave radar and the coordinate system of the camera device coincide, with only the origin position differing.
[0050] To accurately establish the correspondence between the excavator's coordinate system and each pixel in the image, it is necessary to establish an imaging model of the camera device in the excavator's construction area, as shown in Figure 2.
[0051] Figure 2 shows the coordinate system of the excavator, represented by Ow, Xw, Yw, and Zw; the coordinate system of the camera device (e.g., camera or video camera), represented by Oc, Xc, Yc, and Zc; and the image coordinate system of the camera device, represented by the origin Oi. The U and V axes are perpendicular to each other. The image coordinate system is two-dimensional. Point Pw is a point on the material within the excavator's construction area. Pi is the point on the image plane corresponding to point Pw. The line connecting points Pw, Pi, and Oc is the projection line of point Pw on the excavator's construction contact surface on the image plane. This projection line is obtained by connecting points Pw and Oc. The intersection point Pi of the projection line and the image plane is the image point of point Pw on the image plane.
[0052] The transformation relationship between the three-dimensional coordinates (xw, yw, zw) of point Pw in the excavator's coordinate system and the two-dimensional coordinates (ui, vi) in the image coordinate system is given by the following equation.
[0053]
number
[0054] In the formula, Z is the scale factor, f is the focal length of the camera, dX and dY represent the physical lengths of the pixels on the photosensitive plate in the X and Y directions, respectively, (u0,v0) represents the coordinates of the center of the camera's photosensitive plate in the pixel coordinate system, respectively, R is the rotation matrix from the excavator's 3D coordinate system to the camera device's (e.g., camera or video camera) 3D coordinate system, t is the translation vector from the excavator's 3D coordinate system to the camera device's (e.g., camera or video camera) 3D coordinate system, K is the internal reference matrix of the camera device (e.g., camera or video camera), and T is the transformation matrix from the excavator's 3D coordinate system to the camera device's (e.g., camera or video camera) 3D coordinate system.
[0055] In the above equation, to obtain the normalized coordinates of Pw in the camera's coordinate system, we multiply 1 / Z by the three-dimensional excavator coordinates of the Pw point. These coordinates are located in the plane Z=1 in front of the camera. The transformation matrix consisting of R and t represents a rigid body transformation. The transformation from the camera's coordinate system to the image coordinate system is a perspective transformation, and the transformation from the image coordinate system to the pixel coordinate system is an affine transformation.
[0056] The transformation matrix between the lidar coordinate system and the excavator coordinate system is as follows:
[0057]
number
[0058] During the ceremony,
[0059] Rt is the rotation matrix from the lidar coordinate system to the excavator coordinate system.
[0060] Tt is the translation vector from the lidar's coordinate system to the excavator's coordinate system.
[0061] Xw, Yw, and Zw are the coordinate points in the excavator's coordinate system.
[0062] Xt, Yt, and Zt are the positions of coordinate points in the LiDAR coordinate system.
[0063] The transformation matrix between the coordinate system of the four-dimensional millimeter-wave radar and the coordinate system of the excavator is as follows:
[0064]
number
[0065] During the ceremony,
[0066] Rm is the rotation matrix from the coordinate system of the four-dimensional millimeter-wave radar to the coordinate system of the excavator.
[0067] Tm is the translation vector from the coordinate system of the four-dimensional millimeter-wave radar to the coordinate system of the excavator.
[0068] Xw, Yw, and Zw are the coordinate points in the excavator's coordinate system.
[0069] Xm, Ym, and Zm are the positions of coordinate points in the coordinate system of the four-dimensional millimeter-wave radar.
[0070] Based on a multi-sensor simultaneous calibration solution, data conversion from the camera, lidar, and four-dimensional millimeter-wave radar is integrated into the excavator's coordinate system to enable simultaneous calibration of the coordinate systems of multiple sensors. Therefore, the image coordinate system, the lidar coordinate system, and the four-dimensional millimeter-wave radar coordinate system can be converted to and from each other based on the excavator's coordinate system.
[0071] The software algorithm architecture design solution is described below.
[0072] Figure 3 shows a schematic flowchart of an interactive semantic segmentation method according to some embodiments of the present disclosure.
[0073] As shown in Figure 3, the interactive semantic segmentation method of this embodiment includes the following steps:
[0074] In step 31a, the camera device collects image information of the work environment.
[0075] In some embodiments, the camera device collects image information of the work environment based on the operation of control components such as a processor.
[0076] In step 31b, the lidar collects first radar point cloud information of the work environment.
[0077] In some embodiments, the lidar collects first radar point cloud information of the work environment based on the driving of a control component such as a processor.
[0078] In step 31c, the four-dimensional millimeter-wave radar collects a second radar point cloud of the work environment.
[0079] In some embodiments, the four-dimensional millimeter-wave radar collects a second radar point cloud of the working environment based on the operation of a control component such as a processor.
[0080] Steps 31a, 31b, and 31c may be performed in any order. Depending on the requirements of the working conditions, steps 31a, 31b, and 31c may only need to be performed in some cases. Similarly, steps 32a, 32b, and 32c may only need to be performed in some of the corresponding cases.
[0081] In step 32a, the image information is processed. These processes include, but are not limited to, geometric transformations of the image, saturation transformations of the image, image enhancement achieved by using a dust removal algorithm when adverse weather conditions occur, and by using homomorphic filtering to increase the pixel values of low-luminance areas in the image.
[0082] In step 32b, the first radar point cloud information is processed. These processes include, but are not limited to, noise reduction, point cloud generation, filtering, and similar processes.
[0083] In step 32c, the second radar point cloud information is processed. These processes include, but are not limited to, calibration, coordinate transformation, denoising, and static-dynamic separation.
[0084] Steps 32a, 32b, and 32c may be performed in any order.
[0085] In step 33, an interactive semantic segmentation method is performed based on the above information from multiple sensors that has been collected or processed, in order to achieve interactive semantic segmentation of point cloud data and images, and interactive semantic segmentation of the surface shape and size of the environment and the surface texture of the environment, and descriptive information such as the position, height, texture, and precise edges of the object is output.
[0086] Interactive semantic segmentation methods can be implemented using interactive semantic segmentation models. Figure 4 shows a schematic diagram of an interactive semantic segmentation model according to some embodiments of the present disclosure.
[0087] As shown in Figure 4, the interactive semantic segmentation model comprises the following components:
[0088] The image encoder 41 for feature coding of images can be implemented by an attention module. For example, the image encoder may have the same architecture as the visual transformer, so it can be pre-trained on labeled datasets of various scenes of the excavator in operation collected by the excavator itself. The image encoder can acquire input images of any size, unify the size of the data format, convert them into fixed-size image blocks, and perform feature coding on these image blocks, so that images that have passed through the image encoder output data of the same dimension. To convert an input image into a fixed-size image block, for example, first change the size to an equal ratio, change the size of the long side to 1024, then subtract the mean, divide it by the variability, and finally impair with an imputation value of 0.
[0089] The feature embedding layer 42 transforms the feature code data into fixed-size feature representations (feature vectors, also called first feature representations) to facilitate processing and computation (e.g., distance calculation) (performing dimensionality reduction on the feature code data). The primary purpose of embedding is to perform dimensionality reduction on (sparse) features. This dimensionality reduction method can be analogous to a fully connected layer (without activation functionality) where dimensionality is reduced by calculating the weight matrix of the embedding layer.
[0090] The semantic segmentation model 43 can be obtained through pre-training to implement the semantic segmentation function of an image. This can be implemented using an existing model capable of performing semantic segmentation. This semantic segmentation model is typically based on the architecture of a deep convolutional neural network (CNN), and obtains category predictions for each pixel by performing convolution, pooling, and upsampling operations on the input image. This semantic segmentation model achieves semantic segmentation by minimizing the classification loss at the pixel level during the training process. This semantic segmentation model includes, but is not limited to, FCN (Fully Convolutional Networks) models, U-Net models, DeepLab models, SegNet models, GMMSeg models, and similar models.
[0091] An image mask 44 means that masking is performed on an image. This mask means blocking out the image to be processed (whole or local) using a selected image, shape, or object to control the area or process in which the image is processed. Since certain areas on the image are screened using the mask, these areas do not participate in the processing or the calculation of processing parameters, or only the screened areas are processed or counted. Specific to this disclosure, an image mask means that a target area is masked. This target area is obtained from the original image by a mature, pre-trained semantic segmentation model.
[0092] The convolutional layer 45 performs a convolution operation on the input data. In particular, the convolution operation is performed on the mask information of the object.
[0093] The prompt encoder 46 is configured to encode input indicator information or mask information in order to obtain a corresponding feature representation. To encode indicator points, the indicator points are converted from points to vectors. To encode indicator boxes, the indicator boxes are converted from the two corner points of the box to vectors. For mask encoding, it is only necessary that the downsampling of the mask reliably matches the output of the image encoder. To obtain a second feature representation of the object's indicator information, each indicator point represented by the object's indicator information is encoded. To obtain a third feature representation of the object's mask information, convolution and encoding are performed on the object's mask information. The prompt encoder supports both sparse prompts (points, boxes, text) and dense prompts (masks) simultaneously. Sparse prompts are projected and embedded into the image. Dense prompts, on the other hand, are embedded by convolution and added by element-wise image embedding.
[0094] The marking information for the object is determined by the following method.
[0095] (1) Depending on the installation position and angle of the camera, lidar, and four-dimensional millimeter-wave radar on the excavator, transformation matrices are determined between the coordinate system of the excavator and the image coordinate system of the camera device, the coordinate system of the lidar, and the coordinate system of the four-dimensional millimeter-wave radar, respectively. For specific forms of each transformation matrix, refer to the above.
[0096] (2) According to the transformation matrix, a second marker information in the image coordinate system of a first marker information of an object in the first radar point cloud information (i.e., radar point cloud information of the work environment obtained using lidar) is determined, and / or a fourth marker information in the image coordinate system of a third marker information of an object in the second radar point cloud information (i.e., radar point cloud information of the work environment obtained using four-dimensional millimeter-wave radar) is determined.
[0097] The first marker information is, for example, any marker information representing the position of an object in the first radar point cloud information. To obtain the second marker information, the first marker information is converted from the LiDAR coordinate system to the image coordinate system. The third marker information is, for example, any marker information representing the position of an object in the second radar point cloud information. To obtain the fourth marker information, the third marker information is converted from the four-dimensional millimeter-wave radar coordinate system to the image coordinate system.
[0098] A transformation matrix between the lidar coordinate system and the image coordinate system can be determined according to the transformation matrix between the lidar coordinate system and the excavator coordinate system, and the transformation matrix between the excavator coordinate system and the image coordinate system. Based on this transformation matrix, a second marker information in the image coordinate system of the first marker information of an object in the first radar point cloud information can be determined.
[0099] Similarly, a transformation matrix between the four-dimensional millimeter-wave radar coordinate system and the image coordinate system can be determined according to the transformation matrix between the four-dimensional millimeter-wave radar coordinate system and the excavator coordinate system, and the transformation matrix between the excavator coordinate system and the image coordinate system. Based on this transformation matrix, a fourth marker information in the image coordinate system of the third marker information of an object in the second radar point cloud information can be determined.
[0100] (3) In order to determine the object's identifier, the original fifth identifier of the object in the image information in the image coordinate system is fused with the second identifier and / or fourth identifier in the image coordinate system. This includes, for example, the following exemplary fusion method.
[0101] In other words, in order to determine the object's identification information, the second and / or fourth identification information in the image coordinate system are also fused based on the original fifth identification information of the object in the image coordinate system within the image information.
[0102] (3-1) A weighted average is performed on the fifth marker information, the second marker information, and / or the fourth marker information in the image coordinate system, so that the resulting average marker information becomes the marker information of the object. The respective weights for weighting the fifth marker information, the second marker information, and the fourth marker information are configured to be the same, or are determined according to the mapping properties of the camera device, lidar, and four-dimensional millimeter-wave radar under the current working conditions. The fifth marker information is determined according to the mapping properties of the camera device under the current working conditions. The second marker information is determined according to the mapping properties of the lidar under the current working conditions. The fourth marker information is determined according to the mapping properties of the four-dimensional millimeter-wave radar under the current working conditions.
[0103] In other words, the method for determining the labeling information of an object includes any of the following:
[0104] A weighted average is performed on the fifth marker information and the second marker information, so the resulting average marker information becomes the marker information of the object.
[0105] A weighted average is performed on the fifth and fourth pieces of marker information, so the resulting average marker information becomes the marker information of the object.
[0106] A weighted average is performed on the fifth, second, and fourth marker information, so the resulting average marker information becomes the marker information of the object.
[0107] (3-2) Depending on the imaging performance of the camera device, lidar, and four-dimensional millimeter-wave radar under the current operating conditions, the marker information corresponding to the best imaging performance is selected as the object marker information from the fifth marker information, the second marker information, and / or the fourth marker information in the image coordinate system.
[0108] In other words, depending on the imaging performance of the camera device, lidar, and four-dimensional millimeter-wave radar under the current operating conditions, the marker information corresponding to the best imaging performance is selected as the object's marker information from the fifth marker information and the second marker information in the image coordinate system, or from the fifth marker information and the fourth marker information in the image coordinate system, or from the fifth marker information, the second marker information, and the fourth marker information.
[0109] The interpreter 47 performs upsampling on the image embedding through two transposed convolutional layers to obtain more accurate object prediction results, and fuses the amplified image embedding with the output data of the prompt encoder 46. This fusion method includes, but is not limited to, fusion of a first feature representation of the image information with a second feature representation of the object's labeling information and / or a third feature representation of the object's mask information, in order to detect objects in the image information. Both the second and third feature representations may indicate the object's position information. Based on the object's position information, the object in the image information is detected according to the first feature representation of the image information.
[0110] In other words, this fusion method includes any of the following:
[0111] To extract an object from the image information, the first feature representation of the image information is fused with the second feature representation of the object's identification information.
[0112] To extract an object from the image information, the first feature representation of the image information is fused with the third feature representation of the object's mask information.
[0113] To extract an object from the image information, the first feature representation of the image information is merged with the second feature representation of the object's labeling information and the third feature representation of the object's mask information.
[0114] In the following, the interactive semantic segmentation method will be described in relation to the interactive semantic segmentation model. Figure 5 shows a schematic diagram of the interactive semantic segmentation method according to some embodiments of this disclosure.
[0115] In step 51, image information of the work environment and radar point cloud information (including first radar point cloud information and / or second radar point cloud information) are obtained.
[0116] Image information of the work environment is obtained using a camera device. First radar point cloud information of the work environment is obtained using lidar. Second radar point cloud information of the work environment is obtained using four-dimensional millimeter-wave radar.
[0117] In step 52, the first feature representation of the image information is determined.
[0118] The image is converted into a fixed-size image block, and feature encoding is performed on this image block. This step can be implemented using an image encoder 41. Dimensionality reduction is performed on this feature-encoded data to obtain a fixed-size first feature representation. This step can be implemented using a feature embedding layer 42.
[0119] In step 53, a second feature representation of the object's marker information is determined based on the image information and the object's marker information within the radar point cloud information.
[0120] To obtain a second feature representation of the object's labeling information, each label point represented by the object's labeling information is encoded. This step can be implemented using a prompt encoder.
[0121] The marking information for the object is determined by the following method.
[0122] (1) Depending on the installation position and angle of the camera, lidar, and four-dimensional millimeter-wave radar on the excavator, transformation matrices are determined between the coordinate system of the excavator and the image coordinate system of the camera device, the coordinate system of the lidar, and the coordinate system of the four-dimensional millimeter-wave radar, respectively. For specific forms of each transformation matrix, refer to the above description.
[0123] (2) According to the transformation matrix, a second marker information in the image coordinate system of a first marker information of an object in the first radar point cloud information (i.e., radar point cloud information of the work environment obtained using lidar) is determined, and / or a fourth marker information in the image coordinate system of a third marker information of an object in the second radar point cloud information (i.e., radar point cloud information of the work environment obtained using four-dimensional millimeter-wave radar) is determined.
[0124] A transformation matrix between the lidar coordinate system and the image coordinate system can be determined according to the transformation matrix between the lidar coordinate system and the excavator coordinate system, and the transformation matrix between the excavator coordinate system and the image coordinate system. Based on this transformation matrix, a second marker information in the image coordinate system of the first marker information of an object in the first radar point cloud information can be determined.
[0125] Similarly, a transformation matrix between the four-dimensional millimeter-wave radar coordinate system and the image coordinate system can be determined according to the transformation matrix between the four-dimensional millimeter-wave radar coordinate system and the excavator coordinate system, and the transformation matrix between the excavator coordinate system and the image coordinate system. Based on this transformation matrix, a fourth marker information in the image coordinate system of the third marker information of an object in the second radar point cloud information can be determined.
[0126] (3) In order to determine the object's identification information, the original fifth identification information of the object in the image coordinate system within the image information is merged with the second identification information and / or fourth identification information in the image coordinate system. This includes, for example, the following exemplary merging method.
[0127] (3-1) A weighted average is performed on the fifth marker information, the second marker information, and / or the fourth marker information in the image coordinate system, so that the resulting average marker information becomes the marker information of the object. The respective weights for weighting the fifth marker information, the second marker information, and the fourth marker information are configured to be the same, or are determined according to the mapping capabilities of the camera device, lidar, and four-dimensional millimeter-wave radar under the current working conditions.
[0128] (3-2) Depending on the imaging performance of the camera device, lidar, and four-dimensional millimeter-wave radar under the current operating conditions, the marker information corresponding to the best imaging performance is selected as the object marker information from the fifth marker information, the second marker information, and / or the fourth marker information in the image coordinate system.
[0129] In step 54, a third feature representation of the object's mask information is determined based on the object's mask information.
[0130] The object and its mask information are determined based on image information. The object may be extracted using a semantic segmentation model 43. The object's mask information may be determined using an image mask 44. To obtain a third feature representation of the object's mask information, the object's mask information is convolved using a convolutional layer 45 and then encoded using a prompt encoder 46.
[0131] Steps 52, 53, and 54 may be performed in any order, and either or both of steps 53 and 54 may be performed as needed.
[0132] In step 55, in order to extract the object from the image information, the first feature representation of the image information is merged with the second feature representation of the object's labeling information and / or the third feature representation of the object's mask information.
[0133] Both the second and third feature representations can indicate the location information of the object. Based on the location information of the object, the object is detected in the image information according to the first feature representation of the image information.
[0134] While obtaining the three-dimensional coordinates of the object, it is also possible to obtain features such as the image's position, texture, appearance, actual height, and precise edges, thus enabling the inference of the object's type and providing a crucial foundation for decision-making planning of the excavator. Based on an interactive semantic segmentation method, dynamic information of the segmentation area can be accurately provided in real time. This information provides a sufficient cognitive foundation for the automation, intelligent planning, and control of the excavator.
[0135] Figure 6 shows a schematic diagram of a semantic segmentation apparatus according to some embodiments of the present disclosure. As shown in Figure 6, the semantic segmentation apparatus 600 of this embodiment comprises a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to perform a semantic segmentation method according to any one of the embodiments based on instructions stored in the memory 610.
[0136] The device 600 may further include an input / output (I / O) interface 630, a network interface 640, a storage interface 650, and similar interfaces. These interfaces 630, 640, 650, as well as the memory 610 and processor 620, may be connected, for example, via a bus 660 between them.
[0137] The memory 610 may include, for example, system memory, a non-volatile fixed storage medium, or something similar. The system memory stores, for example, the operating system, application programs, a boot loader, and other programs.
[0138] The processor 620 can be implemented as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, as well as discrete hardware components such as discrete gates or transistors.
[0139] The input / output interface 630 provides connection interfaces for input and output devices such as displays, mice, keyboards, and touchscreens. The network interface 640 provides connection interfaces for various network devices. The storage interface 650 provides connection interfaces for external storage devices such as SD cards or USB flash drives. The bus 660 may use any of several bus structures. For example, the bus structures may include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Microchannel Architecture (MCA) bus, and the Peripheral Component Interconnect (PCI) bus.
[0140] Embodiments of this disclosure provide a computer-readable storage medium that stores a computer program. When executed by a processor, this computer program embodies a semantic segmentation method according to any one of the embodiments.
[0141] According to embodiments of this disclosure, a computer program product is provided. This computer program product comprises a computer program, which, when executed by a processor, embodies a semantic segmentation method according to any one of the embodiments.
[0142] Embodiments of the present disclosure provide a semantic segmentation system that may be used for an excavator 14. As shown in Figure 7, the semantic segmentation system 700 comprises a camera device 11 for collecting image information of a work environment; a radar for collecting radar point cloud information of a work environment, for example, a lidar 13 for collecting first radar point cloud information of a work environment and / or a four-dimensional millimeter-wave radar 12 for collecting second radar point cloud information of a work environment; and a semantic segmentation device 600 configured to carry out a semantic segmentation method in any one of the embodiments based on information collected by the plurality of sensors.
[0143] Those skilled in the art will understand that multiple embodiments of this disclosure may be provided as methods, systems, or computer program products. Accordingly, this disclosure may take the form of hardware-only embodiments, software-only embodiments, or embodiments combining both software and hardware aspects. Furthermore, this disclosure may take the form of a computer program product implemented on one or more (non-temporary) computer-readable storage media (including, but not limited to, disk memory, CD-ROM, optical memory, cloud storage, and similar) on which computer program code is embodied.
[0144] This disclosure has been described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to several embodiments of this disclosure. It will be understood that each step and / or block of a flowchart and / or block diagram, as well as combinations of multiple steps and / or blocks of a flowchart and / or block diagram, can be embodied by computer program instructions. These computer program instructions may be provided to the processor of a general-purpose computer, a dedicated computer, an embedded processing machine, or other programmable data processing device for producing a machine, which, when executed by the processor of a computer or other programmable data processing device, produces a device for producing a device that realizes the function specified in one or more steps of a flowchart and / or one or more blocks in a block diagram.
[0145] These computer program instructions can also be stored in computer-readable memory. Computer program instructions can cause a computer or other programmable data processing device to operate in a specific way, so that instructions stored in computer-readable memory give rise to a product with an instruction device. This instruction device implements the functions specified in one or more steps in a flowchart or one or more blocks in a block diagram.
[0146] These computer program instructions may also be loaded into a computer or other programmable data processing device so that a series of operational steps are performed on the computer or other programmable device to produce a process that is embodied by the computer, such that instructions executed on the computer or other programmable device provide multiple steps to realize a function specified in one or more steps of a flowchart and / or one or more blocks in a block diagram.
[0147] The above description represents only a number of preferred embodiments of the Disclosure and is not intended to limit the Disclosure. Any amendments, substitutions of equivalents, improvements, and similar modifications made within the spirit and principles of the Disclosure shall all be within the scope of the Disclosure.
Claims
1. A semantic segmentation method, Steps include acquiring image information and radar point cloud information of the work environment, The steps include determining a first feature representation of the aforementioned image information, A step of determining a second feature representation of the object's marking information in the radar point cloud information based on the image information and the object's marking information, In order to extract the object from the image information, the steps include: fusing the first feature representation of the image information with the second feature representation of the marking information of the object; A semantic segmentation method that includes [this].
2. The step of acquiring image information and radar point cloud information of the aforementioned work environment is: Acquiring the image information of the work environment using a camera device, The method involves using lidar to acquire first radar point cloud information of the work environment, and / or using four-dimensional millimeter-wave radar to acquire second radar point cloud information of the work environment. A semantic segmentation method according to claim 1, comprising:
3. Depending on the installation position and angle of the camera device, lidar, and four-dimensional millimeter-wave radar on the excavator, transformation matrices are determined between the coordinate system of the excavator and the image coordinate system of the camera device, between the coordinate system of the excavator and the coordinate system of the lidar, and between the coordinate system of the excavator and the coordinate system of the four-dimensional millimeter-wave radar, respectively. In accordance with the transformation matrix, the second marker information in the image coordinate system of the first marker information of the object in the first radar point cloud information is determined, and / or the fourth marker information in the image coordinate system of the third marker information of the object in the second radar point cloud information is determined, In order to determine the marking information of the object, the original fifth marking information of the object in the image coordinate system within the image information is fused with the second marking information and / or the fourth marking information in the image coordinate system. The semantic segmentation method according to claim 2, further comprising the step of determining the label information of the object by the method.
4. To determine the label information of the object, merging the original fifth label information of the object in the image coordinate system within the image information with the second label information and / or the fourth label information in the image coordinate system is: In order to obtain averaged label information as the label information of the object, a weighted average is performed on the fifth label information and the second label information in the image coordinate system, In order to obtain averaged label information as the label information of the object, a weighted average is performed on the fifth label information and the fourth label information in the image coordinate system, In order to obtain average label information as the label information of the object, a weighted average is performed on the fifth label information, the second label information, and the fourth label information. It includes any of the following: The respective weights for weighting the fifth, second, and fourth indicator information are configured to be the same, or determined according to the mapping characteristics of the camera device, lidar, and four-dimensional millimeter-wave radar under current operating conditions. The semantic segmentation method according to claim 3.
5. To determine the label information of the object, merging the original fifth label information of the object in the image coordinate system within the image information with the second label information and / or the fourth label information in the image coordinate system is: Depending on the imaging performance of the camera device, lidar, and four-dimensional millimeter-wave radar under current operating conditions, the method includes selecting the marker information corresponding to the best imaging performance of the object from the fifth marker information and the second marker information in the image coordinate system, or from the fifth marker information and the fourth marker information in the image coordinate system, or from the fifth marker information, the second marker information, and the fourth marker information. The semantic segmentation method according to claim 3.
6. The step of determining the first feature representation of the image information is: Convert the image into a fixed-size image block, and perform feature encoding on the image block. Reducing the dimension of feature code data to a fixed-size first feature representation, including, The semantic segmentation method according to claim 1.
7. The step of determining the second characteristic representation of the marking information of the object is: To obtain the second feature representation of the marking information of the object, the method includes encoding each marking point represented by the marking information of the object. The semantic segmentation method according to claim 1.
8. A step of determining the mask information of the object based on the aforementioned image information, A step of determining a third feature representation of the mask information of the object based on the mask information of the object, It further includes, The merging step includes merging the first feature representation of the image information with the second feature representation of the marking information of the object and the third feature representation of the mask information of the object in order to extract the object from the image information. The semantic segmentation method according to claim 1.
9. The step of determining the mask information of the object based on the aforementioned image information is: In order to obtain the image region of the aforementioned object, semantic segmentation is performed based on the image information, Based on the image region of the object, the mask information of the object is generated. including, The semantic segmentation method according to claim 8.
10. The step of determining the third feature representation of the mask information of the object is: To obtain the third feature representation of the mask information of the object, the process includes performing a convolution operation on the mask information of the object and encoding it. The semantic segmentation method according to claim 8.
11. A semantic segmentation apparatus comprising a memory and a processor coupled to the memory, wherein the processor is configured to perform the semantic segmentation method described in any one of claims 1 to 10 based on instructions stored in the memory.
12. A computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, embodies the semantic segmentation method described in any one of claims 1 to 10.
13. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, embodies the semantic segmentation method described in any one of claims 1 to 10.
14. A camera device for collecting image information of the work environment, A radar comprising a lidar for collecting a first radar point cloud information of the work environment, and / or a four-dimensional millimeter-wave radar for collecting a second radar point cloud information of the work environment, A semantic segmentation apparatus configured to carry out the semantic segmentation method described in any one of claims 1 to 10, A semantic segmentation system equipped with [a specific feature / feature].
15. A semantic segmentation system according to claim 14 Excavator.
Citation Information
Patent Citations
Recognition system, vehicle control system, recognition method, and program
JP2021021967A
Multi-sensor object detection fusion system and method using point cloud projection
US11403860B1
Methods and Apparatuses for Object Detection in a Scene Based on Lidar Data and Radar Data of the Scene
US20200301013A1