Image processing method and device for improving robot's environmental perception ability
By combining RGB-D cameras and lidar, and using a deep image fusion network to process heterogeneous depth information, an accurate and dense fused image is generated, which solves the problem of robots having difficulty navigating in dark environments, reduces system cost and complexity, and enhances navigation capabilities.
Patent Information
- Application Number
- CN202410526588.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-04-29
AI Technical Summary
In existing technologies, it is difficult for robots to effectively obtain accurate depth information in dark environments, resulting in navigation difficulties and collision risks, and relying on multiple expensive LiDARs increases system cost and complexity.
By combining RGB-D cameras with LiDAR, the deep image fusion network is used to process heterogeneous depth information to generate accurate and dense fused images, and the pre-trained deep learning model is used to fuse the data of RGB-D cameras and LiDAR.
It improves the robot's navigation ability in dark environments, reduces system cost and complexity, reduces collision risks, and enhances accessibility in confined environments.
Smart Images

Figure CN119274031B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of depth completion in image processing, and in particular to an image processing method, device, electronic device, non-transitory computer-readable storage medium, and computer program product for improving a robot's environmental perception capability. Background Art
[0002] With the development of robotics-related technologies, robots are expected to replace humans in performing tasks. Certain types of robots (for example, those with the ability to autonomously perceive and create three-dimensional maps of unknown environments beyond visual range) can perform some relatively dangerous tasks (for example, detecting unstable structures in underground mines or buildings that may collapse). In this context, a key data source required for robots to perform tasks is depth information (i.e., depth information is a key data support for robots to autonomously locate, perceive, and map in dark environments).
[0003] Current technologies primarily rely on LiDAR (Light Detection and Ranging) and cameras to collect depth information. However, robots face the challenge of camera perception degradation in dark underground environments. Existing technologies often address this challenge by equipping robots with multiple, expensive LiDARs of varying types, but this in turn limits the cost, complexity, and accessibility of robotic systems. Furthermore, the sparse depth information scanned by LiDAR alone not only makes it difficult to capture small geometric defects, but also exposes robots to the risk of collision when navigating complex, unknown environments.
[0004] Therefore, it is necessary to provide a new solution to overcome the above problems. Summary of the Invention
[0005] The present disclosure provides an image processing method, device, electronic device, non-transitory computer-readable storage medium, and computer program product for improving the environmental perception capability of a robot, so as to address the deficiencies in the prior art.
[0006] The present disclosure provides an image processing method for improving the environmental perception capability of a robot, comprising:
[0007] Acquire an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocess the initial point cloud data to obtain a point cloud depth image; wherein, the sensor includes a lidar and a depth camera, and the depth camera is capable of collecting depth information of the target environment area; input the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; wherein, the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein, the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0008] The present disclosure also provides an image processing device for improving the environmental perception capability of a robot, comprising:
[0009] The data acquisition module is configured to: acquire an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and pre-process the initial point cloud data to obtain a point cloud depth image; wherein, the sensor includes a lidar and a depth camera, and the depth camera is capable of collecting depth information of the target environment area; the data processing module is configured to: input the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; wherein, the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein, the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0010] The present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the image processing method for improving the robot's environmental perception capability as described above is implemented.
[0011] The present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the image processing method for improving the robot's environmental perception capability as described in any one of the above is implemented.
[0012] The present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-described image processing methods for improving the environmental perception capability of a robot.
[0013] As described above, by using the image processing method for improving the robot's environmental perception capability provided by the above embodiments of the present disclosure, the heterogeneous depth information of LiDAR and RGB-D cameras is used for environmental perception, and then the heterogeneous depth information is fused using a pre-trained deep image fusion network. It is possible to predict an accurate and dense fused image of the same target environmental area, thereby providing effective data support for the robot's next work (for example, navigation), and thus reducing the risk of collision faced by the robot during navigation in complex and unknown environments to a certain extent.
[0014] In addition, since this solution uses a data acquisition method that combines RGB-D cameras and lidar, compared with the existing solution of "equipping the robot with multiple different types of expensive lidar", it can to a certain extent reduce the cost and complexity of the robot system and improve the robot's accessibility in confined environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a flowchart of an image processing method for improving a robot's environmental perception capability provided by an embodiment of the present disclosure;
[0017] Figure 2 is a structural diagram of the deep image fusion network provided by an embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram of a process for training the deep image fusion network to be trained provided by an embodiment of the present disclosure;
[0019] Figure 4 This is a schematic diagram of the coarse-to-fine point cloud registration process provided by an embodiment of the present disclosure;
[0020] Figure 5 is a schematic structural diagram of an image processing device for improving the environmental perception capability of a robot provided by an embodiment of the present disclosure;
[0021] Figure 6 It is a structural diagram of the electronic device provided by the present disclosure. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of this disclosure more clear, the technical solutions of this disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of this disclosure without making any creative efforts shall fall within the scope of protection of this disclosure.
[0023] Summary of the invention concept:
[0024] After repeated research, the inventors of this disclosure have discovered that RGB-D cameras (which may include an RGB camera and a depth sensor) can be used for sensing in dark environments (e.g., underground spaces) due to their ability to actively image in dark environments. Furthermore, RGB-D cameras can acquire denser environmental depth information than LiDAR in short-range detection. LiDAR, on the other hand, can detect farther away than RGB-D cameras, thus compensating for the RGB-D camera's vulnerability to environmental factors, resulting in data holes and unreliable measurements. Therefore, combining RGB-D cameras with LiDAR, in place of conventional cameras, can address the aforementioned issues associated with the prior art.
[0025] In light of this, the inventors of this disclosure considered fusing these two types of heterogeneous depth information to complement each other's strengths, thereby enhancing the robot's perception capabilities in dark environments. It should be noted that because RGB-D cameras and LiDAR collect depth information in different ways, this disclosure refers to it as heterogeneous depth information (HDI).
[0026] Example:
[0027] The image processing solution disclosed in the present invention for improving the robot's environmental perception capability is described below with reference to the accompanying drawings.
[0028] Figure 1 This is a flow chart of an image processing method for improving the robot's environmental perception capability provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices (e.g., a server or cloud computing platform that is connected to the robot for communication, or a processor directly configured on the robot), such as Figure 1 As shown, the image processing method for improving the robot's environmental perception ability includes the following steps:
[0029] S110, acquiring an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocessing the initial point cloud data to obtain a point cloud depth image;
[0030] The sensor includes a LiDAR (Light Detection and Ranging) and a depth camera, and the depth camera can collect depth information of the target environment area.
[0031] As an optional example, the depth camera is an RGB-D camera. Among them, RGB-D cameras and LiDARs can be used in underground perception systems because of their ability to actively image in dark environments. In related practices, RGB-D cameras are only used to sense terrain changes under the feet of the robot or to detect close obstacles to assist robot navigation. However, the inventors of the present disclosure consider that since RGB-D cameras can obtain denser environmental depth information than LiDARs, RGB-D cameras also have potential in mapping. In particular, consumer-grade RGB-D cameras have the advantages of small size and light weight, and are suitable for being mounted on small robot platforms, which is conducive to improving the accessibility of small, low-cost robots in cluttered environments (for example, dark environments).
[0032] The present disclosure is not limited to consumer-grade RGB-D cameras; for example, it may include but is not limited to Microsoft's Kinect, Intel's RealSense, etc.
[0033] The specific implementation details of "preprocessing the initial point cloud data" will be described below and will not be elaborated here.
[0034] Among them, the initial depth image can be collected by an RGB-D camera; the initial point cloud data can be collected by a lidar.
[0035] S120: Input the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model.
[0036] The fused image includes the initial depth image and depth information in the initial point cloud data.
[0037] Among them, the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0038] As described above, by using the image processing method for improving the robot's environmental perception capability provided by the above embodiments of the present disclosure, the heterogeneous depth information of LiDAR and RGB-D cameras is used for environmental perception, and then the heterogeneous depth information is fused using a pre-trained deep image fusion network. It is possible to predict an accurate and dense fused image of the same target environmental area, thereby providing effective data support for the robot's next work (for example, navigation), and thus reducing the risk of collision faced by the robot during navigation in complex and unknown environments to a certain extent.
[0039] In addition, since this solution uses a data acquisition method that combines RGB-D cameras and lidar, compared with the existing solution of "equipping the robot with multiple different types of expensive lidar", it can to a certain extent reduce the cost and complexity of the robot system and improve the robot's accessibility in confined environments.
[0040] exist Figure 1 Based on the embodiment, as an optional implementation method, refer to Figure 2 , the deep image fusion network may include a sparse convolution unit, a hole convolution unit and a depth measurement confidence filtering unit.
[0041] Specifically, as an optional example, refer to Figure 2 ,The design concepts of the deep image fusion network include:
[0042] 1) Design five sparse convolutional layers to process the input LiDAR depth image, with a stride of 1 and 32 output channels, and use the ReLU activation function to maintain depth non-negativity;
[0043] 2) Design five dilated convolutional layers to process the input RGB-D camera depth image, set the convolution dilation coefficients to k = 1, 2, 4, 8, 8 respectively, set the feature map output to 32, and use batch normalization and ReLU activation;
[0044] 3) The features of the output LiDAR depth image and the RGB-D camera depth image are concatenated in the fifth layer and then fused through four layers of residual blocks. After that, the fused features are also concatenated with the LiDAR depth image to further refine the output;
[0045] 4) Filter the output fused depth image and retain the range in the output depth image (in millimeters) and The depth values are weighted averaged and calculated as follows: Among them, d f represents the network output depth map, The weight representing the LiDAR depth; Represents the weight of LiDAR depth; d l (r) and d o (r) respectively represent the range For the depth values within 6000 mm, δ = 0.005 is a constant. For depth values exceeding 6000 mm, it is directly replaced by the corresponding value in the LiDAR depth.
[0046] exist Figure 1 、 2 Based on the embodiment, as an optional implementation, preprocessing the point cloud data may include the following steps:
[0047] Step 1) Based on the acquired intrinsic parameter matrix K of the depth camera, determine the transformation matrix T between the depth camera and the lidar.
[0048] As an optional example, the RGB-D camera can be calibrated first to obtain the intrinsic parameter matrix K; then the RGB-D camera and LiDAR can be externally calibrated to obtain the transformation matrix T between the sensors.
[0049] Step 2) Using the intrinsic parameter matrix K and the transformation matrix T, calculate the projection value of the point cloud data onto the corresponding depth camera (for example, RGB-D camera) plane to obtain the point cloud depth image.
[0050] As an optional example, two consecutive frames of point clouds L input in time sequence can be transformed according to the intrinsic parameter matrix K and the transformation matrix T. t , L t+1 Project it onto the corresponding camera plane to obtain the LiDAR point cloud depth image
[0051] In the above example, the following calculation formula can be used to determine the point cloud depth image:
[0052]
[0053] Among them, L * Represents the LiDAR point cloud at any moment.
[0054] exist Figure 1 、 2 Based on the embodiment, as an optional implementation method, refer to Figure 3 as well as Figure 4 , the training process of the pre-trained deep learning model includes:
[0055] Step 1: Create a sample data set.
[0056] The sample data set includes multiple sample depth images of the same target area, multiple frames of sample point clouds at consecutive moments, and corresponding multiple frames of sample point cloud depth images at consecutive moments.
[0057] As an optional example, according to the time sequence of the sample point clouds at multiple consecutive frames, a sample depth image and a frame of sample point cloud depth image regarding the same target area are taken as a set of sample data; in this example, the aforementioned implementation method of "preprocessing the point cloud data" can be adopted to preprocess the sample point cloud to obtain the sample point cloud depth image.
[0058] Or, as another alternative example, refer to Figure 3 , a sample depth image and a frame of sample point cloud of the same target area can be directly used as a set of sample data. In this example, the above-mentioned process of "preprocessing the point cloud data" is performed in a pre-trained deep learning model.
[0059] Step 2: Reference Figure 3 , the sample data of the target time t (for example, the sample depth image D t , sample point cloud L t ) and the sample data of the next time instant t+1 directly adjacent to the target time instant (for example, the sample depth image D t+1 , sample point cloud L t+1 ), respectively input the deep image fusion network to be trained to obtain the initial fusion image corresponding to the target moment and the initial fusion image corresponding to the next moment.
[0060] The target time can be any time within the sample data collection time in the sample data set, for example, the time when an RGB-D camera captures a sample depth image, or the time when a lidar captures a frame of point cloud data.
[0061] Step 3: Based on the preset view synthesis rule, the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment are synthesized to obtain a simulated co-view image at the target moment.
[0062] As an alternative example, refer to Figure 3 and Figure 4 , the step 3 may further include:
[0063] Step 31: Obtain the sample point cloud L at the target time in the sample data set t , and the sample point cloud L at the next moment directly adjacent to the target moment t+1 .
[0064] Step 32: Determine the relative motion transformation matrix between the two moments using the sample point cloud at the target moment and the sample point cloud at the next moment directly adjacent to the target moment.
[0065] As an optional example, a point cloud registration module with coarse-to-fine accuracy can be constructed by using fast global registration and point-to-surface iterative closest point (ICP) algorithm to obtain the point cloud registration result from two consecutive frames of point cloud L input in time sequence. t , L t+1 (For example, see Figure 4 , two-frame point cloud L t , L t+1 The relative motion transformation matrix T of the rigid assembly at two moments can be obtained from the original point cloud by voxel downsampling. t,t+1 (i.e., for example Figure 4 The registration results are shown).
[0066] The ICP algorithm is an algorithm used to solve the free-form surface registration problem. Its basic principle is to minimize the distance between corresponding points of the source and target data through continuous iteration to achieve accurate registration.
[0067] Step 33: Using the relative motion transformation matrix, determine a conversion view of the initial fused image at the next moment under the perspective of the initial fused image at the target moment.
[0068] As an optional example, the following formula can be used to determine the conversion view:
[0069]
[0070]
[0071] in, Represents the initial fused image corresponding to the next moment; represents the initial fused image The coordinates of all pixels in the pixel space coordinate system; K represents the internal parameter matrix; X i,j represents the initial fused image All pixels in ; represents the initial fusion image corresponding to the target moment; T t,t+1 represents the relative motion transformation matrix; Indicates the depth values corresponding to all pixels in the initial fused image at the next moment.
[0072] It should be explained that the principle of step 33 is that in the rigid assembly formed by the RGB-D camera and LiDAR through the rigid connection, the relative position relationship between the sensors remains unchanged during the motion process. That is, when the LiDAR is translated or (and) rotated, the RGB-D camera will also respond to the same physical quantity in the same way to maintain the relative position relationship between them. Therefore, the relative motion transformation matrix T can be used in the example of step 33. t,t+1 Synthesize the initial fused image In the initial fusion image The synthetic view under the viewing angle, that is, the transformed view
[0073] Step 34: Determine a simulated co-visual view at the target moment by using the converted view and the initial fused image at the target moment.
[0074] Reference Figure 3 During the view synthesis process, the following formula can be used to calculate the simulated co-visual image at the target moment:
[0075]
[0076] in, representing the transformation view; Represents the initial fused image corresponding to the target moment. The calculation formula represents: Calculate the transformed view With the initial fusion image Pixel-by-pixel correlation of depth images.
[0077] Step 4: Based on the simulated co-visual image at the target moment, calculate the depth consistency loss function value and the smoothing loss function value;
[0078] As an optional example, the depth consistency loss function value is determined using the following calculation formula:
[0079]
[0080] Wherein, Ω represents the pixel set of the simulated co-visual image; Represents the pixel value of pixel v in the initial fused image at the target moment; Represents the pixel value of pixel v in the simulated co-visual map at the target time.
[0081] Constructing temporal-depth consistency loss function based on simulated co-visual image This is done to penalize pixels with inconsistent depths at the same position in the two frames of preliminary fused depth images, thereby optimizing the deep image fusion network to be trained.
[0082] It should be noted that in the above embodiments and examples of this disclosure, normalized depth differences are used rather than absolute distances. This is because it allows points at different absolute depths to be treated equally during optimization. Furthermore, this design makes the function symmetrical, with the output naturally ranging between 0 and 1, which makes training more stable.
[0083] It should be noted that, in the above examples of the present disclosure, the depth consistency loss term is constructed based only on the depth information of the single-view image.
[0084] As another optional example, the smoothing loss function value is determined using the following calculation formula:
[0085]
[0086] Wherein, Ω represents the pixel set of the simulated co-visual image; D f (v) represents the pixel value of the pixel v in the updated image; Δ x Represents the gradient of the predicted depth map in the horizontal direction; Δ y Represents the gradient of the predicted depth map in the vertical direction.
[0087] In this example, the second-order derivative of the fused depth image is used The norm, as another component of the fusion network loss function, can be used to optimize the deep image fusion network to be trained.
[0088] Step 5: Use the initial fusion image at the target time to filter the sample depth image at the target time to obtain the filtered sample depth image D at the target time. c , which can be used as a verification image.
[0089] Among them, in the model training stage, step 5 is recorded as the verification stage.
[0090] Step 6: The filtered sample depth image D at the target moment c , and the sample point cloud depth image at the target time, input the depth image fusion network to be trained, and obtain the second fusion image D corresponding to the target time f , as the updated image.
[0091] Among them, in the model training stage, step 6 is recorded as the update stage.
[0092] Step 7: Calculate the verification loss function value using the verification image and the update image.
[0093] As an optional example, the following formula can be used to determine the verification loss function value:
[0094]
[0095] Where ε=30mm
[0096] Wherein, F represents the mask of the depth point after filtering in step 5; ||x|| τ Indicates truncated Norm function calculation rules.
[0097] In this example, the fused depth image D output by the update stage is calculated f and the filtered depth image D c Truncated Norm function, used to construct the validation loss term As one of the components of the fusion network loss function, it can be used to optimize the deep image fusion network to be trained.
[0098] Step 8: Determine a comprehensive loss function value based on the depth consistency loss function value, the smoothing loss function, and the verification loss function value.
[0099] As an optional example, the following formula can be used to determine the comprehensive loss function value:
[0100]
[0101] Among them, λ1 and λ2 represent weight parameters; Represents the smoothed loss function value; Represents the verification loss function value; Represents the value of the depth consistency loss function.
[0102] Based on the above example, preferably, λ1 is set to 1 and λ2 is set to 0.001.
[0103] It should be noted that after training is completed, the fusion of depth images no longer distinguishes between the verification stage and the update stage.
[0104] Step 9: Use the comprehensive loss function value to train the deep image fusion network to be trained until the preset training completion condition is met, and obtain the deep image fusion network from the deep image fusion network to be trained.
[0105] As an optional example, step 9 may include:
[0106] Step 91: In response to the convergence of the comprehensive loss function value, the comprehensive loss function value is fed back to the deep image fusion network to be trained.
[0107] Step 92: the deep image fusion network to be trained adjusts the network parameters of each network layer according to the comprehensive loss function value;
[0108] Step 93, iteratively execute the steps from step 2 to the step of adjusting the network parameters of each network layer (ie, step 92) until the preset training completion condition is met.
[0109] Among them, the preset training completion condition is: all sample data in the sample data set participate in the training, and the corresponding comprehensive loss function value converges.
[0110] Optionally, in response to the failure of the comprehensive loss function value to converge, the currently trained sample data can be skipped, and the training of the first set of sample data can be continued. Then, after the first round of training is completed, in the second round of training, the sample data that did not converge in the first round can be trained separately.
[0111] Optionally, if there is sample data that has undergone multiple rounds of training but the corresponding comprehensive loss function value still does not converge, then this group of sample data can be discarded because this group of sample data may contain interference factors such as noise, which causes the corresponding comprehensive loss function value to not converge.
[0112] As described above, by utilizing the image processing method for improving the robot's environmental perception capability provided by the above embodiments of the present disclosure, the heterogeneous depth information of LiDAR and RGB-D cameras is used for environmental perception, and then the heterogeneous depth information is fused and processed using a pre-trained deep image fusion network, an accurate and dense fused image of the same target environmental area can be predicted, thereby providing effective data support for the robot's next step (e.g., navigation), thereby reducing the risk of collision faced by the robot during navigation in complex and unknown environments to a certain extent. Moreover, since this solution adopts a data acquisition method combining RGB-D cameras and LiDAR, it can reduce the cost and complexity of the robot system to a certain extent compared to the prior art solution of "equipping the robot with multiple different types of expensive LiDAR."
[0113] In addition, the image processing method for improving the robot's environmental perception ability provided by the above-mentioned embodiment of the present disclosure can also perform unsupervised learning without relying on large-scale true depth images; it can reasonably utilize the effective measurement data of different types of depth sensors to accurately restore the scene's geometric features; the input data can be obtained by consumer-grade depth sensors, which can reduce the cost and size of the robot platform; thereby helping to enhance the perception ability of small, low-cost robots in dark environments.
[0114] The image processing device for improving the robot's environmental perception ability provided by the present disclosure is described below. The image processing device for improving the robot's environmental perception ability described below and the image processing method for improving the robot's environmental perception ability described above can be referenced to each other.
[0115] Figure 5 FIG is a schematic diagram of the structure of an image processing device for improving the robot's environmental perception capability provided by an exemplary embodiment of the present disclosure. Figure 5 As shown, an image processing device for improving the environmental perception capability of a robot includes: a data acquisition module 510, configured to: acquire an initial depth image and initial point cloud data of the same target environmental area collected by a sensor associated with the robot, and pre-process the initial point cloud data to obtain a point cloud depth image; wherein, the sensor includes a lidar and a depth camera, and the depth camera can collect depth information of the target environmental area; a data processing module 520, configured to: input the initial depth image and the point cloud depth image into a pre-trained deep learning model, and obtain a fused image of the same target environmental area output by the pre-trained deep learning model; wherein, the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein, the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0116] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6As shown, the electronic device may include: a processor (processor) 610, a communication interface (Communications Interface) 620, a memory (memory) 630 and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logic instructions in the memory 630 to execute an image processing method for improving the robot's environmental perception capability, the method including: obtaining an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocessing the initial point cloud data to obtain a point cloud depth image; wherein the sensor includes a lidar and a depth camera, and the depth camera can collect depth information of the target environment area; inputting the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; wherein the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0117] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0118] On the other hand, the present disclosure also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image processing method provided by the above methods for improving the robot's environmental perception ability, the method including: obtaining an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocessing the initial point cloud data to obtain a point cloud depth image; wherein the sensor includes a lidar and a depth camera, and the depth camera can collect depth information of the target environment area; the initial depth image and the point cloud depth image are input into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; wherein the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network.
[0119] On the other hand, the present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the image processing method provided by the above-mentioned methods for improving the robot's environmental perception ability, the method comprising: obtaining an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocessing the initial point cloud data to obtain a point cloud depth image; wherein the sensor comprises a lidar and a depth camera, and the depth camera is capable of collecting depth information of the target environment area; the initial depth image and the point cloud depth image are input into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; wherein the fused image contains the depth information in the initial depth image and the initial point cloud data; wherein the pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model comprises a deep image fusion network.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.
Claims
1. An image processing method for improving a robot's environmental perception capability, characterized in that: include: Obtaining an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and preprocessing the initial point cloud data to obtain a point cloud depth image; wherein the sensor includes a laser radar and a depth camera, and the depth camera is capable of collecting depth information of the target environment area; Inputting the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; The fused image includes the depth information in the initial depth image and the initial point cloud data; The pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network; The training process of the pre-trained deep learning model includes: Step 1: Establish a sample data set; wherein the sample data set includes multiple sample depth images of the same target area, multiple frames of sample point clouds at consecutive moments, and corresponding multiple frames of sample point cloud depth images at consecutive moments; wherein, according to the time sequence of the sample point clouds at consecutive moments, a sample depth image and a frame of sample point cloud depth image of the same target area are taken as a set of sample data; Step 2: Input the sample data at the target moment and the sample data at the next moment directly adjacent to the target moment into the deep image fusion network to be trained, respectively, to obtain the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment; Step 3: Based on a preset view synthesis rule, the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment are synthesized to obtain a simulated co-view image at the target moment; Step 4: Based on the simulated co-visual image at the target moment, calculate the depth consistency loss function value and the smoothing loss function value; Step 5: Use the initial fusion image at the target time to filter the sample depth image at the target time to obtain the filtered sample depth image at the target time as the verification image; Step 6: Input the filtered sample depth image at the target moment and the sample point cloud depth image at the target moment into the depth image fusion network to be trained to obtain a second fused image corresponding to the target moment as an updated image; Step 7: Calculate the verification loss function value using the verification image and the update image; Step 8: Determine a comprehensive loss function value based on the depth consistency loss function value, the smoothing loss function, and the verification loss function value; Step 9: Use the comprehensive loss function value to train the deep image fusion network to be trained until the preset training completion condition is met, and obtain the deep image fusion network from the deep image fusion network to be trained.
2. The image processing method according to claim 1, wherein: The deep image fusion network includes a sparse convolution unit, a hole convolution unit and a depth measurement confidence filtering unit.
3. The image processing method according to claim 1, wherein: Preprocessing the point cloud data includes: Based on the acquired intrinsic parameter matrix of the depth camera K , determine the transformation matrix between the depth camera and the lidar T ; Using the internal parameter matrix K and the transformation matrix T , calculate the projection value of the point cloud data onto the corresponding depth camera plane to obtain the point cloud depth image.
4. The image processing method according to claim 1, wherein: The deep image fusion network to be trained is trained using the comprehensive loss function value until a preset training completion condition is met, including: In response to the comprehensive loss function value converging, the comprehensive loss function value is fed back to the deep image fusion network to be trained; The deep image fusion network to be trained adjusts the network parameters of each network layer according to the comprehensive loss function value; Iteratively executing step 2 to the step of adjusting the network parameters of each network layer until a preset training completion condition is met; Among them, the preset training completion condition is: all sample data in the sample data set participate in the training, and the corresponding comprehensive loss function value converges.
5. The image processing method according to claim 1, wherein: Based on the preset view synthesis rules, the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment are synthesized to obtain a simulated co-view image at the target moment, including: Obtaining a sample point cloud at a target moment and a sample point cloud at a next moment directly adjacent to the target moment in the sample data set; Determine a relative motion transformation matrix between the two moments using the sample point cloud at the target moment and the sample point cloud at a next moment directly adjacent to the target moment; Determining a conversion view of the initial fused image at the next moment under the perspective of the initial fused image at the target moment using the relative motion transformation matrix; A simulated co-visual view at the target moment is determined by using the transformed view and the initial fused image at the target moment.
6. The image processing method according to claim 1, wherein: Based on the simulated co-visual image at the target moment, the depth consistency loss function value and the smoothness loss function value are calculated, including: Use the following calculation formula to determine the depth consistency loss function value : in, a set of pixels representing the simulated co-visual image; Represents the pixel in the initial fusion image at the target moment v The pixel value of the position; Represents the pixel in the simulated co-visual map at the target time v The pixel value of the position; Use the following calculation formula to determine the smooth loss function value : in, a set of pixels representing the simulated co-visual image; Represents the pixels in the updated image v The pixel value of the position; Represents the gradient of the predicted depth map in the horizontal direction; Represents the gradient of the predicted depth map in the vertical direction.
7. The image processing method according to claim 1, wherein: Calculating a verification loss function value using the verification image and the update image includes: Use the following calculation formula to determine the verification loss function value : in, F A mask representing the depth point after filtering in step 5; Indicates truncated Norm function calculation rules.
8. The image processing method according to claim 1, wherein: Determine a comprehensive loss function value based on the depth consistency loss function value, the smoothing loss function and the verification loss function value, include: Use the following calculation formula to determine the comprehensive loss function value : in, 、 represents the weight parameter; Represents the smoothed loss function value; Represents the verification loss function value; Represents the value of the depth consistency loss function.
9. An image processing device for improving a robot's environmental perception capability, characterized in that: include: a data acquisition module configured to: acquire an initial depth image and initial point cloud data of the same target environment area collected by a sensor associated with the robot, and pre-process the initial point cloud data to obtain a point cloud depth image; wherein the sensor includes a laser radar and a depth camera, and the depth camera is capable of collecting depth information of the target environment area; A data processing module is configured to: input the initial depth image and the point cloud depth image into a pre-trained deep learning model to obtain a fused image of the same target environment area output by the pre-trained deep learning model; The fused image includes the depth information in the initial depth image and the initial point cloud data; The pre-trained deep learning model is trained using sample depth images and sample point cloud depth images; the pre-trained deep learning model includes a deep image fusion network; The training process of the pre-trained deep learning model includes: Step 1: Establish a sample data set; wherein the sample data set includes multiple sample depth images of the same target area, multiple frames of sample point clouds at consecutive moments, and corresponding multiple frames of sample point cloud depth images at consecutive moments; wherein, according to the time sequence of the sample point clouds at consecutive moments, a sample depth image and a frame of sample point cloud depth image of the same target area are taken as a set of sample data; Step 2: Input the sample data at the target moment and the sample data at the next moment directly adjacent to the target moment into the deep image fusion network to be trained, respectively, to obtain the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment; Step 3: Based on a preset view synthesis rule, the initial fused image corresponding to the target moment and the initial fused image corresponding to the next moment are synthesized to obtain a simulated co-view image at the target moment; Step 4: Based on the simulated co-visual image at the target moment, calculate the depth consistency loss function value and the smoothing loss function value; Step 5: Use the initial fusion image at the target time to filter the sample depth image at the target time to obtain the filtered sample depth image at the target time as the verification image; Step 6: Input the filtered sample depth image at the target moment and the sample point cloud depth image at the target moment into the depth image fusion network to be trained to obtain a second fused image corresponding to the target moment as an updated image; Step 7: Calculate the verification loss function value using the verification image and the update image; Step 8: Determine a comprehensive loss function value based on the depth consistency loss function value, the smoothing loss function, and the verification loss function value; Step 9: Use the comprehensive loss function value to train the deep image fusion network to be trained until the preset training completion condition is met, and obtain the deep image fusion network from the deep image fusion network to be trained.
Citation Information
Patent Citations
Robot fleet management and additive manufacturing for value chain networks
CA3177985A1
Autonomous underwater vehicle combined navigation system
CN102042835A