An image diffusion generation method, system and electronic device based on target perception
The target perception platform acquires multiple sensor data and generates diffuse rendered images, solving the problem of high-definition video transmission in remote driving control, achieving more efficient data transmission and more realistic environmental perception.
Patent Information
- Application Number
- CN202510594248.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The high-definition video scene display of unmanned devices in remote driving control is limited by the camera resolution and wireless communication bandwidth, which makes it difficult for remote driving controllers to obtain accurate environmental and target information, and high-dimensional data transmission causes the burden of wireless data transmission and lack of real-time real-time performance.
A target perception platform is used to obtain a variety of sensor data of unmanned devices, such as visible light, infrared, lidar and millimeter wave radar data, and target detection and perception results are determined. A diffusion rendered image is generated through the target diffusion model to avoid direct transmission of high-dimensional data.
It improves the real-time data transmission and bandwidth utilization, and the generated diffuse rendered images are more realistic and faster, helping drivers accurately perceive the unmanned equipment environment and support correct driving decisions.
Smart Images

Figure CN120107569B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image diffusion generation method, system and electronic device based on target perception. Background Art
[0002] With the development of artificial intelligence technology, intelligent vehicles have become the focus of research in the high-tech field, and the remote control of unmanned devices has become the most important component. The current remote control process directly transmits the video images of the drone camera to the terminal for display through wireless communication. The remote control operator obtains the scene and environment where the unmanned device is located through the terminal display screen, realizes remote perception of the scene, and remotely controls the unmanned device to complete the operation task. However, limited by the resolution of the unmanned device camera and the bandwidth problem of wireless communication transmission, it is impossible to achieve the terminal display of high-definition video scenes. The remote control operator is difficult to obtain accurate environmental information and target information and cannot accurately complete the operation task. In order to enable the remote control operator to have a more comprehensive understanding of the environment around the boat, the unmanned device will use more perspective cameras to remotely transmit data to the shore-based platform. However, the large amount of video data will cause the burden of up and down wireless data transmission, the transmission real-time performance is not strong, and it consumes traffic. Summary of the Invention
[0003] The present disclosure provides an image diffusion generation method, system, electronic device and storage medium based on target perception to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided an image diffusion generation method based on target perception, the method including: a target perception platform obtains perception data of an unmanned device, the perception data of the unmanned device including: visible light images, infrared images, lidar data, and millimeter wave radar data, and the target perception platform is installed on the unmanned device; performing target detection based on the perception data of the unmanned device and determining a perception result of the target, the perception result of the target including a semantic segmentation map, target category and bounding box information, and three-dimensional point cloud data; transmitting the perception result of the target to an image diffusion generation platform through network data or satellite data; the image diffusion generation platform performs diffusion generation on the perception result of the target through a target diffusion model to generate a diffusion rendering image.
[0005] In an implementable manner, the performing target detection based on the perception data of the unmanned device and determining a perception result of the target includes: inputting the visible light image, the infrared image, and the lidar data into a target convolutional network to obtain a multi-scale feature map; generating anchor boxes on the multi-scale feature map and determining a first target detection result; determining a second target detection result according to the millimeter wave radar data; determining a perception result of the target according to the first target detection result and the second target detection result.
[0006] In one implementable manner, determining the second target detection result according to the millimeter-wave radar data includes: determining a unit to be measured, determining a reference unit and a protection unit based on the unit to be measured; determining an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determining a detection threshold based on the average noise level and a preset threshold factor; comparing the millimeter-wave radar data of the unit to be measured with the detection threshold, and determining the second target detection result according to the comparison result.
[0007] In one implementable manner, the image diffusion generation platform performs diffusion generation on a diffusion rendering image through a target diffusion model according to the perception result of the target, including: inputting the perception result of the target into the target diffusion model, where the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, among which 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 Vision Transformer; extracting features from the semantic segmentation map through the encoding blocks of the target diffusion model to generate an intermediate feature representation, and the decoding blocks of the target diffusion model generating a feature tensor according to the intermediate feature representation; encoding the target category and box information through a first image encoder to obtain a first encoding result; encoding the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; generating a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0008] In one implementable manner, the method further includes: calculating the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image input into the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0009] According to a second aspect of the present disclosure, there is provided an image diffusion generation system based on target perception, the system comprising: a target perception platform for acquiring unmanned device perception data, the unmanned device perception data including: visible light images, infrared images, lidar data, and millimeter wave radar data, the target perception platform being installed on the unmanned device; the target perception platform is further configured to perform target detection based on the unmanned device perception data and determine the perception result of the target, the perception result of the target including a semantic segmentation map, target category and bounding box information, and three-dimensional point cloud data; a data transmission module for transmitting the perception result of the target to an image diffusion generation platform through network data or satellite data; an image diffusion generation platform for generating a diffusion rendering image through a target diffusion model according to the perception result of the target.
[0010] In an implementable embodiment, the target perception platform includes: a data processing module for inputting the visible light image, the infrared image, and the lidar data into a target convolutional network to obtain a multi-scale feature map; a first target detection module for generating anchor boxes on the multi-scale feature map and determining a first target detection result; a second target detection module for determining a second target detection result according to the millimeter wave radar data; a determination module for determining the perception result of the target according to the first target detection result and the second target detection result.
[0011] In an implementable embodiment, the second target detection module is specifically configured to determine a unit under test, determine a reference unit and a protection unit based on the unit under test; determine an average noise level based on the amplitude of the millimeter wave radar data of the reference unit; determine a detection threshold based on the average noise level and a preset threshold factor; compare the millimeter wave radar data of the unit under test with the detection threshold, and determine the second target detection result according to the comparison result.
[0012] In an implementable embodiment, the image diffusion generation platform includes: a processing module for inputting the perception result of the target into the target diffusion model, where the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, among which 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 Vision Transformers; a feature extraction module for extracting features from the semantic segmentation map through the encoding blocks of the target diffusion model to generate an intermediate feature representation, and the decoding blocks of the target diffusion model generate a feature tensor according to the intermediate feature representation; a first encoding module for encoding the target category and box information through a first image encoder to obtain a first encoding result; a second encoding module for encoding the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; a generation module for generating a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0013] In an implementable embodiment, the image diffusion generation platform further includes: a loss calculation module for calculating the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectations of the input data distribution and the noise distribution, represents the initial image input into the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0014] According to the third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present disclosure.
[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.
[0019] A method, system, electronic device, and storage medium for image diffusion generation based on target perception according to the present disclosure, wherein a target perception platform is installed on an unmanned device to obtain perception data of the unmanned device, perform target detection based on the perception data of the unmanned device, determine the perception result of the target, and then transmit the perception result of the target to an image diffusion generation platform. The image diffusion generation platform performs diffusion generation through a target diffusion model according to the perception result of the target to generate a diffusion rendering image. Applying this method, the target perception platform obtains perception data and performs target perception to obtain the perception result of the target, and then only transmits the perception result of the target to the image diffusion generation platform for image diffusion rendering, which can avoid the data transmission burden caused by a large amount of data when directly transmitting the original high-dimensional data to the image diffusion generation platform, improve the real-time performance of data transmission, reduce the bandwidth of the transmitted data, and the image diffusion generation platform located at the driving control end uses rich computing power to generate a more realistic and faster diffusion rendering image of the perception result of the target, which is beneficial for the driving control personnel to perceive the real environment of the unmanned device to make a correct driving choice.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, wherein:
[0022] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0023] Figure 1 Shows the implementation process schematic of a method for image diffusion generation based on target perception according to an embodiment of the present disclosure Figure 1 ;
[0024] Figure 2 Shows the implementation process schematic of a method for image diffusion generation based on target perception according to an embodiment of the present disclosure Figure 2 ;
[0025] Figure 3 Shows the implementation process schematic of a method for image diffusion generation based on target perception according to an embodiment of the present disclosureFigure 3 ;
[0026] Figure 4 Shows a schematic diagram of the modules of an image diffusion generation system based on target perception according to an embodiment of the present disclosure;
[0027] Figure 5 Shows a schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0028] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.
[0029] Figure 1 Shows a schematic implementation process of an image diffusion generation method based on target perception according to an embodiment of the present disclosure Figure 1 , including:
[0030] Step 101, the target perception platform obtains the perception data of the unmanned device. The perception data of the unmanned device includes visible light images, infrared images, lidar data, and millimeter-wave radar data. The target perception platform is installed on the unmanned device.
[0031] To meet the target requirements of remote driving control in all-weather complex environments, multiple sensors are installed on the unmanned device to perceive the targets around the unmanned device, such as dual-channel visible light sensors, infrared sensors, lidar sensors, and millimeter-wave radar sensors, etc.; the infrared sensor can make up for the lack of visible light visibility at night, and the lidar sensor can increase the detection range around the target; the millimeter-wave radar sensor can improve the accuracy of small target detection and recognition in harsh environments such as rain and fog environments, and can also provide high-semantic point cloud information.
[0032] Multiple sensors on the unmanned device perceive the environment to obtain the perception data of the unmanned device including visible light images, infrared images, lidar data, and millimeter-wave radar data, etc. The target perception platform located on the unmanned device can obtain this perception data of the unmanned device.
[0033] Step 102, perform target detection based on the perception data of the unmanned device and determine the perception result of the target. The perception result of the target includes a semantic segmentation map, target category and box information, and three-dimensional point cloud data.
[0034] After the target perception platform obtains the perception data of the unmanned device, namely infrared images, visible light images, lidar data, and millimeter wave radar data, it performs target detection based on the perception data of the unmanned device combined with target detection and perception algorithms to obtain the perception results of the target. For example, preprocessing such as denoising and filtering is performed on the perception data of the unmanned device, and a semantic segmentation model is used to perform pixel-level classification on the images and data to obtain a semantic segmentation map. The semantic segmentation map is the result of classifying each pixel, marking the categories of different targets or regions. A target detection model can be used to detect the targets in the images and data, and a point cloud processing algorithm can be used to detect the targets and extract the three-dimensional point cloud data of the targets, etc. The three-dimensional point cloud data includes the three-dimensional spatial position and shape information of the targets.
[0035] Step 103: Transmit the perception results of the target to the image diffusion generation platform through network data or satellite data.
[0036] After the target perception platform located on the unmanned device obtains the perception results of the target, it transmits the perception results of the target to the image diffusion generation platform through network data or satellite data. Network data (4G or 5G) and / or satellite data transmission have the characteristics of high bandwidth, high-speed transmission, good real-time performance, and high stability, and can meet the data transmission requirements in different scenarios.
[0037] Step 104: The image diffusion generation platform performs diffusion generation on the perception results of the target through the target diffusion model to generate a diffusion rendering image.
[0038] In an image generation task, a diffusion model is a generative model based on a diffusion process, which can generate images that meet specific requirements by introducing additional control conditions such as text, images, and radar point clouds. The target diffusion model of this application is a generative model based on a diffusion process. By gradually adding noise, the data or image is transformed from the original distribution to a Gaussian distribution, and then the original data is recovered from the Gaussian distribution through a reverse process. Inputting the perception results of the target into the target diffusion model for diffusion can generate a diffusion rendering image related to the target. The target diffusion platform can generate high-quality diffusion rendering images according to the perception results of the target.
[0039] The target perception platform of this application obtains the perception data of the unmanned device, conducts target detection on the perception data of the unmanned device and determines the perception result of the target. Then, it transmits the perception result of the target to the image diffusion generation platform, and the image diffusion generation platform performs diffusion generation to generate a diffusion rendering image according to the perception result of the target. Applying this method, after the target perception platform obtains the perception result of the target based on the acquired perception data of the unmanned device, it only transmits the perception result of the target to the image diffusion generation platform for image diffusion rendering. This can avoid the data transmission burden caused by a large amount of data when directly transmitting the original high-dimensional data to the image diffusion generation platform, improve the real-time performance of data transmission, reduce the bandwidth of the transmitted data. The image diffusion generation platform located at the driving control end uses rich computing power to generate a more realistic and faster diffusion rendering image of the perception result of the target, which is beneficial for the driving control personnel to perceive the real environment of the unmanned device and make correct driving choices.
[0040] In an implementable manner, as Figure 2 shown, conducting target detection based on the perception data of the unmanned device and determining the perception result of the target includes:
[0041] Step 201: Input the visible light image, infrared image, and lidar data into the target convolutional network to obtain a multi-scale feature map;
[0042] Step 202: Generate anchor boxes on the multi-scale feature map and determine the first target detection result;
[0043] Step 203: Determine the second target detection result according to the millimeter-wave radar data;
[0044] Step 204: Determine the perception result of the target according to the first target detection result and the second target detection result.
[0045] Input the visible light image, infrared image, and lidar image into the target convolutional neural network. The target convolutional network is used to perform target perception based on the perception data of the unmanned device to output the result of target detection and recognition. The target convolutional network adopted in this application is ResNet18. ResNet18 is a classic convolutional neural network, which includes 18 layers of convolutional layers, pooling layers, and fully connected layers. It mainly solves the problem of gradient disappearance in deep networks through residual connections. The infrared image reflects the thermal radiation characteristics of the target. The visible light image is an RGB image, which contains rich color and texture information. The lidar data is usually in the form of point clouds, which contains the three-dimensional space information of the target. Project the point cloud onto a two-dimensional plane to generate a depth map or intensity map.
[0046] The visible light image, the infrared image, and the depth map or intensity map corresponding to the lidar data are respectively input into independent ResNet18 networks for feature extraction. The infrared features of the target are extracted through ResNet18, the visible light features of the target are extracted through ResNet18, and the depth or intensity features of the target are extracted through ResNet18. Through ResNet18, multi-scale feature maps can be obtained, including large-scale feature maps, medium-scale feature maps, and small-scale feature maps. The large-scale feature maps are shallow feature maps with higher resolution, suitable for capturing the overall layout and background information; the medium-scale feature maps are middle-layer feature maps used to balance resolution and semantic information; the small-scale feature maps are deep feature maps with lower resolution, suitable for capturing details and local features. The outputs of the conv2_x, conv3_x, and conv4_x layers of ResNet18 are the feature maps of the above three scales. Then, convolution and pooling operations are respectively performed on the feature maps of each scale to further extract features. It can be understood that a multi-head attention mechanism can be introduced on the feature maps of each scale, and the long-range dependencies in the feature maps are captured through the multi-head attention mechanism to enhance the model's understanding and expression ability of the image content. By fusing the output of the multi-head attention with the original feature map, an enhanced feature map can be obtained.
[0047] The extracted features are fused by means of channel concatenation or weighted summation to generate multi-modal features, and the fused features are classified and regressed through a fully connected layer or a convolutional layer, and the probability of the target category and the bounding box coordinates are output to obtain the target category and box information.
[0048] Corresponding anchor boxes are generated on the feature maps of each scale. Larger anchor boxes are generated on the large-scale feature maps, suitable for detecting large targets; medium-sized anchor boxes are generated on the medium-scale feature maps, suitable for detecting medium targets; smaller anchor boxes are generated on the small-scale feature maps, suitable for detecting small targets. Each anchor box contains classification and regression bounding box offset information to obtain the first target detection result. Then, target detection is performed on the millimeter-wave radar data to obtain the second target detection result. In this application, the second target detection result can be obtained through a constant false alarm rate algorithm.
[0049] The first target detection result and the second target detection result are fused to determine the perception result of the target. Specifically, the perception result of the target is obtained through the following formula: , where is a normalized fusion function, is the first target detection result, is the second target detection result, is the weight of the second target detection result. By the importance degree of the second target detection result can be adjusted, The value range of
[0050] In addition, a bounding box regression algorithm can be used for regression prediction to adjust the position and size of the bounding box and obtain the precise position of the target for subsequent target recognition and detection.
[0051] To enable the target convolutional network to detect, recognize, and segment targets, a multi-task loss function is used to optimize the model:
[0052]
[0053] where indicates that a target appears in the cell, with a value of 1; indicates that no target appears in the cell, with a value of 0; represents the size of the feature map used to predict the target box and the segmentation map; represents the number of candidate target boxes predicted by the prediction points; indicates traversing the predicted target boxes from [0, ; indicates traversing the predicted target boxes from [0, ; represents the training weight for target box loss prediction; represents the loss training weight for non-target boxes; x and y represent the upper-left coordinates of the target, and w and h represent the width and height of the target respectively. , , , respectively represent the ground truth of the upper-left coordinates, width, and height of the training data; c represents the class prediction of the target, represents the ground truth of the class of the training data, p represents the probability output of predicting the target, represents the ground truth of the probability output of the training data; represents the segmentation result output of the predicted target, represents the ground truth of the segmentation result output of the training data; CE is the cross-entropy loss calculation. In this formula, the cross-entropy of the prediction probability at each point is calculated and summed to obtain the final loss calculation.
[0054] By processing infrared images, visible light images, and lidar data through the ResNet18 network, multi-modal target detection and recognition can be achieved. This method makes full use of the complementarity of different modal data and improves the accuracy and robustness of target detection.
[0055] In an implementable embodiment, as shown in Figure 3 , the second target detection result is determined according to the millimeter-wave radar data, including:
[0056] Step 301: Determine the unit to be measured, and determine the reference unit and the protection unit based on the unit to be measured.
[0057] Step 302: Determine the average noise level based on the amplitude of the millimeter-wave radar data of the reference unit.
[0058] Step 303: Determine the detection threshold based on the average noise level and a preset threshold factor.
[0059] Step 304: Compare the millimeter-wave radar data of the unit to be measured with the detection threshold, and determine the second target detection result according to the comparison result.
[0060] Since the millimeter-wave radar point cloud data is sparse and punctate, a Constant False Alarm Rate (CFAR) algorithm can be designed to detect targets in the millimeter-wave radar data. The purpose of CFAR detection is to maintain a constant false alarm rate under the background of noise and clutter, while effectively detecting targets. First, perform a fast Fourier transform on the millimeter-wave radar data to generate a range-Doppler map. The millimeter-wave radar data will be divided into multiple units, and the size of the unit, that is, the number and physical size of the unit, need to be designed according to the specific parameters and application scenarios of the radar. Then select the unit to be measured, and determine the reference unit and the protection unit based on the unit to be measured. The protection unit is the neighbor unit of the unit to be measured, that is, a protection band is formed around the unit to be measured to avoid the influence of strong target echoes on noise estimation. The protection unit is not used for noise estimation. Secondly, select reference units around the unit to be measured. The reference units are used to estimate the level of background noise. Estimate the background noise level through the amplitude of the millimeter-wave radar data of the reference units. In this application, the average value of the millimeter-wave radar data of the reference units is calculated to obtain the average noise level, and this average noise level is used to estimate the background noise.
[0061] Then determine the detection threshold based on the average noise level and a preset threshold factor, that is, take the product of the average noise level and the preset threshold factor as the detection threshold. The specific value of this preset threshold factor can be determined according to the actual situation. Compare the millimeter-wave radar data of the unit to be measured with the detection threshold to obtain a comparison result. If the comparison result is that the amplitude of the millimeter-wave radar data of the detection unit exceeds the detection threshold, it is determined that there is a target in the unit to be measured. If the comparison result is that the amplitude of the millimeter-wave radar data of the detection unit does not exceed the detection threshold, it is determined that there is no target in the unit to be measured.
[0062] Using the Constant False Alarm Rate algorithm to detect targets in the millimeter-wave radar data can reduce excessive false alarms caused by the intensity asymmetry between clutter and echoes, so as to ensure a relatively constant false alarm rate.
[0063] In an implementable manner, the image diffusion generation platform performs diffusion generation and diffusion rendering of an image through a target diffusion model according to the perception result of a target, including:
[0064] Input the perception result of the target into the target diffusion model. The target diffusion model includes 12 encoding blocks, 12 decoding blocks, and 1 intermediate block, where 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 Vision Transformers; extract features from the semantic segmentation map through the encoding blocks of the target diffusion model to generate intermediate feature representations, and the decoding blocks of the target diffusion model generate feature tensors according to the intermediate feature representations; encode the target category and box information through the first image encoder to obtain a first encoding result; encode the three-dimensional point cloud data through the second image encoder to obtain a second encoding result; generate a diffusion rendering image according to the feature tensor, the first encoding result, and the second encoding result.
[0065] Input the perception result of the target into the target diffusion model for diffusion to generate a diffusion rendering image. The target diffusion model is based on the latent diffusion model Stable diffusion and is obtained through pre-training. The target diffusion model includes 12 encoding blocks, 12 decoding blocks, and 1 intermediate block. Input the semantic segmentation map of the perception result of the target into the target diffusion model. The encoding blocks are used to extract features of the semantic segmentation map to generate a latent space representation. The intermediate block is used to perform diffusion in the latent space to gradually generate intermediate features. The decoding blocks are used to restore the intermediate features to a high-resolution image. The target diffusion model includes 17 main blocks, and each main block includes 4 ResNet layers and 2 Vision Transformers (ViT). The ResNet layer includes multiple residual blocks, and each residual block consists of several convolutional layers and skip connections. The ResNet layer is used to extract local features of the image and capture the detailed information of the image; the ViT is used to capture the global structure and context information of the image. Due to its powerful modeling ability, the ViT is commonly used in diffusion models to process high-level semantic information. By combining the ViT with the ResNet layer, local and global feature fusion can be achieved. By combining the two, the target diffusion model can simultaneously process the local details and global structure of the image, thereby generating high-quality images.
[0066] To make the diffusion-generated diffusion rendering image more controllable, the target category and box information in the perception result of the target and the three-dimensional point cloud data can be used as prompts to control the generation of the rendering image. The first encoder encodes the target category and box information to obtain the first encoding result. In this application, the first encoder uses FrozenCLIP, and FrozenCLIP encoding is used to encode the target category and box information into high-dimensional features, with an output dimension of [1, 4, 256, 384]. The first encoding result is scaled to the dimensions of each layer of the decoding block and concatenated with the feature tensor of the decoding block in the channel dimension to control the content and layout of the generated image. The second encoder encodes the three-dimensional point cloud data to obtain the second encoding result. In this application, the second encoder uses a small UNet network, and the small UNet encoding encodes the three-dimensional point cloud data into low-dimensional features, with the output dimension normalized to [1, 4, 32, 48]. The second encoding result is scaled to the dimensions of each layer of the decoding block and concatenated with the feature tensor of the decoding block in the channel dimension to enhance the spatial perception ability of the generated image. Therefore, the diffusion rendering image is generated according to the output of the decoding block, the first encoding result, and the second encoding result. After that, the generated diffusion rendering image can also be processed such as color correction and sharpening to improve the visual effect.
[0067] In an implementable manner, the method further includes: calculating the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectations of the input data distribution and the noise distribution, represents the initial image input to the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0068] Calculate the loss function of the controllable diffusion model through the above formula, and by minimizing the predicted noise and the real noise The differences enable the target diffusion model to gradually remove noise and generate high-quality diffusion rendering images. The style loss term L is used to constrain the consistency in style between the generated diffusion rendering image and the reference image. By minimizing the style loss term, the style of the generated diffusion rendering image can be made close to that of the reference image, where the reference image can be the original image around the unmanned device captured by the sensor of the unmanned device; is the weight coefficient of the style loss term, which is used to balance the importance of the core loss term and the style loss term. If is larger, the target diffusion model will pay more attention to style consistency. If is smaller, the target diffusion model will pay more attention to the quality of image generation. is the feature vector obtained by converting the text information through the first encoder, is the feature vector obtained by converting the 3D point cloud data through the second encoder.
[0069] ϵ is the Gaussian noise added to the image during the diffusion process, representing a certain rendering style during model training. This application adopts the open-source data style of the latent diffusion model (stable Diffusion). is the style loss term. It is hoped that the style of the generated diffusion rendering image is as close as possible to that of the reference image. In this application, the style is represented by the Gram matrix between the same feature maps in the same hidden layer, and the effect trained by simulating the rendering style is obtained. In addition, the context loss is superimposed on the feature maps of the same layer for stronger image style rendering generation and repair.
[0070] Specifically, ; where is the Gram matrix of the generated diffusion rendering image, is the Gram matrix of the reference image, is the style loss, representing the difference in style between the generated diffusion rendering image and the reference image; the Gram matrix is obtained by calculating the inner product of the feature maps, reflecting the correlation between the features, thereby capturing the texture and style information of the image. By minimizing the style loss, the texture and color distribution of the generated diffusion rendering image will gradually approach the style of the reference image. is the feature map corresponding to the generated diffusion rendering image, is the feature map corresponding to the reference image, It represents the context loss, which indicates the difference in content or semantics between the generated diffusion-rendered image and the reference image. It is usually achieved by comparing the semantic information of feature maps. The feature maps of the context loss are typically high-level features extracted from a pre-trained deep neural network such as a convolutional neural network (Visual Geometry Group, VGG). By minimizing the context loss, the generated diffusion-rendered image will include content information such as object shapes and structures in the reference image. The style loss term L is the weighted sum of the style loss and the context loss , which is used to optimize both the style and content of the generated diffusion-rendered image simultaneously. By minimizing the style loss term, the generated diffusion-rendered image can match the target style while retaining the target content.
[0071] Figure 4 The schematic diagram of the modules of an image diffusion generation system based on target perception according to an embodiment of the present disclosure is shown.
[0072] Referring to Figure 4 , according to the second aspect of the present disclosure, an image diffusion generation system based on target perception is provided. The system includes: a target perception platform 401 for acquiring unmanned device perception data, where the unmanned device perception data includes: visible light images, infrared images, lidar data, and millimeter-wave radar data. The target perception platform is installed on the unmanned device; the target perception platform 401 is further configured to perform target detection based on the unmanned device perception data and determine the perception result of the target. The perception result of the target includes a semantic segmentation map, target category and bounding box information, and three-dimensional point cloud data; a data transmission module 402 for transmitting the perception result of the target to the image diffusion generation platform through network data or satellite data; an image diffusion generation platform 403 for generating a diffusion-rendered image through target diffusion according to the perception result of the target.
[0073] In an implementable manner, the target perception platform 401 includes: a data processing module 4011 for inputting visible light images, infrared images, and lidar data into a target convolutional network to obtain multi-scale feature maps; a first target detection module 4012 for generating anchor boxes on the multi-scale feature maps and determining the first target detection result; a second target detection module 4013 for determining the second target detection result according to the millimeter-wave radar data; and a determination module 4014 for determining the perception result of the target according to the first target detection result and the second target detection result.
[0074] In an implementable embodiment, the second target detection module 4013 is specifically configured to determine a unit to be measured, determine a reference unit and a protection unit based on the unit to be measured; determine an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determine a detection threshold based on the average noise level and a preset threshold factor; compare the millimeter-wave radar data of the unit to be measured with the detection threshold, and determine a second target detection result according to the comparison result.
[0075] In an implementable embodiment, the image diffusion generation platform 403 includes: a processing module 4031, configured to input a perception result of a target into a target diffusion model, the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, where 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 Vision Transformer; a feature extraction module 4032, configured to extract features from a semantic segmentation map through the encoding blocks of the target diffusion model to generate an intermediate feature representation, and the decoding blocks of the target diffusion model generate a feature tensor according to the intermediate feature representation; a first encoding module 4033, configured to encode target category and box information through a first image encoder to obtain a first encoding result; a second encoding module 4034, configured to encode three-dimensional point cloud data through a second image encoder to obtain a second encoding result; a generation module 4035, configured to generate a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0076] In an implementable embodiment, the image diffusion generation platform 403 further includes: a loss calculation module 4036, configured to calculate the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image input into the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0077] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0078] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0079] As Figure 5 shown, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0080] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0081] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as an object-aware image diffusion generation method. For example, in some embodiments, an object-aware image diffusion generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the object-aware image diffusion generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute an object-aware image diffusion generation method by any other suitable means (e.g., by means of firmware).
[0082] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0083] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0084] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0085] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0086] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0087] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0088] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.
[0089] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present disclosure, "a plurality" means two or more, unless otherwise specifically defined.
[0090] As described above, the above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An image diffusion generation method based on target perception, characterized in that The method includes: The target perception platform obtains the perception data of the unmanned device, and the perception data of the unmanned device includes: visible light image, infrared image, lidar data and millimeter wave radar data, and the target perception platform is installed on the unmanned device; Based on the perception data of the unmanned device, target detection is performed and the perception result of the target is determined. The perception result of the target includes a semantic segmentation map, target category and bounding box information, and three-dimensional point cloud data; Transmit the perception result of the target to the image diffusion generation platform through network data or satellite data; The image diffusion generation platform performs diffusion to generate a diffusion rendering image through the target diffusion model according to the perception result of the target; The image diffusion generation platform performs diffusion to generate a diffusion rendering image through the target diffusion model according to the perception result of the target, including: inputting the perception result of the target into the target diffusion model. The target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, where 8 are upsampling or downsampling convolutional layers, and 17 are main blocks. Each main block includes 4 residual network layers and 2 Vision Transformers; extracting features from the semantic segmentation map through the encoding blocks of the target diffusion model to generate an intermediate feature representation, and the decoding blocks of the target diffusion model generate a feature tensor according to the intermediate feature representation; encoding the target category and bounding box information through a first image encoder to obtain a first encoding result; encoding the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; generating a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result; The method further includes: calculating the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image input to the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
2. The method according to claim 1, characterized in that The performing target detection based on the perception data of the unmanned device and determining the perception result of the target includes: Input the visible light image, the infrared image and the lidar data into a target convolutional network to obtain a multi-scale feature map; Generate anchor boxes on the multi-scale feature map to determine a first target detection result; Determine a second target detection result according to the millimeter wave radar data; Determine the perception result of the target according to the first target detection result and the second target detection result.
3. The method according to claim 2, characterized in that, The determining the second target detection result according to the millimeter wave radar data includes: Determine a unit to be measured, and determine a reference unit and a protection unit based on the unit to be measured; Determine the average noise level based on the amplitude of the millimeter wave radar data of the reference unit; Determine a detection threshold based on the average noise level and a preset threshold factor; Compare the millimeter wave radar data of the unit to be measured with the detection threshold, and determine the second target detection result according to the comparison result.
4. An image diffusion generation system based on target perception, characterized in that, The system includes: A target perception platform for obtaining the perception data of the unmanned device. The perception data of the unmanned device includes: visible light image, infrared image, lidar data and millimeter wave radar data, and the target perception platform is installed on the unmanned device; The target perception platform is further configured to perform target detection based on the perception data of the unmanned device and determine the perception result of the target. The perception result of the target includes a semantic segmentation map, target category and bounding box information, and 3D point cloud data; The data transmission module is configured to transmit the perception result of the target to the image diffusion generation platform via network data or satellite data; The image diffusion generation platform is configured to generate a diffusion rendering image through diffusion using the target diffusion model according to the perception result of the target; The image diffusion generation platform includes: a processing module configured to input the perception result of the target into the target diffusion model. The target diffusion model includes 12 encoding blocks, 12 decoding blocks, and 1 intermediate block, where 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 Vision Transformers; a feature extraction module configured to extract features from the semantic segmentation map through the encoding blocks of the target diffusion model to generate an intermediate feature representation, and the decoding blocks of the target diffusion model generate a feature tensor according to the intermediate feature representation; a first encoding module configured to encode the target category and bounding box information through a first image encoder to obtain a first encoding result; a second encoding module configured to encode the 3D point cloud data through a second image encoder to obtain a second encoding result; a generation module configured to generate a diffusion rendering image according to the feature tensor, the first encoding result, and the second encoding result; The image diffusion generation platform further includes: a loss calculation module, which is used to calculate the loss corresponding to the target diffusion model through the following formula: ; where represents the total loss corresponding to the target diffusion model, represents the expectations of the input data distribution and the noise distribution, represents the initial image input into the target diffusion model, represents the number of diffusion times, represents the text conversion feature of the target category and box information, represents the conversion feature of the three-dimensional point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the t-th diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
5. The system according to claim 4, wherein The target perception platform includes: A data processing module configured to input the visible light image, the infrared image, and the lidar data into a target convolutional network to obtain a multi-scale feature map; A first target detection module configured to generate anchor boxes on the multi-scale feature map and determine a first target detection result; A second target detection module configured to determine a second target detection result according to the millimeter-wave radar data; A determination module configured to determine the perception result of the target according to the first target detection result and the second target detection result.
6. The system according to claim 5, wherein The second target detection module is specifically configured to, Determine a unit under test, and determine a reference unit and a protection unit based on the unit under test; Determine an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; Determine a detection threshold based on the average noise level and a preset threshold factor; Compare the millimeter-wave radar data of the unit under test with the detection threshold, and determine the second target detection result according to the comparison result.
7. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-3.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-3.
Citation Information
Patent Citations
Perception fusion system, electronic equipment and storage medium
CN116664997A
Text-guided image processing method based on diffusion model
CN116977489A