Image diffusion generation method and system based on target perception, and electronic equipment
By introducing an image diffusion generation method based on target perception in the remote driving control system, the problem of difficult display of high-definition video scenes in remote driving control technology is solved, more efficient data transmission and more realistic environment perception are achieved, and the accuracy of the work tasks is improved.
Patent Information
- Application Number
- CN202510594248.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Due to the limitations of unmanned equipment camera resolution and wireless communication bandwidth, existing remote driving control technology cannot realize terminal display of high-definition video scenes, making it difficult for remote driving controllers to obtain accurate environmental information and target information, affecting the accurate completion of work tasks.
Using the image diffusion generation method based on target perception, the unmanned device perception data is obtained through the target perception platform, the object detection and perception results are determined, the perception results are transmitted to the image diffusion generation platform, and the diffusion rendered image is generated using the target diffusion model.
This method avoids the direct transmission of original high-dimensional data, reduces the burden of data transmission, improves the real-time and bandwidth efficiency of data transmission, and the generated diffuse rendered images are more realistic and faster, helping drivers and controllers more accurately perceive the environment of unmanned equipment and make the correct driving choices.
Smart Images

Figure CN120107569A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to a method, system and electronic device for generating image diffusion based on target perception. Background Art
[0002] With the development of artificial intelligence technology, intelligent vehicles have become the research focus in the field of high-tech, and remote control of unmanned equipment has become the most important component. The current remote control process is to directly transmit the video image of the drone camera to the terminal display through wireless communication. The remote operator obtains the scene and environment of the unmanned equipment through the terminal display, realizes remote perception of the scene, and remotely controls the unmanned equipment to complete the task. However, due to the resolution of the camera of the unmanned equipment and the bandwidth of wireless communication transmission, the terminal display of high-definition video scenes cannot be realized. It is difficult for remote operators to obtain accurate environmental information and target information, and it is impossible to accurately complete the task. In order to allow remote operators to have a more comprehensive understanding of the environment around the boat, the unmanned equipment will use more viewing angle cameras to remotely transmit data to the shore-based platform, but the amount of video data will cause a burden on the uplink and downlink wireless data transmission, the transmission is not real-time, and it consumes traffic. Summary of the invention
[0003] The present disclosure provides a method, system, electronic device and storage medium for generating image diffusion based on target perception, so as to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, a method for generating image diffusion based on target perception is provided, the method comprising: a target perception platform acquiring perception data of an unmanned device, the perception data of the unmanned device comprising: visible light images, infrared images, lidar data and millimeter wave radar data, the target perception platform being installed on the unmanned device; performing target detection based on the perception data of the unmanned device and determining the perception result of the target, the perception result of the target comprising a semantic segmentation map, target category and frame information and three-dimensional point cloud data; transmitting the perception result of the target to an image diffusion generation platform via network data or satellite data; the image diffusion generation platform generating a diffusion rendering image by diffusion based on the perception result of the target through a target diffusion model.
[0005] In one possible implementation, the target detection based on the unmanned equipment perception data and determining the perception result of the target include: inputting the visible light image, the infrared image and the lidar data into a target convolutional network to obtain a multi-scale feature map; generating an anchor frame on the multi-scale feature map to determine a first target detection result; determining a second target detection result based on the millimeter wave radar data; and determining the perception result of the target based on the first target detection result and the second target detection result.
[0006] In one possible implementation manner, determining the second target detection result based on the millimeter-wave radar data includes: determining a unit to be tested, and determining a reference unit and a protection unit based on the unit to be tested; determining an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determining a detection threshold based on the average noise level and a preset threshold factor; comparing the millimeter-wave radar data of the unit to be tested with the detection threshold, and determining the second target detection result based on the comparison result.
[0007] In one possible implementation, the image diffusion generation platform generates a diffusion rendering image by diffusion according to the perception result of the target through a target diffusion model, including: inputting the perception result of the target into the target diffusion model, the target diffusion model including 12 encoding blocks, 12 decoding blocks and 1 intermediate block, of which 8 are upsampling or downsampling convolutional layers, and 17 are main blocks, each main block including 4 residual network layers and 2 vision transformers; extracting features from the semantic segmentation map through the encoding block of the target diffusion model to generate an intermediate feature representation, and the decoding block of the target diffusion model generates a feature tensor according to the intermediate feature representation; encoding the target category and frame information through a first image encoder to obtain a first encoding result; encoding the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; generating a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0008] In one possible implementation, the method further includes: calculating the loss corresponding to the target diffusion model by the following formula: ;in, represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image of the input target diffusion model, represents the number of diffusions, Text transformation features representing target categories and box information, Represents the transformation characteristics of 3D point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the tth diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0009] According to a second aspect of the present disclosure, there is provided an image diffusion generation system based on target perception, the system comprising: a target perception platform, for acquiring perception data of an unmanned device, the perception data of the unmanned device comprising: visible light images, infrared images, lidar data and millimeter wave radar data, the target perception platform being installed on the unmanned device; the target perception platform, further for performing target detection based on the perception data of the unmanned device and determining the perception result of the target, the perception result of the target comprising a semantic segmentation map, target category and frame information and three-dimensional point cloud data; a data transmission module, for transmitting the perception result of the target to the image diffusion generation platform via network data or satellite data; and the image diffusion generation platform, for generating a diffusion rendering image by diffusion based on the perception result of the target through a target diffusion model.
[0010] In one possible implementation, the target perception platform includes: a data processing module, used to input the visible light image, the infrared image and the lidar data into a target convolutional network to obtain a multi-scale feature map; a first target detection module, used to generate an anchor frame on the multi-scale feature map to determine a first target detection result; a second target detection module, used to determine a second target detection result based on the millimeter wave radar data; and a determination module, used to determine a perception result of the target based on the first target detection result and the second target detection result.
[0011] In one possible implementation manner, the second target detection module is specifically used to determine a unit to be tested, and determine a reference unit and a protection unit based on the unit to be tested; determine an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determine a detection threshold based on the average noise level and a preset threshold factor; compare the millimeter-wave radar data of the unit to be tested with the detection threshold, and determine a second target detection result based on the comparison result.
[0012] In one embodiment, the image diffusion generation platform includes: a processing module, which is used to input the perception result of the target into the target diffusion model, the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, of which 8 are upsampling or downsampling convolutional layers, and 17 are main blocks, each main block includes 4 residual network layers and 2 vision transformers; a feature extraction module, which is used to extract features of the semantic segmentation map through the encoding block of the target diffusion model to generate an intermediate feature representation, and the decoding block of the target diffusion model generates a feature tensor according to the intermediate feature representation; a first encoding module, which is used to encode the target category and frame information through a first image encoder to obtain a first encoding result; a second encoding module, which is used to encode the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; a generation module, which is used to generate a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0013] In one possible implementation manner, the image diffusion generation platform further includes: a loss calculation module, configured to calculate the loss corresponding to the target diffusion model by using the following formula: ;in, represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image of the input target diffusion model, represents the number of diffusions, Text transformation features representing target categories and box information, Represents the transformation characteristics of 3D point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the tth diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0015] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method described in the present disclosure.
[0016] The present invention discloses an image diffusion generation method, system, electronic device and storage medium based on target perception, wherein the target perception platform is installed on the unmanned equipment, obtains the perception data of the unmanned equipment, and performs target detection based on the perception data of the unmanned equipment to determine the perception result of the target, and then transmits the perception result of the target to the image diffusion generation platform, and the image diffusion generation platform diffuses and generates a diffusion rendering image according to the perception result of the target through the target diffusion model. Applying this method, the target perception platform obtains the perception data and performs target perception to obtain the perception result of the target, and then only transmits the perception result of the target to the image diffusion generation platform for image diffusion rendering, which can avoid the data transmission burden caused by the large amount of data when the original high-dimensional data is directly transmitted to the image diffusion generation platform, improve the real-time performance of data transmission, and reduce the bandwidth of the transmission data. The image diffusion generation platform at the driving control end uses abundant computing power to diffuse the perception result of the target to generate a more realistic and faster diffusion rendering image, which is conducive to the driver's perception of the real environment of the unmanned equipment to make the correct driving choice.
[0017] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which: In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0019] Figure 1 The schematic diagram shows the implementation process of a method for generating image diffusion based on target perception according to an embodiment of the present disclosure. Figure 1 ; Figure 2 The schematic diagram shows the implementation process of a method for generating image diffusion based on target perception according to an embodiment of the present disclosure. Figure 2 ; Figure 3 The schematic diagram shows the implementation process of a method for generating image diffusion based on target perception according to an embodiment of the present disclosure. Figure 3 ; Figure 4A module schematic diagram of an image diffusion generation system based on target perception according to an embodiment of the present disclosure is shown; Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0020] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0021] Figure 1 The schematic diagram shows the implementation process of a method for generating image diffusion based on target perception according to an embodiment of the present disclosure. Figure 1 ,include: Step 101, the target perception platform obtains the perception data of the unmanned equipment, the perception data of the unmanned equipment includes: visible light image, infrared image, lidar data and millimeter wave radar data, and the target perception platform is installed on the unmanned equipment.
[0022] In order to achieve the goal of remote driving in complex all-weather environments, a variety of sensors are installed on unmanned equipment to perceive targets around the unmanned equipment, such as dual-channel visible light sensors, infrared sensors, lidar sensors and millimeter-wave radar sensors. The use of infrared sensors can make up for the problem of insufficient visibility of visible light at night, and the use of lidar sensors can increase the detection range around the target. The use of millimeter-wave radar sensors can improve the accuracy of small target detection and recognition in harsh environments such as rain and fog, and can also provide high-semantic point cloud information.
[0023] Various sensors on the unmanned equipment perceive the environment and obtain unmanned equipment perception data including visible light images, infrared images, lidar data, and millimeter wave radar data. The target perception platform located on the unmanned equipment can obtain the unmanned equipment perception data.
[0024] Step 102, performing target detection based on the unmanned equipment perception data and determining the perception result of the target, wherein the perception result of the target includes a semantic segmentation map, target category and frame information, and three-dimensional point cloud data.
[0025] After obtaining the unmanned equipment perception data, namely infrared images, visible light images, lidar data and millimeter wave radar data, the target perception platform performs target detection based on the unmanned equipment perception data combined with target detection and perception algorithms to obtain the perception results of the target, such as denoising, filtering and other pre-processing of the unmanned equipment perception data, and using the semantic segmentation model to perform pixel-level classification of images and data to obtain a semantic segmentation map. The semantic segmentation map is the result of classifying each pixel, marking the categories of different targets or areas. The target detection model can be used to detect targets in images and data, and the point cloud processing algorithm can be used to detect targets and extract the three-dimensional point cloud data of the target. The three-dimensional point cloud data includes the three-dimensional spatial position and shape information of the target.
[0026] Step 103: Transmit the perception result of the target to the image diffusion generation platform via network data or satellite data.
[0027] After obtaining the perception results of the target, the target perception platform located in the unmanned equipment transmits the perception results of the target to the image diffusion generation platform through network data or satellite data. Network data (4G or 5G) and / or satellite data transmission has the characteristics of high bandwidth, high-speed transmission, good real-time performance and high stability, which can meet the data transmission requirements in different scenarios.
[0028] Step 104: the image diffusion generation platform generates a diffusion rendering image by performing diffusion according to the perception result of the target through the target diffusion model.
[0029] In the image generation task, the diffusion model is a generation model based on the diffusion process, which can generate images that meet specific requirements by introducing additional control conditions such as text, images, radar point clouds, etc. The target diffusion model of this application is a generation model based on the diffusion process, which gradually adds noise to convert data or images from the original distribution to a Gaussian distribution, and then restores the original data from the Gaussian distribution through the reverse process. The target's perception results are input into the target diffusion model for diffusion to generate a diffusion rendering image related to the target. The target diffusion platform can generate high-quality diffusion rendering images based on the target's perception results.
[0030] The target perception platform of the present application obtains the perception data of the unmanned equipment, performs target detection on the perception data of the unmanned equipment and determines the perception result of the target, and then transmits the perception result of the target to the image diffusion generation platform, which diffuses and generates a diffusion rendering image based on the perception result of the target. Applying this method, after the target perception platform performs target perception based on the acquired perception data of the unmanned equipment to obtain the perception result of the target, it only transmits the perception result of the target to the image diffusion generation platform for image diffusion rendering, which can avoid the data transmission burden caused by the large amount of data when the original high-dimensional data is directly transmitted to the image diffusion generation platform, improves the real-time performance of data transmission, and reduces the bandwidth of the transmitted data. The image diffusion generation platform at the driving control end uses abundant computing power to diffuse the perception result of the target to generate a more realistic and faster diffusion rendering image, which is conducive to the driver's perception of the real environment of the unmanned equipment so as to make the correct driving choice.
[0031] In one possible implementation, Figure 2 As shown, the target is detected based on the perception data of the unmanned equipment and the perception results of the target are determined, including: Step 201, inputting the visible light image, infrared image and lidar data into the target convolutional network to obtain a multi-scale feature map; Step 202, generating an anchor frame on the multi-scale feature map and determining a first target detection result; Step 203, determining a second target detection result according to the millimeter wave radar data; Step 204: Determine a perception result of the target according to the first target detection result and the second target detection result.
[0032] The visible light image, infrared image and lidar image are input into the target convolutional neural network. The target convolutional network is used to perceive the target according to the perception data of the unmanned equipment to output the result of target detection and recognition. The target convolutional network used in this application is ResNet18. ResNet18 is a classic convolutional neural network, which contains 18 layers including convolutional layer, pooling layer and fully connected layer. It mainly solves the gradient vanishing problem of the deep network through residual connection. The infrared image reflects the thermal radiation characteristics of the target. The visible light image is an RGB image, which contains rich color and texture information. The lidar data is usually in the form of a point cloud, which contains the three-dimensional spatial information of the target. The point cloud is projected onto a two-dimensional plane to generate a depth map or intensity map.
[0033] The depth map or intensity map corresponding to the visible light image, infrared image and lidar data is respectively input into an independent ResNet18 network for feature extraction. The infrared features of the target are extracted by ResNet18, the visible light features of the target are extracted by ResNet18, and the depth or intensity features of the target are extracted by ResNet18. Multi-scale feature maps can be obtained through ResNet18, including large-scale feature maps, medium-scale feature maps and small-scale feature maps. The large-scale feature map is a shallow feature map with high resolution, which is suitable for capturing the overall layout and background information; the medium-scale feature map is a medium-level feature map used to balance resolution and semantic information; the small-scale feature map is a deep feature map with low resolution, which is suitable for capturing details and local features. The output of the conv2_x, conv3_x, and conv4_x layers of ResNet18 is the feature map of the above three scales. After that, convolution and pooling operations are performed on the feature map of each scale to further extract features. It is understandable that a multi-head attention mechanism can be introduced on the feature map of each scale. The multi-head attention mechanism can capture the long-distance dependencies in the feature map, enhance the model's understanding and expression capabilities of the image content, and fuse the output of the multi-head attention with the original feature map to obtain an enhanced feature map.
[0034] The extracted features are fused through channel splicing or weighted summation to generate multimodal features. The fused features are classified and regressed through a fully connected layer or a convolutional layer, and the probability of the target category and the bounding box coordinates are output to obtain the target category and box information.
[0035] Generate a corresponding anchor frame on the feature map of each scale. The large-scale feature map generates a larger anchor frame, which is suitable for detecting large targets. The medium-scale feature map generates a medium-sized anchor frame, which is suitable for detecting medium targets. The small-scale feature map generates a smaller anchor frame, which is suitable for detecting small targets. Each anchor frame contains classification and regression bounding box offset information to obtain the first target detection result. Then perform target detection on the millimeter-wave radar data to obtain the second target detection result. The present application can obtain the second target detection result through the constant false alarm rate algorithm.
[0036] The first target detection result and the second target detection result are fused to determine the target perception result. Specifically, the target perception result is obtained by the following formula: ,in, is the normalized fusion function, is the first target detection result, is the second target detection result, is the weight of the second target detection result, through The importance of the second target detection result can be adjusted. The value range of is [0,1]. The second target detection result can serve as additional prior information to enhance the accuracy of classification.
[0037] In addition, the bounding box regression algorithm can be used for regression prediction, the position and size of the bounding box can be adjusted, and the precise position of the target can be obtained for subsequent target recognition and detection.
[0038] In order to enable the target convolutional network to detect, identify and segment targets, a multi-task loss function is used to optimize the model:
[0039] in, Indicates that a target appears in the cell, and the value is 1; Indicates that no target appears in the cell, and the value is 0; Indicates the size of the feature map used to predict the target box and segmentation map; Indicates the number of candidate target boxes predicted by the prediction point; Indicates that from [0, ]Traverse the predicted target box; Indicates that from [0, ]Traverse the predicted target box; Represents the training weights for target box loss prediction; represents the loss training weight without the target box; x, y represent the coordinates of the upper left corner of the target, w, h represent the width and height of the target respectively, , , , They represent the true values of the upper left corner coordinates, width, and height of the training data respectively; c represents the category prediction of the target, Represents the true value of the training data category, p represents the probability output of predicting the target, Represents the true value of the probability output of the training data; Represents the segmentation result output of the predicted target, represents the true value of the segmentation result output of the training data; CE is the cross entropy loss calculation. In this formula, the predicted probability cross entropy of each point is calculated and summed to obtain the final loss calculation.
[0040] By processing infrared images, visible light images and lidar data through the ResNet18 network, multimodal target detection and recognition can be achieved. This method makes full use of the complementarity of different modal data and improves the accuracy and robustness of target detection.
[0041] In one possible implementation, Figure 3 As shown, according to the millimeter wave radar data, determining the second target detection result includes: Step 301, determining a unit to be tested, and determining a reference unit and a protection unit based on the unit to be tested; Step 302, determining an average noise level based on the amplitude of the millimeter wave radar data of the reference unit; Step 303, determining a detection threshold based on an average noise level and a preset threshold factor; Step 304: compare the millimeter-wave radar data of the unit to be tested with the detection threshold, and determine the second target detection result according to the comparison result.
[0042] Since the millimeter-wave radar point cloud data is sparse and point-shaped, a constant false alarm rate algorithm (CFAR) can be designed to detect targets on millimeter-wave radar data. The purpose of CFAR detection is to maintain a constant false alarm rate and effectively detect targets under the background of noise and clutter. First, the millimeter-wave radar data is fast Fourier transformed to generate a range-Doppler map. The millimeter-wave radar data will be divided into multiple units. The size of the unit, that is, the number and physical size of the unit, needs to be designed according to the specific parameters and application scenarios of the radar. Then, the unit to be tested is selected, and the reference unit and the protection unit are determined based on the unit to be tested. The protection unit is the neighboring unit of the unit to be tested, that is, a protection band is formed around the unit to be tested to avoid strong target echoes affecting noise estimation. The protection unit is not used for noise estimation; secondly, a reference unit is selected around the unit to be tested. The reference unit is used to estimate the level of background noise. The background noise level is estimated by the amplitude of the millimeter-wave radar data of the reference unit. This application calculates the average value based on the amplitude of the millimeter-wave radar data of the reference unit to obtain the average noise level, and the average noise level is used to estimate the background noise.
[0043] Then, the detection threshold is determined based on the average noise level and the preset threshold factor, that is, the product of the average noise level and the preset threshold factor is used as the detection threshold, and the specific value of the preset threshold factor can be determined according to the actual situation. The millimeter-wave radar data of the unit to be tested is compared with the detection threshold to obtain a comparison result. If the comparison result is that the amplitude of the millimeter-wave radar data of the detection unit exceeds the detection threshold, it is determined that there is a target in the unit to be tested. If the comparison result is that the amplitude of the millimeter-wave radar data of the detection unit does not exceed the detection threshold, it is determined that there is no target in the unit to be tested.
[0044] Using a constant false alarm rate algorithm to detect targets on millimeter-wave radar data can reduce the excessive false alarms caused by the intensity asymmetry between clutter and echo, thereby ensuring a relatively constant false alarm rate.
[0045] In one possible implementation, the image diffusion generation platform generates a diffusion rendering image by diffusion based on the perception result of the target through the target diffusion model, including: The perception result of the target is input into the target diffusion model, which includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, of which 8 are upsampling or downsampling convolutional layers and 17 are main blocks. Each main block includes 4 residual network layers and 2 vision transformers. The encoding block of the target diffusion model is used to extract features from the semantic segmentation map to generate an intermediate feature representation, and the decoding block of the target diffusion model generates a feature tensor based on the intermediate feature representation. The target category and box information are encoded by the first image encoder to obtain a first encoding result. The three-dimensional point cloud data is encoded by the second image encoder to obtain a second encoding result. A diffusion rendering image is generated based on the feature tensor, the first encoding result and the second encoding result.
[0046] The perception result of the target is input into the target diffusion model for diffusion to generate a diffusion rendering image. The target diffusion model is based on the potential diffusion model Stable diffusion and is pre-trained. The target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block. The semantic segmentation map of the perception result of the target is input into the target diffusion model. The encoding block is used to extract the features of the semantic segmentation map and generate a latent space representation. The intermediate block is used to diffuse in the latent space and gradually generate intermediate features. The decoding block is used to restore the intermediate features to a high-resolution image. The target diffusion model includes 17 main blocks, each of which includes 4 residual network layers ResNet and 2 visual converters Vision Transformer (ViT). The ResNet layer includes multiple residual blocks, each of which is composed of several convolutional layers and jump connections. The ResNet layer is used to extract local features of the image and capture the detailed information of the image; ViT is used to capture the global structure and contextual information of the image. Due to its powerful modeling ability, ViT is often used to process high-level semantic information in the diffusion model. ViT combined with the ResNet layer can realize the fusion of local and global features. By combining the two, the target diffusion model can simultaneously process the local details and global structure of the image, thereby generating high-quality images.
[0047] In order to make the diffusion rendering image generated by diffusion more controllable, the target category and frame information in the perception result of the target and the three-dimensional point cloud data can be used as prompts to control the generation of the rendered image. The first encoding result is obtained by encoding the target category and frame information through the first encoder. The first encoder of the present application adopts FrozenCLIP. FrozenCLIP encoding is used to encode the target category and frame information into high-dimensional features. The output dimension is [1, 4, 256, 384]. The first encoding result is scaled to the dimension of each layer of the decoding block, and the channel is spliced with the feature tensor of the decoding block to control the content and layout of the generated image. The second encoding result is obtained by encoding the three-dimensional point cloud data through the second encoder. The second encoder of the present application adopts a small UNet network. The small UNet encoding encodes the three-dimensional point cloud data into low-dimensional features, and the output dimension is normalized to [1, 4, 32, 48]. The second encoding result is scaled to the dimension of each layer of the decoding block, and the channel is spliced with the feature tensor of the decoding block to enhance the spatial perception ability of the generated image. Therefore, a diffusion rendering image is generated according to the output of the decoding block, the first encoding result, and the second encoding result. Then, the generated diffusion rendering image may be subjected to color correction, sharpening, and other processing to improve the visual effect.
[0048] In one possible implementation, the method further includes: calculating the loss corresponding to the target diffusion model by the following formula: ;in, represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image of the input target diffusion model, represents the number of diffusions, Text transformation features representing target categories and box information, Represents the transformation characteristics of 3D point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the tth diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0049] The loss function of the controllable diffusion model is calculated by the above formula, by minimizing the prediction noise With real noise The difference enables the target diffusion model to gradually remove noise and generate a high-quality diffusion rendering image. The style loss term L is used to constrain the consistency of the generated diffusion rendering image and the reference image in style. By minimizing the style loss term, the style of the generated diffusion rendering image can be close to the reference image, where the reference image can be the original image around the unmanned device taken by the sensor of the unmanned device; Is the weight coefficient of the style loss term, which is used to balance the importance of the core loss term and the style loss term. If If the target diffusion model is larger, it will pay more attention to style consistency. If it is smaller, the target diffusion model will pay more attention to the quality of image generation. is the feature vector after the text information is converted by the first encoder, is the feature vector obtained by converting the three-dimensional point cloud data through the second encoder.
[0050] ϵ is the Gaussian noise added to the image during the diffusion process, which represents a certain rendering style during model training. This application uses the open source data style of the potential diffusion model (stable Diffusion). It is a style loss term. It is hoped that the style of the generated diffusion rendering image is as close to the style of the reference image as possible. This application represents the style through the Gram matrix Gram between the same feature maps in the same hidden layer to simulate the effect of rendering style training. In addition, the context loss is superimposed on the feature maps of the same layer for stronger image style rendering generation and restoration.
[0051] Specifically, ;in, is the Gram matrix of the generated diffuse rendered image, is the Gram matrix of the reference image, is the style loss, which indicates the difference in style between the generated diffuse rendering image and the reference image. The Gram matrix is obtained by calculating the inner product of the feature map, which reflects the correlation between the features, thereby capturing the texture and style information of the image. By minimizing the style loss, the generated diffuse rendering image will gradually approach the style of the reference image in terms of texture and color distribution. is the feature map corresponding to the generated diffusion rendering image, is the feature map corresponding to the reference image, The context loss represents the difference in content or semantics between the generated diffuse rendering image and the reference image. It is usually achieved by comparing the semantic information of the feature map. The feature map of the context loss is usually a high-level feature extracted from a pre-trained deep neural network such as a convolutional neural network (Visual Geometry Group, VGG). By minimizing the context loss, the generated diffuse rendering image will include content information such as object shape and structure in the reference image. The style loss term L is the style loss. and context loss The weighted sum of is used to simultaneously optimize the style and content of the generated diffusion rendering image. By minimizing the style loss term, the generated diffusion rendering image can match the target style while retaining the target content.
[0052] Figure 4 A module schematic diagram of a target-aware image diffusion generation system according to an embodiment of the present disclosure is shown.
[0053] See also Figure 4 According to the second aspect of the present disclosure, a target perception-based image diffusion generation system is provided, the system comprising: a target perception platform 401, used to obtain unmanned equipment perception data, the unmanned equipment perception data comprising: visible light images, infrared images, lidar data and millimeter wave radar data, the target perception platform is installed on the unmanned equipment; the target perception platform 401 is also used to perform target detection based on the unmanned equipment perception data and determine the perception result of the target, the perception result of the target comprising a semantic segmentation map, target category and frame information and three-dimensional point cloud data; a data transmission module 402 is used to transmit the perception result of the target to the image diffusion generation platform via network data or satellite data; an image diffusion generation platform 403 is used to generate a diffusion rendering image by diffusion according to the perception result of the target through a target diffusion model.
[0054] In one embodiment, the target perception platform 401 includes: a data processing module 4011, which is used to input visible light images, infrared images and lidar data into a target convolutional network to obtain a multi-scale feature map; a first target detection module 4012, which is used to generate an anchor frame on the multi-scale feature map and determine a first target detection result; a second target detection module 4013, which is used to determine a second target detection result based on millimeter wave radar data; and a determination module 4014, which is used to determine a perception result of the target based on the first target detection result and the second target detection result.
[0055] In one possible implementation mode, the second target detection module 4013 is specifically used to determine the unit to be tested, determine the reference unit and the protection unit based on the unit to be tested; determine the average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determine the detection threshold based on the average noise level and a preset threshold factor; compare the millimeter-wave radar data of the unit to be tested with the detection threshold, and determine the second target detection result according to the comparison result.
[0056] In one embodiment, the image diffusion generation platform 403 includes: a processing module 4031, which is used to input the perception result of the target into the target diffusion model, the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, of which 8 are upsampling or downsampling convolutional layers, and 17 are main blocks, each of which includes 4 residual network layers and 2 vision transformers; a feature extraction module 4032, which is used to extract features of the semantic segmentation map through the encoding block of the target diffusion model to generate an intermediate feature representation, and the decoding block of the target diffusion model generates a feature tensor according to the intermediate feature representation; a first encoding module 4033, which is used to encode the target category and frame information through a first image encoder to obtain a first encoding result; a second encoding module 4034, which is used to encode the three-dimensional point cloud data through a second image encoder to obtain a second encoding result; a generation module 4035, which is used to generate a diffusion rendering image according to the feature tensor, the first encoding result and the second encoding result.
[0057] In one embodiment, the image diffusion generation platform 403 further includes: a loss calculation module 4036, which is used to calculate the loss corresponding to the target diffusion model by the following formula: ;in, represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image of the input target diffusion model, represents the number of diffusions, Text transformation features representing target categories and box information, Represents the transformation characteristics of 3D point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the tth diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
[0058] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0059] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0060] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0061] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0062] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as a method for generating image diffusion based on target perception. For example, in some embodiments, a method for generating image diffusion based on target perception may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for generating image diffusion based on target perception described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured in any other appropriate manner (eg, by means of firmware) to execute a target-aware based image diffusion generation method.
[0063] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0064] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0065] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0066] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0067] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0068] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0069] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0070] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0071] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A method for generating image diffusion based on target perception, characterized in that: The method comprises: The target perception platform acquires the perception data of the unmanned equipment, wherein the perception data of the unmanned equipment includes: visible light images, infrared images, laser radar data and millimeter wave radar data, and the target perception platform is installed on the unmanned equipment; Performing target detection based on the unmanned equipment perception data and determining a perception result of the target, wherein the perception result of the target includes a semantic segmentation map, target category and frame information, and three-dimensional point cloud data; Transmitting the perception result of the target to the image diffusion generation platform via network data or satellite data; The image diffusion generation platform generates a diffusion rendering image by diffusion based on the perception result of the target through a target diffusion model.
2. The method according to claim 1, characterized in that The performing target detection based on the unmanned equipment perception data and determining the perception result of the target includes: Inputting the visible light image, the infrared image and the lidar data into a target convolutional network to obtain a multi-scale feature map; Generating an anchor frame on the multi-scale feature map to determine a first target detection result; Determining a second target detection result according to the millimeter-wave radar data; A perception result of the target is determined according to the first target detection result and the second target detection result.
3. The method according to claim 2, characterized in that The determining, according to the millimeter-wave radar data, a second target detection result includes: Determine a unit to be tested, and determine a reference unit and a protection unit based on the unit to be tested; determining an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determining a detection threshold based on the average noise level and a preset threshold factor; The millimeter-wave radar data of the unit to be tested is compared with the detection threshold, and a second target detection result is determined according to the comparison result.
4. The method according to claim 1, characterized in that: The image diffusion generation platform generates a diffusion rendering image by diffusion according to the perception result of the target through a target diffusion model, including: Inputting the perception result of the target into the target diffusion model, the target diffusion model includes 12 encoding blocks, 12 decoding blocks and 1 intermediate block, of which 8 are upsampling or downsampling convolutional layers, 17 are main blocks, and each main block includes 4 residual network layers and 2 vision transformers; Performing feature extraction on the semantic segmentation map through the encoding block of the target diffusion model to generate an intermediate feature representation, and the decoding block of the target diffusion model generates a feature tensor according to the intermediate feature representation; Encoding the target category and the frame information by a first image encoder to obtain a first encoding result; Encoding the three-dimensional point cloud data by a second image encoder to obtain a second encoding result; A diffusion rendering image is generated according to the feature tensor, the first encoding result, and the second encoding result.
5. The method according to claim 4, characterized in that The method further comprises: The loss corresponding to the target diffusion model is calculated by the following formula: ; in, represents the total loss corresponding to the target diffusion model, represents the expectation of the input data distribution and the noise distribution, represents the initial image of the input target diffusion model, represents the number of diffusions, Text transformation features representing target categories and box information, Represents the transformation characteristics of 3D point cloud data, is the Gaussian noise added to the image during the diffusion process, is the noise predicted by the target diffusion model, represents the image after the tth diffusion, || || is the core loss term of the target diffusion model, is the style loss term of the target diffusion model, is the weight coefficient of the style loss term.
6. An image diffusion generation system based on target perception, characterized in that: The system comprises: A target perception platform, used to obtain perception data of unmanned equipment, wherein the perception data of unmanned equipment includes: visible light images, infrared images, laser radar data and millimeter wave radar data, and the target perception platform is installed on the unmanned equipment; The target perception platform is further used to perform target detection based on the unmanned equipment perception data and determine the perception result of the target, wherein the perception result of the target includes a semantic segmentation map, target category and frame information, and three-dimensional point cloud data; A data transmission module, used for transmitting the perception result of the target to the image diffusion generation platform via network data or satellite data; The image diffusion generation platform is used to generate a diffusion rendering image by diffusion through a target diffusion model according to the perception result of the target.
7. The system according to claim 6, characterized in that The target perception platform includes: A data processing module, used for inputting the visible light image, the infrared image and the lidar data into a target convolution network to obtain a multi-scale feature map; A first object detection module, used to generate an anchor frame on the multi-scale feature map and determine a first object detection result; A second target detection module, used to determine a second target detection result according to the millimeter wave radar data; A determination module is used to determine a perception result of a target based on the first target detection result and the second target detection result.
8. The system according to claim 7, characterized in that The second target detection module is specifically used to: Determine a unit to be tested, and determine a reference unit and a protection unit based on the unit to be tested; determining an average noise level based on the amplitude of the millimeter-wave radar data of the reference unit; determining a detection threshold based on the average noise level and a preset threshold factor; The millimeter-wave radar data of the unit to be tested is compared with the detection threshold, and a second target detection result is determined according to the comparison result.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Perception fusion system, electronic equipment and storage medium
CN116664997A
Text-guided image processing method based on diffusion model
CN116977489A
Method of generating partial area of image by using generative model and electronic device for performing the method
US20250078366A1
Cited By
Distributed photovoltaic power station surveying method and system based on unmanned aerial vehicle
CN120320711A