Data processing method and device

By establishing a closed-loop mechanism between scene perception and generation and combining the simulation framework of perception model and diffusion model, the technical problems between scene perception and generation are solved, high-precision and high-realism scene generation and perception are achieved, the performance of perception model is improved, and powerful tools and new possibilities are provided.

CN120689873APending Publication Date: 2025-09-23BEIJING PHIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463609.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, scene perception and scene generation are two independent aspects. The technical problem that is difficult to effectively solve with existing technologies is how to achieve mutual promotion between scene perception and scene generation, especially in processing tasks beyond data distribution. The technical problem of how to enhance scene perception is how to perceive the scene.

Method used

A technical means is adopted to establish a closed-loop mechanism between scene perception and generation, combine the perception model and the diffusion model, form a simulation framework, and realize the integration of scene perception and generation. Through the collaborative training of the perception model and the diffusion model in the simulation framework, high-precision and high-realism scenes are generated.

Benefits of technology

It achieves efficient conversion from text to real scenes, improves the performance of perception models, generates high-quality 3D realistic videos, and provides powerful tools and new possibilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689873A_ABST
    Figure CN120689873A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device. And acquiring the description text, acquiring image features of the noisy image, and acquiring the noisy image according to the image features of the noisy image. Obtaining a three-dimensional semantic occupancy grid of the noisy image based on the noisy image by using a perception model in the simulation framework, and denoising the image features of the noisy image based on the description text and the three-dimensional semantic occupancy grid of the noisy image by using a diffusion model in the simulation framework to obtain the noisy image. Image features of the denoised image are obtained, the content theme of the denoised image comprises the description text, and the denoised image is obtained according to the image features of the denoised image. The diffusion model priori integrates fine and rich semantic information and geometric information in a three-dimensional semantic occupancy grid of a noisy image for conditional processing, and assists the diffusion model to generate a high-precision, high-integrity and high-reality image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a data processing method and device. Background Art

[0002] At present, scene perception and scene generation are two independent aspects. Scene perception can assist scene generation to obtain higher quality scenes, and scene generation can also assist in enhancing the generalization ability of scene perception. Especially in processing tasks beyond the data distribution, the role of enhancing the generalization ability of scene perception is particularly obvious. Summary of the Invention

[0003] This application provides a data processing method and device.

[0004] In a first aspect, the present application provides a data processing method, the method comprising:

[0005] Get the description text and the image features of the input image of the Nth round, where N is a positive integer;

[0006] Obtaining an input image for the Nth round according to the image features of the input image for the Nth round;

[0007] Obtain a three-dimensional semantic occupancy grid of the input image of the Nth round based on the input image of the Nth round using the perception model in the simulation framework;

[0008] Denoising image features of the input image of the Nth round based on the description text and a three-dimensional semantic occupancy grid of the input image of the Nth round using a diffusion model in a simulation framework to obtain image features of an output image of the Nth round, wherein the content theme of the output image of the Nth round includes the description text;

[0009] The output image of the Nth round is obtained according to the image features of the output image of the Nth round.

[0010] In an optional implementation, N is equal to 1;

[0011] The obtaining of image features of the input image of the Nth round includes:

[0012] Get random noise characteristics;

[0013] The image features of the input image in the first round are obtained based on the random noise features.

[0014] In an optional implementation, N is greater than or equal to 2;

[0015] The obtaining of image features of the input image of the Nth round includes:

[0016] Obtaining image features of the output image of the N-1th round; the image features of the output image of the N-1th round are obtained by denoising the image features of the input image of the N-1th round based on the description text and the three-dimensional semantic occupancy grid of the input image of the N-1th round using a diffusion model; the three-dimensional semantic occupancy grid of the input image of the N-1th round is obtained based on the input image of the N-1th round using a perception model;

[0017] The image features of the input image of the Nth round are obtained according to the image features of the output image of the N-1th round.

[0018] In an optional implementation, the method further includes:

[0019] Acquire multiple sample data, where each sample data includes a noise-free image, a description text of the noise-free image, and a three-dimensional semantic occupancy grid of the noise-free image; the description text of the noise-free image is used to describe the content theme of the noise-free image;

[0020] Use multiple sample data to train the perception model and diffusion model in the simulation framework;

[0021] The perception model is used to obtain a three-dimensional semantic occupancy grid of a noisy image based on a noisy image. The noisy image is obtained based on image features of the noisy image. The image features of the noisy image are obtained by superimposing random noise features on image features of the noise-free image.

[0022] The diffusion model is used to denoise image features of the noisy image based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image, so as to obtain image features of the denoised image.

[0023] In an optional implementation, the training of the perception model and the diffusion model in the simulation framework using the plurality of sample data includes:

[0024] Dividing the plurality of sample data into a plurality of first sample data and a plurality of second sample data;

[0025] Using the plurality of first sample data to train the diffusion model in the simulation framework separately, and in the process of using the plurality of first sample data to train the diffusion model in the simulation framework separately, keeping the parameters of the perception model unchanged;

[0026] The plurality of second sample data are used to collaboratively train the perception model and the diffusion model in the simulation framework.

[0027] In an optional implementation, the collaborative training of the perception model and the diffusion model in the simulation framework using the plurality of second sample data includes:

[0028] Acquire image features of a noise-free image in the second sample data;

[0029] The image features of the noise-free image are superimposed with random noise features to obtain the image features of the noisy image;

[0030] obtaining a noisy image according to image features of the noisy image;

[0031] Obtaining a 3D semantic occupancy grid of a noisy image based on the noisy image using a perception model;

[0032] Denoising the image features of the noisy image using a diffusion model based on the description text of the noise-free image in the second sample data and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image;

[0033] Obtaining a denoised image according to image features of the denoised image;

[0034] The perception model and the diffusion model in the simulation framework are collaboratively trained based on the denoised image, the 3D semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free 3D semantic occupancy grid in the second sample data.

[0035] In an optional implementation, the collaborative training of the perception model and the diffusion model in the simulation framework based on the denoised image, the three-dimensional semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free three-dimensional semantic occupancy grid in the second sample data includes:

[0036] Obtain the loss value corresponding to the diffusion model based on the noise-free image and the denoised image;

[0037] Obtaining a loss value corresponding to the perception model based on the 3D semantic occupancy grid of the noisy image and the 3D semantic occupancy grid without noise;

[0038] According to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model, the parameters in the perception model and the parameters in the diffusion model are collaboratively optimized.

[0039] In an optional implementation, the collaborative optimization of parameters in the perception model and parameters in the diffusion model according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model includes:

[0040] Obtain the comprehensive loss value of the simulation framework according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model;

[0041] The parameters in the perception model and the parameters in the diffusion model are collaboratively optimized according to the comprehensive loss value of the simulation framework.

[0042] In an optional implementation, obtaining the comprehensive loss value of the simulation framework according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model includes:

[0043] Determine the weight of the loss value corresponding to the perception model according to the content of random noise features in the image features of the noisy image; the weight is greater than 0 and less than 1;

[0044] The higher the content of random noise features in the image features of the noisy image, the smaller the weight of the loss value corresponding to the perception model; or the lower the content of random noise features in the image features of the noisy image, the greater the weight of the loss value corresponding to the perception model;

[0045] Calculating the product between the weight and the loss value corresponding to the perception model;

[0046] The sum of the product and the loss value corresponding to the diffusion model is calculated to obtain the comprehensive loss value of the entire framework.

[0047] In an optional implementation, the method further includes:

[0048] Obtain a three-dimensional semantic occupancy grid of the output image of the Nth round based on the output image of the Nth round using the perception model in the simulation framework;

[0049] A training data pair is generated according to the three-dimensional semantic occupancy grid of the output image of the Nth round and the output image of the Nth round, and the training data pair is used to at least train the perception model.

[0050] In a second aspect, the present application provides a data processing device, comprising:

[0051] The first acquisition module is used to acquire the description text, and the second acquisition module is used to acquire the image features of the input image of the Nth round, where N is a positive integer;

[0052] A third acquisition module is used to acquire the input image of the Nth round according to the image features of the input image of the Nth round;

[0053] a fourth acquisition module, configured to acquire a three-dimensional semantic occupancy grid of the input image of the Nth round based on the input image of the Nth round using the perception model in the simulation framework;

[0054] a denoising module, configured to denoise image features of the input image of the Nth round using a diffusion model in the simulation framework based on the description text and a three-dimensional semantic occupancy grid of the input image of the Nth round, to obtain image features of an output image of the Nth round, wherein the content theme of the output image of the Nth round includes the description text;

[0055] The fifth acquisition module is used to acquire the output image of the Nth round according to the image features of the output image of the Nth round.

[0056] In an optional implementation, N is equal to 1;

[0057] The second acquisition module includes:

[0058] A first acquisition unit is used to acquire random noise characteristics;

[0059] The second acquisition unit is used to acquire the image features of the input image of the first round according to the random noise features.

[0060] In an optional implementation, N is greater than or equal to 2;

[0061] The second acquisition module includes:

[0062] a third acquisition unit, configured to acquire image features of the output image of the N-1th round; the image features of the output image of the N-1th round are obtained by denoising the image features of the input image of the N-1th round using a diffusion model based on the description text and the three-dimensional semantic occupancy grid of the input image of the N-1th round; the three-dimensional semantic occupancy grid of the input image of the N-1th round is acquired based on the input image of the N-1th round using a perception model;

[0063] The fourth acquiring unit is configured to acquire the image features of the input image of the Nth round according to the image features of the output image of the N-1th round.

[0064] In an optional implementation, the apparatus further includes:

[0065] a sixth acquisition module, configured to acquire a plurality of sample data, wherein each sample data includes a noise-free image, a description text of the noise-free image, and a three-dimensional semantic occupancy grid of the noise-free image; the description text of the noise-free image is used to describe the content theme of the noise-free image;

[0066] A training module, used to train the perception model and the diffusion model in the simulation framework using multiple sample data;

[0067] The perception model is used to obtain a three-dimensional semantic occupancy grid of a noisy image based on a noisy image. The noisy image is obtained based on image features of the noisy image. The image features of the noisy image are obtained by superimposing random noise features on image features of the noise-free image.

[0068] The diffusion model is used to denoise image features of the noisy image based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image, so as to obtain image features of the denoised image.

[0069] In an optional implementation, the training module includes:

[0070] a dividing unit, configured to divide the plurality of sample data into a plurality of first sample data and a plurality of second sample data;

[0071] A first training unit is configured to separately train the diffusion model in the simulation framework using the plurality of first sample data, and to keep parameters of the perception model unchanged during the separate training of the diffusion model in the simulation framework using the plurality of first sample data;

[0072] The second training unit is used to collaboratively train the perception model and the diffusion model in the simulation framework using a plurality of second sample data.

[0073] In an optional implementation, the second training unit includes:

[0074] A first acquisition subunit is used to acquire image features of a noise-free image in the second sample data;

[0075] a superposition subunit, configured to superimpose random noise features on image features of a noise-free image to obtain image features of a noisy image;

[0076] a second acquisition subunit, configured to acquire a noisy image according to image features of the noisy image;

[0077] a third acquisition subunit, configured to acquire a three-dimensional semantic occupancy grid of the noisy image based on the noisy image using the perception model;

[0078] a denoising subunit, configured to denoise image features of the noisy image using a diffusion model based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image in the second sample data, to obtain image features of the denoised image;

[0079] a fourth acquisition subunit, configured to acquire a denoised image according to image features of the denoised image;

[0080] A training subunit is used to collaboratively train the perception model and the diffusion model in the simulation framework based on the denoised image, the three-dimensional semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free three-dimensional semantic occupancy grid in the second sample data.

[0081] In an optional implementation, the training subunit is specifically used to: obtain a loss value corresponding to the diffusion model based on the noise-free image and the denoised image; obtain a loss value corresponding to the perception model based on the three-dimensional semantic occupancy grid of the noisy image and the noise-free three-dimensional semantic occupancy grid; and collaboratively optimize the parameters in the perception model and the parameters in the diffusion model based on the loss value corresponding to the diffusion model and the loss value corresponding to the perception model.

[0082] In an optional implementation, the training subunit is specifically used to: obtain the comprehensive loss value of the simulation framework based on the loss value corresponding to the diffusion model and the loss value corresponding to the perception model; and collaboratively optimize the parameters in the perception model and the parameters in the diffusion model based on the comprehensive loss value of the simulation framework.

[0083] In an optional implementation, the training subunit is specifically used to: determine the weight of the loss value corresponding to the perception model based on the content of random noise features in the image features of the noisy image; the weight is greater than 0 and less than 1; wherein, the higher the content of random noise features in the image features of the noisy image, the smaller the weight of the loss value corresponding to the perception model, or, the lower the content of random noise features in the image features of the noisy image, the greater the weight of the loss value corresponding to the perception model; calculate the product between the weight and the loss value corresponding to the perception model; calculate the sum of the product and the loss value corresponding to the diffusion model to obtain the comprehensive loss value of the overall framework.

[0084] In an optional implementation, the apparatus further includes:

[0085] a seventh acquisition module, configured to acquire a three-dimensional semantic occupancy grid of the output image of the Nth round based on the output image of the Nth round using the perception model in the simulation framework;

[0086] A generation module is used to generate a training data pair based on the three-dimensional semantic occupancy grid of the output image of the Nth round and the output image of the Nth round, where the training data pair is at least used to train the perception model.

[0087] In a third aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.

[0088] In a fourth aspect, the present application shows a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.

[0089] In a fifth aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method as described in any one of the above aspects.

[0090] The technical solution provided by this application may have the following beneficial effects:

[0091] In the solution of the present application, descriptive text can be obtained, and image features of a noisy image can be obtained. The descriptive text is used to describe the content theme of the image to be generated, and the noisy image is obtained based on the image features of the noisy image. A perception model in a simulation framework is used to obtain a three-dimensional semantic occupancy grid of the noisy image based on the noisy image. A diffusion model in the simulation framework is used to denoise the image features of the noisy image based on the descriptive text and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image. The content theme of the denoised image includes the descriptive text, and the denoised image is obtained based on the image features of the denoised image.

[0092] In the present application, the perception model in the simulation framework can obtain a high-quality three-dimensional semantic occupancy grid of the noisy image based on the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image is rich in clearer and more comprehensive information (position information and geometric information, etc.) in the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image can constrain and guide the diffusion model in the simulation framework, so that the diffusion model in the simulation framework can integrate the fine and rich semantic information and geometric information in the high-quality three-dimensional semantic occupancy grid of the noisy image a priori for conditioning, so as to assist the diffusion model in the simulation framework to generate high-precision, high-integrity, and high-realism (visually realistic) images, so that the images generated by the diffusion model in the simulation framework are also highly consistent with the content of the descriptive text at the semantic level, and even generate novel, coherent and realistic videos (such as three-dimensional realistic videos, etc.).

[0093] Through this application, efficient conversion from text to real scenes can be achieved, and the performance of the perception model is also improved to the cutting-edge level of the industry. It not only provides a powerful tool for scene creation (including three-dimensional scenes) and visual research, but also opens up new possibilities for future mixed reality applications, automated design and planning and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 It is a flowchart of the steps of a data processing method of the present application.

[0095] Figure 2 It is a schematic diagram of a data processing method of the present application.

[0096] Figure 3It is a flowchart of the steps of a data processing method of the present application.

[0097] Figure 4 It is a schematic diagram of a data processing method of the present application.

[0098] Figure 5 It is a flowchart of the steps of a data processing method of the present application.

[0099] Figure 6 This is a structural block diagram of a data processing device of the present application.

[0100] Figure 7 This is a block diagram of an electronic device of the present application.

[0101] Figure 8 This is a block diagram of an electronic device of the present application. DETAILED DESCRIPTION

[0102] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0103] Currently, scene perception and scene generation are usually viewed independently. However, this limits the potential for mutual promotion between scene perception and scene generation, and in practical applications it may be difficult to meet the requirements for high precision and realism in perception, as well as the requirements for high precision and realism in generation.

[0104] Recently, diffusion models have made significant progress in generating high-fidelity images, providing new ideas for meeting the needs of high accuracy and realism.

[0105] However, how to incorporate the semantic information of the scene into the generation of the scene, and how to use the generated scene to feed back the perception model, remain underexplored challenges.

[0106] To this end, the solution of this application is proposed. The solution of this application can be used in computer vision and artificial intelligence. This application proposes a simulation framework that organically combines the perception model and the diffusion model, integrating the task of scene perception with the task of scene generation to form a closed-loop mutual improvement mechanism. It aims to provide a unified and efficient solution for multiple fields such as virtual reality, augmented reality, urban planning, and interior design by integrating the fine scene perception capability of the perception model and the high-quality scene generation capability of the diffusion model.

[0107] Before introducing the technical solution of this application, the technical terms that may be involved in the technical solution of this application are first explained.

[0108] A 3D semantic occupancy grid is a data structure used to represent the position, shape, and semantic information of objects in 3D space. It divides the 3D space into a series of voxels, each of which can contain the following information:

[0109] Occupancy state: Indicates whether the voxel is occupied by an object, usually represented by 0 / 1 or a probability value.

[0110] Semantic information: indicates the object category to which the voxel belongs, such as road, vehicle, pedestrian, etc.

[0111] Other attributes: such as velocity information (OccupancyFlow), used to describe the motion state of the voxel.

[0112] This representation method can describe the shapes and spatial relationships of objects in three-dimensional scenes more finely, and is suitable for environmental perception in fields such as autonomous driving and robot navigation.

[0113] Perception model, used to generate a 3D semantic occupancy grid for the image. Perception models can include VoxFormer, OccupancyDETR, and TPVFormer.

[0114] VoxFormer, Sparse Voxel Transformer, is a vision-based 3D semantic scene completion method that generates complete voxel features through depth estimation and voxel occupancy prediction, suitable for autonomous driving scenarios.

[0115] OccupancyDETR, Occupancy Detection Transformer, combines the object detection module and the 3D occupancy decoder to simplify the 3D semantic occupancy perception task and improve the small object detection performance.

[0116] TPVFormer, Tri-Perspective View Transformer, proposes a three-perspective view (TPV) representation to describe a three-dimensional scene through three orthogonal two-dimensional planes, providing more comprehensive and fine-grained three-dimensional scene understanding.

[0117] Diffusion models, for example, U-Net for Denoising Diffusion Probabilistic Models

[0118] U-Net is a specialized convolutional neural network architecture, originally proposed by Olaf Ronneberger et al. in 2015, primarily for biomedical image segmentation. In diffusion models, U-Net is widely used for denoising, which involves gradually recovering a clean image from noisy data. Specifically, U-Net is often used to implement denoisers. Specifically, U-Net learns to extract useful features from noisy data and gradually generates a clean image. Its symmetrical encoder-decoder structure and skip connection design enable it to effectively process high-dimensional features while preserving detailed information from the original image.

[0119] CLIP (Contrastive Language-Image Pre-training) is a multimodal model proposed by OpenAI that uses contrastive learning methods to pre-train on large-scale text-image datasets. It aims to learn a model that maps text and images into the same feature space, thereby achieving image-text matching.

[0120] The VAE encoder is an important component of the Variational Autoencoder (VAE). Its main function is to map input data (such as images, text, etc.) to the distribution parameters of the latent space. Specifically, the encoder compresses the input data into the mean (μ) and variance (σ2) in the latent space through a neural network. These two parameters define a Gaussian distribution that represents the position and uncertainty of the input data in the latent space. The core goal of the encoder is to learn the probability distribution of the input data, rather than simply mapping it to a fixed low-dimensional representation.

[0121] The VAE decoder is another key component of the variational autoencoder. Its primary task is to reconstruct the original data from samples in the latent space. The decoder is typically a neural network that accepts latent variables as input and generates outputs that are similar to the original input data. The decoder defines a conditional distribution pθ(x|z), which represents the generative distribution of data x given the latent variable z.

[0122] Figure 1 A flowchart showing the steps of a data processing method of the present application is shown. Figure 2 A schematic diagram of a data processing method of the present application is shown. The method is used to generate an image based on a trained perception model and a diffusion model in a simulation framework. The method can be applied to electronic devices, wherein the method can specifically include the following steps:

[0123] In step S101, a description text is obtained, and image features of an input image of the Nth round are obtained, where N is a positive integer.

[0124] In one embodiment of the present application, when N is equal to 1, random noise features are obtained, such as Gaussian random noise, and then image features of the input image of the first round are obtained based on the random noise features. For example, the random noise features are directly determined as the image features of the input image of the first round.

[0125] The description text is used to describe the content theme of the image to be generated. The description text can also be regarded as prompt text. The description text can be input by the user to the electronic device according to the content theme of the image to be generated.

[0126] Alternatively, in another embodiment of the present application, when N is greater than or equal to 2, the image features of the output image of the N-1th round can be obtained; the image features of the output image of the N-1th round are obtained by denoising the image features of the input image of the N-1th round based on the description text and the three-dimensional semantic occupancy grid of the input image of the N-1th round using a diffusion model; the three-dimensional semantic occupancy grid of the input image of the N-1th round is obtained based on the input image of the N-1th round using a perception model; the image features of the input image of the Nth round are obtained based on the image features of the output image of the N-1th round; for example, the image features of the output image of the N-1th round are directly determined as the image features of the input image of the Nth round.

[0127] In step S102 , the input image of the Nth round is obtained according to the image features of the input image of the Nth round.

[0128] In this application, a VAE decoder can be used to decode the image features of the input image of the Nth round to obtain the input image of the Nth round.

[0129] In step S103 , a three-dimensional semantic occupancy grid of the input image of the Nth round is obtained based on the input image of the Nth round using the perception model in the simulation framework.

[0130] In the present application, the input image of the Nth round can be input into the perception model in the simulation framework, so that the perception model in the simulation framework processes the input image of the Nth round, obtains a three-dimensional semantic occupancy grid of the input image of the Nth round, and outputs the three-dimensional semantic occupancy grid of the input image of the Nth round. The electronic device can obtain the three-dimensional semantic occupancy grid of the input image of the Nth round output by the perception model.

[0131] In step S104, the diffusion model in the simulation framework is used to denoise the image features of the input image of the Nth round based on the descriptive text and the three-dimensional semantic occupancy grid of the input image of the Nth round to obtain the image features of the output image of the Nth round, and the content theme of the output image of the Nth round includes the descriptive text.

[0132] In the present application, the description text, the three-dimensional semantic occupancy grid of the input image of the Nth round, and the image features of the input image of the Nth round can be input into a diffusion model, so that the diffusion model denoises the image features of the input image of the Nth round based on the description text and the three-dimensional semantic occupancy grid of the input image of the Nth round, obtains the image features of the output image of the Nth round, the content theme of the output image of the Nth round includes the description text, and outputs the image features of the output image of the Nth round. The electronic device can obtain the image features of the output image of the Nth round output by the diffusion model.

[0133] It should be noted that the training process of the perception model and diffusion model in the simulation framework can be found in the following Figure 3 and Figure 4 The embodiment shown will not be described in detail here.

[0134] Among them, the three-dimensional semantic occupancy grid of the input image of the Nth round is three-dimensional data, for example, a three-dimensional matrix. If the image features of the input image of the Nth round are two-dimensional data, for example, a two-dimensional matrix, then the three-dimensional semantic occupancy grid of the input image of the Nth round needs to be aligned with the image features of the input image of the Nth round.

[0135] In this application, the three-dimensional semantic occupancy grid of the input image of the Nth round can be aligned with the image features of the input image of the Nth round through the camera parameter P.

[0136] For example, to ensure the consistency of cross-view videos, we can integrate the camera parameters P and the three-dimensional semantic occupancy grid of the input image of the Nth round based on CSE (Camera-aware Sequential Encoding) to achieve perception of the camera trajectory.

[0137] In addition, in order to align the three-dimensional semantic occupancy grid of the input image of the Nth round with the image features of the input image of the Nth round, the SOE (Sequential Occupancy Encoding) of the convolutional GRU can be used, which is used to process the semantic occupancy along the depth dimension of the three-dimensional semantic occupancy grid of the input image of the Nth round for context aggregation.

[0138] Among them, for cross-view camera coding, the three-dimensional semantic occupancy grid generated by the perception model is sufficient to describe a large scene (for example, 51.2m*51.2m*6.4m in SemanticKITTI).

[0139] To save computational resources, we can only utilize the 3D semantic occupancy grid predicted from the first keyframe and slide the camera viewpoint to generate features of the 3D semantic occupancy grid that is aware of the camera trajectory.

[0140] Given an input image containing N video frames, the image features L∈C L *N*H L *W L and the 3D semantic occupancy grid X of the input image predicted from the first keyframe OCC ∈1*D*H OCC *W OCC , the goal is to use the camera parameters P i For the i th Viewpoint (i∈(0, N-1)) encodes features based on the 3D semantic occupancy grid

[0141] Thus, the features of N frames based on 3D semantic occupancy grids It is combined with the corresponding camera trajectory.

[0142] Specifically, to encode different camera parameters with the semantic occupancy of keyframes.

[0143] The camera parameters P i (including internal and external parameters) are input into the {parameter encoder}, and the formula is as follows:

[0144]

[0145] Conv and FC represent convolutional and fully connected layers, respectively. σ and Reshape represent the sigmoid function and reshape operation, respectively.

[0146] After that, the interpolation semantics takes up X OCC Align the image feature L with the input image of the Nth round in the spatial dimension, and use deformable 3D convolution to generate the i-th th The dynamic occupancy volume of the camera view is then compared with the encoded camera parameters Multiply for camera perception:

[0147]

[0148] K w represents the number of points in the deformable sampling process, wk represents the spatial feature weight, △P kis an additional offset in the sampling grid, adaptively adjusting the sampling position p+p k .

[0149] so, It encodes the features of specific semantic 3D occupancy grids and is camera parameter aware.

[0150] To facilitate the generation of N frames of video, the camera parameters of different viewpoints are encoded separately, and different deformable 3D convolutions without shared weights are used to generate the features of the corresponding 3D semantic occupancy grid. The image features L of the input image corresponding to the Nth round along the time axis.

[0151] The same method can be used to generate a single-view image, and N=1 can be set.

[0152] Among them, for occupancy sequence encoding, the features of the 3D semantic occupancy grid of the input image Contains D semantic graphs spliced ​​in the depth dimension. In order to make it th The image feature L of the input image of the video frame i Alignment, which can sequentially process the features of the 3D semantic occupancy grid of the input image along the depth axis

[0153] For example, a series of stacked convolutional GRUs are employed to integrate temporal context along the depth dimension.

[0154] Features of the 3D semantic occupancy grid of the input image First, it is encoded along the depth dimension and compared with the image feature L of the input image. i Connection formation

[0155] Right now Use this value to update the hidden state through the convolutional GRU

[0156] Among them, [] represents concatenation, and ⊙ represents element-by-element multiplication. occ Represents the features of the 3D semantic occupancy grid of the input image using a 3D convolutional layer Encode. L Represents the image feature L of the input image using the self-attention block i to encode.

[0157] Among them, in cross-view video generation, the self-attention module is the implemented spatiotemporal attention mechanism. and Represent the normalized features of the previous step and the updated features of the current step, respectively. is the update gate mapping that determines whether to update the output of the current step, and r k It is a decision Reset gate mapping for the influence of the current update. and is the weight matrix of the GRU unit. Based on the hidden state h k , we decode the aggregated latent features L via residual connections and zero-convolution agg , like ControlNet, to preserve the inherent ability to generate models:

[0158]

[0159] In this way, the sequence encoding architecture achieves the alignment of the features of the 3D semantic occupancy grid of the input image with the image features of the input image by integrating contextual information along the depth axis.

[0160] In step S105 , the output image of the Nth round is obtained according to the image features of the output image of the Nth round.

[0161] In the present application, a VAE decoder may be used to decode the image features of the output image of the Nth round to obtain the output image of the Nth round.

[0162] The output image of the Nth round is an image obtained after denoising based on the input image of the Nth round, and can be used in downstream applications. Alternatively, the output image of the Nth round is continued to be used as the input image of the N+1th round and the denoising process of the N+1th round is performed according to the process of steps S101 to S105 to obtain the output image of the N+1th round, and the output image of the N+1th round can be used in downstream applications, etc. Alternatively, the output image of the N+1th round is continued to be used as the input image of the N+2th round and the denoising process of the N+2th round is performed according to the process of steps S101 to S105 until a denoised image that meets the requirements is obtained.

[0163] By using the output of the previous round of denoising as the input data for the next round of denoising, the algorithm continues to iterate. As the number of iterations increases, the denoised image becomes clearer and richer in information, and the three-dimensional semantic occupancy grid becomes more complete and accurate.

[0164] The 3D semantic occupancy grid provides more specific semantic and geometric context, thereby constraining and refining the denoising process, which can improve the quality of the denoised image.

[0165] In the solution of the present application, descriptive text can be obtained, and image features of a noisy image can be obtained. The descriptive text is used to describe the content theme of the image to be generated, and the noisy image is obtained based on the image features of the noisy image. A perception model in a simulation framework is used to obtain a three-dimensional semantic occupancy grid of the noisy image based on the noisy image. A diffusion model in the simulation framework is used to denoise the image features of the noisy image based on the descriptive text and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image. The content theme of the denoised image includes the descriptive text, and the denoised image is obtained based on the image features of the denoised image.

[0166] In the present application, the perception model in the simulation framework can obtain a high-quality three-dimensional semantic occupancy grid of the noisy image based on the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image is rich in clearer and more comprehensive information (position information and geometric information, etc.) in the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image can constrain and guide the diffusion model in the simulation framework, so that the diffusion model in the simulation framework can integrate the fine and rich semantic information and geometric information in the high-quality three-dimensional semantic occupancy grid of the noisy image a priori for conditioning, so as to assist the diffusion model in the simulation framework to generate high-precision, high-integrity, and high-realism (visually realistic) images, so that the images generated by the diffusion model in the simulation framework are also highly consistent with the content of the descriptive text at the semantic level, and even generate novel, coherent and realistic videos (such as three-dimensional realistic videos, etc.).

[0167] Through this application, efficient conversion from text to real scenes can be achieved, and the performance of the perception model is also improved to the cutting-edge level of the industry. It not only provides a powerful tool for scene creation (including three-dimensional scenes) and visual research, but also opens up new possibilities for future mixed reality applications, automated design and planning and other fields.

[0168] Furthermore, in another embodiment of the present application, after obtaining the output image of the Nth round, a perception model in the simulation framework can be used to obtain a 3D semantic occupancy grid of the output image of the Nth round based on the output image of the Nth round. Then, a training data pair is generated based on the 3D semantic occupancy grid of the output image of the Nth round and the output image of the Nth round. The training data pair is used at least to train the perception model. For example, the 3D semantic occupancy grid of the output image of the Nth round and the output image of the Nth round are combined into a training data pair. For example, the output image of the Nth round is used as sample data and the 3D semantic occupancy grid of the output image of the Nth round is used as annotation data to train other models, such as training other perception models.

[0169] Figure 3A flowchart showing the steps of a data processing method of the present application is shown. Figure 4 A schematic diagram of a data processing method of the present application is shown. The method is used to train a perception model and a diffusion model in a simulation framework, so as to subsequently generate an image or denoise an image based on the trained perception model and diffusion model in the simulation framework. The method can be applied to electronic devices, wherein the method can specifically include the following steps:

[0170] In step S201, a plurality of sample data are obtained, wherein one sample data includes a noise-free image, a description text of the noise-free image, and a 3D semantic occupancy grid of the noise-free image. The description text of the noise-free image is used to describe the content theme of the noise-free image.

[0171] The noise-free images in different sample data are different, so that the description texts of the noise-free images in different sample data are different, and the three-dimensional semantic occupancy grids of the noise-free images in different sample data are different.

[0172] Noise-free images can be collected in advance.

[0173] The description text of the noise-free image is obtained based on the noise-free image. For example, the description text of the noise-free image can be obtained using a CLIP model.

[0174] The three-dimensional semantic occupancy grid of the noise-free image can be obtained based on the noise-free image. For example, the three-dimensional semantic occupancy grid of the noise-free image can be obtained using a perception model (which can be the perception model of the present application or other perception models already available on the market).

[0175] In step S202 , a plurality of sample data are used to train the perception model and the diffusion model in the simulation framework.

[0176] The perception model is used to obtain a 3D semantic occupancy grid of the noisy image based on the noisy image. The noisy image is obtained based on the image features of the noisy image, which are obtained by superimposing random noise features on the image features of the noise-free image. The diffusion model is used to denoise the image features of the noisy image based on the descriptive text of the noise-free image and the 3D semantic occupancy grid of the noisy image, thereby obtaining the image features of the denoised image.

[0177] Once the parameters in the perception model and the diffusion model in the simulation framework converge, the training can be terminated and the subsequent simulation framework can be put into use.

[0178] In another embodiment of the present application, see Figure 5 , step S202 includes:

[0179] In step S301 , a plurality of sample data are divided into a plurality of first sample data and a plurality of second sample data.

[0180] The plurality of first sample data and the plurality of second sample data do not overlap.

[0181] In step S302 , the diffusion model in the simulation framework is trained separately using the plurality of first sample data, and in the process of training the diffusion model in the simulation framework separately using the plurality of first sample data, the parameters of the perception model are kept unchanged.

[0182] In the present application, in the initial stage, the parameters in the diffusion model in the simulation framework are initialized, so that the diffusion model in the simulation framework has no fitting ability, and the parameters in the perception model in the simulation framework are not initialized, but are obtained after training, so that the perception model in the simulation framework has fitting ability. The present application can first keep the parameters of the perception model unchanged, train the diffusion model in the simulation framework, so that the diffusion model in the simulation framework has a certain fitting ability, and then collaboratively train the perception model and the diffusion model in the simulation framework to improve the collaborative training effect, increase the stability of the training, increase the robustness, and thus improve the generalization ability of the diffusion model and the generalization ability of the perception model finally obtained. Secondly, it also improves the diversity and authenticity of scene generation, and also improves the performance of the perception model in the semantic occupancy prediction task, opening up new possibilities for the generation and analysis of scenes (including three-dimensional scenes).

[0183] In step S303, the perception model and the diffusion model in the simulation framework are collaboratively trained using a plurality of second sample data.

[0184] During the collaborative training process, the perception model and the diffusion model promote each other synergistically.

[0185] During the collaborative training process, the parameters of the perception model are not fixed (may change according to the actual training situation) and / or the parameters of the extended model are not fixed (may change according to the actual training situation).

[0186] This step can be implemented through the following process, including:

[0187] 3031. Obtain image features of a noise-free image in the second sample data.

[0188] A VAE encoder may be used to obtain image features of a noise-free image in the second sample data.

[0189] 3032. Superimpose random noise features on image features of the noise-free image to obtain image features of the noisy image.

[0190] For example, Gaussian random noise is superimposed on the image feature of a noise-free image to obtain another image feature. The other image feature is a noisy feature, which can correspond to an image, which is a noisy image. In this way, the other image feature can be regarded as the image feature of the noisy image.

[0191] 3033. Obtain a noisy image according to image features of the noisy image.

[0192] In this application, a VAE decoder can be used to decode image features of a noisy image to obtain a noisy image.

[0193] 3034. Obtaining a 3D semantic occupancy grid of a noisy image based on a noisy image using a perceptual model.

[0194] In the present application, a noisy image can be input into a perception model in a simulation framework, so that the perception model in the simulation framework processes the noisy image to obtain a three-dimensional semantic occupancy grid of the noisy image and outputs the three-dimensional semantic occupancy grid of the noisy image. An electronic device can obtain the three-dimensional semantic occupancy grid of the noisy image output by the perception model.

[0195] 3035. Denoise the image features of the noisy image using a diffusion model based on the description text of the noise-free image in the second sample data and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image.

[0196] In the present application, a descriptive text of a noise-free image, a three-dimensional semantic occupancy grid of a noisy image, and image features of the noisy image can be input into a diffusion model, so that the diffusion model denoises the image features of the noisy image based on the descriptive text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image, obtains image features of the denoised image, and outputs the image features of the denoised image. The electronic device can obtain the image features of the denoised image output by the diffusion model.

[0197] 3036. Obtain a denoised image based on the image features of the denoised image.

[0198] In this application, a VAE decoder can be used to decode the image features of the denoised image to obtain the denoised image.

[0199] 3037. Co-train the perception model and the diffusion model in the simulation framework based on the denoised image, the three-dimensional semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free three-dimensional semantic occupancy grid in the second sample data.

[0200] In step 3037, a loss value corresponding to the diffusion model may be obtained based on the noise-free image and the denoised image. A loss value corresponding to the perception model may be obtained based on the 3D semantic occupancy grid of the noisy image and the noise-free 3D semantic occupancy grid.

[0201] According to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model, the parameters in the perception model and the parameters in the diffusion model are collaboratively optimized.

[0202] In one example, a comprehensive loss value of the simulation framework can be obtained based on the loss value corresponding to the diffusion model and the loss value corresponding to the perception model. Then, parameters in the perception model and the diffusion model can be collaboratively optimized based on the comprehensive loss value of the simulation framework.

[0203] For example, in one example, the sum of the loss value corresponding to the diffusion model and the loss value corresponding to the perception model can be calculated to obtain the comprehensive loss value of the simulation framework.

[0204] However, during the training of the diffusion model and simulation model in the simulation framework, the input to the simulation model is a noisy image. The noise in the noisy image can affect the accuracy of the simulation model's output. Therefore, when calculating the comprehensive loss value of the simulation framework, the fact that the simulation model is fed with a noisy image needs to be taken into account. For example, to mitigate the impact of noise, in another example, the weight of the loss value corresponding to the perception model can be determined based on the content of random noise features in the image features of the noisy image (the content can be understood as the degree of change in the image features of the noise-free image to random noise features). The weight is greater than 0 and less than 1. The higher the content of random noise features in the image features of the noisy image, the smaller the weight of the loss value corresponding to the perception model. Alternatively, the lower the content of random noise features in the image features of the noisy image, the larger the weight of the loss value corresponding to the perception model. This weight can emphasize supervision under low noise levels and reduce the impact under high noise levels. The product of this weight and the loss value corresponding to the perception model is calculated. The sum of this product and the loss value corresponding to the diffusion model is calculated to obtain the comprehensive loss value of the entire framework. This allows the training scheme to balance the noise components in the input image and ensure the stability and effectiveness of supervised training.

[0205] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions involved are not necessarily required by this application.

[0206] Reference Figure 6 , shows a structural block diagram of a data processing device of the present application, the device comprising:

[0207] The first acquisition module 11 is used to acquire the description text, and the second acquisition module 12 is used to acquire the image features of the input image of the Nth round, where N is a positive integer;

[0208] A third acquisition module 13 is configured to acquire an input image for the Nth round according to image features of the input image for the Nth round;

[0209] A fourth acquisition module 14 is configured to acquire a three-dimensional semantic occupancy grid of the input image of the Nth round based on the input image of the Nth round using the perception model in the simulation framework;

[0210] a denoising module 15 configured to use a diffusion model in a simulation framework to denoise image features of the input image of the Nth round based on the description text and the three-dimensional semantic occupancy grid of the input image of the Nth round, thereby obtaining image features of the output image of the Nth round, wherein the content theme of the output image of the Nth round includes the description text;

[0211] The fifth acquisition module 16 is configured to acquire the output image of the Nth round according to the image features of the output image of the Nth round.

[0212] In an optional implementation, N is equal to 1;

[0213] The second acquisition module includes:

[0214] A first acquisition unit is used to acquire random noise characteristics;

[0215] The second acquisition unit is used to acquire the image features of the input image of the first round according to the random noise features.

[0216] In an optional implementation, N is greater than or equal to 2;

[0217] The second acquisition module includes:

[0218] a third acquisition unit, configured to acquire image features of the output image of the N-1th round; the image features of the output image of the N-1th round are obtained by denoising the image features of the input image of the N-1th round using a diffusion model based on the description text and the three-dimensional semantic occupancy grid of the input image of the N-1th round; the three-dimensional semantic occupancy grid of the input image of the N-1th round is acquired based on the input image of the N-1th round using a perception model;

[0219] The fourth acquiring unit is configured to acquire the image features of the input image of the Nth round according to the image features of the output image of the N-1th round.

[0220] In an optional implementation, the apparatus further includes:

[0221] a sixth acquisition module, configured to acquire a plurality of sample data, wherein each sample data includes a noise-free image, a description text of the noise-free image, and a three-dimensional semantic occupancy grid of the noise-free image; the description text of the noise-free image is used to describe the content theme of the noise-free image;

[0222] A training module, used to train the perception model and the diffusion model in the simulation framework using multiple sample data;

[0223] The perception model is used to obtain a three-dimensional semantic occupancy grid of a noisy image based on a noisy image. The noisy image is obtained based on image features of the noisy image. The image features of the noisy image are obtained by superimposing random noise features on image features of the noise-free image.

[0224] The diffusion model is used to denoise image features of the noisy image based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image, so as to obtain image features of the denoised image.

[0225] In an optional implementation, the training module includes:

[0226] a dividing unit, configured to divide the plurality of sample data into a plurality of first sample data and a plurality of second sample data;

[0227] A first training unit is configured to separately train the diffusion model in the simulation framework using the plurality of first sample data, and to keep parameters of the perception model unchanged during the separate training of the diffusion model in the simulation framework using the plurality of first sample data;

[0228] The second training unit is used to collaboratively train the perception model and the diffusion model in the simulation framework using a plurality of second sample data.

[0229] In an optional implementation, the second training unit includes:

[0230] A first acquisition subunit is used to acquire image features of a noise-free image in the second sample data;

[0231] a superposition subunit, configured to superimpose random noise features on image features of a noise-free image to obtain image features of a noisy image;

[0232] a second acquisition subunit, configured to acquire a noisy image according to image features of the noisy image;

[0233] a third acquisition subunit, configured to acquire a three-dimensional semantic occupancy grid of the noisy image based on the noisy image using the perception model;

[0234] a denoising subunit, configured to denoise image features of the noisy image using a diffusion model based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image in the second sample data, to obtain image features of the denoised image;

[0235] a fourth acquisition subunit, configured to acquire a denoised image according to image features of the denoised image;

[0236] A training subunit is used to collaboratively train the perception model and the diffusion model in the simulation framework based on the denoised image, the three-dimensional semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free three-dimensional semantic occupancy grid in the second sample data.

[0237] In an optional implementation, the training subunit is specifically used to: obtain a loss value corresponding to the diffusion model based on the noise-free image and the denoised image; obtain a loss value corresponding to the perception model based on the three-dimensional semantic occupancy grid of the noisy image and the noise-free three-dimensional semantic occupancy grid; and collaboratively optimize the parameters in the perception model and the parameters in the diffusion model based on the loss value corresponding to the diffusion model and the loss value corresponding to the perception model.

[0238] In an optional implementation, the training subunit is specifically used to: obtain the comprehensive loss value of the simulation framework based on the loss value corresponding to the diffusion model and the loss value corresponding to the perception model; and collaboratively optimize the parameters in the perception model and the parameters in the diffusion model based on the comprehensive loss value of the simulation framework.

[0239] In an optional implementation, the training subunit is specifically used to: determine the weight of the loss value corresponding to the perception model based on the content of random noise features in the image features of the noisy image; the weight is greater than 0 and less than 1; wherein, the higher the content of random noise features in the image features of the noisy image, the smaller the weight of the loss value corresponding to the perception model, or, the lower the content of random noise features in the image features of the noisy image, the greater the weight of the loss value corresponding to the perception model; calculate the product between the weight and the loss value corresponding to the perception model; calculate the sum of the product and the loss value corresponding to the diffusion model to obtain the comprehensive loss value of the overall framework.

[0240] In an optional implementation, the apparatus further includes:

[0241] a seventh acquisition module, configured to acquire a three-dimensional semantic occupancy grid of the output image of the Nth round based on the output image of the Nth round using the perception model in the simulation framework;

[0242] A generation module is used to generate a training data pair based on the three-dimensional semantic occupancy grid of the output image of the Nth round and the output image of the Nth round, where the training data pair is at least used to train the perception model.

[0243] In the solution of the present application, descriptive text can be obtained, and image features of a noisy image can be obtained. The descriptive text is used to describe the content theme of the image to be generated, and the noisy image is obtained based on the image features of the noisy image. A perception model in a simulation framework is used to obtain a three-dimensional semantic occupancy grid of the noisy image based on the noisy image. A diffusion model in the simulation framework is used to denoise the image features of the noisy image based on the descriptive text and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image. The content theme of the denoised image includes the descriptive text, and the denoised image is obtained based on the image features of the denoised image.

[0244] In the present application, the perception model in the simulation framework can obtain a high-quality three-dimensional semantic occupancy grid of the noisy image based on the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image is rich in clearer and more comprehensive information (position information and geometric information, etc.) in the noisy image. The high-quality three-dimensional semantic occupancy grid of the noisy image can constrain and guide the diffusion model in the simulation framework, so that the diffusion model in the simulation framework can integrate the fine and rich semantic information and geometric information in the high-quality three-dimensional semantic occupancy grid of the noisy image a priori for conditioning, so as to assist the diffusion model in the simulation framework to generate high-precision, high-integrity, and high-realism (visually realistic) images, so that the images generated by the diffusion model in the simulation framework are also highly consistent with the content of the descriptive text at the semantic level, and even generate novel, coherent and realistic videos (such as three-dimensional realistic videos, etc.).

[0245] Through this application, efficient conversion from text to real scenes can be achieved, and the performance of the perception model is also improved to the cutting-edge level of the industry. It not only provides a powerful tool for scene creation (including three-dimensional scenes) and visual research, but also opens up new possibilities for future mixed reality applications, automated design and planning and other fields.

[0246] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0247] Optionally, an embodiment of the present application also provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0248] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the various processes of the above-mentioned method embodiments are implemented and the same technical effects are achieved. To avoid repetition, the details are not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0249] Figure 7 8 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0250] Reference Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0251] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0252] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0253] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0254] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0255] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0256] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0257] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0258] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0259] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0260] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0261] Figure 8 1 is a block diagram of an electronic device 1900 shown in the present application. For example, the electronic device 1900 can be provided as a server.

[0262] Reference Figure 8 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0263] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0264] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0265] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0266] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0267] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0268] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0269] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0270] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0271] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0272] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0273] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: The method comprises: Get the description text and the image features of the input image of the Nth round, where N is a positive integer; Obtaining an input image for the Nth round according to the image features of the input image for the Nth round; Obtain a three-dimensional semantic occupancy grid of the input image of the Nth round based on the input image of the Nth round using the perception model in the simulation framework; Denoising image features of the input image of the Nth round based on the description text and a three-dimensional semantic occupancy grid of the input image of the Nth round using a diffusion model in a simulation framework to obtain image features of an output image of the Nth round, wherein the content theme of the output image of the Nth round includes the description text; The output image of the Nth round is obtained according to the image features of the output image of the Nth round.

2. The method according to claim 1, characterized in that N is equal to 1; The obtaining of image features of the input image of the Nth round includes: Get random noise characteristics; The image features of the input image in the first round are obtained based on the random noise features.

3. The method according to claim 1, characterized in that N is greater than or equal to 2; The obtaining of image features of the input image of the Nth round includes: Obtaining image features of the output image of the N-1th round; the image features of the output image of the N-1th round are obtained by denoising the image features of the input image of the N-1th round based on the description text and the three-dimensional semantic occupancy grid of the input image of the N-1th round using a diffusion model; the three-dimensional semantic occupancy grid of the input image of the N-1th round is obtained based on the input image of the N-1th round using a perception model; The image features of the input image of the Nth round are obtained according to the image features of the output image of the N-1th round.

4. The method according to claim 1, wherein The method further comprises: Acquire multiple sample data, where each sample data includes a noise-free image, a description text of the noise-free image, and a three-dimensional semantic occupancy grid of the noise-free image; the description text of the noise-free image is used to describe the content theme of the noise-free image; Use multiple sample data to train the perception model and diffusion model in the simulation framework; The perception model is used to obtain a three-dimensional semantic occupancy grid of a noisy image based on a noisy image. The noisy image is obtained based on image features of the noisy image. The image features of the noisy image are obtained by superimposing random noise features on image features of the noise-free image. The diffusion model is used to denoise image features of the noisy image based on the description text of the noise-free image and the three-dimensional semantic occupancy grid of the noisy image, so as to obtain image features of the denoised image.

5. The method according to claim 4, characterized in that The method of using multiple sample data to train the perception model and the diffusion model in the simulation framework includes: Dividing the plurality of sample data into a plurality of first sample data and a plurality of second sample data; Using the plurality of first sample data to train the diffusion model in the simulation framework separately, and in the process of using the plurality of first sample data to train the diffusion model in the simulation framework separately, keeping the parameters of the perception model unchanged; The plurality of second sample data are used to collaboratively train the perception model and the diffusion model in the simulation framework.

6. The method according to claim 5, characterized in that The collaborative training of the perception model and the diffusion model in the simulation framework using the plurality of second sample data includes: Acquire image features of a noise-free image in the second sample data; The image features of the noise-free image are superimposed with random noise features to obtain the image features of the noisy image; obtaining a noisy image according to image features of the noisy image; Obtaining a 3D semantic occupancy grid of a noisy image based on the noisy image using a perception model; Denoising the image features of the noisy image using a diffusion model based on the description text of the noise-free image in the second sample data and the three-dimensional semantic occupancy grid of the noisy image to obtain image features of the denoised image; Obtaining a denoised image according to image features of the denoised image; The perception model and the diffusion model in the simulation framework are collaboratively trained based on the denoised image, the 3D semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free 3D semantic occupancy grid in the second sample data.

7. The method according to claim 6, characterized in that The collaborative training of the perception model and the diffusion model in the simulation framework based on the denoised image, the three-dimensional semantic occupancy grid of the noisy image, the noise-free image in the second sample data, and the noise-free three-dimensional semantic occupancy grid in the second sample data includes: Obtain the loss value corresponding to the diffusion model based on the noise-free image and the denoised image; Obtaining a loss value corresponding to the perception model based on the 3D semantic occupancy grid of the noisy image and the 3D semantic occupancy grid without noise; According to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model, the parameters in the perception model and the parameters in the diffusion model are collaboratively optimized.

8. The method according to claim 7, characterized in that The collaborative optimization of parameters in the perception model and parameters in the diffusion model according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model includes: Obtain the comprehensive loss value of the simulation framework according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model; The parameters in the perception model and the parameters in the diffusion model are collaboratively optimized according to the comprehensive loss value of the simulation framework.

9. The method according to claim 8, characterized in that The obtaining of the comprehensive loss value of the simulation framework according to the loss value corresponding to the diffusion model and the loss value corresponding to the perception model includes: Determine the weight of the loss value corresponding to the perception model according to the content of random noise features in the image features of the noisy image; the weight is greater than 0 and less than 1; The higher the content of random noise features in the image features of the noisy image, the smaller the weight of the loss value corresponding to the perception model; or the lower the content of random noise features in the image features of the noisy image, the greater the weight of the loss value corresponding to the perception model; Calculating the product between the weight and the loss value corresponding to the perception model; The sum of the product and the loss value corresponding to the diffusion model is calculated to obtain the comprehensive loss value of the entire framework.

10. The method according to claim 1, characterized in that The method further comprises: Obtain a three-dimensional semantic occupancy grid of the output image of the Nth round based on the output image of the Nth round using the perception model in the simulation framework; A training data pair is generated according to the three-dimensional semantic occupancy grid of the output image of the Nth round and the output image of the Nth round, and the training data pair is used for at least training the perception model.

11. A data processing device, characterized in that: The device comprises: The first acquisition module is used to acquire the description text, and the second acquisition module is used to acquire the image features of the input image of the Nth round, where N is a positive integer; A third acquisition module is used to acquire the input image of the Nth round according to the image features of the input image of the Nth round; a fourth acquisition module, configured to acquire a three-dimensional semantic occupancy grid of the input image of the Nth round based on the input image of the Nth round using the perception model in the simulation framework; a denoising module, configured to denoise image features of the input image of the Nth round using a diffusion model in the simulation framework based on the description text and a three-dimensional semantic occupancy grid of the input image of the Nth round, to obtain image features of an output image of the Nth round, wherein the content theme of the output image of the Nth round includes the description text; The fifth acquisition module is used to acquire the output image of the Nth round according to the image features of the output image of the Nth round.

12. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by the processor.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the method according to any one of claims 1 to 10 when executed by a processor.