Image processing method and device, equipment and storage medium
By generating hybrid modal maps and using a dual-branch diffusion model, the problem of insufficient processing capabilities for complex user input and multimodal data in the prior art is solved, and efficient three-dimensional scene generation is achieved, with high spatial consistency and geometric details.
Patent Information
- Application Number
- CN202510129261.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The existing three-dimensional scene generation technology mainly relies on text input and is difficult to effectively process complex user input, especially data in multiple modes.
By acquiring input data of multiple modalities, a hybrid modal graph is generated, which includes the relationship edges between multiple nodes and nodes. Then, based on the hybrid modal graph, a three-dimensional scene is generated using a double-branch diffusion model, containing the geometric shape and spatial position information of the object.
It realizes flexible processing and integration of data in different modalities, improves the adaptability to complex input data, and the generated three-dimensional scenes have high spatial consistency and geometric details.
Smart Images

Figure CN120047619A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] The technology of generating three-dimensional scenes indoors is an important research direction in virtual reality, interior design, and intelligent environment construction. The goal of this technology is to be able to generate highly realistic and semantically compliant three-dimensional scenes based on the data input by users.
[0003] In practical applications, most of the three-dimensional scene generation technologies only support text input, that is, generating corresponding three-dimensional scenes based on the text data input by users. Therefore, they are unable to cope when dealing with complex user inputs and have limitations. Summary of the Invention
[0004] Embodiments of the present application provide an image processing method, apparatus, device, and storage medium to support users in inputting data of different modalities and flexibly process data of different modalities, improving the adaptability to complex input data.
[0005] In a first aspect, an embodiment of the present application provides an image processing method, including:
[0006] Obtain input data, where the input data includes data of multiple modalities and is used to generate a three-dimensional scene;
[0007] Generate a hybrid modality graph based on the input data, where the hybrid modality graph includes multiple nodes and multiple edges between the multiple nodes. Each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge;
[0008] Generate the three-dimensional scene based on the hybrid modality graph, where the content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0009] Optionally, the multiple nodes include a first node, a second node, a third node, and a fourth node. The features of the first node include the object category feature and text feature of the first node, the features of the fourth node include the object category feature, text feature, and visual feature of the fourth node, and there is no edge between the second node and the third node;
[0010] The generating a hybrid modality graph based on the input data includes:
[0011] Extract features from the input data to obtain a first hybrid modality graph, where the first hybrid modality graph includes each node, the features of each node, each edge, and the features of each edge;
[0012] Enhance the text features of the first node in the first hybrid modality graph to obtain a second hybrid modality graph, where the enhancement process is used to convert the text features of the first node into visual features;
[0013] Perform relationship prediction on the features corresponding to the second node and the third node in the second hybrid modality graph to obtain the hybrid modality graph, where the relationship prediction is used to generate the features of the edge between the second node and the third node to predict the relationship between the second node and the third node.
[0014] Optionally, the input data includes text data and image data, and the extracting features from the input data to obtain a first hybrid modality graph includes:
[0015] Extract features from the text data and the image data through a vision-language model to obtain the multiple nodes, and extract features from the text data and the image features through the vision-language model to obtain each node, the features included in each node respectively, each edge, and the features of each edge;
[0016] Generate the first hybrid modality graph based on each node, the features included in each node respectively, each edge, and the features of each edge.
[0017] Optionally, the enhancing the text features of the first node in the first hybrid modality graph to obtain a second hybrid modality graph includes:
[0018] Encode the text features of the first node through an encoder to obtain a latent vector;
[0019] Quantize the latent vector through a codebook to obtain a quantized latent vector;
[0020] Decode the quantized latent vector through a decoder to obtain the visual features of the first node;
[0021] Generate the second hybrid modality graph based on the visual features of the first node and the first hybrid modality graph.
[0022] Optionally, the performing relationship prediction on the features corresponding to the second node and the third node in the second hybrid modality graph to obtain the hybrid modality graph includes:
[0023] Construct a triple, which sequentially includes the features of the second node, the relationship between the second node and the third node, and the features of the third node, and the relationship between the second node and the third node is filled with zeros;
[0024] Input the triple into a relationship prediction model, and output the features of the edge between the second node and the third node through the relationship prediction model;
[0025] Generate the hybrid modality graph based on the features of the edge between the second node and the third node and the second hybrid modality graph.
[0026] Optionally, generating the three-dimensional scene based on the hybrid modality graph includes:
[0027] Using the hybrid modality graph as the condition of the shape branch in the dual-branch diffusion model, and performing sampling and denoising processing through the shape branch to obtain the geometric shape information of the object corresponding to each node;
[0028] Using the hybrid modality graph as the condition of the layout branch in the dual-branch diffusion model, and performing sampling and denoising processing through the layout branch to obtain the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0029] Optionally, the loss function of the shape branch is used to minimize the deviation between the true noise of the object shape in the three-dimensional scene and the predicted noise of the shape branch, the loss function of the layout branch is used to minimize the deviation between the true noise of the object layout in the three-dimensional scene and the predicted noise of the layout branch, and the loss function of the dual-branch diffusion model is determined based on the weight corresponding to the loss function of the shape branch, the loss function of the shape branch, the weight of the loss function of the layout branch, and the loss function of the layout branch.
[0030] In a second aspect, an embodiment of the present application provides an image processing apparatus, including:
[0031] An input acquisition module, configured to acquire input data, where the input data includes data of multiple modalities, and the input data is used to generate a three-dimensional scene;
[0032] A hybrid modality graph generation module, configured to generate a hybrid modality graph based on the input data, where the hybrid modality graph includes multiple nodes and multiple edges between the multiple nodes, each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge;
[0033] A three-dimensional scene generation module for generating the three-dimensional scene based on the hybrid modality graph, where the content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor, a memory, and a system bus;
[0035] The processor and the memory are connected through the system bus;
[0036] The memory is used to store a program, and the program includes instructions, and when the instructions are executed by the processor, the processor is caused to execute any implementation step of the above image processing method.
[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on an electronic device, the electronic device is caused to execute any implementation step of the above image processing method.
[0038] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0039] In the embodiments of the present application, first, input data can be obtained. The input data includes data of multiple modalities and is used to generate a three-dimensional scene. Then, a hybrid modality graph can be generated based on the input data. The hybrid modality graph includes multiple nodes and multiple edges between the multiple nodes. Each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge. In this way, a three-dimensional scene can be generated based on the hybrid modality graph, where the content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene. Since the hybrid modality graph is generated based on data of multiple modalities, that is to say, this solution can support a user to input data of different modalities, and then convert the data of different modalities into a hybrid modality graph for three-dimensional scene generation. Therefore, it can flexibly process data of different modalities, improve the adaptability to the input data, and thus significantly improve the support ability for complex data input by the user. In addition, the image generated based on the hybrid modality graph can comprehensively process the spatial layout and geometric shape of the object under complex input conditions, so as to generate a three-dimensional scene with high spatial consistency and geometric details. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0041] Figure 2A flowchart for generating a three-dimensional scene by a dual-branch diffusion model provided by an embodiment of the present application;
[0042] Figure 3 A schematic structural diagram of an image processing apparatus provided by an embodiment of the present application. Detailed implementation manners
[0043] As described above, in practical applications, most of the three-dimensional scene generation technologies only support text input, that is, generating a corresponding three-dimensional scene based on the text data input by the user, and it is difficult to integrate the image data or mixed-modal data (such as a combination of text data and image data) input by the user. Therefore, it is unable to handle complex user inputs effectively and has limitations.
[0044] Based on this, to solve the above problems, an embodiment of the present application provides an image processing method, which includes: First, input data can be obtained. The input data includes data of multiple modalities and is used to generate a three-dimensional scene. Then, a mixed-modal graph can be generated based on the input data. The mixed-modal graph includes multiple nodes and multiple edges between the multiple nodes. Each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge. In this way, a three-dimensional scene can be generated based on the mixed-modal graph, where the content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0045] Since the mixed-modal graph is generated based on data of multiple modalities, that is to say, this solution can support the user to input data of different modalities, and then convert the data of different modalities into a mixed-modal graph for three-dimensional scene generation. Therefore, it can flexibly process data of different modalities, improve the adaptability to the input data, and thus significantly improve the support ability for complex data input by the user. In addition, the image generated based on the mixed-modal graph can comprehensively process the spatial layout and geometric shape of the object under complex input conditions, so as to generate a three-dimensional scene with high spatial consistency and geometric details.
[0046] It should be noted that the embodiment of the present application does not limit the execution subject of this image processing method. For example, the image processing method in the embodiment of the present application can be applied to an image processing device such as a terminal device or a server. Among them, the terminal device can be an electronic device such as a smart phone, a computer, a personal digital assistant (PDA), or a tablet computer. The server can be an independent server, a cluster server, or a cloud server.
[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0048] Figure 1 This is a flowchart of an image processing method provided by an embodiment of this application. In combination with Figure 1 As shown, for the image processing method provided by an embodiment of this application, the corresponding image processing device is used as the execution subject to describe the specific implementation of the solution. This image processing method may include the following steps S101 to step S103.
[0049] S101: Obtain input data, where the input data includes data of multiple modalities and is used to generate a three-dimensional scene.
[0050] In the embodiments of this application, the data of multiple modalities are text data and image data. Correspondingly, in specific implementation, the image processing device may receive the data of multiple modalities input by the user through the terminal device as the input data and generate a three-dimensional scene therefrom.
[0051] S102: Generate a mixed-modal graph based on the input data. The mixed-modal graph includes multiple nodes and multiple edges between the multiple nodes. Each of the multiple nodes is used to represent an object, and each of the multiple edges is used to represent the relationship between the two nodes corresponding to each edge.
[0052] In the embodiments of this application, the mixed-modal graph is the core structure for integrating the above-mentioned data of multiple modalities. By representing the objects and relationship information in the input data as corresponding nodes and edges, it can provide a unified expression framework for generating a three-dimensional scene, thereby realizing the integration of the data of multiple modalities. For ease of understanding, the following describes the generation process of the mixed-modal graph in combination with a possible implementation manner.
[0053] As a possible implementation manner, the above-mentioned multiple nodes may include a first node, a second node, a third node, and a fourth node. Among them, the features of the first node include the object category feature and text feature of the first node, and the features of the fourth node include the object category feature, text feature, and visual feature of the fourth node. That is to say, the visual graph feature of the first node is missing in the features of the first node. And there is no edge between the second node and the third node. That is to say, the corresponding relationship is missing between the second node and the third node.
[0054] In addition, among multiple nodes, a fifth node may also be included, and the features of the fifth node include the visual features of the fifth node. Therefore, based on the features of the nodes, the above-mentioned multiple nodes can be divided into three types. One type is the nodes that only include visual features, such as the fifth node; another type is the nodes that include object category features and text features, such as the first node; and yet another type is the nodes that include object category features, text features, and visual features, such as the fourth node.
[0055] It should be noted that the embodiments of the present application only use the first node, the second node, the third node, the fourth node, and the fifth node as examples to illustrate the features respectively included in the multiple nodes, and do not specifically limit the number of the first node, the second node, the third node, the fourth node, and the fifth node.
[0056] Based on this, the process of generating the hybrid modality graph based on the input data may include the following steps 21 - step 23.
[0057] Step 21: Extract features from the input data to obtain a first hybrid modality graph, where the first hybrid modality graph includes each node, the features of each node, each edge, and the features of each edge.
[0058] As mentioned above, the input data includes text data and image data. Correspondingly, in specific implementation, first, the text data and the image data can be subjected to feature extraction through a vision - language model to obtain multiple nodes. And the text data and the image features can be subjected to feature extraction through the vision - language model to obtain the features respectively included in each node, each edge, and the features of each edge.
[0059] Among them, the multiple nodes can form a node set V m , and the multiple edges can form an edge set E m , V m = {v 1 , v 2 , …, v N}, where N is the number of multiple nodes, and E m = {e i,j | i, j ∈ {1, …, N}, i ≠ j}, and e i,j represents the edge between node i and node j, and the features of each edge include the similarity relationship and the spatial position relationship between the objects corresponding to each node.
[0060] Taking node i and node j as an example, when node i is the first node, the features of node i may include the object category feature and the text feature The features of node j may include the object category feature and the text feature and the visual feature That is, the features of node i are represented as The feature of node j is represented as The edge between node i and node j is represented as where r i,j represents the relationship between node i and node j, which may include semantic similarity relationship and spatial position relationship. Among them, the similarity relationship can describe the similarity between the objects corresponding to node i and node j, such as whether they are objects with similar styles or similar materials, etc.; the spatial position relationship can describe the spatial position relationship between the objects corresponding to node i and node j, such as the relative distance and relative orientation between them.
[0061] It should be noted that for the acquisition methods of the above three features of nodes, the embodiments of the present application may not specifically limit them. For example, the category feature can be extracted through an embedding layer, the text feature can be extracted through a text encoder, such as CLIP or LLaMa, etc., and the visual feature can be extracted through an image feature encoder, such as CLIP or DINO, etc.
[0062] In addition, it should be noted that if a certain node lacks object category features, text features or visual features, zero padding can be used to ensure the consistency of feature categories.
[0063] Next, based on each node, the features included in each node respectively, each edge and the features of each edge, a first hybrid modality graph can be generated. This first hybrid modality graph G m can be formally defined as G m =(V m , E m ).
[0064] In this way, generating the first hybrid modality graph based on input data of multiple modalities can significantly improve the flexibility and scalability of the data. It can support text, image and their combined inputs, adapt to diverse user needs, and also provide high-quality input features for subsequent enhancement processing and relationship prediction processing.
[0065] Step 22: Perform enhancement processing on the text features of the first node in the first hybrid modality graph to obtain a second hybrid modality graph, and the enhancement processing is used to convert the text features of the first node into visual features.
[0066] As mentioned above, the features of the first node lack the visual features of the first node. Based on this, in the embodiments of the present application, enhancement processing can be performed on the nodes lacking visual features to provide high-quality visual features, thereby solving the deficiencies existing in the existing 3D scene generation solutions in terms of geometric control ability.
[0067] As an example, taking the first node as node i, correspondingly, the features of the first node can be enhanced by a visual enhancement module, which can adopt the Q-VAE architecture (a variant of the variational autoencoder). Among them, the Q-VAE architecture includes an encoder, a decoder, and a codebook.
[0068] Correspondingly, first, the text features of the first node can be encoded by the encoder to obtain a latent vector, as shown in the following formula (1):
[0069]
[0070] Among them, is the latent vector, E is the encoder, is the text feature of the first node.
[0071] Next, the obtained latent vector can be quantized by the codebook to obtain a quantized latent vector. Among them, the codebook can be expressed as This means that the codebook C includes K embedding vectors e k . Correspondingly, the process of quantization by the codebook can be reflected as determining the n embedding vectors closest to from these K embedding vectors, as shown in the following formula (2):
[0072]
[0073] Among them, is the quantized latent vector, is the j-th embedding vector, is the latent vector, is the l-th embedding vector.
[0074] Combined with the above formula (2), during the quantization process, the square value of the Euclidean distance between the latent vector and each of the K embedding vectors can be calculated first, and then n embedding vectors are selected from the K embedding vectors based on the square value corresponding to each embedding vector. The sum of the square values corresponding to the selected n embedding vectors is the smallest, so that the quantized latent vector
[0075] After that, the quantized latent vector can be decoded by the decoder to obtain the visual features of the first node, as shown in the following formula (3):
[0076]
[0077] Among them, is the visual feature of the first node, D is the decoder, is the quantized latent vector.
[0078] In this way, based on the visual features of the above-mentioned first node and the first hybrid modality graph G m , a second hybrid modality graph can be generated That is, through the initial first hybrid modality graph, a second hybrid modality graph with enhanced vision is further generated, providing strong support for generating high-quality and geometrically accurate three-dimensional objects.
[0079] In addition, in the embodiment of the present application, the training objective of the above-mentioned vision enhancement module is to maximize the evidence lower bound (ELBO) of the data to optimize the generation quality of visual features. Accordingly, the objective function of this vision enhancement module can be shown as the following formula (4):
[0080]
[0081] where is the value of the objective function, q E (h|t) is the latent vector distribution, p D (u|h) is the likelihood probability of visual feature generation, D KL is the Kullback-Leibler divergence, which is used to measure the difference between the latent distribution and the Gaussian prior distribution p(h), and β is an important parameter used to weigh the divergence term. And, to solve the non-differentiability problem of the quantization process, the Gumbel-Softmax relaxation technique is used to optimize the ELBO objective function.
[0082] Step 23: Perform relationship prediction on the features corresponding to the second node and the third node in the second hybrid modality graph to obtain a hybrid modality graph. The relationship prediction is used to generate the features of the edge between the second node and the third node to predict the relationship between the second node and the third node.
[0083] As mentioned above, there is a lack of corresponding relationship between the second node and the third node. Based on this, in the embodiment of the present application, relationship prediction can be performed for each pair of nodes lacking a relationship to generate a more reasonable scene layout, thereby solving the problem that the existing three-dimensional scene generation scheme lacks the inference ability of relationship information, resulting in inaccurate scene layout generation results.
[0084] As an example, relationship prediction can be performed on the features of the second node and the third node through a relationship prediction module. Here, in the embodiment of the present application, the structure of the relationship prediction module may not be specifically limited, and any existing or future possible network capable of extracting triple features can be used.
[0085] Accordingly, in the embodiment of the present application, taking node m as an example for the second node and node n as an example for the third node, first, a triple can be constructed, which can be used as the input of the relationship prediction module. Among them, the triple can sequentially include the features of the second node, the relationship between the second node and the third node, and the features of the third node, that is Since the relationship between the second node and the third node is missing, the relationship between the second node and the third node is filled with zeros. That is to say, the triple can be expressed as Thus, the consistency of the feature space is ensured.
[0086] Next, the triple can be input into the relationship prediction model, and the features of the edge between the second node and the third node are output through the relationship prediction model. Here, the features of the edge between the second node and the third node output by the relationship prediction module can represent the relationship between the second node and the third node.
[0087] In this way, based on the features of the edge between the second node and the third node and the second hybrid modality graph a hybrid modality graph is generated That is, a hybrid modality graph with the missing relationship complemented is further generated through the second hybrid modality graph, so as to capture the relationship between nodes and effectively improve the overall consistency and rationality of the scene layout.
[0088] In addition, in the embodiment of the present application, the above relationship prediction module can be trained by the cross-entropy loss function. Correspondingly, the cross-entropy loss function can be shown in the following formula (5):
[0089]
[0090] Among them, is the cross-entropy loss value, N is the number of node pairs, C is the number of relationship categories, y ic is the one-hot encoding of the true relationship, is the probability of the relationship category predicted by the relationship prediction module. By minimizing this loss function, the relationship prediction module can accurately predict the relationship between nodes, thereby efficiently complementing the missing relationship.
[0091] S103: Generate a three-dimensional scene based on the hybrid modality graph. The content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0092] In the embodiment of the present application, a two-branch diffusion model can be used to generate a three-dimensional scene. Its main goal is to use the hybrid modality graph as a condition to generate a three-dimensional scene that meets this condition. The three-dimensional scene can include the spatial layout and geometric shape of the object. Correspondingly, combined with Figure 2As shown, the two-branch diffusion model may include two branches, namely, a layout branch and a shape branch. Among them, the layout branch is responsible for generating the spatial position information of the object, such as the center position, size, and rotation angle of the object, while the shape branch is responsible for generating the three-dimensional geometric shape of the object.
[0093] Furthermore, in order to effectively achieve information exchange and relationship modeling between nodes, the two-branch diffusion model can use a graph encoder E g , such as a Graph Convolutional Network (GCN), to update the latent representation of the hybrid modality graph. For ease of understanding, the process of generating a three-dimensional scene based on the hybrid modality graph will be exemplarily described below in conjunction with a possible implementation manner.
[0094] As a possible implementation manner, the hybrid modality graph can be used as the condition for the above-mentioned shape branch, and through sampling and denoising processing by the shape branch, the geometric shape information of the object corresponding to each node can be obtained. And, using the hybrid modality graph as the condition for the above-mentioned layout branch, through sampling and denoising processing by the layout branch, the spatial position information of the object corresponding to each node in the three-dimensional scene can be obtained.
[0095] In specific implementation, first, the hybrid modality graph can be used as the condition and input into the shape branch and the layout branch. In this way, the hybrid modality graph can enable the shape branch to know the initial shape of the object to be generated and enable the layout branch to know the initial layout of the object to be generated. Then, sampling data can be obtained by sampling from a Gaussian distribution or other suitable noise distributions, and this sampling data can reflect the initial geometric shape information and initial spatial position information of the object. Subsequently, using the sampling data as the input data and performing step-by-step denoising through the shape branch and the layout branch, the sampled data can be converted into the final shape and layout, that is, the geometric shape information of the object corresponding to each node and the spatial position information in the three-dimensional scene.
[0096] Furthermore, in the model training stage, first, real data can be obtained, that is, the geometric shape information and spatial position information of the object in a real three-dimensional scene. Then, noise addition processing can be performed on the real data. In this way, the two-branch diffusion model can use the hybrid modality graph as the condition to predict the noise added to the real data, and measure the difference between the accuracy of the model prediction and the real noise through a loss function. In this way, by continuously optimizing the parameters of the model to minimize the loss function, a two-branch diffusion model with accurate noise prediction can be obtained.
[0097] More specifically, the shape branch uses the Truncated Signed Distance Field (TSDF) to represent the object shape and encodes it through a pre-trained and frozen VQ-VAE model to generate an initial latent representation. Subsequently, the final shape is gradually generated through a denoising process.
[0098] During the training phase of this shape branch, for each denoising time step t ∈ {1, 2, …, T}, the graph encoder E g can process the latent representation and the latent graph to generate an updated latent representation and the latent graph where the latent graph is obtained based on the above mixed-modal graph and the nodes in the above updated latent graph are denoted as Next, the nodes in the above updated latent graph θ can serve as the conditional input to the denoiser ∈ θ which can be implemented using a 3D UNet (a deep learning architecture for image segmentation) or DiT (a new model combining Transformer architecture and diffusion model), etc.
[0099] In practical applications, the training objective of this shape branch is to minimize the deviation between the true noise ∈ and the model-predicted noise, that is, the loss function of the shape branch is used to minimize the deviation between the true noise of the object shape in the 3D scene and the predicted noise of the shape branch. Accordingly, the loss function of the shape branch can be shown as the following formula (6):
[0100]
[0101] where is the loss value of the shape branch, is the updated latent representation, t is the denoising time step, ∈ is the true noise, ∈ θ is the denoiser, is the updated latent graph and the nodes in it.
[0102] The layout branch is based on the representation of the object's bounding box, and each bounding box includes the position dimensions and the rotation angle
[0103] During the training phase of this layout branch, to improve the scale and numerical stability of the training process, the rotation angle can be represented in the form of trigonometric functions to improve numerical stability, that is, the rotation angle can be expressed as Moreover, and can be normalized.
[0104] In the graph structure, the layout branch performs relationship modeling through the graph encoder E g and can generate an updated latent layout representation and node embeddings Next, the above-mentioned updated node embeddings and latent layout representation can be used as conditional inputs to the denoiser ψ θ , and the denoiser ψ θ gradually generates the final spatial layout. The denoiser ψ θ can be implemented using a 1D UNet or Transformer (a deep learning model based on self-attention mechanism), etc.
[0105] In practical applications, the training objective of this layout branch is to minimize the deviation between the real noise ψ and the model-predicted noise, that is, the loss function of the layout branch is used to minimize the deviation between the real noise of the object layout in the 3D scene and the predicted noise of the layout branch. Correspondingly, the loss function of the layout branch can be shown as the following formula (7):
[0106]
[0107] where is the loss value of the layout branch, is the updated latent layout representation, t is the denoising time step, ψ is the real noise, and ψ θ is the denoiser, is the updated latent graph and the node embeddings in it.
[0108] Based on this, in the embodiments of this application, the overall optimization objective of the above-mentioned double-branch diffusion model can be reflected by the weighted combination of the losses of the shape branch and the layout branch, that is, the loss function of the double-branch diffusion model is determined based on the weight corresponding to the loss function of the shape branch, the loss function of the shape branch, the weight of the loss function of the layout branch, and the loss function of the layout branch. Specifically, it can be shown as the following formula (8):
[0109]
[0110] where is the loss value of the double-branch diffusion model, and α 1 is the weight of the loss function of the layout branch, is the loss value of the layout branch, α 2 is the weight of the loss function of the shape branch, is the loss value of the shape branch.
[0111] It can be seen that based on the relevant content of the above steps S101 - S103, in the embodiment of the present application, first, input data can be obtained. The input data includes data of multiple modalities and is used to generate a three - dimensional scene. Then, a hybrid modality graph can be generated based on the input data. The hybrid modality graph includes multiple nodes and multiple edges between the multiple nodes. Each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge. In this way, a three - dimensional scene can be generated based on the hybrid modality graph. The content of the three - dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three - dimensional scene. Since the hybrid modality graph is generated based on data of multiple modalities, that is to say, this solution can support users to input data of different modalities, and then convert the data of different modalities into a hybrid modality graph for generating a three - dimensional scene. Therefore, it can flexibly process data of different modalities, improve the adaptability to the input data, and thus significantly improve the support ability for complex data input by users. In addition, the image generated based on the hybrid modality graph can comprehensively process the spatial layout and geometric shape of the object under complex input conditions, so as to generate a three - dimensional scene with high spatial consistency and geometric details.
[0112] Furthermore, based on the image processing method provided in the above embodiment, the embodiment of the present application can also provide an image processing device. The following will describe this image processing device in combination with the embodiment and the drawings respectively.
[0113] Figure 3 is a schematic structural diagram of an image processing device provided by an embodiment of the present application. Combining Figure 3 as shown, the image processing device 300 provided by the embodiment of the present application may include:
[0114] An input acquisition module 301, configured to acquire input data, where the input data includes data of multiple modalities and is used to generate a three - dimensional scene;
[0115] A hybrid modality graph generation module 302, configured to generate a hybrid modality graph based on the input data. The hybrid modality graph includes multiple nodes and multiple edges between the multiple nodes. Each node in the multiple nodes is used to represent an object, and each edge in the multiple edges is used to represent the relationship between the two nodes corresponding to each edge;
[0116] A three-dimensional scene generation module 303 is configured to generate the three-dimensional scene based on the hybrid modality graph. The content of the three-dimensional scene includes the geometric shape information of the object corresponding to each node and the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0117] Optionally, the multiple nodes include a first node, a second node, a third node, and a fourth node. The features of the first node include the object category feature and text feature of the first node. The features of the fourth node include the object category feature, text feature, and visual feature of the fourth node. There is no edge between the second node and the third node.
[0118] The hybrid modality graph generation module 302 includes:
[0119] A feature extraction module is configured to perform feature extraction on the input data to obtain a first hybrid modality graph, where the first hybrid modality graph includes each node, the features of each node, each edge, and the features of each edge.
[0120] An enhancement processing module is configured to perform enhancement processing on the text feature of the first node in the first hybrid modality graph to obtain a second hybrid modality graph. The enhancement processing is used to convert the text feature of the first node into a visual feature.
[0121] A relationship prediction module is configured to perform relationship prediction on the features corresponding to the second node and the third node in the second hybrid modality graph to obtain the hybrid modality graph. The relationship prediction is used to generate the features of the edge between the second node and the third node to predict the relationship between the second node and the third node.
[0122] Optionally, the input data includes text data and image data. The feature extraction module is specifically configured to:
[0123] Perform feature extraction on the text data and the image data through a vision-language model to obtain the multiple nodes, and perform feature extraction on the text data and image features through the vision-language model to obtain each node, the features included in each node, each edge, and the features of each edge.
[0124] Generate the first hybrid modality graph based on each node, the features included in each node, each edge, and the features of each edge.
[0125] Optionally, the enhancement processing module is specifically configured to:
[0126] Encode the text feature of the first node through an encoder to obtain a latent vector.
[0127] Quantize the potential vector through a codebook to obtain a quantized potential vector;
[0128] Decode the quantized potential vector through a decoder to obtain the visual features of the first node;
[0129] Generate the second hybrid modality graph based on the visual features of the first node and the first hybrid modality graph.
[0130] Optionally, the relationship prediction module is specifically configured to:
[0131] Construct a triple, which sequentially includes the features of the second node, the relationship between the second node and the third node, and the features of the third node, and the relationship between the second node and the third node is filled with zeros;
[0132] Input the triple into a relationship prediction model, and output the features of the edge between the second node and the third node through the relationship prediction model;
[0133] Generate the hybrid modality graph based on the features of the edge between the second node and the third node and the second hybrid modality graph.
[0134] Optionally, the three-dimensional scene generation module 303 is specifically configured to:
[0135] Use the hybrid modality graph as the condition of the shape branch in the dual-branch diffusion model, and perform sampling and denoising processing through the shape branch to obtain the geometric shape information of the object corresponding to each node;
[0136] Use the hybrid modality graph as the condition of the layout branch in the dual-branch diffusion model, and perform sampling and denoising processing through the layout branch to obtain the spatial position information of the object corresponding to each node in the three-dimensional scene.
[0137] Optionally, the loss function of the shape branch is used to minimize the deviation between the true noise of the object shape in the three-dimensional scene and the predicted noise of the shape branch, the loss function of the layout branch is used to minimize the deviation between the true noise of the object layout in the three-dimensional scene and the predicted noise of the layout branch, and the loss function of the dual-branch diffusion model is determined based on the weight corresponding to the loss function of the shape branch, the loss function of the shape branch, the weight of the loss function of the layout branch, and the loss function of the layout branch.
[0138] Furthermore, an embodiment of the present application also provides an electronic device, including: a processor, a memory, and a system bus;
[0139] The processor and the memory are connected through the system bus;
[0140] The memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any implementation step of the above image processing method.
[0141] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on an electronic device, any implementation step of the above image processing method is enabled.
[0142] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application. It should be noted that the embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.
[0143] For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0144] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0145] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image processing method, characterized in that: include: Acquiring input data, wherein the input data includes data of multiple modalities, and the input data is used to generate a three-dimensional scene; Generate a mixed modal graph based on the input data, the mixed modal graph comprising a plurality of nodes and a plurality of edges between the plurality of nodes, each of the plurality of nodes being used to represent an object, and each of the plurality of edges being used to represent a relationship between two nodes corresponding to each edge; The three-dimensional scene is generated based on the mixed modal graph, and the content of the three-dimensional scene includes geometric shape information of the object corresponding to each node and spatial position information of the object corresponding to each node in the three-dimensional scene.
2. The image processing method according to claim 1, characterized in that: The multiple nodes include a first node, a second node, a third node, and a fourth node, the features of the first node include an object category feature and a text feature of the first node, the features of the fourth node include an object category feature, a text feature, and a visual feature of the fourth node, and there is no edge between the second node and the third node; Generating a mixed mode diagram based on the input data comprises: Performing feature extraction on the input data to obtain a first mixed mode graph, wherein the first mixed mode graph includes each node, a feature of each node, each edge, and a feature of each edge; Performing enhancement processing on the text features of the first node in the first mixed modal graph to obtain a second mixed modal graph, wherein the enhancement processing is used to convert the text features of the first node into visual features; Relationship prediction is performed on features respectively corresponding to the second node and the third node in the second mixed modal graph to obtain the mixed modal graph, wherein the relationship prediction is used to generate features of the edge between the second node and the third node to predict the relationship between the second node and the third node.
3. The image processing method according to claim 2, characterized in that: The input data includes text data and image data, and the feature extraction of the input data to obtain a first mixed modal graph includes: Extracting features from the text data and the image data through a visual language model to obtain the multiple nodes, and extracting features from the text data and the image features through the visual language model to obtain each node, features respectively included in each node, each edge, and features of each edge; The first mixed modal graph is generated based on each node, the features respectively included in each node, each edge and the features of each edge.
4. The image processing method according to claim 2, characterized in that: The step of enhancing the text features of the first node in the first mixed modal graph to obtain a second mixed modal graph includes: Encoding the text features of the first node by an encoder to obtain a latent vector; Quantizing the latent vector using a code book to obtain a quantized latent vector; Decoding the quantized latent vector by a decoder to obtain a visual feature of the first node; The second mixed modal graph is generated based on the visual features of the first node and the first mixed modal graph.
5. The image processing method according to claim 2, characterized in that: The performing relationship prediction on the features respectively corresponding to the second node and the third node in the second mixed modal graph to obtain the mixed modal graph includes: constructing a triplet, wherein the triplet sequentially includes a feature of the second node, a relationship between the second node and the third node, and a feature of the third node, wherein the relationship between the second node and the third node is filled with zeros; Inputting the triple into a relationship prediction model, and outputting features of an edge between the second node and the third node through the relationship prediction model; The mixed modal graph is generated based on the feature of the edge between the second node and the third node and the second mixed modal graph.
6. The image processing method according to any one of claims 1 to 5, characterized in that: The generating the three-dimensional scene based on the mixed modal graph comprises: Using the mixed mode graph as a condition of a shape branch in a dual-branch diffusion model, sampling and denoising are performed through the shape branch to obtain geometric shape information of the object corresponding to each node; The mixed mode graph is used as a condition of the layout branch in the dual-branch diffusion model, and sampling and denoising processing are performed through the layout branch to obtain the spatial position information of the object corresponding to each node in the three-dimensional scene.
7. The image processing method according to claim 6, characterized in that: The loss function of the shape branch is used to minimize the deviation between the real noise of the object shape of the three-dimensional scene and the predicted noise of the shape branch, and the loss function of the layout branch is used to minimize the deviation between the real noise of the object layout of the three-dimensional scene and the predicted noise of the layout branch. The loss function of the dual-branch diffusion model is determined based on the weight corresponding to the loss function of the shape branch, the loss function of the shape branch, the weight of the loss function of the layout branch, and the loss function of the layout branch.
8. An image processing device, characterized in that: include: An input acquisition module, used to acquire input data, wherein the input data includes data of multiple modes, and the input data is used to generate a three-dimensional scene; A hybrid modal graph generating module, configured to generate a hybrid modal graph based on the input data, wherein the hybrid modal graph includes a plurality of nodes and a plurality of edges between the plurality of nodes, each of the plurality of nodes being used to represent an object, and each of the plurality of edges being used to represent a relationship between two nodes corresponding to each edge; A three-dimensional scene generation module is used to generate the three-dimensional scene based on the mixed modal graph, and the content of the three-dimensional scene includes geometric shape information of the object corresponding to each node and spatial position information of the object corresponding to each node in the three-dimensional scene.
9. An electronic device, characterized in that: The device includes: a processor, a memory, and a system bus; The processor and the memory are connected via the system bus; The memory is used to store a program, wherein the program includes instructions, and when the instructions are executed by the processor, the processor executes the steps of the image processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a terminal device, the steps of the image processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
A task-oriented text generation image network model
CN111858954A
Multi-modal medical image multi-organ positioning method based on one-to-one target query Transform
CN114359642A
Indoor single-view scene semantic reconstruction method and system
CN116385660A
Three-dimensional scene generation method based on pre-training language model and related components
CN117475089A
Three-dimensional model reconstruction and image generation method and device, storage medium and program product
CN118298127A