Image generation method and device, equipment, storage medium and product
Through the image generation model, data fusion of multi-frame and multi-view real images is generated to generate simulated images of multi-view and multi-frame numbers, solving the problem of limited labeling data sets in autonomous driving technology, and achieving efficient data generation and improvement of autonomous driving technology.
Patent Information
- Application Number
- CN202510095073.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In autonomous driving technology, it is costly to obtain continuous synthetic labeled data with multiple perspective angles and multiple frames, and the existing data sets are limited, making it difficult to effectively improve the performance of autonomous driving technology.
By acquiring multi-frame, multi-view real images and labeled data, the image generation model is used to fusion data to generate simulated images with multiple-view and multi-frame numbers. The method includes adding noise to the image generation model, fusing the real image and the labeled data, performing multiple fusion processing until T simulated images are generated.
It realizes the generation of continuous and real simulated images based on limited multi-frame and multi-view real images, which reduces the cost of data annotation, and integrates information from different perspectives and time dimensions, improving the performance of autonomous driving technology.
Smart Images

Figure CN120014089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data synthesis, and in particular to an image generation method and device, equipment, storage medium, and product. Background Art
[0002] Autonomous driving technology has attracted extensive attention and research due to its great potential in improving travel safety, efficiency and economy. The development of autonomous driving technology is inseparable from various subtasks such as target detection and recognition, semantic segmentation, target tracking and trajectory prediction. Although the implementation methods and roles of these tasks are different, they all rely on labeled data. Obtaining accurate labeled data to train and improve the performance of these tasks plays a vital role in the field of autonomous driving. However, the acquisition of massive multi-scene and multi-view labeled data faces high costs in terms of time and manpower. Therefore, how to generate multi-view and multi-frame continuous synthetic labeled data to improve autonomous driving technology under the condition of limited existing autonomous driving labeled data sets is a technical problem that needs to be solved urgently. Summary of the invention
[0003] One of the objects of the present invention is to provide an image generation method to solve the problem of how to generate multi-perspective and multi-frame continuous synthetic annotation data when the existing autonomous driving annotation data set is limited; the second object is to provide an image generation device; the third object is to provide an electronic device; the fourth object is to provide a computer-readable storage medium; and the fifth object is to provide a computer program product.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A method for generating an image, the method comprising:
[0006] Obtain multi-frame multi-view real images and annotation data of real images;
[0007] Based on the added noise of T time steps included in the image generation model, each real image in the multi-frame multi-view real image and the annotated data of each real image are fused to obtain T first fusion results of each real image; T is an integer greater than 0;
[0008] Based on the image generation model, the first fusion result of the real image of the same frame from other perspectives except any one perspective is fused into the first fusion result of the real image of any perspective to obtain a second fusion result of the real image of any perspective;
[0009] Based on the image generation model, the second fusion results of the remaining frames of real images of the same perspective except any frame are fused into the second fusion result of any frame of real image to obtain a processing result of any frame of real image;
[0010] Until the T first fusion results of each real image are processed, T simulation images of each real image are obtained based on the T processing results of each real image in the multi-frame multi-view real images.
[0011] According to the above technical means, T simulated images corresponding to each real image can be generated based on a limited number of multi-frame multi-perspective real images, thereby achieving the purpose of reducing costs; in addition, by fusing real images of different perspectives of the same frame, complementary information from different perspectives can be integrated; further, by fusing real images of different frames of the same perspective, information in the time dimension can be further integrated; therefore, the generated simulated images are continuous and realistic between different perspectives and different frame numbers, thereby improving autonomous driving technology.
[0012] Furthermore, based on the added noise of T time steps included in the image generation model, each real image in the multi-frame multi-view real image and the annotation data of each real image are fused to obtain T first fusion results of each real image, including: encoding each real image to obtain a first encoding result of each real image; encoding the annotation data of each real image to obtain a second encoding result of each real image; performing noise processing on the first encoding result of each real image based on the added noise of the T time steps to obtain T noise-added results of each real image; fusing the second encoding result of each real image into the T noise-added results of each real image to obtain T first fusion results of each real image.
[0013] Furthermore, the annotated data includes relevant data and scene data of dynamic targets; the encoding of the annotated data of each real image to obtain a second encoding result of each real image includes: encoding the relevant data of the dynamic targets included in each real image to obtain a dynamic target encoding result; encoding the scene data of each real image to obtain a scene encoding result; the second encoding result includes the dynamic target encoding result and the scene encoding result.
[0014] Furthermore, the second encoding result of each real image is integrated into the T noise-added results of each real image to obtain T first fusion results of each real image, including: using a convolution layer to process the T noise-added results of each real image to obtain T convolution results of each real image; integrating the dynamic target encoding result and the scene encoding result of each real image into the T convolution results of each real image to obtain T first fusion results of each real image.
[0015] According to the above technical means, the dynamic target encoding result and the scene encoding result of each real image are fused into T convolution results of each real image, so that the first fusion result of each real image is consistent with the dynamic target and the scene.
[0016] Furthermore, the method of fusing the dynamic target coding result and the scene coding result of each real image into T convolution results of each real image to obtain T first fusion results of each real image includes: fusing the scene coding result of each real image into the T convolution results of each real image to obtain T third fusion results of each real image; fusing the dynamic target coding result of each real image into the T third fusion results of each real image to obtain T first fusion results of each real image.
[0017] Furthermore, the relevant data of the dynamic target includes its category and three-dimensional coordinates; encoding the relevant data of the dynamic target included in each real image to obtain a dynamic target encoding result includes: encoding the category of the dynamic target included in each real image to obtain a category encoding result; encoding the three-dimensional coordinates of the dynamic target included in each real image to obtain a position encoding result; fusing the category encoding result and the position encoding result to obtain the dynamic target encoding result.
[0018] Furthermore, the scene data includes map data and context data; encoding the scene data of each real image to obtain a scene coding result includes: encoding the map data of each real image to obtain a map coding result; encoding the context data of each real image to obtain a context coding result; splicing the map coding result and the context coding result to obtain the scene coding result.
[0019] Furthermore, based on the image generation model, the second fusion results of the real images of the remaining frames except any frame of the same perspective are integrated into the second fusion result of any frame of the real image to obtain the processing result of any frame of the real image, including: determining the similarity between the second fusion result of any frame of the real image in multiple frames of the same perspective and the second fusion results of the real images of the remaining frames before any frame; integrating the second fusion result of the real image of the target frame corresponding to the maximum similarity and the second fusion result of the real image of the previous frame of any frame into the second fusion result of any frame of the real image to obtain the processing result of any frame of the real image.
[0020] According to the above technical means, the second fusion result of the target frame real image most similar to any frame and the second fusion result of the previous frame real image of any frame are fused into the second fusion result of any frame real image to obtain the processing result of any frame real image. In this way, the comprehensiveness and accuracy of the processing result of any frame real image can be enhanced.
[0021] Furthermore, the image generation method also includes: obtaining multiple frames of multi-view sample images and annotation data of each sample image; based on the added noise of T time steps included in the diffusion model, fusing each sample image in the multiple frames of multi-view sample images and the annotation data of each sample image to obtain T first fusion results of each sample image; based on the diffusion model, fusing the first fusion result of the sample images of the remaining perspectives of the same frame except any one perspective into the first fusion result of the sample image of any one perspective to obtain a second fusion result of the sample image of any one perspective; based on the diffusion model, fusing the second fusion result of the sample images of the remaining frames of the same perspective except any one frame into the second fusion result of the sample image of any one frame to obtain a processing result of the sample image of any one frame; until the T first fusion results of each sample image are processed, the T processing results of each sample image are subjected to noise prediction processing to obtain T predicted noises of each sample image; based on the T added noises and T predicted noises of each sample image in the multiple frames of multi-view sample images, the diffusion model is trained to obtain the trained image generation model.
[0022] An image generating device, the image generating device comprising:
[0023] An acquisition unit, used for acquiring multi-frame multi-view real images and annotation data of the real images;
[0024] A processing unit, configured to fuse each real image and the annotated data of each real image in the multiple frames of multi-view real images based on the added noise of T time steps included in the image generation model, to obtain T first fusion results of each real image; T is an integer greater than 0;
[0025] The processing unit is further used to fuse the first fusion result of the real image of the other perspectives of the same frame except any perspective into the first fusion result of the real image of any perspective based on the image generation model to obtain a second fusion result of the real image of any perspective;
[0026] The processing unit is further configured to fuse the second fusion results of the remaining frames of real images of the same perspective except any frame into the second fusion result of any frame of real image based on the image generation model to obtain a processing result of any frame of real image;
[0027] The processing unit is also used to obtain T simulation images of each real image based on the T processing results of each real image in the multi-frame multi-view real image until the T first fusion results of each real image are processed.
[0028] An electronic device comprises: a processor and a memory configured to store a computer program that can be run on the processor, wherein the processor is configured to execute the steps of the above method when running the computer program.
[0029] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.
[0030] A computer program product comprises a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the steps of the aforementioned method are implemented.
[0031] Beneficial effects of the present invention:
[0032] (1) The present invention can generate T simulated images corresponding to each real image based on a limited number of multi-frame multi-perspective real images, thereby achieving the purpose of reducing costs; in addition, by fusing real images of different perspectives of the same frame, complementary information from different perspectives can be integrated; further, by fusing real images of different frames of the same perspective, information in the time dimension can be further integrated; therefore, the generated simulated images are continuous and realistic between different perspectives and different frame numbers, thereby improving autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 1 ;
[0034] Figure 2 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 2 ;
[0035] Figure 3 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 3 ;
[0036] Figure 4 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 4 ;
[0037] Figure 5 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 5 ;
[0038] Figure 6 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 6 ;
[0039] Figure 7 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 7 ;
[0040] Figure 8 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 8 ;
[0041] Fig. 9 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 9 ;
[0042] Fig.10 An overall framework diagram of an image generation method provided by an embodiment of the present invention;
[0043] Fig.11 Schematic diagram of the noise addition and denoising process of the diffusion module in an embodiment of the present invention;
[0044] Fig.12 A frame of six-viewing angle simulation image as an example of an embodiment of the present invention;
[0045] Fig.13 A schematic diagram of the composition structure of an image generating device according to an embodiment of the present invention;
[0046] Fig.14 Schematic diagram of the composition structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present invention, the implementation of the embodiments of the present invention is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not intended to limit the embodiments of the present invention.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein are only for the purpose of describing the present embodiment and are not intended to limit the present invention.
[0049] In the following description, references to “some embodiments,” “this embodiment,” “this embodiment,” and examples, etc., describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0050] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the present embodiment described here can be implemented in an order other than that illustrated or described here.
[0051] In this embodiment, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, object A and / or object B may represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0052] The embodiment of the present invention provides an image generation method. Figure 1 A schematic diagram of a process of an image generation method provided by an embodiment of the present invention Figure 1 , applied to electronic devices, including but not limited to desktop computers, smart phones, and tablet computers.
[0053] like Figure 1 As shown, the image generation method includes the following steps:
[0054] Step S101: Acquire multiple frames of multi-view real images and annotation data of the real images.
[0055] In the embodiment of the present invention, it can be applied to intelligent driving scenarios. Cameras are installed at the left front, front, right front, left rear, rear and right rear of the car respectively. Then, driving images taken by the six cameras at multiple times can be obtained, that is, multi-frame six-view driving images or real images. The annotated data of the real image refers to the data processed to mark the key information or features in the real image.
[0056] Step S102: Based on the added noise of T time steps included in the image generation model, each real image in the multi-frame multi-view real image and the annotation data of each real image are fused to obtain T first fusion results of each real image; T is an integer greater than 0.
[0057] In the embodiment of the present invention, based on the addition of noise at T time steps, each real image is subjected to noise addition processing to obtain T noise addition results of each real image, and then the annotation data of each real image is fused into the T noise addition results of each real image to obtain T first fusion results of each real image. Specifically, based on the addition of noise at the first time step, each real image is subjected to noise addition processing to obtain the first noise addition result of each real image, and then the annotation data of each real image is fused into the first noise addition result of each real image to obtain the first first fusion result of each real image; based on the addition of noise at the second time step, the first noise addition result of each real image is subjected to noise addition processing to obtain the second noise addition result of each real image, and then the annotation data of each real image is fused into the second noise addition result of each real image to obtain the second first fusion result of each real image; based on the addition of noise at the third time step, each real image is subjected to noise addition processing to obtain the first noise addition result of each real image, and then the annotation data of each real image is fused into the second noise addition result of each real image to obtain the second first fusion result of each real image. Based on the noise added at the T-th time step, the (T-1)th noise addition result of each real image is subjected to noise addition to obtain the Tth noise addition result, and the annotation data of each real image is fused into the third noise addition result of each real image to obtain the third first fusion result of each real image. Based on the noise added at the T-th time step, the (T-1)th noise addition result of each real image is subjected to noise addition to obtain the Tth noise addition result, and the annotation data of each real image is fused into the T-th noise addition result of each real image to obtain the Tth first fusion result of each real image. That is, T first fusion results of each real image are obtained.
[0058] Alternatively, the annotated data of each real image is fused into each real image, and then noise is added based on T time steps to perform noise addition processing to obtain T first fusion results of each real image. The specific process can be referred to the previous paragraph and will not be repeated here.
[0059] Step S103: Based on the image generation model, the first fusion result of the real image of the same frame from other perspectives except any one perspective is fused into the first fusion result of the real image of any perspective to obtain the second fusion result of the real image of any perspective.
[0060] In the embodiment of the present invention, the cross attention mechanism is used to fuse the first fusion result of the real image of the same frame except for any one perspective into the first fusion result of the real image of any one perspective to obtain the second fusion result of the real image of any one perspective. In this way, the real images of different perspectives in the same frame can be kept consistent.
[0061] Step S104: Based on the image generation model, the second fusion results of the real images of the remaining frames except any frame at the same perspective are fused into the second fusion result of any frame of the real image to obtain the processing result of any frame of the real image.
[0062] In the embodiment of the present invention, the temporal attention mechanism is used to fuse the second fusion results of the real images of the remaining frames except any frame at the same perspective into the second fusion result of any frame of the real image to obtain the processing result of any frame of the real image. In this way, the continuity between the real images of different frames at the same perspective can be maintained.
[0063] Step S105: until the T first fusion results of each real image are processed, then T simulation images of each real image are obtained based on the T processing results of each real image in the multi-frame multi-view real image.
[0064] In an embodiment of the present invention, by fusing real images of different perspectives of the same frame, complementary information from different perspectives can be integrated; further, by fusing real images of different frames of the same perspective, information in the time dimension can be further integrated; therefore, the generated simulated image is continuous and realistic between different perspectives and different frame numbers, thereby improving autonomous driving technology.
[0065] In some embodiments of the present invention, based on the added noise of T time steps included in the image generation model, each real image in the multi-frame multi-view real image and the annotation data of each real image are fused to obtain T first fusion results of each real image, including the following steps S201 to S204:
[0066] Step S201: Encode each real image to obtain a first encoding result of each real image.
[0067] In the embodiment of the present invention, a variable self-decomposing encoder is used to encode each real image in a multi-frame multi-view real image to obtain a first encoding result of each real image.
[0068] Step S202: Encode the labeled data of each real image to obtain a second encoding result of each real image.
[0069] In the embodiment of the present invention, a correlation encoder is used to encode the annotation data of each real image in multiple frames of multi-view real images to obtain a second encoding result of each real image.
[0070] Step S203: performing noise addition processing on the first encoding result of each real image based on the added noise of T time steps to obtain T noise addition results of each real image.
[0071] In the embodiment of the present invention, noise is continuously added to the first encoding result of each real image to make it become a standard Gaussian distribution, and after T time steps, T noise-added results of each real image are obtained.
[0072] Specifically, when the time step is 1, the first noise is added to the first encoding result of each real image to obtain the first noise result; when the time step is 2, the second noise is added to the first first noise result of each real image to obtain the second noise result; when the time step is 3, the third noise is added to the second first noise result of each real image to obtain the third noise result; and when the time step is T, the Tth noise is added to the (T-1)th first noise result of each real image to obtain the Tth noise result. That is, T noise results of each real image are obtained.
[0073] Step S204: merging the second encoding result of each real image into the T noise-added results of each real image to obtain T first fusion results of each real image.
[0074] In the embodiment of the present invention, the second encoding result of each real image is fused into the T noise-added results of each real image using the spatial attention mechanism to obtain T first fusion results of each real image. In this way, the first fusion result of each real image is kept consistent with the labeled data.
[0075] In some embodiments of the present invention, the annotation data includes relevant data and scene data of the dynamic target.
[0076] Here, the dynamic targets included in the real image include but are not limited to other vehicles, pedestrians, and pets. The scene data of the real image refers to the relevant information describing the real image.
[0077] Based on this, the annotated data of each real image is encoded to obtain a second encoding result of each real image, including the following steps S301 to S302:
[0078] Step S301: Encode the relevant data of the dynamic target included in each real image to obtain the dynamic target encoding result.
[0079] Step S302: Encode the scene data of each real image to obtain a scene encoding result; the second encoding result includes a dynamic object encoding result and a scene encoding result.
[0080] Further, in some embodiments of the present invention, the second encoding result of each real image is fused into the T noise addition results of each real image to obtain T first fusion results of each real image, including the following steps S401 to S402:
[0081] Step S401: using a convolution layer to process the T noise addition results of each real image to obtain T convolution results of each real image.
[0082] In the embodiment of the present invention, the convolution layer is used to realize feature extraction and restoration. Specifically, the convolution layer can be used to extract key information from the noise results of each real image through layer-by-layer convolution and pooling operations, while reducing the dimension of the noise result to facilitate subsequent processing and analysis. Furthermore, in the restoration stage, the dimension of the extracted features is gradually restored through operations such as deconvolution or upsampling, so that the dimension of the obtained convolution result is consistent with the dimension of the noise result.
[0083] Step S402: Fusing the dynamic object coding result and the scene coding result of each real image into the T convolution results of each real image to obtain T first fusion results of each real image.
[0084] In some embodiments of the present invention, the dynamic object encoding result and the scene encoding result of each real image are fused into T convolution results of each real image to obtain T first fusion results of each real image, including the following steps S501 to S502:
[0085] Step S501: Fusing the scene encoding results of each real image into the T convolution results of each real image to obtain T third fusion results of each real image.
[0086] In the embodiment of the present invention, the scene encoding result of each real image is fused into each convolution result by using the spatial attention mechanism to obtain the corresponding second fusion result, so that the third fusion result of each real image is consistent with the scene.
[0087] Step S502: Fusing the dynamic target encoding result of each real image into the T third fusion results of each real image to obtain T first fusion results of each real image.
[0088] In the embodiment of the present invention, the spatial attention mechanism is used to fuse the dynamic target encoding results of each real image into each third fusion result to obtain the corresponding first fusion result, so that the first fusion result of each real image is coordinated and consistent with the dynamic target and scene.
[0089] In other embodiments of the present invention, the dynamic target coding result and the scene coding result of each real image are fused into T convolution results of each real image to obtain T first fusion results of each real image, including: fusing the dynamic target coding result of each real image into the T convolution results of each real image to obtain T third fusion results of each real image; fusing the scene coding result of each real image into the T third fusion results of each real image to obtain T first fusion results of each real image.
[0090] In some embodiments of the present invention, the relevant data of the dynamic target includes the category and the three-dimensional coordinates.
[0091] Here, the dynamic target includes but is not limited to vehicles, pets, and people. The three-dimensional coordinates of the dynamic target are the three-dimensional coordinates of the eight vertices of the dynamic target frame.
[0092] Based on this, the relevant data of the dynamic target included in each real image is encoded to obtain the dynamic target encoding result, including the following steps S601 to S603:
[0093] Step S601: Encode the category to which the dynamic target included in each real image belongs to, and obtain a category encoding result.
[0094] In the embodiment of the present invention, the contrastive language-image pre-training encoder (CLIP) encoder is composed of two sub-networks, one is a text encoder and the other is an image encoder. Here, the CLIP encoder is used to encode the category of the dynamic target included in each real image to obtain a category encoding result.
[0095] Step S602: Encode the three-dimensional coordinates of the dynamic targets included in each real image to obtain a position encoding result.
[0096] In the embodiment of the present invention, a Fourier encoder is used to encode the three-dimensional coordinates of the dynamic target included in each real image to obtain a position encoding result.
[0097] Step S603: Fusing the category coding result and the position coding result to obtain a dynamic target coding result.
[0098] In an embodiment of the present invention, the three-dimensional coordinates of the dynamic target and the category to which the dynamic target belongs are both information related to the dynamic target. Therefore, the encoding result of the category to which the dynamic target belongs, i.e., the category encoding result, and the encoding result of the three-dimensional coordinates of the dynamic target, i.e., the position encoding result, are fused to obtain the dynamic target encoding result.
[0099] In some embodiments of the present invention, the scene data includes map data and context data.
[0100] Here, the map data includes road information. The scenario data includes but is not limited to weather and time text description information in the real image.
[0101] Based on this, the scene data of each real image is encoded to obtain a scene encoding result, including the following steps S701 to S703:
[0102] Step S701: Encode the map data of each real image to obtain a map encoding result.
[0103] In the embodiment of the present invention, the map data of each real image is encoded using a CLIP encoder to obtain a map encoding result.
[0104] Step S702: Encode the context data of each real image to obtain a context encoding result.
[0105] In the embodiment of the present invention, the CLIP encoder is used to encode the context data of each real image to obtain a context encoding result. Specifically, the CLIP encoder is used to encode the weather text description information of each real image to obtain a weather encoding result. The CLIP encoder is used to encode the time text description information of each real image to obtain a time encoding result. The context encoding result includes a weather encoding result and a time encoding result.
[0106] Step S703: concatenate the map coding result and the context coding result to obtain the scene coding result.
[0107] In the embodiment of the present invention, the map data and the scenario data are two independent but related data types, and therefore, the map encoding result and the scenario encoding result are concatenated to obtain the scenario encoding result.
[0108] In some embodiments of the present invention, based on the image generation model, the second fusion results of the remaining frames of real images of the same perspective except any frame are fused into the second fusion result of any frame of real image to obtain the processing result of any frame of real image, including the following steps S801 to S802:
[0109] Step S801: Determine the similarity between the second fusion result of any real image in multiple frames of the same viewing angle and the second fusion results of the real images of the remaining frames before any frame.
[0110] Step S802: The second fusion result of the target frame real image corresponding to the maximum similarity and the second fusion result of the previous frame real image of any frame are fused into the second fusion result of any frame real image to obtain the processing result of any frame real image.
[0111] Exemplarily, the similarity between the second fusion result of the sixth frame of the real image in the ten frames of the same perspective and the second fusion result of the fifth frame of the real image is determined to obtain a first similarity. The similarity between the second fusion result of the sixth frame of the real image in the same perspective and the second fusion result of the fourth frame of the real image is determined to obtain a second similarity. The similarity between the second fusion result of the sixth frame of the real image in the same perspective and the second fusion result of the third frame of the real image is determined to obtain a third similarity. The similarity between the second fusion result of the sixth frame of the real image in the same perspective and the second fusion result of the second frame of the real image is determined to obtain a fourth similarity. The similarity between the second fusion result of the sixth frame of the real image in the same perspective and the second fusion result of the first frame of the real image is determined to obtain a fifth similarity. If the fourth similarity is the maximum value, the second fusion result of the second frame of the real image and the second fusion result of the fourth frame of the real image are fused into the second fusion result of the fifth frame of the real image to obtain the processing result of the fifth frame of the real image.
[0112] Specifically, the second fusion result of the second frame real image and the second fusion result of the fourth frame real image are fused into the second fusion result of the fifth frame real image by using the temporal attention mechanism to obtain the processing result of the fifth frame real image.
[0113] Based on this, the second fusion result of the target frame real image most similar to any frame and the second fusion result of the previous frame real image of any frame are fused into the second fusion result of any frame real image to obtain the processing result of any frame real image. In this way, the comprehensiveness and accuracy of the processing result of any frame real image can be enhanced.
[0114] In some embodiments of the present invention, based on the T processing results of each real image in the multi-frame multi-perspective real image, T simulated images of each real image are obtained, including: decoding the T processing results of each real image in the multi-frame multi-perspective real image to obtain T simulated images of each real image.
[0115] In some embodiments of the present invention, the training process of the image generation model includes the following steps S901 to S906:
[0116] Step S901: Acquire multiple frames of multi-view sample images and annotation data of each sample image.
[0117] In the embodiment of the present invention, the multi-frame multi-view sample images are training sample images. The labeled data of the sample images refers to data obtained by processing the sample images to mark key information or features in the sample images.
[0118] Step S902: Based on the added noise of T time steps included in the diffusion model, each sample image and the annotation data of each sample image in the multi-frame multi-view sample image are fused to obtain T first fusion results of each sample image.
[0119] Step S903: Based on the diffusion model, the first fusion results of the view sample images of the same frame except for any view are fused into the first fusion result of any view sample image to obtain the second fusion result of any view sample image.
[0120] Step S904: Based on the diffusion model, the second fusion results of the sample images of the frames other than any frame at the same viewing angle are fused into the second fusion result of any frame sample image to obtain the processing result of any frame sample image.
[0121] Step S905: until the processing of the T first fusion results of each sample image is completed, noise prediction processing is performed on the T processing results of each sample image to obtain T predicted noises of each sample image.
[0122] Step S906: Based on the T added noises and T predicted noises of each sample image in the multi-frame multi-view sample images, the diffusion model is trained to obtain a trained image generation model.
[0123] In the embodiment of the present invention, based on T added noises and T predicted noises of each sample image in the multi-frame multi-view sample images, a loss value is calculated using a pre-established loss function, and a set of model parameters corresponding to the minimum loss is used as the optimal model parameters. Based on the optimal model parameters, the parameters in the diffusion model are modified to obtain a trained image generation model.
[0124] In an embodiment of the present invention, by introducing multi-frame multi-view sample images and annotated data of each sample image, and learning the spatial association between the multi-view images of the same frame and the temporal association between the multi-frame images of the same view based on the diffusion model, the trained image generation model learns the spatial association between the multi-view images of the same frame and the temporal association between the multi-frame images of the same view. Therefore, the simulated image generated by the image generation model is continuous and realistic between different viewpoints and different frame numbers, thereby improving the autonomous driving technology.
[0125] Based on the above embodiments, an embodiment of the present invention provides an overall framework diagram of an image generation method. Fig.10 An overall framework diagram of an image generation method provided by an embodiment of the present invention, combined with Fig.10 , described as follows:
[0126] Step S1: Obtain camera recorded images of the six perspectives of the vehicle in the historical driving process, namely, the sample images, of all frames, namely, the front view, front left, front right, rear view, rear left, and rear right. And obtain the annotation data of each sample image, including the category to which the dynamic target belongs (such as a car), the three-dimensional coordinates of the eight vertices of the dynamic target frame, map data, and context data, such as the weather and time text description information in the sample image (such as "on a sunny day, there is a car on the street"), etc.
[0127] Step S2: Encode the video data of the six viewing angles of all frames during the vehicle driving process through a variable self-decoding encoder to obtain a video data encoding vector. The encoding process is expressed by formula (1):
[0128] z0=Encoder v (x) (1)
[0129] Among them, x is the original video data, Encoder v is a variable self-decomposing encoder, which is composed of a fully connected layer, and z0 is a video data encoding vector. The original video data is a multi-frame multi-view sample image. The video data encoding vector is the first encoding result of each sample image in the multi-frame multi-view sample image.
[0130] Step S3: Encode the weather and time text description information in each image using a text encoder to obtain an external information encoding vector. The encoding process is expressed by formula (2):
[0131] h text =Encoder text (L) (2)
[0132] Among them, Encoder text is the CLIP encoder, L is the weather and time text description information, h text is the external information encoding vector, which is the result of context encoding.
[0133] The map encoder is used to encode the map data in the image into a map encoding vector (i.e., the map encoding result). The encoding process is expressed by formula (3):
[0134] h map =Encoder map (M) (3)
[0135] Among them, M is the map data of the video data, Encoder map is the map data encoder, h map Encode the map vector.
[0136] The external information encoding vector and the map encoding vector are concatenated to obtain the scene encoding vector. This process is expressed by formula (4):
[0137] h scene =[h text ,h map ] (4)
[0138] Step S4: Encode the category to which the dynamic target belongs using the CLIP encoder to obtain a category encoding vector (i.e., category encoding result). This process is expressed by formula (5):
[0139] h class =Encoder text (L class ) (5)
[0140] Among them, L class is the text describing the dynamic target category, h class is the category encoding vector.
[0141] The Fourier encoder is used to encode the coordinate information of the eight vertices of the dynamic target 3D box to obtain the position encoding result. The process is expressed by formula (6):
[0142] h position =MLP(Fourier(b))
[0143] Fourier(b)=(sin(2 0 πb),cos(2 0 πb),...,sin(2 L-1 πb),cos(2 L-1 πb)) (6)
[0144] Among them, b is the coordinates of the vertices of the dynamic target 3D box, L is the dimension, and h is position The position encoding result.
[0145] The category encoding result and the position encoding result are fused to obtain the dynamic target encoding vector (i.e., the dynamic target encoding result). This process is expressed by formula (7):
[0146] h box =[h class ,h position ] (7)
[0147] Among them, h box Encode the vector for the dynamic target.
[0148] Step S5: Input the video data encoding vector into the diffusion model, such as Fig.11As shown in the figure, in the forward process, real noise is continuously added to the video data coding vector to make it become a standard Gaussian distribution. After t time steps (a total of T time steps, T is an integer greater than 0), the video data coding vector z after adding noise can be obtained. t (i.e., the video data encoding vector after adding noise corresponding to one time step). The video data encoding vector after adding noise is the noise addition result. T time steps correspond to T groups of convolutional layers, spatial attention, cross attention, and temporal attention. Fig.10 z in T Represents the T noise-added video data encoding vectors z corresponding to T time steps t .
[0149] Among them, Fig.11 As shown, in the forward process, the first noise is added to an image x0 to obtain q(x1 / x0); , , , ; and q(x t / x t-1 ) adds the tth noise to get q(x t+1 / x t ),,,, after adding the Tth noise, we get q(x T / x T-1 ). Accordingly, in the reverse process, the denoising process is performed and finally q(x0 / x1) is obtained.
[0150] Step S6: Use the convolution layer to convert the dimension of the video data encoding vector after adding noise, that is, extract and restore the features. This process is expressed by formula (7):
[0151] z v =Conv(z t ) (7)
[0152] Among them, z t is the encoding vector of the video data after t time steps, Conv(·) is the convolutional neural network, z v is the encoding vector of the video data after dimension conversion (i.e., the convolution result).
[0153] Step S7: Use the spatial attention mechanism to fuse the scene encoding vector into the video data encoding after dimension conversion (i.e., the convolution result), so that the generated video data (i.e., the third fusion result) is consistent with the scene. The third fusion result is obtained by multiplying the video data encoding after dimension conversion by formula (8), which is expressed as:
[0154]
[0155] in, is the parameter matrix, d is the output feature dimension, z vis the video coding vector after dimension change, h scence Encodes the scene information vector.
[0156] Step S8: Use the spatial attention mechanism to fuse the dynamic target encoding vector into the video data encoding vector after the fusion scene information (i.e., the third fusion result), and keep the generated video data (i.e., the first fusion result) and the dynamic target coordinated and consistent. The first fusion result is obtained by multiplying the video data encoding after the fusion scene information by formula (9), which is expressed as:
[0157]
[0158] in, is the parameter matrix, d is the input feature dimension, z s are the video data encoding vectors after integrating scene information, h box Encode the vector for the dynamic target.
[0159] Step S9: Use the cross attention mechanism to model the relationship between each view and the other five views of the same frame to maintain consistency between different views. The specific method can be expressed by formula (10):
[0160]
[0161] in, is the parameter matrix, d is the input feature dimension, is the embedding representation of the current view (i.e., the first fusion result of the current view), are the other five views corresponding thereto (ie, the first fusion result of the other five views).
[0162] Step S10: Calculate the similarity between the current frame and the embedded representations of all previous frames. This can be expressed by formula (11):
[0163]
[0164] in, is the embedded representation of the current frame (i.e., the second fusion result), is the number of frames before the current frame, m<i.
[0165] The temporal attention mechanism is used to calculate the relationship between the current frame and the frame most similar to the current frame and the previous frame of the current frame, which can be expressed by formula (12):
[0166]
[0167] in, is the parameter matrix, is the embedding representation for the current time step, is the embedding representation of the frame that is most similar to the embedding representation of the current frame, is the embedded representation of the previous frame of the current frame.
[0168] Repeat steps S6-S10 to finally obtain T processing results for each sample image.
[0169] Step S11: performing noise prediction processing on the T processing results of each sample image to obtain T predicted noises of each sample image.
[0170] Step S12: Based on the T real noises and T predicted noises of each sample image in the multi-frame multi-view sample images, the following loss function is used to train and optimize the model parameters of the diffusion model. The loss function can be expressed by formula (13):
[0171]
[0172] Among them, S = {M, B, L}, M is the map data, B is the dynamic target 3D box information, L is the weather and time text description information. ∈ is the real noise, The content in brackets indicates the processing result for predicting noise. The real noise is the noise added in the forward process of the diffusion model.
[0173] Calculate the minimum value of the loss function to obtain a set of optimal model parameters.
[0174] The parameters of the diffusion model are modified based on the optimal model parameters to obtain a trained image generation model.
[0175] Furthermore, we obtain multi-frame multi-view real images and real image annotation data, and combine Fig.10 , perform encoding processing, and then input it into the trained image generation model. Next, decode the output result of the image generation model to obtain a multi-frame multi-view simulation image. Fig.12 The figure shows a frame of six-view simulation image of an example.
[0176] An embodiment of the present invention provides an image generating device. Fig.13 FIG. 1 is a schematic diagram of the structure of the image generating device in an embodiment of the present invention. Fig.13 As shown, the image generating device 130 includes:
[0177] An acquisition unit 1301 is used to acquire multiple frames of multi-view real images and annotation data of the real images;
[0178] The processing unit 1302 is configured to fuse each real image and the annotation data of each real image in the multiple frames of multi-view real images based on the added noise of T time steps included in the image generation model to obtain T first fusion results of each real image; T is an integer greater than 0;
[0179] The processing unit 1302 is further configured to fuse the first fusion result of the real image of the other perspectives of the same frame except for any perspective into the first fusion result of the real image of any perspective based on the image generation model to obtain a second fusion result of the real image of any perspective;
[0180] The processing unit 1302 is further configured to fuse the second fusion results of the real images of the remaining frames except any frame at the same viewing angle into the second fusion result of any frame of the real image based on the image generation model to obtain a processing result of any frame of the real image;
[0181] The processing unit 1302 is further used to obtain T simulation images of each real image based on the T processing results of each real image in the multi-frame multi-view real image until the T first fusion results of each real image are processed.
[0182] In some embodiments of the present invention, the processing unit 1302 is also used to encode each real image to obtain a first encoding result of each real image; encode the annotation data of each real image to obtain a second encoding result of each real image; perform noise processing on the first encoding result of each real image based on the added noise of the T time steps to obtain T noisy results of each real image; and fuse the second encoding result of each real image into the T noisy results of each real image to obtain T first fusion results of each real image.
[0183] In some embodiments of the present invention, the annotation data includes relevant data and scene data of dynamic targets; the processing unit 1302 is also used to encode the relevant data of the dynamic targets included in each real image to obtain a dynamic target encoding result; encode the scene data of each real image to obtain a scene encoding result; the second encoding result includes the dynamic target encoding result and the scene encoding result.
[0184] In some embodiments of the present invention, the processing unit 1302 is also used to process the T noise addition results of each real image using a convolution layer to obtain T convolution results of each real image; and fuse the dynamic target encoding results and scene encoding results of each real image into the T convolution results of each real image to obtain T first fusion results of each real image.
[0185] In some embodiments of the present invention, the processing unit 1302 is also used to fuse the scene encoding result of each real image into the T convolution results of each real image to obtain T third fusion results of each real image; and fuse the dynamic target encoding result of each real image into the T third fusion results of each real image to obtain T first fusion results of each real image.
[0186] In some embodiments of the present invention, the relevant data of the dynamic target includes the category and three-dimensional coordinates; the processing unit 1302 is also used to encode the category of the dynamic target included in each real image to obtain a category coding result; encode the three-dimensional coordinates of the dynamic target included in each real image to obtain a position coding result; and fuse the category coding result and the position coding result to obtain the dynamic target coding result.
[0187] In some embodiments of the present invention, the processing unit 1302 is further used to encode the map data of each real image to obtain a map encoding result; encode the context data of each real image to obtain a context encoding result; and splice the map encoding result and the context encoding result to obtain the scene encoding result.
[0188] In some embodiments of the present invention, the processing unit 1302 is also used to determine the similarity between the second fusion result of the real image of any frame in multiple frames of the same perspective and the second fusion results of the real images of the remaining frames before any frame; the second fusion result of the real image of the target frame corresponding to the maximum similarity and the second fusion result of the real image of the previous frame of any frame are fused into the second fusion result of the real image of any frame to obtain the processing result of any frame of the real image.
[0189] The embodiment of the present invention further provides another electronic device, Fig.14 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Fig.14 As shown, the electronic device 140 includes: a processor 1401 and a memory 1402 configured to store a computer program that can be run on the processor;
[0190] The processor 1401 is configured to execute the method steps in the aforementioned embodiment when running a computer program.
[0191] Of course, in practical applications, Fig.14 As shown, the components in the electronic device 140 are coupled together via a bus system 1403. It is understood that the bus system 1403 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1403 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.14 Various buses are labeled as bus system 1403.
[0192] In practical applications, the processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a controller, a microcontroller, and a microprocessor. It is understandable that for different devices, the electronic device used to implement the functions of the processor may also be other, and the embodiments of the present invention do not specifically limit this.
[0193] The above-mentioned memory can be a volatile memory (volatile memory), such as a random access memory (RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above-mentioned types of memory, and provide instructions and data to the processor.
[0194] In an exemplary embodiment, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program.
[0195] Optionally, the computer-readable storage medium can be applied to any one of the methods in the embodiments of the present invention, and the computer program enables the computer to execute the corresponding processes implemented by the processor in each method in the embodiments of the present invention. For the sake of brevity, they are not described here.
[0196] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0197] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0198] In addition, all functional units in the embodiments of the present invention may be integrated into one processing module, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art may understand that all or part of the steps of implementing the above method embodiments may be completed by hardware related to program instructions, and the above program may be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the above storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0199] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0200] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0201] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0202] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An image generation method, characterized in that: The image generation method comprises: Obtain multi-frame multi-view real images and annotation data of real images; Based on the added noise of T time steps included in the image generation model, each real image in the multi-frame multi-view real image and the annotated data of each real image are fused to obtain T first fusion results of each real image; T is an integer greater than 0; Based on the image generation model, the first fusion result of the real image of the same frame from other perspectives except any one perspective is fused into the first fusion result of the real image of any perspective to obtain a second fusion result of the real image of any perspective; Based on the image generation model, the second fusion results of the remaining frames of real images of the same perspective except any frame are fused into the second fusion result of any frame of real image to obtain a processing result of any frame of real image; Until the T first fusion results of each real image are processed, T simulation images of each real image are obtained based on the T processing results of each real image in the multi-frame multi-view real images.
2. The image generation method according to claim 1, characterized in that: The adding of noise in T time steps based on the image generation model, fusing each real image in the multi-frame multi-view real image and the annotated data of each real image, to obtain T first fusion results of each real image, includes: Encoding each real image to obtain a first encoding result of each real image; Encoding the labeled data of each real image to obtain a second encoding result of each real image; Performing noise addition processing on the first encoding result of each real image based on the added noise of the T time steps to obtain T noise addition results of each real image; The second encoding result of each real image is fused into the T noise-added results of each real image to obtain T first fusion results of each real image.
3. The image generation method according to claim 2, characterized in that: The annotated data includes relevant data and scene data of the dynamic target; the encoding of the annotated data of each real image to obtain a second encoding result of each real image includes: Encoding the relevant data of the dynamic target included in each real image to obtain the dynamic target encoding result; The scene data of each real image is encoded to obtain a scene encoding result; the second encoding result includes the dynamic object encoding result and the scene encoding result.
4. The image generation method according to claim 3, characterized in that: The step of fusing the second encoding result of each real image into the T noise-added results of each real image to obtain T first fusion results of each real image includes: The T noise-added results of each real image are processed by a convolution layer to obtain T convolution results of each real image; The dynamic object coding result and the scene coding result of each real image are fused into the T convolution results of each real image to obtain T first fusion results of each real image.
5. The image generation method according to claim 4, characterized in that: The step of fusing the dynamic target encoding result and the scene encoding result of each real image into the T convolution results of each real image to obtain T first fusion results of each real image includes: The scene encoding result of each real image is fused into the T convolution results of each real image to obtain T third fusion results of each real image; The dynamic target encoding result of each real image is fused into the T third fusion results of each real image to obtain T first fusion results of each real image.
6. The image generation method according to claim 3, characterized in that: The relevant data of the dynamic target includes the category and the three-dimensional coordinates; the relevant data of the dynamic target included in each real image is encoded to obtain the dynamic target encoding result, including: Encoding the category of the dynamic target included in each real image to obtain a category encoding result; Encoding the three-dimensional coordinates of the dynamic targets included in each real image to obtain a position encoding result; The category coding result and the position coding result are fused to obtain the dynamic target coding result.
7. The image generation method according to claim 3, characterized in that: The scene data includes map data and context data; the scene data of each real image is encoded to obtain a scene encoding result, including: Encoding the map data of each real image to obtain a map encoding result; Encoding the context data of each real image to obtain a context encoding result; The map coding result and the context coding result are concatenated to obtain the scene coding result.
8. The image generation method according to any one of claims 1 to 7, characterized in that: The second fusion result of the remaining frames of real images of the same perspective except any frame is fused into the second fusion result of any frame of real image based on the image generation model to obtain the processing result of any frame of real image, including: Determine the similarity between the second fusion result of any real image frame in multiple frames of the same viewing angle and the second fusion results of the real images of the remaining frames before any frame; The second fusion result of the target frame real image corresponding to the maximum similarity and the second fusion result of the previous frame real image of any frame are fused into the second fusion result of any frame real image to obtain the processing result of any frame real image.
9. The image generation method according to any one of claims 1 to 7, characterized in that: The image generation method further comprises: Obtaining multi-frame multi-view sample images and annotation data of each sample image; Based on the added noise of T time steps included in the diffusion model, each sample image and the annotation data of each sample image in the multi-frame multi-view sample image are fused to obtain T first fusion results of each sample image; Based on the diffusion model, the first fusion results of the sample images of the other viewing angles of the same frame except for any viewing angle are fused into the first fusion result of the sample image of any viewing angle to obtain a second fusion result of the sample image of any viewing angle; Based on the diffusion model, the second fusion results of the sample images of the remaining frames except any frame at the same viewing angle are fused into the second fusion result of any frame of the sample image to obtain a processing result of any frame of the sample image; After the T first fusion results of each sample image are processed, the T processing results of each sample image are subjected to noise prediction processing to obtain T predicted noises of each sample image; Based on the T added noises and T predicted noises of each sample image in the multi-frame multi-view sample images, the diffusion model is trained to obtain the trained image generation model.
10. An image generating device, characterized in that: The image generating device comprises: An acquisition unit, used for acquiring multi-frame multi-view real images and annotation data of the real images; A processing unit, configured to fuse each real image and the annotated data of each real image in the multiple frames of multi-view real images based on the added noise of T time steps included in the image generation model, to obtain T first fusion results of each real image; T is an integer greater than 0; The processing unit is further used to fuse the first fusion result of the real image of the other perspectives of the same frame except any perspective into the first fusion result of the real image of any perspective based on the image generation model to obtain a second fusion result of the real image of any perspective; The processing unit is further configured to fuse the second fusion results of the remaining frames of real images of the same perspective except any frame into the second fusion result of any frame of real image based on the image generation model to obtain a processing result of any frame of real image; The processing unit is also used to obtain T simulation images of each real image based on the T processing results of each real image in the multi-frame multi-view real image until the T first fusion results of each real image are processed.
11. An electronic device, characterized in that: The electronic device comprises: a processor and a memory configured to store a computer program capable of running on the processor, Wherein, the processor is configured to execute the steps of the image generation method according to any one of claims 1 to 9 when running the computer program.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image generating method according to any one of claims 1 to 9 are implemented.
13. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the image generation method according to any one of claims 1 to 9 are implemented.