Image generation method and related device

By extracting the target object from the original image and synthesizing it with the scene environment map generated by artificial intelligence, the problems of high manual operation cost and low accuracy in the existing technology are solved, and an efficient and low-cost image generation method is realized, and the generated image is more realistic.

CN120599088AInactive Publication Date: 2025-09-05PORSCHE (SHANGHAI) DIGITAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410245307.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies require a lot of manual work in image generation, resulting in high costs and low efficiency. In addition, the images generated by artificial intelligence are not accurate enough, especially in terms of the size and appearance design details of the target object, which are difficult to accurately define.

Method used

By extracting the image of the target object from the original image and combining it with the scene environment map generated by the artificial intelligence model, it is synthesized using image segmentation and recognition technology to ensure the accuracy and lighting effects of the target object in the synthesized image.

Benefits of technology

The accurate restoration of the target object in the synthetic image is achieved, the efficiency of image generation is improved and the cost is reduced, and the generated image is more realistic and meets expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599088A_ABST
    Figure CN120599088A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and a related image generation device. The method comprises the following steps: extracting an image of a first target object from a first image; on the basis of an artificial intelligence model, according to the information describing the environment scene input by the user, a second image reflecting the information describing the environment scene is generated, the information describing the environment scene at least comprises information associated with the target object, and the second image at least comprises a second target object and a scene environment graph; and synthesizing the extracted image of the first target object and the scene environment map in the second image to obtain a synthesized image comprising the extracted first target object. According to the invention, the accuracy of the related target object in the obtained composite image is ensured, so that a more vivid or accurate composite image about the first target object is obtained. In addition, by adopting the scheme of the invention, the efficiency of obtaining the expected image about the target object can be improved, and the cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation, and more particularly, to an image generation method, a related apparatus, a computer device, and a computer-readable storage medium. Background Art

[0002] In applications related to image generation, the existing technology usually uses software such as Photoshop to edit and produce related images, but this requires designers to perform manual operations, consumes a lot of manual labor, and results in high costs and low efficiency.

[0003] In addition, with the further development of artificial intelligence (AI) technology, it can be used to generate images for reasons such as reducing costs and improving efficiency. For example, deep learning technology can be used to train an AI model by inputting a large number of relevant images, and based on the trained AI model, better images can be generated. However, due to the limitations of AI technology, the accuracy of the generated images cannot be guaranteed. For example, it is impossible to accurately define the size and appearance design details related to the target object in the generated image.

[0004] Therefore, there is a need for an improved image generation solution. Summary of the Invention

[0005] In order to solve at least one of the above technical problems, the present invention provides an image generation method, a related image generation device, a computer device and a computer-readable storage medium.

[0006] According to a first aspect of the present invention, there is provided an image generation method, the method comprising: extracting an image of a first target object from a first image; generating a second image embodying the information describing the environmental scene based on information describing the environmental scene input by a user based on an artificial intelligence model, wherein the information describing the environmental scene includes at least information associated with the target object, and wherein the second image includes at least the second target object and a scene environment map; and synthesizing the extracted image of the first target object and the scene environment map in the second image to obtain a synthesized image including the extracted first target object.

[0007] In one embodiment, the method further comprises: segmenting the second image into an image of the second target object and the scene environment map by using an image segmentation technique, wherein the scene environment map includes a shadow of the second target object.

[0008] In one embodiment, the artificial intelligence model includes an artificial intelligence model that supports image generation from Sora, ChatGPT, midjourney, stablediffusion, or Tongyi Qianwen.

[0009] In one embodiment, the method further includes: determining the position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image through image recognition technology, and based on the position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image, using the extracted image of the first target object to replace the image of the second target object to achieve the synthesis of the image of the first target object and the scene environment map in the second image.

[0010] In one embodiment, the first target object is a first vehicle and the second target object is a second vehicle.

[0011] According to a second aspect of the present invention, an image generating device is provided, comprising: an image extraction module, configured to extract an image of a first target object from a first image having the first target object; an image generating module based on an artificial intelligence model, configured to generate a second image embodying information describing an environmental scene according to information describing an environmental scene input by a user based on an artificial intelligence model, wherein the information describing the environmental scene includes at least information associated with the target object, and wherein the second image includes at least the second target object and a scene environment map; and an image synthesis module, configured to synthesize the extracted image of the first target object and the scene environment map in the second image to obtain a synthesized image including the extracted first target object.

[0012] In one embodiment, the image generating device further includes an image segmentation module, which is configured to segment the second image into an image of the second target object and the scene environment map through image segmentation technology, wherein the scene environment map includes a shadow of the second target object.

[0013] In one embodiment, the image synthesis module is further configured to determine the position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image through image recognition technology, and based on the position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image, use the extracted image of the first target object to replace the image of the second target object to achieve the synthesis of the image of the first target object and the scene environment map in the second image.

[0014] According to a third aspect of the present invention, a computer device is provided, comprising a memory and a processor, wherein the memory stores computer instructions, which, when executed by the processor, cause the image generation method described above to be performed.

[0015] According to a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the image generation method described above to be performed.

[0016] The solution of the present invention uses an original first image with a desired first target object and a second image with a second target object generated by artificial intelligence. The image of the first target object extracted from the first image is then synthesized with a scene environment map in the second image generated based on an artificial intelligence model, thereby obtaining a composite image including the extracted first target object. Advantageously, the accuracy of the first target object in the obtained composite image and its integration into the scene are guaranteed. In other words, the relevant target object in the composite image can be accurately restored to obtain a more realistic or accurate composite image. In addition, by adopting the solution of the present invention, the efficiency of obtaining the desired image of the target object can be improved and the cost can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Non-limiting and non-exhaustive embodiments of the present invention are described with reference to the following drawings, in which:

[0018] Figure 1 A schematic flow chart illustrating an image generation method according to one embodiment of the present invention is provided;

[0019] Figure 2 A schematic block diagram illustrating an image generating apparatus according to an embodiment of the present invention; and

[0020] Figure 3A-Figure 3E A schematic diagram illustrating an image generating method according to an embodiment of the present invention is shown.

[0021] Those skilled in the art will appreciate that the elements in the drawings are illustrated for simplicity and clarity and are not necessarily drawn to scale. Additionally, common but well-understood elements that are useful or necessary in a commercially feasible embodiment are often not depicted in order to less obstruct viewing of the various embodiments of the present invention. DETAILED DESCRIPTION

[0022] In the following description, many specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that it is not necessary to adopt specific details to practice the present invention. In other cases, well-known materials or methods are not described in detail to avoid obscuring the present invention.

[0023] Reference throughout this specification to "one embodiment," "another embodiment," "an example," or "another example" means that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention, and they are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics may be combined in any suitable combinations and / or subcombinations in one or more embodiments or examples.

[0024] Figure 1 The schematic flowchart of the image generation method 100 according to one embodiment of the present invention is illustrated. The method 100 may include steps S102, S104, and S106.

[0025] In step S102, an image of the first target object is extracted from the first image. The first image may refer to an original image or a real image of the first target object (for example, a high-definition image of a person or product used for promotion or display). The first target object generally refers to any object contained in the first image for which subsequent image synthesis is intended. For example, the first target object includes but is not limited to products (such as vehicles, smart terminal devices, or any other items), people, animals, or other natural objects (such as the sun, flowers, trees, and other natural scenes).

[0026] In step S104, based on the artificial intelligence model and according to the information describing the environmental scene input by the user, a second image embodying the information describing the environmental scene is generated, wherein the information describing the environmental scene includes at least information associated with the target object, and wherein the second image includes at least the second target object and the scene environment map.

[0027] In one embodiment, the artificial intelligence model includes but is not limited to Sora, ChatGPT, midjourney, stablediffusion or Tongyi Qianwen's artificial intelligence model that supports image generation. With the help of these artificial intelligence models, artificial intelligence pictures with desired scene environment pictures can be generated at low cost and high efficiency. In order to make the scene environment picture in the generated artificial intelligence picture more in line with expectations and to make the second target object therein closer to the above-mentioned first target object, the user can select appropriate information describing the scene when using the artificial intelligence model and use the information of the first target object to limit the target object in the artificial intelligence picture, such as keywords that can include more accurate and key scene description information and related information such as product model, character feature description or animal species description of the first target object.

[0028] In step S106 , the extracted image of the first target object and the scene environment map in the second image are synthesized to obtain a synthesized image including the extracted first target object.

[0029] Advantageously, the accuracy of the first target object in the obtained composite image is guaranteed. In addition, by adopting the solution of the present invention, the efficiency of obtaining the composite image can be improved and the cost can be reduced.

[0030] In one embodiment, the method 100 may further include step S105 (not shown), which may be included after step S104 and before step S106. In step S105, the second image is segmented into an image of the second target object and the scene environment map using image segmentation technology, wherein the scene environment map includes a shadow of the second target object. The shadow generally refers to the shadow cast by the relevant target object (e.g., the second target object) under lighting conditions (such as sunlight, outdoor or indoor lighting, etc.).

[0031] In one embodiment, step S106 of method 100 may further include step S1062. In step S1062, the position, size, and / or positional relationship of the second target object relative to the shadow of the second target object in the second image are determined using image recognition technology, and based on the position, size, and / or positional relationship of the second target object relative to the shadow of the second target object in the second image, the image of the second target object is replaced with the extracted image of the first target object to achieve synthesis of the image of the first target object and the scene environment map in the second image.

[0032] For example, based on the position and size of the second target object in the second image and / or the positional relationship of the second target object's shadow relative to the second target object, the extracted image of the first target object is directly synthesized with the scene environment map segmented and extracted from the second image (which includes the shadow corresponding to the second target object). In another example, based on the position and size of the second target object in the second image and / or the positional relationship of the second target object's shadow relative to the second target object, the extracted image of the first target object is overlaid on the image of the second target object to achieve synthesis of the image of the first target object and the scene environment map in the second image including the corresponding shadow.

[0033] Advantageously, the present invention can synthesize the extracted accurate original image of the first target object into the position of the second target object in the image generated by artificial intelligence, so as to generate an image with accurate light and shadow related to the target object (such as the first target object) (such as the vehicle promotional image mentioned below). In other words, it can be guaranteed that the relevant target objects in the synthetic image are 100% restored to obtain a more realistic or accurate synthetic image. For example, this can ensure that in the final synthetic image obtained after image synthesis, the relevant target objects can obtain the correct shadows under the lighting conditions in the image, so that in addition to the appearance of the first target object in the synthetic image remaining authentic and consistent with the expected promotion or display, the correct light and shadow imaging requirements of the relevant target object can also be met, thereby improving the viewing experience and aesthetics.

[0034] In one embodiment, the first target object may be a first vehicle, and the second target object may be a second vehicle. Figure 3A-Figure 3E Described in detail.

[0035] It should be understood that although Figure 1 The steps in the flowchart are shown in the order indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0036] Figure 2The schematic block diagram of an image generation device 20 according to an embodiment of the present invention is illustrated. The image generation device 20 may include an image extraction module 202, an image generation module 204 based on an artificial intelligence model, and an image synthesis module 206. The image extraction module 202 is configured to extract an image of the first target object from a first image having the first target object. The image generation module 204 is configured to generate a second image reflecting the information describing the environmental scene based on the information describing the environmental scene input by the user based on the artificial intelligence model, wherein the information describing the environmental scene includes at least information associated with the target object, and wherein the second image includes at least the second target object and a scene environment map. The image synthesis module 206 is configured to synthesize the extracted image of the first target object and the scene environment map in the second image to obtain a synthesized image including the extracted first target object.

[0037] In one embodiment, the image generating device 20 further includes an image segmentation module 205 (not shown), which is configured to segment the second image into an image of the second target object and the scene environment map through image segmentation technology, wherein the scene environment map includes the shadow of the second target object.

[0038] In one embodiment, the image synthesis module 20 is further configured to determine the position, size and / or the positional relationship of the second target object relative to the shadow of the second target object in the second image through image recognition technology, and based on the position, size and / or the positional relationship of the second target object relative to the shadow of the second target object in the second image, use the extracted image of the first target object to replace the image of the second target object to achieve the synthesis of the image of the first target object and the scene environment map in the second image.

[0039] The image generating device herein can execute the image generating method discussed in any of the above implementation schemes or examples, and the functions or related features of the image generating device and its various modules can be expanded similarly to the above image generating method, which will not be described in detail here.

[0040] Figure 3A-Figure 3E A schematic diagram illustrates an image generation method according to one embodiment of the present invention. As described above, in this embodiment, the target objects may be vehicles, with the first target object being a first vehicle and the second target object being a second vehicle. For example, the generated composite image includes an accurate image of the vehicle being promoted. The vehicle exhibits accurate lighting and shading within the composite image's scene environment, blending naturally with the scene environment.

[0041] Figure 3A The example illustrates extracting an image of a first vehicle, where the extracted image of the first vehicle can be an image of the vehicle's exterior. Specifically, by inputting an existing image with an accurate vehicle exterior, image recognition technology can be used to identify the vehicle exterior image itself, excluding environmental information within the image. Image segmentation technology can then be used to extract the vehicle exterior image, ensuring that the extracted vehicle exterior image accurately reproduces the first vehicle's model and details. The extracted vehicle exterior image is then stored in a memory.

[0042] Figure 3B A second image generated based on AI is illustrated, which includes at least a vehicle (e.g., the second vehicle herein), a shadow created by the vehicle, and other relevant environmental backgrounds; Figure 3C The example illustrates segmenting and extracting the second vehicle and the scene environment map (including the vehicle shadow part and other related environmental background) in the second image by image segmentation technology; Figure 3D The example illustrates identifying the location and size of a vehicle in a second image generated by AI using image recognition technology; and Figure 3E The example illustrates synthesizing an accurate extracted original image of a vehicle (e.g., the first vehicle in this article) into the location of a vehicle (e.g., the second vehicle in this article) in an AI-generated scene environment map, thereby generating a composite image containing the original image of the first vehicle with accurate lighting and shadows.

[0043] It is conceivable that those skilled in the art may obtain a composite image having an accurate image of the target object (such as the first vehicle) and its light and shadow through other techniques related to image segmentation, extraction, recognition and synthesis based on the implementation scheme or examples described in the context.

[0044] refer to Figure 3A-Figure 3E As described above, the generated synthetic image can be used to create promotional posters for vehicles. For example, when promoting new models, organizing test drives, and organizing events like owner clubs, the marketing and sales departments of a car company need to create and release promotional posters for new cars. Advantageously, using the above solution, a synthetic image can be generated that accurately depicts the appearance of a vehicle (such as a new car) (such as its dimensions and exterior design details) and a rich, anticipated scene environment. In other words, the appearance of the vehicle in the synthetic image can be consistent with the appearance of the actual vehicle.

[0045] In one embodiment, when an artificial intelligence model is used to generate an AI image, the descriptive information associated with the vehicle may include, for example, the make and model of the vehicle, or other relevant information that can characterize the appearance of the vehicle. Based on the above input by the user, the appearance of the second vehicle in the second image is substantially the same as the appearance of the first vehicle. For another example, the information input by the user to describe the environmental scene may be, for example, the following sentence: "Behind the ancient temple, petals falling under the cherry blossom trees are like snowflakes, and the entire street is covered with the pink of cherry blossoms, quiet and beautiful", to form the scene environment map in the second image.

[0046] Advantageously, the present invention can ensure that not only the appearance of the second vehicle in the composite image is consistent with the appearance of the corresponding real first vehicle, but also the light and shadow imaging requirements of the first vehicle in a simulated real environment can be met, thereby improving the viewing experience and aesthetics, providing customers or users with good visual experience, and better used for vehicle promotion.

[0047] Another aspect of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores computer instructions which, when executed by the processor, cause the image generation method described above to be executed. The computer device can be broadly defined as a server, an in-vehicle terminal, or any other electronic device having the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, a memory, a network interface, a communication interface, etc. connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and an internal memory. An operating system, a computer program, etc. may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the computer device can be used to connect to and communicate with external devices via a network. When the computer program is executed by the processor, it can implement the steps of any of the above-mentioned image generation methods of the present invention.

[0048] Another aspect of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the image generation method described in any of the above implementation schemes or examples is implemented.

[0049] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned image generation method can be performed by instructing related hardware such as a computer device or a processor through a computer program, and the computer program can be stored in a non-transitory computer-readable storage medium. When the computer program is executed, the steps of the image generation method of the present invention are performed. Depending on the circumstances, any reference to memory, storage, database or other media in this document may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0050] The above description of the illustrated embodiments or examples of the present invention, including what is described in the Abstract, is not intended to be exhaustive or to limit the precise forms disclosed. Although specific embodiments and examples of the invention are described herein for illustrative purposes, various equivalent modifications are possible without departing from the broader spirit and scope of the invention.

Claims

1. An image generation method, characterized in that: The method comprises: extracting an image of a first target object from the first image; generating, based on the artificial intelligence model and according to the information describing the environmental scene input by the user, a second image embodying the information describing the environmental scene, wherein the information describing the environmental scene includes at least information associated with the target object, and wherein the second image includes at least the second target object and a scene environment map; and The extracted image of the first target object and the scene environment map in the second image are synthesized to obtain a synthesized image including the extracted first target object.

2. The image generation method according to claim 1, wherein: The method further comprises: The second image is segmented into an image of the second target object and the scene environment map by using an image segmentation technology, wherein the scene environment map includes a shadow of the second target object.

3. The image generation method according to claim 2, wherein: The artificial intelligence models include Sora, ChatGPT, midjourney, stablediffusion or Tongyi Qianwen's artificial intelligence models that support image generation.

4. The image generation method according to claim 2, wherein: The method further comprises: The position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image are determined by image recognition technology, and based on the position, size and / or positional relationship of the second target object relative to the shadow of the second target object in the second image, the image of the second target object is replaced with the extracted image of the first target object to achieve synthesis of the image of the first target object and the scene environment map in the second image.

5. The image generation method according to any one of claims 1 to 4, characterized in that: The first target object is a first vehicle, and the second target object is a second vehicle.

6. An image generating device, characterized in that: The device comprises: an image extraction module, configured to extract an image of the first target object from a first image having the first target object; an image generation module based on an artificial intelligence model, the image generation module being configured to generate, based on the artificial intelligence model and according to information describing an environmental scene input by a user, a second image embodying the information describing the environmental scene, wherein the information describing the environmental scene at least includes information associated with a target object, and wherein the second image at least includes a second target object and a scene environment map; and An image synthesis module is configured to synthesize the extracted image of the first target object and the scene environment map in the second image to obtain a synthesized image including the extracted first target object.

7. The image generating device according to claim 6, wherein: The image generating device further includes an image segmentation module configured to segment the second image into an image of the second target object and the scene environment map using an image segmentation technique, wherein the scene environment map includes a shadow of the second target object.

8. The image generating device according to claim 7, wherein: The artificial intelligence models include Sora, ChatGPT, midjourney, stablediffusion or Tongyi Qianwen's artificial intelligence models that support image generation.

9. The image generating device according to claim 7, wherein: The image synthesis module is further configured to determine the position, size and / or the positional relationship of the second target object relative to the shadow of the second target object in the second image through image recognition technology, and based on the position, size and / or the positional relationship of the second target object relative to the shadow of the second target object in the second image, use the extracted image of the first target object to replace the image of the second target object to achieve the synthesis of the image of the first target object and the scene environment map in the second image.

10. The image generating device according to any one of claims 6 to 9, characterized in that: The first target object is a first vehicle, and the second target object is a second vehicle.

11. A computer device comprising a memory and a processor, wherein the memory stores computer instructions, wherein: The computer instructions, when executed by the processor, result in the image generation method according to any one of claims 1 to 5 being performed.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it causes the image generation method according to any one of claims 1 to 5 to be performed.

Citation Information

Patent Citations

  • Method for enhancing environment scene and automatic driving vehicle testing system

    CN116805294A

  • Vehicle preview generation method and device, equipment and medium

    CN117078835A

  • Image generation method and device and storage medium

    CN117475031A