Methods, apparatus, devices, and programs for generating images

By determining multiple layers from vehicle environments and incorporating user input, the method generates personalized images that align with the user's needs and vehicle state, improving the driving experience.

JP2026090210APending Publication Date: 2026-06-02MOBILITY ASIA SMART TECH CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MOBILITY ASIA SMART TECH CO LTD
Filing Date
2025-11-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current AI-powered image generation methods in vehicles fail to produce personalized images that connect with the user's current mood and vehicle usage scenarios, often resulting in images unrelated to the user's needs.

Method used

A method and apparatus for generating images by determining multiple layers based on the vehicle's internal and external environments, incorporating spatial relationships, and using an image generation model that considers the vehicle's state and user input information to create personalized images.

Benefits of technology

The method generates images that better suit the user's preferences, enhancing the driving and riding experience by combining vehicle data, environmental information, and user needs, resulting in personalized and precise image adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026090210000001_ABST
    Figure 2026090210000001_ABST
Patent Text Reader

Abstract

The present invention provides methods, apparatus, devices, and products for generating images. [Solution] The method includes determining a number of corresponding layers based on images of the vehicle's interior and exterior environment. The spatial relationships of the layers are different. The method further includes generating corresponding images based on the layers using an image generation model, based on the vehicle's state and / or user input information. This method allows for the generation of images that match user intent and the vehicle environment by combining user input, vehicle state, and the vehicle's surrounding environment, thereby enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of vehicle technologies, and more specifically, to a method, an apparatus, a device, and a product for generating an image.

Background Art

[0002] With the rapid development of vehicle technologies and artificial intelligence, automobiles have already become an indispensable part of people's daily lives. During the driving process, users often encounter many beautiful scenes and often want to record the current mood and feelings of the journey. At the same time, technologies such as text-to-text, text-to-image, and image-to-image have begun to be widely applied in vehicles. Users' expectations for the entertainment functions of in-vehicle systems are increasing, and for example, users are further pursuing smart and personalized smart image generation services.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Embodiments of the present disclosure propose a method, an apparatus, a device, and a product for generating an image.

Means for Solving the Problems

[0004] In a first aspect of the present disclosure, a method for generating an image is provided. The method includes determining a corresponding plurality of layers based on an image of the internal and external environments of a vehicle, where the spatial relationships of the plurality of layers are different. The method further includes generating a corresponding image based on the plurality of layers by an image generation model based on the state of the vehicle and / or the input information of the user.

[0005] A second aspect of the present disclosure provides an apparatus for generating images. The apparatus includes a layer determination module configured to determine a plurality of corresponding layers based on images of the vehicle's interior and exterior environment, wherein the spatial relationships of the plurality of layers are different. The apparatus further includes an image generation module configured to generate a corresponding image based on the plurality of layers by an image generation model based on the vehicle's state and / or user input information.

[0006] A third aspect of this disclosure provides an electronic device, which includes one or more processors and a memory device for storing one or more programs, which, when executed by one or more processors, cause one or more processors to implement the method provided in the first aspect of this disclosure.

[0007] A fourth aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions, which are executed by a processor to implement the method provided in the first aspect of this disclosure.

[0008] A fifth aspect of this disclosure provides a computer program product which includes machine-executable instructions that are physically stored on a computer-readable medium and, when the machine-executable instructions are executed, implements the method of the first aspect of this disclosure.

[0009] It should be understood that the contents described in the summary section of the invention are not intended to limit the essential or important features of the embodiments of this disclosure, nor to limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description. [Brief explanation of the drawing]

[0010] The above and other features, advantages, and aspects of each embodiment of this disclosure will become clearer when viewed in conjunction with the following detailed description. In the drawings, the same or similar reference numerals represent the same or similar elements. [Figure 1] A schematic diagram of an exemplary environment capable of realizing multiple embodiments of this disclosure is shown. [Figure 2] A schematic diagram of the flow for generating an image according to some embodiments of this disclosure is shown. [Figure 3] A flowchart of a method for generating an image according to some embodiments of this disclosure is shown. [Figure 4] A schematic diagram showing the determination of multiple layers according to some embodiments of this disclosure is shown. [Figure 5] A schematic diagram illustrating the adjustment of multiple layers according to some embodiments of this disclosure is shown. [Figure 6] A block diagram of an apparatus for generating images according to some embodiments of the present disclosure is shown. [Figure 7] A block diagram of equipment capable of realizing multiple embodiments of this disclosure is shown. [Modes for carrying out the invention]

[0011] The embodiments of this disclosure will be described in more detail below with reference to the drawings. While the drawings show several embodiments of this disclosure, this disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to provide a clearer and more complete understanding of this disclosure. The drawings and embodiments of this disclosure are for illustrative purposes only and should not be understood as limiting the scope of protection of this disclosure.

[0012] In the descriptions of the embodiments of this disclosure, the term “including” and its synonyms should be understood as open inclusion, i.e., “including, but not limited to.” The term “based on” should be understood as “based at least in part.” The terms “one embodiment” or “this embodiment” should be understood as “at least one embodiment.” The terms “first,” “second,” etc., may refer to different or the same subject. The following may include other explicit and implicit definitions.

[0013] As mentioned earlier, with advancements in vehicle technology and artificial intelligence, users are no longer satisfied with simply taking photos to capture beautiful scenery along the way or to record their current mood and emotions. It is understandable that users now desire to be able to process images according to their needs, generate more personalized images, and preserve their current mood and emotions.

[0014] Currently, artificial intelligence (AI) applications are already enabling AI-powered drawing capabilities in vehicles. For example, in-vehicle systems can directly call traditional "text-to-image" APIs, thereby fulfilling the need to generate images based on user voice commands within the vehicle. However, images generated by current methods are not "personalized," lack connection to the user and current vehicle usage scenarios, and often fail to meet the user's image requirements, frequently resulting in users receiving images unrelated to themselves.

[0015] Therefore, embodiments of the present disclosure propose a method for generating images. In embodiments of the present disclosure, the method includes determining a plurality of corresponding layers based on images of the vehicle's interior and exterior environment, wherein the spatial relationships of the plurality of layers are different. The method further includes generating corresponding images based on the plurality of layers by an image generation model based on the vehicle's state and user input information.

[0016] This method combines multi-view images of the vehicle's interior and exterior with information about the vehicle's state, providing rich material and reference grounds for image generation. Furthermore, user input can be considered during the image generation process, allowing for further customization of the image to better suit user needs and preferences. Because multiple layers provide material for image generation from various dimensions and directions, the presence of multiple layers also makes image generation or adjustment more flexible and precise.

[0017] Figure 1 shows a schematic diagram of an exemplary environment 100 that can realize several embodiments of the present disclosure. As shown in Figure 1, the exemplary environment 100 includes a vehicle 102, a user terminal 104, and images 106 of the vehicle's interior and exterior environment acquired by the vehicle 102. The exemplary environment 100 further includes a processor 108 that is communicated to the vehicle 102. According to embodiments of the present disclosure, the vehicle 102 is any type of powered or unpowered vehicle that is mobile and capable of carrying people and / or objects. As shown in Figure 1, the vehicle 102 is depicted as an automobile. Although the vehicle 102 is depicted as an automobile in Figure 1, this is merely illustrative and not limited to examples, and it should be understood that examples may further include buses, trucks, motorcycles, electric vehicles, and the like. In some embodiments of the present disclosure, the processor 108 may be realized as various service terminals having computing capabilities. Here, the service terminals may be servers, large computing devices, etc., provided by various service providers.

[0018] In the exemplary environment 100, various types of sensors may be deployed inside and outside the vehicle 102 to sense the internal and external environments of the vehicle. Here, the sensors may include, but are not limited to, visual sensors (e.g., cameras, camera lenses, etc.), millimeter-wave radar, laser radar, infrared sensors, global positioning systems, ultrasonic sensors, inertial measurement units, and other sensors. Here, the visual sensors can capture images of the environment around the vehicle 102 and the environment inside the vehicle 102 to obtain internal and external vehicle images. For example, the visual sensors can capture images of the environment around the vehicle 102 and the environment inside the vehicle to obtain image 106.

[0019] In some embodiments of this disclosure, image 106 may include multiple layers, such as an interior vehicle environment layer, an exterior vehicle environment layer, and a person inside the vehicle layer. Different spatial relationships may exist between the different layers, and the different layers may originate from different visual sensors. Image 106 can describe rich information about the interior and exterior vehicle environments through a combination or overlay of multiple layers. The spatial relationships between the layers may be used to represent the actual layout of the interior and exterior of the vehicle and the activity of people. Here, the interior vehicle environment layer may be used to represent the interior environment of the vehicle, such as the seats, steering wheel, and dashboard inside the vehicle 102. This interior vehicle environment layer may be laid out relative to the interior space of the vehicle 102, and the positional relationships between each element in the layer can reflect information such as the actual layout inside the vehicle. The exterior vehicle environment layer may include information about the environment in different directions, such as the front, left, and right of the vehicle 102. Each element in the exterior vehicle environment layer is associated with the exterior space of the vehicle and is used to reflect the relative positional relationship between the vehicle and the exterior environment. For example, the vehicle's external environment image may include buildings, pedestrians, and other vehicles located to the left of the vehicle.

[0020] In some embodiments of the present disclosure, the in-vehicle person image may be a corresponding image including the driver and passengers of the vehicle, for example, an image including the driver obtained by collecting with a camera located on the steering wheel. In other embodiments, the image 106 may further be an image obtained by collecting with a terminal device corresponding to the driver and passengers. For example, an image obtained by the terminal device photographing the external environment of the vehicle or an image obtained by the driver and passengers taking selfies using the terminal device.

[0021] In some embodiments of the present disclosure, a plurality of layers may be input to the processor 108, and the processor 108 may generate a corresponding image 118 based on the plurality of layers. Here, a generation model 116 may be deployed in the processor 108, and the generation model 116 may be an image generation model. The image generation model may be a generation model for generating images that has completed training through a large amount of training data. The image generation model may include various types of models, for example, generative adversarial networks (GANs), variational autoencoders (VAEs), and diffusion models, etc., but is not limited thereto. These models can capture the distribution of image data through different mechanisms, such as adversarial training, probabilistic modeling, or progressive noise removal, etc., and generate realistic new images. That is, the plurality of layers can be used as the basis for image generation model inference.

[0022] In some embodiments of the present disclosure, various reference information may be input to the processor 108 to personalize and diversify the generated image 118 and adapt it to the actual driving state of the vehicle, and may also be used as another basis for the image generation model 116 to generate an image. Here, the various reference information may include, but is not limited to, the vehicle state 114, the user input information 112, weather information (such as sunny day, rainy day, temperature or humidity), geographical location information (such as the current geographical location of the vehicle), etc. Here, the user input information may be used to represent the user's needs. For example, the user may input their needs, such as the style of the image to be generated, specific elements to be added, mood, etc., via the user terminal 104 according to the actual situation. In one example, the user may input their needs by various input methods such as voice, text, gesture, scanning, and touch.

[0023] In some embodiments of the present disclosure, the vehicle state 114 may include the driving data of the vehicle 102 (such as driving speed, acceleration, power consumption, etc.), driving data (such as the rotation angle of the steering wheel, the depression force of the brake pedal, etc.), the data of the vehicle control unit of the vehicle (such as whether the window / door is open, the color of the ambient light, the temperature / wind speed of the air conditioner, etc.), and media resource data (such as music or video played by the in-vehicle system).

[0024] In some embodiments of the present disclosure, the image generation model 116 may directly generate an image 118 that meets the user's needs based on various reference information and a plurality of layers. In other embodiments, the image generation model 116 may generate a candidate image 110 based on a plurality of layers. The candidate image 110 is adjusted based on various reference information to generate the corresponding image 118. Here, the adjustment may be to adjust the candidate image 110 or to adjust some layers of the candidate image 110. This layer may correspond to the original plurality of layers.

[0025] This method makes full use of images collected by the vehicle, combining them with the vehicle's own data, current environmental information, and the user's needs to generate images that better suit the user's preferences. The generated images are no longer simple, unrelated images, but personalized images tailored to the user's requests, enhancing the user's driving and riding experience, strengthening the connection between the vehicle and the user, and making every journey more enjoyable and exciting. Furthermore, the presence of layers allows for more precise adjustment of the generated images, thereby improving the efficiency and accuracy of image generation.

[0026] Figure 2 shows a schematic diagram of the flow for generating images according to some embodiments of the present disclosure. As shown in Figure 2, the in-vehicle / out-of-vehicle visual acquisition module 202 can acquire images including the in-vehicle and out-of-vehicle environment, such as in-vehicle images and out-of-vehicle images (the sensors in Figure 2 are cameras, but are not limited to cameras), based on multiple sensors deployed inside and outside the vehicle. Here, the out-of-vehicle sensors may be deployed around the vehicle to acquire images of the area above, below, left, right, front, and rear of the vehicle, and to obtain external vehicle images, such as environmental images. In-vehicle sensors may be deployed in important locations inside the vehicle, such as the steering wheel, drive recorder, dashboard, and doors. The in-vehicle sensors can acquire internal vehicle images, such as images of people inside the vehicle and images of the internal environment, for the environment inside the vehicle. It can be understood that the deployment locations of the sensors may be set according to the acquisition demand and the image generation demand. In some embodiments, images may be acquired by a user terminal to enrich the image content and to assist in image generation. Here, the images acquired by the user terminal may include external vehicle images and user selfies.

[0027] In some embodiments of this disclosure, image recognition and processing may be performed on images of the interior and exterior environments of a vehicle to obtain an image 216 containing multiple layers, enabling the generation of more personalized images. Here, the multi-image fusion module 210 performs image recognition and processing on images of the interior and exterior of the vehicle, separating the person and the environment from the images, distinguishing between the interior and exterior environments, and obtaining a number of corresponding layers (e.g., an exterior layer, an interior layer, a person layer, etc.). Different layers contain information about different structures at different locations. For example, the person layer may be determined based on a user's selfie obtained by a user terminal and interior images of the vehicle collected by in-vehicle sensors, and includes the person's features. Here, the person's features may include, but are not limited to, facial expressions, limb movements, hairstyle, skin color, clothing, and gender. In some embodiments of this disclosure, the person region and the vehicle region in image 216 may be separated to obtain the person layer. For example, person recognition may be performed on the image, and the layer corresponding to the person included in the image may be determined. After determining the person layer, the person's facial features and posture features may be identified and obtained.

[0028] In some embodiments of this disclosure, each layer can be distinguished by adding tags or annotations to elements of different layers, ensuring that the relative position of each image in the overall scene conforms to the laws of the real world, and facilitating subsequent adjustments of the corresponding layers. For example, depending on the camera source corresponding to image 216, spatial position calibration can be performed for each image in image 216. For example, multiple layers (e.g., an exterior layer and an interior layer) can be determined by calibrating the different orientations (front, rear, left, right) of the exterior image and the spatial relationships and relative positions of the interior image.

[0029] In one example, image 1 can be collected and acquired using camera 1 (installed on the left side of the vehicle), image 2 can be collected and acquired using camera 2 (installed in front of the vehicle), image 3 can be collected and acquired using camera 3 (installed on the right side of the vehicle), and image 4 can be collected and acquired using camera 4 (installed on the steering wheel inside the vehicle). Based on the identification information of cameras 1, 2, 3, and 4, it can be determined that the boundaries of image 1 and image 2 are adjacent, the boundaries of image 2 and image 3 are adjacent, and image 4 is located inside the vehicle. In other embodiments, edge detection technology can also be used to identify important contours in the image, including edges of the external scene, boundaries of the interior structure, and contours of people. Based on this information, the relative positional and spatial relationships between the images can be determined.

[0030] In some embodiments of this disclosure, after determining edge information corresponding to multiple layers, the multiple layers may be merged to obtain a fused image. For example, a corresponding fused image may be obtained by combining edge information and tags corresponding to layers and performing stitching fusion on adjacent images. The fused image may be input to an image generation model 234, and the image generation model 234 may generate a corresponding image 236 based on this fused image. This fused image may substantially be the "base image" of the generated image.

[0031] In some embodiments of this disclosure, various reference information may be input to the image generation model 234 to improve the accuracy of the generated image and to make the generated image more user-friendly, and the image generation model 234 may generate an image 236 that meets the user's needs based on the various reference information and the fused image. The image generation model 234 may adjust the generated image based on the various reference information to finally obtain an image 236 that meets the requirements.

[0032] In one embodiment, the reference information may be the vehicle status. Here, the vehicle status may include vehicle data (information directly related to the vehicle's operating state, such as power consumption and vehicle speed) and vehicle control unit information 204. For example, the dynamics of the generated image can be adjusted according to the difference in vehicle speed. When the vehicle is traveling at high speed, the generated image may be more dynamic and convey a sense of speed, while when the vehicle is traveling at low speed or stopped, the generated image may be more static and detailed. The image generation model can also adjust the image according to the vehicle's power consumption or fuel level. For example, when the vehicle's fuel level or power consumption is relatively low, the generated image may be relatively dark and slow. In this way, the user can be reminded to some extent to charge or refuel in a timely manner. In other embodiments, the vehicle status may further include vehicle control unit data, media resource data, and driving operation data. For example, if the steering wheel rotation angle is relatively large, the image generation model 234 can add curves to the generated image to better fit the generated image to the actual situation, thereby providing the driver and passengers with a more immersive travel experience.

[0033] In some embodiments of this disclosure, the reference information may further include environmental data 206 in which the vehicle is located. Herein, the environmental data 206 may be used to reflect various data relating to the current environment or current conditions in which the vehicle is located. For example, it may include, but is not limited to, weather data, seasonal data, time data, and geolocation data (the geolocation data may include characteristic data of the location where the vehicle is currently located, e.g., urban / rural, landmarks, elevation, etc.). For example, if the vehicle is traveling at a winter evening, elements of a sunset and snow-covered mountains may be added to the generated image, or if the vehicle is traveling in the center of a city, relatively dense urban buildings may be added to the generated image.

[0034] In some embodiments of this disclosure, user sensibilities are paramount, and in order for the generated image to satisfy the user's preferences, user requests must be incorporated into the image generation process as a basis for adjusting or generating the image. For example, the generated image or a single layer of the image may be adjusted based on user input information 208. For example, the features and imagery of people or environments in the image may be adjusted according to the user's requests, ultimately generating an image unique to the user. In some embodiments, user requests may be determined based on user input information 208, for example, the user may input their requests (e.g., desired image style, user's mood, etc.) in a manner such as voice, text, gesture, or touch. In one example, the user may directly express their emotions and feelings at that moment. Alternatively, the user may input specific adjustment commands and modify specific layers. For example, the user may request that the generated image display a specific natural landscape, e.g., "Flowers blooming from the dry tree trunks on this path I walk," or the user may request, "Make the background color a little darker." This input information directly influences the final generated image content and presentation.

[0035] In some embodiments of this disclosure, in order to make full use of the collected images, a generative model (e.g., an exterior semantic understanding module 212 and an interior semantic understanding module 214) may be used to analyze the collected images and determine an image description corresponding to these images. For example, the image description 218 may be that the vehicle is driving on Sakura Street at the time, with cherry trees with lush branches and leaves planted on both sides of the road, petals fluttering in the wind, and the sky being pink. In other embodiments, by extracting features from the person layer, various features 220 corresponding to the person layer may be further determined, such as gender, facial expression, hairstyle, clothing, and location (location inside the vehicle, whether or not the person is in the driver's seat, etc.).

[0036] It can be understood that the exterior semantic understanding module 212, the interior semantic understanding module 214, the vehicle data and vehicle control information module 204, the environmental information data module 206, and the user instruction module 208 may input the generated information to the prompt engineering module 228, and the prompt word engineering module 228 may integrate the information generated by each of the aforementioned modules into a prompt word 232 that is input to the image generation model 234. For example, the prompt word 232 may be "a girl with long hair blowing in the wind, cherry blossoms fluttering outside the car, cherry trees on both sides, the driver's side window is open, and the girl has a cheerful expression."

[0037] In some embodiments of this disclosure, the generation of prompt words 232 in the prompt engineering module 228 may include, but is not limited to, automated generation and predefined styles. Here, automated generation may involve pre-training a single generative text large-scale language model and automatically generating prompt words that match the scene based on input contextual information. In some embodiments of this disclosure, different image styles may be pre-set depending on the actual situation in order to improve the efficiency of prompt word generation. Based on user input, an image style that matches this input may be selected from a plurality of image styles. For example, the user may input general command information such as "make the image colors brighter" or "make the scenery the main focus in the image, with people as secondary" according to their needs. This command information may be used to help the image generation model 234 generate an image having the style of a specific request.

[0038] In some embodiments of this disclosure, a user may modify an image multiple times or adjust it by continuously proposing new user commands based on previous adjustments. Dialogue history management can be used in a prompt engineering module 228 that interacts with a dialogue history database 230 and stores and manages information about the user's multi-round interactions. If a user wishes to continue adjusting an image generated after a single interaction, fine-tuning can be done based on the dialogue history. By adjusting at least one layer based on the command information 232 generated by the prompt engineering module 228, the resulting final image 236 includes all features specified by the user and fully reflects the user's personalized needs, such as the facial features of a female user, the sense of speed in a driving scene, color coordination, and mood expression. All of these details can be clearly shown in image 236.

[0039] This method allows for the generation of images that meet user needs by combining user input, vehicle status, and the surrounding environment during the image generation process. Because the information underlying the generated images is richer and more diverse, the generated images are better suited to the current situation and possess greater richness and hierarchicality, thereby enhancing the user's driving and riding experience.

[0040] The following flowchart describes a method 300 for generating an image according to an embodiment of the present disclosure, with reference to Figure 3. Note that the method 300 according to an embodiment of the present disclosure may be implemented, for example, by the processor 108 shown in Figure 1, and more specifically by the image generation model 116. Furthermore, with the continued development and innovative upgrades of smart cars and smartphones, an increasing number of vehicles and mobile devices have computing capabilities (e.g., smart cars and smartphones). In other words, the method 300 according to an embodiment of the present disclosure is not limited to being implemented in the processor 108, and may also be implemented in the vehicle 102, and the present disclosure does not limit this.

[0041] As shown in Figure 3, in block 302, method 300 includes determining a plurality of corresponding layers based on images of the vehicle's interior and exterior environments, wherein the spatial relationships of the plurality of layers differ. For example, images of the vehicle's interior environment and images of the vehicle's exterior environment may be collected and obtained using image acquisition equipment. The image acquisition equipment may be deployed in the vehicle, for example, on the left side, right side, front, rear of the vehicle, and on the dashboard, air conditioning vents, etc., inside the vehicle. Image acquisition equipment deployed in different locations may be used to collect and obtain different images. For example, image acquisition equipment not located inside the vehicle can collect and obtain images of the vehicle's interior environment. Depending on the source of the images (e.g., from which image acquisition equipment they originate), the collected images may be easily divided into a plurality of layers, for example, an exterior layer and an interior layer, and the positional relationships between the layers may be determined based on this.

[0042] In some embodiments of this disclosure, the images may include various subjects, such as people and backgrounds (e.g., vehicle interior structure, traffic signs in the surrounding environment, lanes, etc.). Based on this, further image recognition can be performed on the collected images to separate the people layer. Here, the people layer may include the driver and passengers (e.g., the driver, passengers in the front and rear seats, etc.).

[0043] In block 304, method 300 includes generating a corresponding image based on multiple layers using an image generation model, based on the vehicle state and / or user input information. Here, the vehicle state may be used to characterize various current states of the vehicle, such as the vehicle's driving state, the user's driving operations on the vehicle, and the state of the vehicle's control components. The vehicle's driving state may include, but is not limited to, the driving speed, acceleration, driving direction, and driving mode. The state of the vehicle control components may include the state of the air conditioning, the state of the ambient lighting, or the state of the windows.

[0044] In some embodiments of this disclosure, user input information includes to some extent the user's needs (e.g., preferred style, elements to be added), user's feelings, or user's experience. For example, the user may input information via the vehicle's in-vehicle system or a terminal device communicating with the vehicle. The user may input information in a variety of ways, including, but not limited to, voice input, touchscreen operation, gesture input, and biometric authentication (e.g., fingerprint or facial recognition). In other embodiments, user input information may further include the user's editing or adjustment commands for the generated image or layer. For example, the user may request to display a specific natural scene or add / remove specific elements from the generated image.

[0045] In some embodiments of this disclosure, the image generation model may be a model trained on a large amount of training data that can complete a specific task. In some embodiments, the image generation model may be fine-tuned to complete a variety of natural language processing tasks, such as text generation, code generation, video generation, text question answering, image generation, paper writing, etc. In some embodiments, the image generation model may be an image-to-image model that can substantially generate a new photorealistic image based on an input image. In some examples, the image generation model may include, but is not limited to, generative adversarial networks, variational autoencoders, and diffusion models.

[0046] In some embodiments of this disclosure, an image generation model may generate an image that meets the requirements by combining the state of the vehicle and / or user input based on multiple layers. If the image input to the image generation model is in the form of multiple layers, it can be understood that the generated image may include multiple layers. The multiple layers and the multiple original layers input may correspond to each other, and the corresponding generated image can be obtained by fusing the multiple layers. In other embodiments, a fused image may be obtained after fusing the multiple original layers, and the image generation model may generate a corresponding image based on this fused image.

[0047] In the image generation process, it is understood that the vehicle state and user input may be used as generation or adjustment criteria to adapt the generated image to the vehicle state and to meet the user's needs. In some embodiments, the image generation model may directly generate an image that meets the requirements based on multiple original layers, the vehicle state, and user input, while in other embodiments, the image generation model may generate multiple layers based on multiple original layers and adjust specific layers based on the vehicle state and user input. An image that meets the requirements is then generated based on the adjusted layers.

[0048] This method combines multi-view images of the vehicle's interior and exterior with information about the vehicle's state, providing rich material and reference grounds for image generation. Furthermore, user input can be considered during the image generation process, allowing for further customization of the image to better suit user needs and preferences. Because multiple layers provide material for image generation from various dimensions and directions, the presence of multiple layers also makes image generation or adjustment more flexible and precise.

[0049] In some embodiments of this disclosure, after collecting multiple images, the images may be classified according to their source, and different types of layers may be determined. For example, the images may be divided into an exterior layer, an interior layer, and a person layer. The relative positions of the different layers may differ.

[0050] Figure 4 shows a schematic diagram of determining multiple layers according to some embodiments of the present disclosure. As shown in Figure 4, cameras 410 (installed at the rear of the vehicle) and camera 412 (installed in the drive recorder inside the vehicle), which are deployed at different locations on the vehicle, can be tagged to give the collected images identification information corresponding to the camera. Based on the identification information, the location of each of the multiple sub-environmental images in the environmental image and the spatial positional relationship between the images can be determined, ensuring that the relative position of each image in the overall scene conforms to real-world laws when generating subsequent layers, and laying the foundation for subsequent image fusion. Multiple images obtained by collecting multiple cameras can be input to a layer identification device 416, and the layer identification device 416 can determine the corresponding multiple layers 418 based on the multiple images. For example, images collected by cameras deployed outside the vehicle may be used as the external layer 424. Since the outside of the vehicle has various orientations relative to the vehicle (e.g., front, rear, left, right, etc.), external layer a located on the left side of the vehicle, external layer b located on the right side of the vehicle, and external layer c located in front of the vehicle can be determined from the camera's identification information. It can be understood that external layer a, external layer b, and external layer c may jointly constitute the overall external layer. For example, external layer 424 may be obtained by fusing adjacent external layer a and external layer c, and adjacent external layer b and external layer c.

[0051] As shown in Figure 4, in some embodiments of this disclosure, multiple images obtained by cameras deployed inside the vehicle may constitute an interior layer 422. Since the driver and passengers are seated in the vehicle, the interior layer 422 may also contain people, which can be obtained by using a layer identification device 416 to identify the contours of people's bodies and the boundaries of the interior structure in the interior layer 422. Based on this, people can be separated from the interior layer 422 to obtain a people layer 420 containing only people.

[0052] This method effectively processes vehicle interior and exterior images captured by multiple image acquisition devices, decomposes them to generate corresponding layers, and makes it easier to more accurately identify target layers in subsequent layer adjustment and image fusion. Furthermore, it can meet users' demand for personalized and unique images, avoid the repetition of a single image, and ensure that each generated image is full of individuality and creativity.

[0053] In some embodiments of this disclosure, multiple original layers may be merged to obtain a merged image before inputting them into an image generation model. For example, an exterior landscape and a person inside a vehicle may be merged into a single image. In one example, three layers, such as an exterior landscape layer, an interior environment layer, and a person layer, can be aligned while ensuring that their edges are clear and natural. In the process of generating multiple layers, identification information may be added to the original layers based on the source of the original layers, and this identification information may be used to uniquely represent the type of original layer. For example, different layers may be marked using different colors or tags. In other words, the multiple original layers included in the merged image also have their own identification information. The image generation model may generate an image that has the meaning of "real space" based on the merged image. This image may include various layers, and different tags may be added to different layers. In this way, the image generation model can be guided to adjust the content of a particular layer by different adjustment suggestions. For example, a prompt word such as "change exterior landscape" may be used to cause the model to modify only the exterior landscape layer in the generated image without affecting other layers.

[0054] In some embodiments of this disclosure, the image generation model may generate corresponding new layers based on a plurality of input original layers. For example, the image generation model may generate a new realistic "person layer" based on an input original person layer, or a new realistic "exterior layer" based on an input exterior layer. The style and elements of this person layer are similar to those of the original person layer. In one example, the image generation model may determine the features of the original person layer (e.g., facial features, clothing style, background environment, etc.) and generate a new person layer based on these features. The hairstyle, facial expression, or clothing of the person in this person layer corresponds to the hairstyle, facial expression, or clothing of the person in the original layer. For example, a feature extraction model such as a convolutional neural network (CNN) or a deep learning model may be used to extract features from the input person layer. These features are used as the basis for adjusting the generated person layer.

[0055] In some embodiments of this disclosure, a semantic understanding of the original layers may be used as the basis for adjustment in order to further improve the efficiency and accuracy of layer adjustment. For example, a Vision-Language Pre-training (VLP) model may be used to interpret the input layer and determine the semantic description corresponding to this layer. Here, the image-text conversion model may include, but is not limited to, a Contrastive Language-Image Pre-training (CLIP) model or a Bootstrapping Language Image Pre-training (BLIP) model. In some embodiments of this disclosure, this semantic description may include a description of the objects contained in this layer, the corresponding scene, and the relationships between the objects (e.g., spatial and logical relationships). For example, in one example, the semantic description of an exterior layer may be "There is a lake to the left of the vehicle, and there is a swan next to the lake" and "The distant mountain peaks are covered with white snow." This semantic description may also be used to describe the overall layout, composition, or perspective relationships of the layer. Based on this, this semantic description may also serve as the basis for the external layer generated by adjusting it using an image generation model, for example, by adding a swan element to the generated external layer.

[0056] This method not only helps to increase the efficiency of layer adjustment by using a highly accurate semantic understanding of the original input layer as the basis for adjusting the generated layer, but also greatly guarantees the accuracy of the adjustment, making the adjusted layer conform to people's aesthetic expectations and visual perceptions.

[0057] In some embodiments of this disclosure, the semantic understanding of the interior layer and the person layer may also serve as a basis for adjusting the generated interior or person layer. For example, the person layer may be interpreted, and a corresponding semantic description may be generated based on the features of this person layer. This semantic description may be "a little girl is sitting in the driver's seat, with long, scattered hair blowing in the wind, and she is holding a French bulldog." It can be understood that the semantic description corresponding to the original layer has higher priority and serves as an important input for subsequent image generation, ensuring that the generated image or content is highly relevant to the actual features of the user.

[0058] Figure 5 shows a schematic diagram of adjusting multiple layers according to some embodiments of the present disclosure. As shown in Figure 5, the unadjusted layer 502 may include an exterior layer 504, a person layer 506, and an interior layer 508. The target layer to be adjusted may be determined based on the vehicle's state and user input information. The image generation model may adjust the target layer based on the vehicle's state and user input information to obtain an adjusted layer 510 including an adjusted exterior layer 512, a person layer 514, and an interior layer 516. For example, in some embodiments, if the vehicle is traveling at a high speed (windows are open), the generated image may be given a sense of speed by representing the effect of wind speed on the person and the environment. Based on this, the target layer that needs adjustment may be determined to be the person layer 506. The image generation model determines the adjusted person layer 514 by adjusting the features of the person in the person layer 506. For example, in the adjusted person layer 514, the girl's long hair is flowing behind her, and such an effect can represent the sense of speed of the vehicle at this time. In other embodiments, the sense of speed of the vehicle may be expressed by simultaneously adjusting some features in the exterior layer 504, the person layer 506, and the interior layer 508. For example, the tree canopy in the exterior layer may be made to sway backward to create an afterimage effect, and flower petals from the scenery outside the vehicle may enter the vehicle and be carried backward by the wind.

[0059] In some embodiments of this disclosure, the user may input their needs according to the actual situation, such as the style of the image they want to generate, specific elements they want to add, or their mood. If the user inputs a need such as "I want the weather to be warmer," it can be understood that the target layer to be adjusted is the exterior layer 504. The image generation model can determine the adjusted exterior layer 512 by adding a sun to the exterior layer 504 and softening the overall color tone of the layer. Depending on the user's needs, it can also be understood that the layer to be adjusted is the interior layer 516. The adjusted layer can be further adapted to the user's needs by adding some fluffy stuffed animals, cute decorations, or changing the overall color tone of the interior environment. In this manner, images can be adjusted according to vehicle data and user needs, allowing the user to participate more significantly in the image generation process and generating images with more personal characteristics.

[0060] In some embodiments of this disclosure, the history image is generally an image that the user finds satisfactory. For example, the style or elements included in the history image are all things that the user likes and that suit the user's preferences. Based on this, the image generation model may generate a corresponding image using the history image as an example. The history image may be used as reference content to cause the image generation model to generate an image that is similar to the history image in terms of the driving scene, specific elements, layout, typesetting, style, etc. For example, if the history image contains a large number of butterfly elements, the generated image may also contain butterfly elements. Or, if the person in the history image has long hair, the person in the generated image may also have long hair.

[0061] In some embodiments of this disclosure, during the image generation process, the user may adjust aspects of the generated image that they are not satisfied with, such as modifying filters or adjusting human features. The user's adjustment methods reflect, to some extent, the user's preferences and recording or sharing habits. Therefore, based on this, the image generation model may generate a corresponding image based on the history of user-model interactions. Here, the history of interactions may include inputs from multiple interactions between the user and the model, generated history images, and associated prompt words. In some embodiments of this disclosure, if the user wants to adjust an image generated after a single interaction, the image generation model may fine-tune the previously generated image based on the interaction history. For example, if the command entered by the user is "slightly brighten the background based on the previous image," the image generation model may generate a new image by making the relevant adjustments to the previously generated image.

[0062] In some embodiments of this disclosure, environmental data of the vehicle's current location may also be used as a basis for the image generation model to generate images or adjust layers. Here, environmental data may be weather data (e.g., sunny day, rainy day, snowy day, temperature, humidity, etc.). For example, the image generation model may adjust the visual representation of the generated image based on the weather data. For example, on a sunny day, a bright sun may be added to the generated image; on a rainy day, a raindrop effect may be added to the generated image; and on a snowy day, a snowflake effect may be added to the generated image.

[0063] In other embodiments, the environmental data may further include seasonal data, where the seasonal data is the weather background for the vehicle's current time of day and may include, but is not limited to, the four seasons of spring, summer, autumn, and winter. Different seasons have different feature elements, and the image generation model may add scene elements to the image that are suitable for the seasonal features during the image generation process. For example, in the case of spring, the generated image may include a scene with bright sunshine and flowers in full bloom; in the case of summer, the generated image may include a bright and vibrant scene; and in the case of autumn, the generated image may include fallen leaves and the image may have a warm color tone.

[0064] In some embodiments of this disclosure, the environmental data may further include temporal data, where the temporal data may be the time of day in which the vehicle is located at the current time, for example, the morning, afternoon, or evening time. Different time periods correspond to different light illumination effects. For example, during sunrise and sunset, the light is soft and objects on the ground cast long shadows, while during the evening, there is no sunlight but instead elements such as moonlight, stars, or artificial lighting, resulting in an overall darker tone. Based on this, it is possible to generate lighting effects that are consistent with the real world based on the temporal data.

[0065] In some embodiments of this disclosure, the environmental data may further include geolocation information in which the vehicle is located at the current time or within a predetermined time frame. It can be understood that the places the vehicle passes through have different geographical features, for example, rural and urban landscapes are completely different, or the geographical features of areas at different elevations are completely different. Based on this, an image or a layer in an image generated based on the vehicle's geolocation information can be adjusted. In one example, an external layer in the image may be adjusted depending on a landmark or scenic spot that the vehicle passes through. For example, this landmark may be added to the external layer as a background element.

[0066] In some embodiments of this disclosure, various adjustments may be made to various layers in various ways, for example, by adjusting certain elements (e.g., adding new elements) or features in a layer (e.g., adjusting the hue or transparency of a layer). The various adjusted layers and the original layers still correspond to each other, meaning that each adjusted layer retains its original pixel and position information. Therefore, based on the relationships between the layers, an image that meets the user's needs can be generated after fusing multiple adjusted layers. For example, based on the edge features or boundaries of a layer, the edges of adjacent images can be smoothed to generate an image with a matching style.

[0067] In some embodiments of this disclosure, when fusing multiple adjusted layers, the style of the resulting image may be adjusted by adjusting the transparency of different layers. In one example, the multiple adjusted layers may include an exterior scenery layer and a person layer. The person layer may be placed on top of the exterior scenery layer, and by adjusting the transparency, size, and position of the person layer, it can be harmonized with the background of the exterior scenery layer, and the overall style of the fusing image can be made more consistent. In some embodiments, the edges between the person layer and the exterior scenery layer can be smoothed to reduce abruptness in the fusing. The overall style of the composite image can be made more consistent by adjusting parameters such as color and brightness.

[0068] Figure 6 shows a block diagram of an image generating apparatus 600 according to some embodiments of the present disclosure. As shown in Figure 6, the apparatus includes a layer determination module 602 configured to determine a plurality of corresponding layers based on images of the vehicle interior and exterior environment, wherein the spatial relationships of the plurality of layers differ. The apparatus 600 further includes an image generation module 604 configured to generate a corresponding image based on the plurality of layers by an image generation model based on the vehicle state and / or user input information.

[0069] In some embodiments, the layer determination module 602 is further configured to use multiple image acquisition devices to acquire environmental images of the surrounding environment of the vehicle and interior images corresponding to the interior of the vehicle, to determine a person image corresponding to a user inside the vehicle based on the image corresponding to the interior of the vehicle, and to generate a plurality of corresponding layers based on the environmental image, the interior image and the person image corresponding to the user.

[0070] In some embodiments, the layer determination module 602 is further configured to determine a plurality of sub-environmental images for a plurality of directions in the environmental image, including the front, rear, left, and right sides of the vehicle, based on the identification information of the plurality of image acquisition devices, and to determine a plurality of corresponding layers based on the plurality of sub-environmental images, the interior image and the person image.

[0071] In some embodiments, the image generation module 604 is further configured to generate a plurality of corresponding candidate layers based on the plurality of layers, determine a target layer to be adjusted in the plurality of candidate layers based on the vehicle state and / or user input information, and adjust the target layer by an image generation model to generate a corresponding image.

[0072] In some embodiments, the apparatus 600 further includes a layer adjustment module configured to determine a type and image features corresponding to the person image, and to adjust the target layer, which includes at least one of the person layer and the environment layer, based on the type and the image features.

[0073] In some embodiments, the layer adjustment module is further configured to adjust the person features of the person layer in at least one layer based on the image features, wherein the person features corresponding to the adjusted person layer match the image features, and the image features include the person's facial features and posture features; and to adjust the environmental features in the in-vehicle environment layer based on the image features.

[0074] In some embodiments, the layer adjustment module is further configured to determine corresponding key elements and image styles based on the user input information, and to adjust the target layer based on the key elements and image styles.

[0075] In some embodiments, the layer adjustment module is further configured to acquire environmental information corresponding to the vehicle, including weather information, time information, and geolocation information, and to adjust the target layer based on the environmental information, the vehicle's status, and user input information.

[0076] In some embodiments, the layer determination module 602 is further configured to acquire vehicle driving data, operating data, media resource data and corresponding vehicle control unit data, and to determine the state of the vehicle based on the driving data, operating data, media resource data and vehicle control unit data.

[0077] In some embodiments, the image generation module 604 is further configured to determine a corresponding fused image by performing multi-layer fusion on the adjusted layers, and to generate a corresponding image by an image generation model based on the fused image.

[0078] In some embodiments, the image generation module 604 is further configured to determine edge information corresponding to each layer in the adjusted layer, determine the x adjacent layers based on the edge information, and determine the corresponding fused image by fusing the adjacent layers.

[0079] In some embodiments, the layer adjustment module is further configured to generate text descriptions corresponding to images of the vehicle's interior and exterior environments using a language model, and to adjust the target layer based on the text descriptions.

[0080] In some embodiments, the layer adjustment module is further configured to acquire history input information for the user's history images and to adjust the target layer based on the history images and the history input information.

[0081] It can be understood that by using the apparatus 600 of this disclosure, at least one of the many advantages that can be achieved by the methods or processes described above can be realized.

[0082] Figure 7 shows a schematic block diagram of an exemplary device 700 that may be used to carry out embodiments of the present disclosure. As shown in the figure, the device 700 includes a computing unit 701 which can perform various appropriate operations and processes based on computer program instructions stored in read-only memory (ROM) 702 or computer program instructions loaded from storage unit 708 into random access memory (RAM) 703. The RAM 703 may further store various programs and data necessary for the operation of the device 700. The computing unit 701, ROM 702 and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0083] Multiple components in the device 700, including, for example, an input unit 706 such as a keyboard and mouse, an output unit 707 such as various types of displays and speakers, a storage unit 708 such as a magnetic disk and an optical disk, and a communication unit 709 such as a network card, modem, and wireless communication transceiver, are connected to the I / O interface 705. The communication unit 709 allows the device 700 to exchange information / data with other devices, for example, via the Internet computer network and / or various telecommunications networks.

[0084] The computing unit 701 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs each of the methods and processes described above, for example, method 300. For example, in some embodiments, method 300 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, for example, a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 700 via ROM 702 and / or a communication unit 709. Once the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of method 300 described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform method 300 by any other suitable means (for example, by firmware).

[0085] In this specification, the functions described above may be performed by at least partially one or more hardware logic components. For example, without limitation, typical types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), dedicated integrated circuits (ASICs), dedicated standard products (ASSPs), systems on a chip (SOCs), and complex programmable logic devices (CPLDs).

[0086] Program code for carrying out the methods of this disclosure can be written using any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, it performs the functions / operations defined in the flowchart and / or block diagrams. The program code may be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine as a standalone software package, or fully on a remote machine or server.

[0087] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in combination with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any appropriate combination of the above. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any appropriate combination of the above. While each operation is described using specific procedures, this should not be understood as requiring such operations to be performed by the specific or sequential procedures shown, or as requiring all illustrated operations to be performed to obtain the desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while the above discussion includes some specific implementation details, these should not be construed as limiting the scope of this disclosure. Some features described in the context of a single embodiment may be combined to be implemented in a single implementation. Conversely, various features described in the context of a single implementation may be implemented in multiple implementations, either individually or in any suitable subcombination.

[0088] Although this topic has already been described using language specific to structural features and / or methodological logic, it should be understood that the topic limited to the claims in the attached file is not necessarily limited to the specific features or behaviors described above. Conversely, the specific features and behaviors described above are merely illustrative forms of realizing the claims.

Claims

1. The process involves determining a number of corresponding layers based on images of the vehicle's interior and exterior environment, where the spatial relationships of the said layers are different. A method for generating an image, comprising generating a corresponding image based on a plurality of layers using an image generation model, based on the vehicle's condition and / or user input information.

2. Determining the corresponding multiple layers based on images of the vehicle's interior and exterior environment is, Using multiple image acquisition devices, environmental images of the surrounding environment of the vehicle and interior images corresponding to the interior of the vehicle are acquired. Based on the image corresponding to the interior of the vehicle, a person image corresponding to a user inside the vehicle is determined, The method according to claim 1, comprising generating a plurality of corresponding layers based on the environmental image, the interior image, and the person image corresponding to the user.

3. Determining the corresponding multiple layers is Based on the identification information of the aforementioned multiple image acquisition devices, multiple sub-environmental images are determined for multiple directions in the environmental image, including the front, rear, left, and right sides of the vehicle. The method according to claim 2, further comprising determining a plurality of corresponding layers based on a plurality of sub-environmental images, the interior image of the vehicle, and the person image.

4. The image generation model generates corresponding images based on the multiple layers, Based on the aforementioned multiple layers, a corresponding number of candidate layers are generated, Based on the vehicle's status and / or user input information, the target layer to be adjusted among the multiple candidate layers is determined. The method according to claim 2, further comprising adjusting the target layer using an image generation model to generate a corresponding image.

5. Adjusting the aforementioned target layer means To determine the type and image features corresponding to the aforementioned person image, The method according to claim 4, comprising adjusting the target layer, which includes at least one of a person layer and an environment layer, based on the type and the image features.

6. Adjusting the aforementioned target layer means Based on the aforementioned image features, the person features of the person layer in at least one layer are adjusted, wherein the person features corresponding to the adjusted person layer match the aforementioned image features, and the aforementioned image features include the person's facial features and posture features. The method according to claim 5, further comprising adjusting the environmental characteristics in the in-vehicle environment layer based on the aforementioned image characteristics.

7. Adjusting the aforementioned target layer means Based on the user's input information, the corresponding key elements and image styles are determined. The method according to claim 4, further comprising adjusting the target layer based on the key elements and image style.

8. Adjusting the aforementioned target layer means To acquire environmental information corresponding to the vehicle, including weather information, time information, and geographic location information, The method according to claim 4, further comprising adjusting the target layer based on the environmental information, the vehicle status, and user input information.

9. To acquire vehicle driving data, operation data, media resource data, and corresponding vehicle control unit data, The method according to claim 1, further comprising determining the state of the vehicle based on the driving data, the operating data, the media resource data, and the vehicle control unit data.

10. To generate the corresponding image, By performing multi-layer fusion on the adjusted layers, the corresponding fused image is determined, The method according to claim 4, further comprising generating a corresponding image using an image generation model based on the fused image.

11. Determining the corresponding fused image is Determining the edge information corresponding to each layer in the adjusted layer, Based on the aforementioned edge information, the layer adjacent to the edge is determined, The method according to claim 10, further comprising determining a corresponding fused image by merging the edges with adjacent layers.

12. Adjusting the aforementioned target layer means The language model generates text descriptions corresponding to images of the vehicle's interior and exterior environment, The method according to claim 4, comprising adjusting the target layer based on the text description.

13. Adjusting the aforementioned target layer means To obtain the user's history input information for their history images, The method according to claim 4, further comprising adjusting the target layer based on the history image and the history input information.

14. A layer determination module configured to determine a plurality of corresponding layers based on images of the vehicle's interior and exterior environment, wherein the spatial relationships of the plurality of layers are different, An image generating apparatus, comprising an image generation module configured to generate corresponding images based on a plurality of layers by an image generation model based on the vehicle status and / or user input information.

15. It is an electronic device, At least one processor, An electronic device comprising a memory coupled to at least one processor and having instructions stored thereon, wherein, when the instructions are executed by the at least one processor, the device causes the device to perform the method according to any one of claims 1 to 13.

16. A computer program that is physically stored on a computer-readable medium and, when executed, implements the method described in any one of claims 1 to 13.