Image generation method, apparatus, controller, and vehicle

By generating hybrid feature vectors in a vehicle environment, retrieving historical images in parallel, and dynamically updating a lightweight model, combined with heterogeneous computing to optimize the inference process, the high latency and high resource consumption problems in vehicle-mounted image generation technology are solved, achieving fast, low-power, high-quality image generation.

CN122676004APending Publication Date: 2026-09-01ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610828401.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

Existing vehicle-mounted text-to-image generation technology suffers from high latency, low cache hit rate, and high computational resource consumption in resource-constrained environments, making it difficult to improve inference speed while ensuring generation quality.

Method used

By generating a hybrid feature vector that integrates image generation instructions and contextual information, and retrieving historical images in parallel from a multi-level caching system based on a similarity threshold, the lightweight generation model is dynamically updated. Combined with heterogeneous computing, the inference process is optimized, reducing computational resource consumption.

Benefits of technology

It significantly improves cache hit rate, reduces inference latency and computational resource consumption, while ensuring the semantic accuracy and aesthetic quality of generated images, achieving a balance between efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676004A_ABST
    Figure CN122676004A_ABST
Patent Text Reader

Abstract

The application provides an image generation method and device, a controller and a vehicle. The method comprises the following steps: receiving an image generation instruction and acquiring context information associated with the image generation instruction; generating a mixed feature vector according to the image generation instruction and the context information; based on the mixed feature vector, searching for a historical image meeting a preset condition from a multi-level cache system in parallel, wherein the preset condition comprises that the similarity between the mixed feature vector and a historical feature vector corresponding to the historical image is greater than a preset similarity threshold; if no historical image meeting the preset condition is searched from the multi-level cache system, acquiring an image generation element corresponding to the image generation instruction and an adaptation parameter corresponding to the image generation element; dynamically updating an initial generation model according to the adaptation parameter to obtain a target generation model; and generating a target output image through the target generation model according to the mixed feature vector. The application improves the cache hit rate and response speed of text-to-image, and reduces the computing power consumption and system power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to an image generation method, apparatus, controller, and vehicle. Background Technology

[0002] With the rapid development of Artificial Intelligence Generated Content (AIGC) technology, text-to-image (TPI) technology (i.e., the technology of automatically generating matching images based on user-input text descriptions) has been gradually applied to in-vehicle smart cockpits. For example, it can generate personalized dashboard backgrounds, wallpapers, or children's entertainment illustrations based on voice commands. Currently, mainstream TPI methods typically rely on large-scale diffusion models (such as Stable Diffusion XL, Flux, etc.), directly calling the model to complete inference by inputting image generation commands.

[0003] However, in resource-constrained environments such as automotive systems, existing technologies still have the following shortcomings: First, the limited video memory and system memory of automotive chips make it difficult to directly load and run diffusion models with billions of parameters, easily leading to memory overflow or excessively high inference latency. Second, to improve response efficiency in repetitive scenarios, some solutions introduce caching mechanisms, but most are based on exact text matching or hash matching. Due to the diversity of user expressions, semantically similar but differently expressed instructions are prone to cache mismatch, resulting in a low hit rate. The system still frequently triggers the complete local inference process, failing to effectively reduce the computational load. Third, the diffusion model requires multiple iterative calculations to generate images, each step involving a large number of attention mechanism operations and text encoding operations, resulting in repetitive computation, high energy consumption, and high heat generation. This not only affects system stability and user experience but may also shorten the lifespan of hardware. Summary of the Invention

[0004] In view of the above, it is necessary to propose an image generation method, device, controller and vehicle to solve the technical problems of high latency, high cache hit rate and high computational resource consumption in vehicle-mounted image generation in the prior art.

[0005] In a first aspect, this application provides an image generation method, the method comprising: receiving an image generation instruction and obtaining context information associated with the image generation instruction; generating a hybrid feature vector based on the image generation instruction and the context information; retrieving historical images that meet preset conditions in parallel from a multi-level cache system based on the hybrid feature vector, wherein the preset conditions include the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image being greater than a preset similarity threshold; if no historical image meeting the preset conditions is found in the multi-level cache system, obtaining image generation elements corresponding to the image generation instruction and adaptation parameters corresponding to the image generation elements; dynamically updating an initial generation model based on the adaptation parameters to obtain a target generation model; and generating a target output image based on the hybrid feature vector and the target generation model.

[0006] In the image generation method of this application embodiment, by generating a hybrid feature vector that integrates image generation instructions and contextual information, and retrieving historical images in parallel from a multi-level caching system based on a similarity threshold, it can effectively identify user requests with similar semantics but diverse expressions, significantly improving the cache hit rate. When a cache hit occurs, historical images can be directly reused, avoiding the time-consuming local inference process, thereby significantly improving the overall response speed and reducing computational resource consumption. Secondly, when the cache misses, by first extracting image generation elements and corresponding adaptation parameters, and then dynamically updating the initial generation model based on these parameters, a lightweight target generation model is obtained. This avoids running the complete original large model, reducing the scale of parameters and computational steps involved in inference, thus significantly reducing memory usage and inference latency while ensuring the quality of the generated image, and alleviating energy consumption and heat generation issues. Finally, the integration of contextual information into the hybrid feature vector makes the generated target output image more consistent with the user's true intent and scenario requirements. Based on this, combined with the dynamically updated target generation model, it is possible to reduce computational overhead without sacrificing the semantic accuracy and aesthetic quality of the image, achieving a balance between efficiency and quality. In summary, this application effectively solves the problems of high image generation latency, high cache hit rate, and high computational resource consumption in existing vehicle-mounted text-to-image generation technologies, and achieves the technical effect of significantly improving inference speed and reducing resource consumption while ensuring generation quality.

[0007] In some embodiments of this application, the multi-level caching system includes an in-vehicle cache, an edge cloud cache, and a central cloud cache. The method further includes: if a historical image that meets the preset conditions is retrieved in at least one level of the multi-level caching system, a target output image is determined from the historical images that meet the preset conditions using a preset image scheduling strategy.

[0008] In some embodiments of this application, generating a hybrid feature vector based on the image generation instruction and the context information includes: using a preset text encoder to perform semantic encoding on the image generation instruction to obtain a text semantic vector; generating a visual feature vector based on the context information using a preset visual feature encoder, wherein the context information includes the user's visual preferences and / or scene information of the vehicle's location; and fusing the text semantic vector and the visual feature vector to obtain the hybrid feature vector.

[0009] In some embodiments of this application, the historical feature vector includes a historical text feature vector and a historical visual feature vector. The method further includes: obtaining a first similarity between the text semantic vector and the historical text feature vector, and a second similarity between the visual feature vector and the historical visual feature vector; weighting the first similarity according to a preset first weight coefficient to obtain a weighted first similarity, and weighting the second similarity according to a preset second weight coefficient to obtain a weighted second similarity; and using the sum of the weighted first similarity and the weighted second similarity as the similarity between the mixed feature vector and the historical feature vector.

[0010] In some embodiments of this application, the method further includes: during the first denoising iteration, caching the key matrix and value matrix output by the text encoder; and during subsequent denoising iterations, reusing the cached key matrix and value matrix for attention operations.

[0011] In some embodiments of this application, the method further includes: assigning the computational task of the attention operation to a neural network processing unit or a digital signal processing unit for execution; and assigning the decoding task of converting the latent space features output by the target generation model into pixel images through a preset decoder to a graphics processing unit for execution.

[0012] In some embodiments of this application, the adaptation parameters include a low-rank weight matrix. Obtaining the image generation elements corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation elements includes: identifying the target entity and / or style descriptor in the image generation instruction to obtain the image generation elements; obtaining the low-rank weight matrix corresponding to the image generation elements from the multi-level caching system; and dynamically updating the linear layer weights corresponding to the initial generation model based on the low-rank weight matrix to obtain the target generation model.

[0013] In some embodiments of this application, the method further includes: storing the target output image and the corresponding hybrid feature vector as new cache entries in the multi-level cache system.

[0014] Secondly, this application also provides an image generation apparatus, the apparatus comprising: an information acquisition module, configured to receive an image generation instruction and acquire context information associated with the image generation instruction; a feature generation module, configured to generate a hybrid feature vector based on the image generation instruction and the context information; an image retrieval module, configured to retrieve historical images that meet preset conditions in parallel from a multi-level caching system based on the hybrid feature vector, wherein the preset conditions include the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image being greater than a preset similarity threshold; a parameter acquisition module, configured to acquire image generation elements corresponding to the image generation instruction and adaptation parameters corresponding to the image generation elements if no historical image meeting the preset conditions is retrieved in the multi-level caching system; a model update module, configured to dynamically update an initial generation model based on the adaptation parameters to obtain a target generation model; and an image generation module, configured to generate a target output image based on the hybrid feature vector and the target generation model.

[0015] Thirdly, this application also provides a controller, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the image generation method described in the above embodiments.

[0016] Fourthly, this application also provides a vehicle, the vehicle including a controller as described in the above embodiments, the controller being configured to perform the image generation method as described in the above embodiments.

[0017] Understandably, the image generation apparatus of the second aspect, the controller of the third aspect, and the vehicle of the fourth aspect provided above all correspond to the image generation method of the first aspect. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding image generation methods provided above, and will not be repeated here. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the application environment of an image generation method provided in an embodiment of this application.

[0019] Figure 2 This is a schematic flowchart of an image generation method provided in an embodiment of this application.

[0020] Figure 3 This is a detailed flowchart illustrating step S20 of an image generation method provided in an embodiment of this application.

[0021] Figure 4 This is a detailed flowchart illustrating step S30 of an image generation method provided in an embodiment of this application.

[0022] Figure 5 This is a detailed flowchart illustrating step S60 of an image generation method provided in an embodiment of this application.

[0023] Figure 6 This is a sequence diagram corresponding to an image generation method provided in an embodiment of this application.

[0024] Figure 7 This is a schematic diagram of the functional modules of an image generation apparatus provided in an embodiment of this application.

[0025] Explanation of main component symbols Vehicle 1 Controller 10 Memory 11 Processor 12 Image generating device 100 Information Acquisition Module 110 Feature generation module 120 Image retrieval module 130 Parameter acquisition module 140 Model update module 150 Image generation module 160 The following detailed description, in conjunction with the accompanying drawings, will further illustrate this application. Detailed Implementation

[0026] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] To provide a clearer understanding of the embodiments of the present invention, the invention will be described in detail below with reference to the accompanying drawings and specific examples. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0028] Please see Figure 1 This is a schematic diagram illustrating an application scenario of the image generation method provided in an embodiment of this application.

[0029] The image generation method of this application embodiment can be applied to vehicle 1 (Permanent Magnet Synchronous Motor, PMSM) and executed by one or more controllers 10.

[0030] Specifically, the controller 10 is typically integrated into the intelligent cockpit system of the vehicle 1, serving as the core computing unit to support the real-time operation of the in-vehicle text and image generation function.

[0031] In some embodiments of this application, the controller 10 can be deployed within the smart cockpit domain controller, utilizing the existing high-performance computing chip resources of the smart cockpit domain controller to perform image generation tasks, thereby avoiding additional hardware costs. Alternatively, in luxury vehicles with higher computing power requirements, the controller 10 can also be an independently configured high-performance AI computing module, specifically responsible for complex generative AI inference tasks.

[0032] In other embodiments, the controller 10 may be in the form of an in-vehicle infotainment host or even a single computing node in a cloud server cluster. As long as it has the ability to store and execute programs to implement the image generation method of this application, it falls within the protection scope of this application.

[0033] In terms of application scenarios, when vehicle 1 is in parking rest mode, camping mode or children's entertainment mode, the user can initiate an image generation command through the voice assistant or touch screen, and the controller 10 will then call the image generation method provided in this application to quickly generate dashboard backgrounds, center console wallpapers, story illustrations, etc. that match the current atmosphere.

[0034] Based on this, this application not only solves the resource bottleneck problem of large model inference in the vehicle environment, but also provides intelligent vehicles with a differentiated emotional interaction experience, making vehicle 1 no longer just a means of transportation, but an intelligent mobile space that can understand user intentions and provide personalized visual services.

[0035] Please see Figure 2 This is a schematic flowchart of an image generation method provided in an embodiment of this application.

[0036] Specifically, the image generation method includes the following steps. Depending on different needs, the order of some steps in the flowchart can be changed, and some steps can be omitted.

[0037] Step S10: Receive the image generation instruction and obtain the context information associated with the image generation instruction.

[0038] Image generation instructions are used to instruct the generation of a target output image that matches the user's description. These instructions can consist of text descriptions input by the user via voice, touch, or gestures. Image generation instructions include, but are not limited to, image generation elements such as entity objects, attribute features, and style elements. For example, an image generation instruction could be "Draw a speeding red Ferrari," which includes the entity object "sports car," the color attribute "red," and the style element "Ferrari."

[0039] Contextual information refers to implicit data, besides explicit image generation instructions, that helps understand the user's true intentions or the current environmental state. Specifically, contextual information can include, but is not limited to, the following two categories: the user's historical generation preferences, such as previously selected image tones (e.g., preference for cool or warm tones), previously selected image styles (e.g., preference for anime or realistic styles), and commonly used wallpaper types (e.g., landscape, science fiction, abstract, etc.). The controller can obtain the user's historical generation preferences by analyzing the user's historical interaction records or by loading personalized settings bound to the user's in-vehicle account; and the vehicle's current state and scene information, such as the current color of the ambient lighting, the vehicle's speed, and external weather conditions. The controller can collect the vehicle's current state and scene information in real time through in-vehicle sensors (e.g., vision sensors, speed sensors, etc.).

[0040] Based on this, this application introduces contextual information so that subsequent image generation no longer relies solely on isolated text instructions, but can perceive user preferences and the current scene, thereby providing a data foundation for generating images that meet the user's potential expectations.

[0041] Step S20: Generate a hybrid feature vector based on the image generation instructions and context information.

[0042] Traditional text-based image retrieval often relies solely on text hashing or exact matching, which can lead to requests that are semantically similar but have different literal expressions being judged as unrelated. For example, "a Ferrari speeding" and "a red race car" are considered completely different requests, resulting in an extremely low cache hit rate.

[0043] Based on this, this embodiment generates a hybrid feature vector by fusing textual semantic information and contextual information, forming a high-dimensional, unified composite cache retrieval key.

[0044] Specifically, when the controller generates a hybrid feature vector based on the image generation instructions and context information, it uses a preset text encoder to semantically encode the image generation instructions to obtain a text semantic vector; based on the context information, it uses a preset visual feature encoder to generate a visual feature vector, where the context information includes the user's visual preferences and / or the scene information of the vehicle's location; and it fuses the text semantic vector and the visual feature vector to obtain a hybrid feature vector.

[0045] It should be noted that how to generate the mixed feature vector based on image generation instructions and context information will be discussed later. Figure 3 The embodiments shown are described in detail, and will not be repeated here to avoid repetition.

[0046] Based on the above embodiments, this application can not only capture explicit semantic content, but also characterize implicit generation preferences and environmental adaptability, thereby effectively bridging the semantic gap between user expression and visual results and significantly improving the recall accuracy of subsequent searches.

[0047] Step S30: Based on the hybrid feature vector, retrieve historical images that meet the preset conditions in parallel from the multi-level cache system.

[0048] Specifically, after generating the hybrid feature vector, in order to reuse the historical generation results as much as possible and avoid triggering high-cost model inference with each request, the controller can first try to search the multi-level cache system to see if there is a historical image that meets the preset conditions.

[0049] The multi-level caching system includes, but is not limited to, on-board cache (L1), edge cloud cache (L2), and central cloud cache (L3). On-board cache (L1) is typically located in the vehicle's local high-speed memory or solid-state drive, closest to the controller, and has the lowest access latency. It primarily stores frequently used personalized data or recently generated images, belonging to the hot data layer. Edge cloud cache (L2) is deployed on roadside units or regional base station servers, serving fleets within a specific geographical area. It stores commonly generated content within this area, providing larger storage capacity than local cache while maintaining low network latency. Central cloud cache (L3) is located in a remote data center, possessing massive storage space and a complete historical generation record. Although network access latency is relatively high, it provides the highest recall coverage.

[0050] It should be understood that although this embodiment uses a three-level cache as an example, in actual applications, the number of cache levels can be adjusted according to specific deployment requirements. For example, in offline mode, only the vehicle-mounted cache L1 can be retained, or in a strong network environment, the edge cloud cache L2 can be skipped and the central cloud cache L3 can be queried directly. As long as the selection logic based on maximizing similarity is followed, it falls within the protection scope of this invention.

[0051] The preset conditions include, but are not limited to, the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image being greater than a preset similarity threshold. The preset similarity threshold can be dynamically adjusted according to the actual application scenario. For example, the threshold can be appropriately increased in scenarios with extremely high image quality requirements, while the threshold can be appropriately decreased in fast preview scenarios to obtain a higher hit rate.

[0052] "Parallel retrieval" refers to the controller simultaneously sending query requests to various levels of the multi-level caching system, rather than the traditional step-by-step source lookup mode. This concurrency mechanism avoids the serial waiting latency caused by a cache miss at a certain level, ensuring that the globally optimal matching result can be obtained as quickly as possible, even in a weak network environment.

[0053] Based on this, this application, through this hierarchical storage mechanism, fully considers the multiple constraints between network bandwidth, response latency and storage capacity in the vehicle environment, and can flexibly schedule under different network conditions and resource conditions to achieve a dynamic balance between latency and capacity.

[0054] In traditional image retrieval methods, the index key is usually based on the hash value of the text image generation instruction or an exact string match of the user's input. However, in the context of in-vehicle scenarios, users' expressions are significantly diverse. For example, a user might first input "draw a speeding Ferrari," and then later input "draw a red race car" or "draw a fast sports car." Although these text expressions differ greatly in string space, their semantic intent and desired visual effect are highly similar. Traditional methods often suffer from extremely low cache hit rates because they cannot recognize this semantic similarity. Furthermore, users' potential aesthetic preferences (e.g., a preference for streamlined or boxy car models) are implicit in visual features, which pure text matching cannot capture. To address these issues, this embodiment proposes a similarity retrieval strategy based on hybrid feature vectors, specifically: When the controller obtains the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image, the historical feature vector includes historical text feature vector and historical visual feature vector. It obtains the first similarity between the text semantic vector and the historical text feature vector, and the second similarity between the visual feature vector and the historical visual feature vector. The first similarity is weighted according to a preset first weight coefficient to obtain a weighted first similarity, and the second similarity is weighted according to a preset second weight coefficient to obtain a weighted second similarity. The sum of the weighted first similarity and the weighted second similarity is taken as the similarity between the hybrid feature vector and the historical feature vector.

[0055] It should be noted that how to obtain the similarity between the mixed feature vector and the historical feature vectors corresponding to historical images will be discussed later. Figure 4 The embodiments shown are described in detail, and will not be repeated here to avoid repetition.

[0056] Based on the hybrid retrieval strategy of the above embodiments, even if the currently input text image generation instruction is not completely consistent with the original text image generation corresponding to the historical image in the multi-level caching system, as long as the semantic content and visual style of the two are similar, the controller can calculate a high similarity through the weighted fusion formula, thereby accurately recalling the high-quality images generated in the past, greatly improving the cache hit rate, and enabling a large number of requests to directly reuse historical images, saving the time-consuming local model inference process.

[0057] Step S40: Determine whether a historical image that meets the preset conditions has been retrieved in the multi-level caching system.

[0058] Specifically, when performing a retrieval, the controller employs a parallel query mechanism to simultaneously send requests to the three-level cache system, rather than the traditional sequential source-following mode. For example, the controller uses asynchronous input / output (I / O) interfaces to simultaneously send retrieval requests containing hybrid feature vectors to the vehicle-mounted cache L1, edge cloud cache L2, and central cloud cache L3, and calculates the similarity between the local data stored in each cache system and the hybrid feature vectors. If the similarity of any historical image exceeds a preset similarity threshold, that historical image is determined to meet the preset conditions.

[0059] Based on this, this concurrency mechanism ensures that even if a certain level of cache responds slowly or misses, it will not block the retrieval of results at other levels, thereby maximizing the use of distributed cache resources and significantly reducing the tail latency of the overall retrieval.

[0060] In some embodiments of this application, if no historical image meeting the preset conditions is found in the multi-level caching system, it indicates that there are not enough similar high-quality samples in the cache, and the controller continues to execute step S50.

[0061] In some embodiments of this application, if a historical image that meets preset conditions is retrieved in at least one level of the cache in a multi-level cache system, the controller proceeds to step S80.

[0062] Based on this, the judgment logic enables the reuse of historical images if they are matched, and the intelligent diversion using the model's precise generation if they are not matched, thus avoiding the waste of resources caused by blindly calling large models.

[0063] Step S50: Obtain the image generation elements corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation elements.

[0064] Specifically, the controller first performs structured parsing on the received image generation instructions to extract image generation elements. Image generation elements include, but are not limited to, entity objects, attribute features, style elements, etc. For example, an image generation instruction could be "Draw a speeding red Ferrari sports car," which includes the entity object "sports car," the color attribute "red," and the style element "Ferrari," among other image generation elements.

[0065] Next, the controller obtains the corresponding adaptation parameters based on the parsed image generation features. These adaptation parameters can be implemented using a lightweight Low-Rank Adaptation (LoRA) plugin. Each type of entity object or style feature is pre-trained with a very small LoRA weight file (typically only a few MB), which contains the offset parameters needed to adapt the initial generation model to the specific image features. The controller, through its internal LoRA manager, downloads the corresponding LoRA adapter from the local cache or cloud as needed based on the identified image generation features, and injects it into the initial generation model.

[0066] The initial generated model refers to the refined basic model pre-deployed in the vehicle's L1 cache. This model has undergone knowledge distillation and model quantization, has a small number of parameters (e.g., 1 / 5 of the original Stable Diffusion), and can reside in the vehicle's chip memory to adapt to hardware resource limitations.

[0067] Step S60: Dynamically update the initial generation model based on the adaptation parameters to obtain the target generation model.

[0068] "Dynamic update" refers to the model adaptation process that occurs in real time during inference, rather than offline training or reloading the entire model.

[0069] Based on this, this embodiment updates the base model by adapting parameters, which can temporarily endow the model with the ability to generate specific entities or styles without changing the base model weight file. This solves the practical constraint that the vehicle terminal cannot store massive full models, and enables small parameter models to have the ability to generate customized models on demand. This ensures the generation quality of the target generated model while significantly reducing model loading and storage overhead.

[0070] Specifically, the adaptation parameters include a low-rank weight matrix. When the controller dynamically updates the initial generation model based on the adaptation parameters to obtain the target generation model, it identifies the target entities and / or style descriptors in the image generation instructions to obtain image generation elements; it retrieves the low-rank weight matrix corresponding to the image generation elements from the multi-level caching system; and it dynamically updates the linear layer weights corresponding to the initial generation model based on the low-rank weight matrix to obtain the target generation model.

[0071] It should be noted that how to obtain the image generation elements corresponding to the image generation command and the corresponding adaptation parameters of the image generation elements will be discussed later. Figure 5 The embodiments shown are described in detail, and will not be repeated here to avoid repetition.

[0072] Step S70: Generate the target output image using the target generation model based on the mixed feature vector.

[0073] In this process, the hybrid feature vector serves not only as a key for retrieving the composite cache, but also as a conditional input for the diffusion generation process. The controller maps the hybrid feature vector to the key-value pairs required by the cross-attention layer in the target generation model, thereby guiding the target output model to focus on visual features related to the image generation instructions and contextual information in each denoising iteration.

[0074] For example, if the image generation element of "red sports car" and the contextual information of "preferring cool tones" are incorporated into the blended feature vector, the generation process will prioritize producing a sports car image with a red main tone and warm-cool lighting, rather than random colors or warm tones, ensuring that the newly generated image is highly consistent with the user's original intent in terms of semantics and style.

[0075] In summary, in steps S10 to S70 of the image generation method in this embodiment, by generating a hybrid feature vector that integrates image generation instructions and contextual information, and by retrieving historical images in parallel from a multi-level cache system based on a similarity threshold, it is possible to effectively identify user requests with similar semantics but diverse expressions, significantly improving the cache hit rate. When a cache hit occurs, historical images can be directly reused, avoiding the time-consuming local inference process, thereby significantly improving the overall response speed and reducing computational resource consumption. Secondly, when a cache miss occurs, by first extracting image generation elements and corresponding adaptation parameters, and then dynamically updating the initial generation model based on these parameters, a lightweight target generation model is obtained. This avoids running the complete original large model, reducing the scale of parameters and computational steps involved in inference, thereby significantly reducing memory usage and inference latency while ensuring the quality of the generated image, and alleviating energy consumption and heat generation issues. Finally, the integration of contextual information into the hybrid feature vector makes the generated target output image more consistent with the user's true intent and scenario requirements. Based on this, combined with the dynamically updated target generation model, it is possible to reduce computational overhead without sacrificing the semantic accuracy and aesthetic quality of the image, achieving a balance between efficiency and quality. In summary, this application effectively solves the problems of high image generation latency, high cache hit rate, and high computational resource consumption in existing vehicle-mounted text-to-image generation technologies, and achieves the technical effect of significantly improving inference speed and reducing resource consumption while ensuring generation quality.

[0076] Furthermore, in some embodiments of this application, considering the multi-step iterative characteristics of the diffusion model, a key-value cache (KV Cache) mechanism is introduced during the inference process to eliminate redundant computation. During image generation by the target generation model, dozens or even hundreds of denoising iterations are typically required to gradually recover clear latent space features from pure noise. Throughout the denoising loop, the mixed feature vector, which serves as the conditional input, remains constant. However, traditional methods re-encode the same mixed feature vector using the text encoder in each iteration and recalculate the key matrix (K) and value matrix (V) required for cross-attention. This redundant computation causes significant resource waste in scenarios with limited onboard computing power, not only prolonging the generation time of a single image but also causing the chip to operate under continuous high load, leading to problems such as overheating and frequency throttling.

[0077] To address the aforementioned issues, the image generation method further includes: during the first denoising iteration, caching the key matrix and value matrix output by the text encoder; and during subsequent denoising iterations, reusing the cached key matrix and value matrix for attention operations.

[0078] Specifically, in the first iteration (Step 1) of the denoising process, the controller drives the text encoder to perform a complete forward propagation calculation on the mixed feature vectors, obtaining the corresponding key matrix K and value matrix V, and writes these two matrices into a preset cache. This cache can be a dedicated tensor region in the graphics processing unit (GPU) memory, or a shared buffer in random access memory (RAM) mapped using zero-copy technology.

[0079] Subsequent iterations: From the second iteration (Step 2) until the final iteration (Step N), the controller no longer triggers the recalculation logic of the text encoder. Instead, it directly reads the pre-stored key matrix K and value matrix V from the cache and injects them into the cross-attention layer of the U-shaped Network (UNet) or Transformer backbone network for computation. Since the query matrix (Q) changes dynamically with each denoising step, only Q needs to be calculated in real time, and the reuse of K and V makes the computational cost on the text encoding side of each iteration approach zero.

[0080] In other real-time examples, when recalculating attention weights in each iteration, the controller can also scale the cached K and V based on the dynamically changing guidance strength in each iteration to adapt to different generation requirements.

[0081] Based on the above embodiments, this application significantly reduces the inference computing power requirements of the target generation model on the vehicle side through this caching mechanism without changing the model structure or losing the generation quality, which helps to achieve low power consumption and low heat generation of real-time images.

[0082] Furthermore, in some embodiments of this application, based on the aforementioned caching mechanism to reduce redundant computation, this application also utilizes the heterogeneous computing characteristics of the in-vehicle computing platform to optimize the operator execution efficiency in a single inference task. Specifically, the image generation method further includes: assigning the attention operation computation task to a neural network processing unit or a digital signal processing unit for execution; and assigning the decoding task of converting the latent space features output by the target generation model into pixel images through a preset decoder to a graphics processing unit for execution.

[0083] Specifically, during the inference process of the target generation model, attention operations (including self-attention and cross-attention) account for the vast majority of the computational load, essentially large-scale matrix multiplication operations. These operations exhibit high parallelism and regular memory access patterns, and are relatively insensitive to numerical precision, making them suitable for execution on Neural Processing Units (NPUs) or Digital Signal Processors (DSPs). NPUs and DSPs typically employ systolic arrays or dedicated vector processing architectures, achieving significantly higher throughput per unit power consumption than general-purpose GPUs when performing low-precision matrix operations such as INT8 or FP16. Therefore, the controller offloads the attention operation computation task to the NPU or DSP for execution, thereby significantly reducing overall power consumption and heat generation during the inference process.

[0084] Furthermore, the process of converting the latent space features output by the target generation model into the final pixel image is typically performed by the decoder of a Variational Autoencoder (VAE). The VAE decoding process involves numerous deconvolution, upsampling, and non-linear activation operations, requiring high precision in floating-point operations and exhibiting complex memory access patterns. Forcing its execution on a low-precision NPU could lead to significant image quality loss or artifacts. GPUs, on the other hand, possess powerful floating-point capabilities and mature texture processing units, enabling them to perform these graphics rendering-related tasks with high precision and efficiency. Therefore, the controller retains the decoding task for execution on the GPU, ensuring the visual quality of the generated image while avoiding the additional overhead caused by precision conversion.

[0085] In other embodiments, RAM serves as the physical storage medium for the vehicle-mounted cache L1 and the model loading space, used to store intermediate data and support the smooth operation of the above-mentioned calculation process.

[0086] Based on the above embodiments, this application utilizes the heterogeneous computing characteristics typically found in vehicle computing platforms to map different types of operators to the hardware units that are best suited for them. This effectively solves the performance bottleneck of a single hardware unit when handling mixed loads, and significantly improves the response speed and generation quality of vehicle-mounted text and images while ensuring the quality of the generated images.

[0087] Furthermore, in some embodiments of this application, based on the collaborative work of the NPU and GPU, the controller drives the diffusion sampler to perform multi-step denoising iterations to gradually generate a latent space image; subsequently, the VAE decoder decodes the latent space features into the final visible RGB image, and performs post-processing as needed (e.g., size adaptation, contrast fine-tuning). The final target output image can be displayed on in-vehicle central control screens, dashboards, etc., to meet the needs of personalized wallpaper generation, children's entertainment illustrations, or atmosphere enhancement.

[0088] Based on this, this application uses the hybrid feature vector as both the retrieval key and the generation condition, and combines the KV caching mechanism with the heterogeneous computing offloading strategy to achieve fast, accurate, and low-power high-quality image generation under the constraints of vehicle resources.

[0089] S80: Using a preset image scheduling strategy, determine the target output image from historical images that meet preset conditions.

[0090] In some embodiments of this application, considering the differences in response latency between different cache levels (for example, the on-board cache L1 typically returns in milliseconds, while the edge cloud cache L2 or the central cloud cache L3 may require tens to hundreds of milliseconds due to network fluctuations), this application introduces a reinforcement learning adaptive scheduler. The adaptive scheduler monitors network latency, available bandwidth, computing load, battery level, and hardware temperature in real time, and combines the monitoring results with image scheduling strategies to determine the target output image, as detailed below: Strategy 1: If at least one historical image meeting preset conditions is found in the vehicle-mounted cache L1, regardless of network conditions, the controller immediately retrieves the historical image with the highest similarity from the multiple historical images in the vehicle-mounted cache L1 as the target output image. Since the vehicle-mounted cache L1 requires no network transmission and has extremely low access latency, it places no additional load on the hardware and is considered the optimal response.

[0091] Strategy 2: If the on-board cache L1 does not retrieve any historical images that meet the preset conditions, but the edge cloud cache L2 or the central cloud cache L3 contains at least one historical image that meets the preset conditions, the adaptive scheduler will make a comprehensive decision based on real-time monitoring data: 2.1 If the current network status is good and the hardware temperature and battery power are normal, the controller pulls the historical image with the highest similarity from remote caches such as edge cloud cache L2 or central cloud cache L3 as the target output image.

[0092] 2.2 If the current network status is detected to be poor (e.g., high latency, low bandwidth) but the hardware temperature and battery level are normal, the controller abandons pulling from remote caches such as edge cloud cache L2 or central cloud cache L3, and switches to local model inference, i.e., steps S50~S70, to avoid users waiting for a long time.

[0093] 2.3 If the current hardware temperature is detected to be too high and / or the battery power is too low, even if the network status is good, the controller will abandon local model inference and instead force the pull of the most similar historical image from remote caches such as edge cloud cache L2 or central cloud cache L3 as the target output image or return a downgrade prompt, in order to protect the hardware and extend the battery life.

[0094] Based on the above embodiments, by distinguishing between local cache hits and remote cache hits, and combining a reinforcement learning-based adaptive scheduler to monitor network latency, available bandwidth, computing load, battery level, and hardware temperature in real time, the controller can ensure both response speed and hardware health. Specifically, local cache hits return the fastest, without waiting; remote cache is only fetched when the network is good and the hardware is in normal condition, avoiding long waiting times for users; when the network is poor, the controller actively switches to local inference to prevent failures due to transmission timeouts; when the hardware temperature is too high or the battery level is too low, even if local inference might be faster, the controller will force a remote fetch or return a downgrade prompt, thereby avoiding high-load computation from exacerbating hardware wear or shortening battery life. Through multi-level, multi-condition adaptive scheduling, the controller achieves rapid response, hardware friendliness, and quality priority in resource-constrained vehicle environments, effectively solving the pain point in existing technologies where cache hits are not efficiently utilized due to network or hardware issues.

[0095] In an optional example, to balance response speed and generation quality, the controller can pre-set a response duration threshold for retrieval time (e.g., 50 milliseconds). During periods when the retrieval time is less than the response duration threshold: If the similarity of the historical image returned by L1 cached by the vehicle terminal exceeds the preset similarity threshold (e.g., 0.95), then this historical image is selected as the target output image, and the response from other levels is no longer waited for, so as to ensure image quality as much as possible while taking into account output speed.

[0096] If the vehicle-mounted cache L1 does not return a result with a similarity exceeding the preset similarity threshold within the response time threshold, but there is at least one historical image in the edge cloud cache L2 or the central cloud cache L3 that meets the preset conditions, then the adaptive scheduler makes a comprehensive decision based on real-time monitoring data. See the above for details, and will not be repeated here to avoid repetition.

[0097] For example, regarding strategy 2.3, if the on-board cache L1 does not retrieve a historical image that meets the preset conditions, but a historical image A with a similarity of 0.96 exists in remote caches such as edge cloud cache L2 or central cloud cache L3, and the network latency is only 20ms, but the current GPU temperature has reached 85℃, which is abnormal. After comprehensive consideration, the adaptive scheduler, to avoid running local inference or high-load network fetching under high temperature (the fetching process requires decoding, which still has a certain load), chooses to fetch historical image A from remote caches such as edge cloud cache L2 or central cloud cache L3 as the target output image and processes it at a limited rate or returns a simple placeholder image.

[0098] Based on this, the balance between response latency and generation quality was optimized to prevent the main process from being blocked due to waiting for slow remote responses.

[0099] S90: Store the target output image and the corresponding hybrid feature vector as new cache entries in the multi-level cache system.

[0100] Specifically, once the target output image is generated, the controller does not discard the intermediate calculation results directly. Instead, it binds the target output image with the blended feature vector on which the image was generated, forming a complete cache entry. The blended feature vector serves as the index key for subsequent searches, and the target output image serves as the return value when a search is successful.

[0101] Based on this, a feedback loop is formed in the technical solution of this application, which has the self-evolutionary ability to improve performance with the increase of usage time. This ensures that when the user makes the same or similar semantic / visual style request again, the controller can directly hit the cache entry through the parallel retrieval mechanism in the aforementioned embodiment, thereby skipping the time-consuming model inference process and achieving millisecond-level response.

[0102] In some embodiments of this application, for newly generated cache entries, the controller performs a write operation according to a preset storage strategy, specifically: Prioritize writing to the vehicle-side cache L1: The target output image and its corresponding blended feature vector are immediately written to the vehicle-side cache L1 so that the same or similar requests can be hit with minimal latency the next time they occur.

[0103] Synchronous or asynchronous updates to edge cloud cache L2: Based on the current status (e.g., network bandwidth, load) and cache hotness policy, the controller synchronously or asynchronously updates cache entries to edge cloud cache L2 to expand the cache coverage for vehicles in the same area to share.

[0104] Selective updates to the central cloud cache L3: For frequently accessed or high-value cache entries (e.g., similarity hit counts exceeding the hit count threshold), the controller can further synchronize this cache entry to the central cloud cache L3 for global sharing.

[0105] Based on the above embodiments, the cache update mechanism of "prioritizing local and spreading step by step" enables newly generated target output images to quickly participate in subsequent retrieval, continuously improves the cache hit rate, and realizes adaptive collaboration of the three-level cache of end-edge-cloud.

[0106] Please see Figure 3 This is a detailed flowchart illustrating step S20 of an image generation method provided in an embodiment of this application.

[0107] This embodiment is a detailed description of the foregoing embodiment, further illustrating how to generate a hybrid feature vector based on image generation instructions and context information. Specifically, it includes the following steps: Step S21: Using a preset text encoder, semantically encode the image generation instructions to obtain a text semantic vector.

[0108] The text encoder can be a pre-trained language model such as Contrastive Language-Image Pre-Training (CLIP) or Bidirectional Encoder Representations from Transformers (BERT). The input is natural language instructions that have been segmented, and the output is a high-dimensional text semantic vector that represents the explicit semantic content.

[0109] Step S22: Based on contextual information, generate visual feature vectors using a preset visual feature encoder.

[0110] Specifically, the input to the visual feature encoder is not direct image pixels, but rather a structured or unstructured description derived from contextual information. For example, when the contextual information includes the scene information "the vehicle is currently in motion mode and the ambient lighting is red," the controller first maps it to a standard sequence of visual cues or an embedded representation, and then inputs it to the visual feature encoder (such as the projection layer of CLIP's visual encoder or a dedicated scene embedding network) to generate a visual feature vector that can characterize the current environmental atmosphere and potential visual style.

[0111] Step S23: Fuse the text semantic vector and the visual feature vector to obtain a hybrid feature vector.

[0112] Specifically, the controller fuses text semantic vectors and visual feature vectors using either weighted concatenation or weighted summation. If weighted concatenation is used, the resulting hybrid feature vector is obtained through feature concatenation: in, This represents the feature vector of the currently input text; This represents the current input visual feature vector.

[0113] Based on the above embodiments, by converting non-visual contextual information into visual feature vectors and fusing them into hybrid feature vectors, subsequent retrieval can not only match text descriptions, but also automatically adapt to the current cabin environment and user aesthetic habits, effectively solving the semantic gap problem caused by expression differences in traditional methods.

[0114] Please see Figure 4 This is a detailed flowchart illustrating step S30 of an image generation method provided in an embodiment of this application.

[0115] This embodiment is a detailed description of the foregoing embodiments, further illustrating how to obtain the similarity between the mixed feature vector and the historical feature vector corresponding to the historical image. Specifically, it includes the following steps: Step S31: Obtain the first similarity between the text semantic vector and the historical text feature vector, and the second similarity between the visual feature vector and the historical visual feature vector.

[0116] Among them, the historical text feature vector and the historical visual feature vector are pre-stored when the historical image is generated or cached and written, and correspond one-to-one with the historical image.

[0117] Specifically, similarity can be calculated using various metrics. For example, the formula for calculating cosine similarity is as follows: in, For similarity function, and Let represent the two feature vectors to be compared, and let represent the range of values ​​for the cosine similarity. The larger the value, the more similar the two numbers are.

[0118] In other embodiments, similarity can also be calculated using the dot product (applicable to normalized vectors) or the reciprocal of the Euclidean distance (the smaller the distance, the higher the similarity).

[0119] Based on this, the first similarity is denoted as The second similarity is denoted as ,in, In a multi-level cache system, the first... The historical text feature vector corresponding to each historical image Indicates the first The historical visual feature vector corresponding to each historical image.

[0120] Step S32: Weight the first similarity according to the preset first weight coefficient to obtain the weighted first similarity, and weight the second similarity according to the preset second weight coefficient to obtain the weighted second similarity.

[0121] Specifically, the weighted first similarity is The weighted second similarity is .

[0122] in, As the first weighting coefficient, As the second weighting coefficient, it usually satisfies The normalization constraint is used to adjust the influence ratio of text modality and visual modality on the final similarity.

[0123] It should be understood that and These are not fixed hyperparameters, but rather strategy variables that can be dynamically adjusted based on the actual application scenario or user settings. For example, when the vehicle is traveling at high speed, the user is more sensitive to response speed than to the accuracy of image style; in this case, the hyperparameters can be appropriately increased. (That is, it relies more on text semantic matching to return results quickly); when the vehicle is stationary, it can increase (That is, it relies more on visual preference matching in order to pursue higher generation quality).

[0124] Step S33: The sum of the weighted first similarity and the weighted second similarity is used as the similarity between the mixed feature vector and the historical feature vector.

[0125] Specifically, the weighted fusion formula is as follows: in, This indicates the similarity between the mixed feature vector and the historical feature vector; In a multi-level cache system, the first... The historical text feature vector corresponding to each historical image Indicates the first The historical visual feature vector corresponding to each historical image is the index key in the multi-level caching system.

[0126] Based on the hybrid retrieval strategy described in the above embodiments, even if the currently input text-image generation instruction is not completely consistent with the original text-image generation corresponding to the historical image in the multi-level caching system, as long as their semantic content and visual style are similar, the controller can calculate a high similarity using a weighted fusion formula. This allows for accurate retrieval of historically generated high-quality images, significantly improving the cache hit rate and enabling a large number of requests to directly reuse historical images, saving time-consuming local model inference processes. Furthermore, through this dynamic weighting mechanism, the effectiveness of cache hits can be maximized by utilizing contextual information while ensuring semantic relevance, avoiding the rigidity problem caused by single-dimensional matching, and significantly enhancing the intelligent perception capability and user experience of in-vehicle text-to-image generation.

[0127] Please see Figure 5 This is a detailed flowchart illustrating step S60 of an image generation method provided in an embodiment of this application.

[0128] This embodiment is a detailed description of the foregoing embodiments, further illustrating how to dynamically update the initial generation model based on adaptation parameters to obtain the target generation model. Specifically, it includes the following steps: Step S61: Identify the target entity and / or style descriptor in the image generation instruction to obtain image generation elements.

[0129] Specifically, before executing the update process, the controller first needs to perform structured parsing of the user's natural language instructions. The controller uses natural language processing technology to accurately extract key information from the image generation instructions. For example, when a user inputs "draw a speeding red Ferrari," the controller will identify image generation elements including the entity "sports car," the color attribute "red," and the style element "Ferrari."

[0130] Step S62: Obtain the low-rank weight matrix corresponding to the image generation elements from the multi-level caching system.

[0131] Specifically, based on the identified image generation elements, the corresponding low-rank weight matrix is ​​retrieved in a multi-level caching system. This reuses the advantages of the edge-cloud three-level architecture from the previous embodiment. Since the low-rank weight matrix is ​​very small (typically only a few MB), the controller can prioritize searching for the low-rank weight matrix of commonly used styles or entity objects in the on-board cache L1; if it is not found locally, it can be retrieved from the edge cloud cache L2 or the central cloud cache L3 in milliseconds, without the significant waiting delay that occurs when downloading the full model.

[0132] The low-rank weight matrix is ​​implemented as a LoRA plugin. Each type of entity or style feature is pre-trained with a very small LoRA weight file (usually only a few MB), which contains the offset parameters needed to adapt the initial generative model to the specific image features.

[0133] Based on this, this design enables the vehicle-mounted device to access an unlimited model library in the cloud while maintaining the low latency characteristics of local inference.

[0134] Step S63: Based on the low-rank weight matrix, dynamically update the linear layer weights corresponding to the initial generative model to obtain the target generative model.

[0135] Specifically, the controller decomposes the low-rank weight matrix into the product of two low-rank matrices and adds it to the original weights of the corresponding linear layer of the initial generation model. That is, the target weight equals the base weight plus the low-rank incremental weight. This fusion process can be completed in one step during the model loading phase or calculated in real-time during inference via a side branch. Regardless of the method used, the result is to instantly endow the model with the ability to generate specific entities or follow specific styles without changing the general capabilities of the base model. After the generation task is completed, the injected low-rank weights can be unloaded or replaced, and the base model returns to its original state, thus achieving on-demand assembly and flexible switching of model capabilities.

[0136] Based on the above embodiments, by introducing LoRA technology, the abstract dynamic update process is visualized as a lightweight parameter injection mechanism. In scenarios where on-board computing resources are limited, it is impractical to pre-configure a complete full model for every possible generation requirement. Not only is storage space insufficient, but the loading time during model switching is also unacceptable. However, the low-rank weight matrix used in this embodiment typically has a file size that is only one-thousandth or even one-ten-thousandth of the full model. This allows the controller to maintain a massive style and entity adaptation library with extremely low storage and network costs, effectively resolving the contradiction between on-board storage capacity and generation diversity.

[0137] Please see Figure 6 This is a sequence diagram of an image generation method provided in an embodiment of this application.

[0138] To better understand the vehicle-mounted image generation method of this application, through Figure 6 The following is a simplified sequence diagram to illustrate the in-vehicle image generation method. Specifically: like Figure 6 As shown, the user inputs the image generation command "Draw a speeding red sports car". The controller generates a hybrid feature vector based on the image generation command and context information, and initiates a parallel search in the multi-level cache system to detect the existence of similar historical images. If cache entry A (the original image generation command is "speeding red sports car") is found in the vehicle-side cache L1, and its similarity to the currently input image generation command is as high as 0.94, which is less than the preset similarity threshold of 0.95, i.e., the preset condition is not met, the controller decides to follow the local model inference path: input the hybrid feature vector, directly call the target generation model for inference, and skip the multi-level cache system. After the inference is completed, the target output image is generated, and the target output image and the corresponding hybrid feature vector are stored in the multi-level cache system. Finally, the generated "speeding red sports car" image is displayed to the user.

[0139] Please see Figure 7 This is a schematic diagram of the functional modules of an image generation apparatus 100 provided in an embodiment of this application.

[0140] In this embodiment, based on the above... Figure 2 The image generation method in the illustrated embodiments follows the same concept. This application also provides an image generation apparatus 100, which can be used to perform the above-described image generation method. For ease of explanation, the schematic diagram of the image generation apparatus 100 embodiment only shows the parts related to the embodiments of this application. Those skilled in the art will understand that the illustrated structure does not constitute a limitation on the image generation apparatus 100, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0141] Specifically, the image generation apparatus 100 provided in this application embodiment includes an information acquisition module 110, a feature generation module 120, an image retrieval module 130, a parameter acquisition module 140, a model update module 150, and an image generation module 160. The information acquisition module 110 receives an image generation instruction and acquires context information associated with the image generation instruction; the feature generation module 120 generates a hybrid feature vector based on the image generation instruction and the context information; the image retrieval module 130 retrieves historical images that meet preset conditions in parallel from a multi-level cache system based on the hybrid feature vector, wherein the preset conditions include a similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image being greater than a preset similarity threshold; the parameter acquisition module 140 acquires the image generation elements corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation elements if no historical image meeting the preset conditions is found in the multi-level cache system; the model update module 150 dynamically updates the initial generation model according to the adaptation parameters to obtain a target generation model; and the image generation module 160 generates a target output image based on the hybrid feature vector and the target generation model.

[0142] For specific limitations regarding the image generation apparatus 100, please refer to the limitations on the image generation method above, which will not be repeated here. Each module in the image generation apparatus 100 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the controller, or stored in software in the memory of the controller, so that the processor can call and execute the operations corresponding to each module.

[0143] Following the example above, please refer to Figure 1 A controller is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0144] The controller 10 provided in this application embodiment includes, but is not limited to, a memory 11, a processor 12, and a computer program stored in the memory 11 and executable on the processor 12, such as an image generation program. When the computer program is executed by the processor 12, it implements the image generation method as described in the above embodiment.

[0145] Figure 1Only the controller 10, which includes a memory 11 and a processor 12, is shown. Those skilled in the art will understand that... Figure 1 The structure shown does not constitute a limitation on the controller 10 and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0146] In some embodiments of this application, the controller 10 can be communicatively connected to devices such as desktop computers, laptops, handheld computers, and cloud servers.

[0147] In some embodiments of this application, the controller 10 can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0148] In some embodiments of this application, the controller 10 may further include network devices and / or client devices. These network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, and a cloud server based on cloud computing, consisting of a large number of hosts or network servers.

[0149] In some embodiments of this application, the network where the controller 10 is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0150] In some embodiments of this application, memory 11 stores multiple computer-readable instructions to implement an image generation method, and processor 12 can execute multiple instructions to implement the image generation method as described in the above embodiments.

[0151] Specifically, the processor 12's implementation method for the above instructions can be found in the description of the relevant steps in the above embodiments, and will not be repeated here.

[0152] Those skilled in the art will understand that the schematic diagram is merely an example of the controller 10 and does not constitute a limitation on the controller 10. The controller 10 can be a bus topology or a star topology. The controller 10 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, the controller 10 may also include an input / output controller 10, a network access device, etc.

[0153] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 1 The symbol is represented by only one arrow, but this does not mean that there is only one bus or one type of bus. The bus is configured to implement communication between memory 11 and processor 12, etc.

[0154] It should be noted that controller 10 is only an example. Other existing or future electronic products that are suitable for this application should also be included within the scope of protection of this application and are incorporated herein by reference.

[0155] In some embodiments of this application, the processor 12 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 12 is the control core of the controller 10, connecting various components of the controller 10 via various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., executing an image generation program) and calls data stored in the memory 11 to perform various functions of the controller 10 and process data.

[0156] The processor 12 executes the operating system of the controller 10 and various installed applications. The processor 12 executes these applications to implement the steps described in each of the above-described image generation method embodiments, for example... Figure 1 The steps are shown.

[0157] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 11 and executed by processor 12 to complete this application. One or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in controller 10. For example, the computer program may be divided into the modules shown in the above embodiments.

[0158] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute a portion of an image generation method according to various embodiments of this application.

[0159] If the modules / units integrated into controller 10 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0160] Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory, and other types of memory.

[0161] This application also provides a computer-readable storage medium (not shown), which stores computer-readable instructions. These computer-readable instructions are executed by a processor in a controller 10 to implement an image generation method according to the above embodiments.

[0162] Specifically, computer-readable storage media can be non-volatile or volatile. Computer-readable storage media include flash memory, portable hard drives, multimedia cards, card-type memories (e.g., SD memory, DX memory, etc.), magnetic storage, magnetic disks, optical disks, etc. In some embodiments, memory 11 can be an internal storage unit of controller 10, such as the portable hard drive of controller 10. In other embodiments, memory 11 can also be an external storage device of controller 10, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on controller 10. Memory 11 can be used not only to store application software and various types of data installed on controller 10, such as the code of an image generation program, but also to temporarily store data that has been output or will be output.

[0163] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, applications required for the functions, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0164] In the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, "multiple" means two or more.

[0165] In the embodiments of this application, it should be noted that, unless otherwise expressly specified and limited, the word "for example" is used to indicate an example, illustration, or description. Any embodiment or design scheme described as "for example" in the embodiments of this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of the word "for example" is intended to present the relevant concepts in a specific manner.

[0166] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0167] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more features.

[0168] In the description of this application, it should be noted that, unless otherwise explicitly stated and limited, "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Furthermore, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.

[0169] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if a method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if a method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.

[0170] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

[0171] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation.

[0172] In the various embodiments of this application, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0173] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the specification may also be implemented by a single unit or device through software or hardware.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.

Claims

1. An image generation method characterized by, The method includes: Receive an image generation instruction and obtain the context information associated with the image generation instruction; Generate a hybrid feature vector based on the image generation instructions and the context information; Based on the hybrid feature vector, historical images that meet preset conditions are retrieved in parallel from the multi-level cache system. The preset conditions include that the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image is greater than a preset similarity threshold. If no historical image meeting the preset conditions is found in the multi-level caching system, obtain the image generation elements corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation elements. Based on the adaptation parameters, the initial generation model is dynamically updated to obtain the target generation model; The target output image is generated by the target generation model based on the hybrid feature vector.

2. The image generation method of claim 1, wherein, The multi-level caching system includes an on-board cache, an edge cloud cache, and a central cloud cache. The method further includes: If a historical image that meets the preset conditions is found in at least one level of the multi-level caching system, a target output image is determined from the historical images that meet the preset conditions using a preset image scheduling strategy.

3. The image generation method of claim 1, wherein, The step of generating a hybrid feature vector based on the image generation instruction and the context information includes: Using a preset text encoder, the image generation instructions are semantically encoded to obtain a text semantic vector; Based on the context information, a visual feature vector is generated using a preset visual feature encoder, wherein the context information includes the user's visual preferences and / or scene information of the vehicle's location. The text semantic vector and the visual feature vector are fused to obtain the hybrid feature vector.

4. The image generation method as described in claim 3, characterized in that, The historical feature vector includes historical text feature vectors and historical visual feature vectors, and the method further includes: Obtain the first similarity between the text semantic vector and the historical text feature vector, and the second similarity between the visual feature vector and the historical visual feature vector; The first similarity is weighted according to a preset first weighting coefficient to obtain a weighted first similarity, and the second similarity is weighted according to a preset second weighting coefficient to obtain a weighted second similarity. The sum of the weighted first similarity and the weighted second similarity is taken as the similarity between the mixed feature vector and the historical feature vector.

5. The image generation method as described in claim 3, characterized in that, The method further includes: During the first denoising iteration, the key matrix and value matrix output by the text encoder are cached; In the denoising iteration following the first denoising iteration, the cached key matrix and value matrix are reused for attention operations.

6. The method as described in claim 5, characterized in that, The method further includes: The computational task of the attention operation is dispatched to the neural network processing unit or the digital signal processing unit for execution; The task of decoding the latent space features output by the target generation model into pixel images through a preset decoder is assigned to the graphics processing unit for execution.

7. The image generation method as described in claim 1, characterized in that, The adaptation parameters include a low-rank weight matrix. Obtaining the image generation features corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation features includes: Identify the target entities and / or style descriptors in the image generation instruction to obtain the image generation elements; Obtain the low-rank weight matrix corresponding to the image generation elements from the multi-level caching system; Based on the low-rank weight matrix, the linear layer weights corresponding to the initial generation model are dynamically updated to obtain the target generation model.

8. The image generation method as described in claim 1, characterized in that, The method further includes: The target output image and the corresponding hybrid feature vector are stored as new cache entries in the multi-level cache system.

9. An image generation apparatus, characterized in that, The device includes: The information acquisition module is used to receive image generation instructions and acquire the context information associated with the image generation instructions; The feature generation module is used to generate a hybrid feature vector based on the image generation instruction and the context information; The image retrieval module is used to retrieve historical images that meet preset conditions in parallel from a multi-level cache system based on the hybrid feature vector, wherein the preset conditions include the similarity between the hybrid feature vector and the historical feature vector corresponding to the historical image being greater than a preset similarity threshold. The parameter acquisition module is used to acquire the image generation elements corresponding to the image generation instruction and the adaptation parameters corresponding to the image generation elements if no historical image that meets the preset conditions is found in the multi-level caching system. The model update module is used to dynamically update the initial generated model according to the adaptation parameters to obtain the target generated model; An image generation module is used to generate a target output image based on the hybrid feature vector using the target generation model.

10. A controller, characterized in that, The controller includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the image generation method as described in any one of claims 1 to 8.

11. A vehicle, characterized in that, The vehicle includes a controller as described in claim 10, the controller being configured to perform the image generation method as described in any one of claims 1 to 8.