Method and system for generating intelligent cabin image based on AIGC
By generating an intelligent cockpit image generation model through generative adversarial training, a high-quality intelligent cockpit image set was generated, which improved the visual quality and logical rationality of the generated images. This solved the technical bottleneck of existing image generation methods in the field of intelligent cockpit image recognition, and achieved high-quality, high-generalization, and highly consistent image generation with the real cockpit environment.
Patent Information
- Application Number
- CN202511434533.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-30
AI Technical Summary
Existing image generation methods in the field of intelligent cockpit image recognition suffer from poor semantic consistency in generated images. Existing technologies struggle to generate high-quality, highly generalizable images that are highly consistent with the real cockpit environment. Furthermore, existing technologies fail to address the challenges of generating high-quality, highly generalizable images that are highly consistent with the real cockpit environment.
The Lora method is adopted, which is based on generative adversarial network to fine-tune the generative adversarial training to generate image generation methods. By introducing the Florence2 visual understanding model with attention adaptive mechanism, the image to be optimized is repaired and optimized to generate image generation models. Through semantic repair and optimization, a high-quality intelligent cockpit image generation method is generated.
It has achieved the generation of high-quality intelligent cockpit image sets, improved the visual quality and logical rationality and consistency of the generation technology, solved the technical bottleneck of existing image generation methods in the field of intelligent cockpit image recognition, and realized the technical application in the fields of advanced driver assistance systems (ADAS) and intelligent human-machine interaction.
Smart Images

Figure CN121236232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart cockpit technology, and relates to, but is not limited to, a method and system for generating smart cockpit images based on AIGC. Background Technology
[0002] With the widespread application of image generation models in image recognition tasks of Occupant Monitoring Systems (OMS), the scarcity of high-quality labeled data has become increasingly prominent. To achieve accurate detection of occupant behavior and effective identification of lost items, the training of intelligent cockpit image recognition models requires large-scale, high-quality image datasets. However, existing image generation methods face numerous technical bottlenecks in practical applications, severely hindering the development of intelligent cockpit image recognition technology. Summary of the Invention
[0003] Based on the above problems, this application provides a method and system for generating smart cockpit images based on Artificial Intelligence Generated Content (AIGC), aiming to generate a high-quality smart cockpit image set.
[0004] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a method for generating smart cockpit images based on AIGC, the method comprising: Obtain text description information used to describe the scene inside the cockpit; The text description information is input into the image generation model to obtain the initial smart cockpit image set; the image generation model is obtained by fine-tuning the generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using the LoRa method. The quality of each image in the initial intelligent cockpit image set is evaluated to obtain the quality evaluation value of each image. When the quality assessment value of the image to be optimized is less than a preset threshold, the Florence2 visual understanding model with an attention adaptive mechanism is used to extract and analyze the visual content of the image to be optimized, and obtain the semantic information of the image to be optimized; wherein, the image to be optimized is an image in the initial intelligent cockpit image set; Using inpaint technology, based on semantic information, the images to be optimized in the initial intelligent cockpit image set are repaired and optimized to obtain the target intelligent cockpit image set.
[0005] In some embodiments, the process of constructing an image generation model includes: Data augmentation is performed on multiple labeled images containing different cockpit environments and user behaviors to obtain a training sample set; Using the training sample set as training data, the generative adversarial network is fine-tuned using the LoRa method to obtain an image generation model. In the process of fine-tuning the pre-trained image generation model using the LoRa method, the Adam optimizer with specific parameters is used for parameter optimization, and the cosine annealing learning rate scheduler is used to adjust the learning rate.
[0006] In some embodiments, the multiple labeled images containing different cabin environments and user behaviors include at least: multiple labeled sub-images with different lighting conditions, multiple labeled sub-images with different seat positions in the cabin, multiple labeled sub-images with different cabin interior configurations, and multiple labeled sub-images with different user postures in the cabin.
[0007] In some embodiments, before inputting textual description information into an image generation model to obtain an initial smart cockpit image set, the method further includes: Based on the preset scenario requirements, the model parameters of the image generation model are adjusted to obtain the adjusted image generation model; The text description information is input into the image generation model to obtain an initial set of smart cockpit images, including: The text description information is input into the adjusted image generation model to obtain the initial smart cockpit image set.
[0008] In some embodiments, the model parameters of the image generation model are adjusted according to preset scenario requirements to obtain an adjusted image generation model, including: Obtain the artistic style and image quality requirements for the preset scene; Based on the artistic style and image quality requirements, the generation parameters of the image generation model are adjusted to obtain the adjusted generation parameters; Based on the adjusted generation parameters, the image generation model is reconfigured to obtain the adjusted image generation model.
[0009] In some embodiments, the quality of each image in the initial smart cockpit image set is evaluated to obtain a quality evaluation value for each image, including: Construct a library of prompt word templates corresponding to image quality concerns in cockpit scenarios; Each image in the initial intelligent cockpit image set is sequentially input into the preset Visual Language Model (VLM), and engages in multi-turn dialogue with the prompt word templates in the prompt word template library to generate multi-turn text responses corresponding to each image; Quantify the multi-round text responses corresponding to each image to obtain the multi-round quality scores for each image; The quality scores of each image are fused from multiple rounds to obtain the quality assessment value of each image.
[0010] In some embodiments, the Florence2 visual understanding model with an attention adaptation mechanism includes: a multimodal encoder with a self-attention layer, a decoder consisting of a masked sub-attention layer and a cross-attention layer.
[0011] In some embodiments, the inpaint technique is used to repair and optimize the images to be optimized in the initial smart cockpit image set based on semantic information. After obtaining the target smart cockpit image set, the method further includes: The target smart cockpit image set is used as training data to iteratively train the preset smart cockpit image recognition model until the output of the target smart cockpit image recognition model meets the preset conditions.
[0012] Secondly, embodiments of this application provide a system for generating smart cockpit images based on AIGC, the system comprising: The acquisition module is used to acquire text description information that describes the scene inside the cockpit; The image generation module is used to input text description information into the image generation model to obtain an initial intelligent cockpit image set. The image generation model is obtained by fine-tuning the generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using the LoRa method. The quality assessment module is used to assess the quality of each image in the initial intelligent cockpit image set and obtain the quality assessment value of each image. The extraction and analysis module is used to extract and analyze the visual content of the image to be optimized when the quality assessment value of the image to be optimized is less than a preset threshold, using the Florence2 visual understanding model with an attention adaptive mechanism to obtain the semantic information of the image to be optimized; wherein, the image to be optimized is an image in the initial intelligent cockpit image set; The repair and optimization module is used to repair and optimize the images to be optimized in the initial smart cockpit image set based on semantic information using inpaint technology, so as to obtain the target smart cockpit image set.
[0013] The beneficial effects of the technical solutions provided in this application include at least the following: The method and system for generating smart cockpit images based on AIGC provided in this application embodiment firstly acquire textual description information describing the scene inside the cockpit; then input the textual description information into an image generation model to obtain an initial smart cockpit image set; wherein, the image generation model is obtained by fine-tuning a generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using the LoRa method; then, the quality of each image in the initial smart cockpit image set is evaluated to obtain a quality evaluation value for each image; and if the quality evaluation value of the image to be optimized is less than a preset threshold, the Florence2 visual understanding model with an attention adaptive mechanism is used to extract and analyze the visual content of the image to be optimized to obtain the semantic information of the image to be optimized; wherein, the image to be optimized is an image in the initial smart cockpit image set; finally, the inpaint technique is used to repair and optimize the image to be optimized in the initial smart cockpit image set based on the semantic information to obtain a target smart cockpit image set. In this way, on the one hand, by using the LoRa method, an image generation model is generated based on multiple labeled images containing different cockpit environments and user behaviors to generate an initial intelligent cockpit image set. This ensures from the source that the generated images conform to the professional specifications of intelligent cockpits in terms of style, layout, and basic elements. On the other hand, based on the quality assessment of each image, the Florence2 visual understanding model with an attention-adaptive mechanism is used to extract and analyze the visual content of the images to be optimized, obtain the semantic information of the images to be optimized, and use the semantic information to repair and optimize the images to be optimized in the initial intelligent cockpit image set. Thus, while being able to perform quality screening on the images in the initial intelligent cockpit image set to ensure the visual quality of the subsequent output image set, the introduction of the Florence2 visual understanding model with an attention-adaptive mechanism achieves a leap from pixel-level evaluation to semantic-level evaluation of the images to be optimized, thereby discovering and defining the deep semantic information in the images to be optimized and realizing closed-loop automatic optimization. Furthermore, it can diagnose problems on its own (Florence2 visual understanding model) and perform precise repairs on its own (inpaint), thus ensuring that the resulting target intelligent cockpit image set is not only visually high-quality, but also highly consistent and reasonable in terms of semantics and logic.In other words, this application forms a closed loop by using four stages: fine-tuning technology (Lora method), automatic quality screening strategy, semantic-level diagnosis (introducing the Florence2 visual understanding model with attention adaptive mechanism), and semantic-guided repair (inpaint). This transforms the traditional single generation process into an iterative, optimizable, and controllable intelligent system, while achieving synergistic effects. Together, these ensure that the final generated target intelligent cockpit image set reaches extremely high levels in terms of visual quality, domain professionalism, and logical rationality. This greatly improves the reliability and practicality of AIGC technology in vertical fields, providing high-quality synthetic data for the training of subsequent intelligent cockpit image recognition models. This, in turn, promotes the further development of intelligent cockpit technology in the fields of Advanced Driving Assistance Systems (ADAS) and Intelligent Human-Machine Interaction (HMI).
[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the technical solutions provided in the embodiments of this application. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 This application provides a flowchart illustrating a method for generating smart cockpit images based on AIGC. Figure 2 A schematic diagram illustrating adversarial training of the generator and discriminator in a generative adversarial network during the training phase; Figure 3 A schematic diagram illustrating the main process of fine-tuning a generative adversarial network to obtain an image generation model; Figure 4 A schematic diagram of the test environment corresponding to the method for generating smart cockpit images based on AIGC provided in the embodiments of this application; Figure 5 This is a schematic diagram of the composition of a system for generating smart cockpit images based on AIGC, as provided in an embodiment of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0018] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0019] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0020] In practical applications, applying image generation models to the generation of cockpit image training data still faces many technical bottlenecks, such as: 1. Poor Model Fine-Tuning Quality: Current mainstream image generation model fine-tuning methods based solely on Low-Rank Adaptation (Lora) technology perform poorly in terms of semantic consistency and detail reproduction of generated images. The generated images cannot accurately reproduce the complex textures, lighting distributions, and passenger postures of real cockpit scenes, resulting in significant differences between the generated data and real data in the feature space. This makes it difficult to meet the high data quality requirements for training intelligent cockpit image recognition models.
[0021] 2. Image Distortion and Artifacts: Existing image generation models commonly produce images with distortion, artifacts, disproportionate human figures, and multiple hands, among other anomalies. These generated images violate fundamental physical laws of the natural world and human anatomy, resulting in a low signal-to-noise ratio (SNR) for the generated data. This makes it impossible to provide high-quality training samples for intelligent cockpit image recognition models, thus affecting the generalization ability and recognition accuracy of these models.
[0022] 3. Poor Image Generalization: Existing image generation solutions on the market exhibit extremely limited generalization capabilities. The generated images lack diversity and cannot effectively cover the complex scenarios faced by intelligent cockpit image recognition models in their recognition tasks, including different lighting conditions, passenger postures, seat positions, and cockpit interior configurations. This lack of generalization means that the generated data cannot provide sufficient positive and negative samples for training the intelligent cockpit image recognition model, limiting its robustness in practical applications.
[0023] 4. Difficulty in Reproducing Realistic Cockpit Environment Light and Shadow Interactions and Lens Parameter Simulation: Existing image generation models have significant shortcomings in reproducing realistic cockpit environments. Specifically, generated images struggle to accurately simulate the light and shadow interactions between people or objects and the cockpit environment, such as ambient light reflection, occlusion shadows, and changes in material reflectivity. Furthermore, existing image generation models cannot precisely control the simulation of lens parameters (such as focal length, aperture, and distortion) in the generated images, resulting in significant visual discrepancies between the generated images and actual cockpit scenes. This makes it impossible to provide training data that is highly consistent with real-world scenarios for intelligent cockpit image recognition models.
[0024] 5. Lack of Post-Generation Closed-Loop Testing and Data Management Mechanisms: After image generation, existing methods lack a reliable visual referee model to automatically evaluate and filter the quality of the generated images. Furthermore, there are no automated testing methods or iterative correction processes for erroneous datasets. This makes it difficult to effectively guarantee the quality of the generated data, and erroneous data may be introduced into the model training process, further affecting the training effect and convergence speed of the intelligent cockpit image recognition model.
[0025] In summary, existing image generation methods face numerous technical bottlenecks in the application of intelligent cockpit image recognition. There is an urgent need to develop an image generation method that can generate high-quality, highly generalizable images that are highly consistent with the real cockpit environment, and to establish a sound post-generation closed-loop testing and data management mechanism.
[0026] Based on the above description, this application provides a method for generating smart cockpit images based on AIGC. Please refer to [link to relevant documentation]. Figure 1 As shown, the method includes the following steps: Step 101: Obtain text description information to describe the scene inside the cockpit.
[0027] In some embodiments of this application, the textual description information used to describe the scene inside the cockpit can be: inside the smart cockpit, the driver is viewing navigation information through a high-definition display interface, and the passenger in the front seat is listening to music with headphones; or inside the smart cockpit, several occupants, including the driver, are chatting happily, etc. This application does not limit the specific content of the textual description information used to describe the scene inside the cockpit.
[0028] Step 102: Input the text description information into the image generation model to obtain the initial smart cockpit image set.
[0029] The image generation model is obtained by fine-tuning the generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using the Lora method.
[0030] In some embodiments of this application, the image generation model is an Explainable Artificial Intelligence Image Generation (XAI ImageGen) model, which is an advanced model that fills image regions based on text descriptions. It can generate high-quality images and can be flexibly invoked and configured through image processing tools (such as ComfyUI).
[0031] In some embodiments of this application, the image generation model may be obtained by fine-tuning a generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors. Here, the LoRa method is a technique for efficiently fine-tuning large deep learning models (especially large language models). Specifically, it efficiently fine-tunes large models through low-rank decomposition, freezes the original model weights, and simulates weight updates by introducing and training two extremely small side-path matrices, thereby achieving excellent fine-tuning performance at extremely low cost.
[0032] It should be noted that Generative Adversarial Networks (GANs) are deep learning models that generate high-quality data through adversarial training. A detailed diagram can be found in [link to diagram]. Figure 2 Furthermore, GANs are inspired by zero-sum games in game theory, and their core idea is to simultaneously train two competing neural network models that progress together. Figure 2 (Generator 203 and Discriminator 204 in the text) 1. Generator (G): Its goal is to learn the distribution of real data and generate fake data (such as images and sounds) that are as realistic as possible. It receives... Figure 2 The 201 shown in the figure is a random noise vector (usually denoted as z) as input, and a fake data sample as output, denoted as G(z).
[0033] 2. Discriminator (D): Its goal is to distinguish whether the input data comes from the real dataset (denoted as x), i.e. Figure 2 The training set 202 shown is still the fake data G(z) generated by generator 203. It outputs a scalar probability value (e.g., between 0 and 1) representing the probability that it considers the input data to be "real".
[0034] The game proceeds as follows: Generator G tries to deceive discriminator D, hoping that discriminator D will give a high probability of truth for its generated false data G(z). Discriminator D strives to improve its judgment ability, classifying real data as "true" and generated data as "false" as accurately as possible. The two continuously compete and iterate to optimize until a Nash equilibrium is reached: the data generated by generator G is so realistic that discriminator D cannot distinguish between true and false (i.e., for any input, discriminator D's judgment probability is 50%, equivalent to random guessing).
[0035] In some embodiments of this application, the image generation model can be constructed through the following process: The first step is to perform data augmentation on multiple labeled images containing different cockpit environments and user behaviors to obtain a training sample set.
[0036] In some embodiments of this application, to ensure the effectiveness and generalization of generative adversarial network training, multiple labeled images containing different cockpit environments and user behaviors can be used as the training dataset, i.e., a large-scale intelligent cockpit image dataset can be used. The intelligent cockpit image dataset may contain the following characteristics: Data size: The training set contains 5,000 labeled images of different cockpit environmental behaviors.
[0037] Data diversity: The images cover different lighting conditions (e.g., natural light, in-vehicle lighting), different human postures (e.g., sitting, semi-reclining, standing), different seat positions (e.g., front row, rear row), and different cabin interior configurations (e.g., different car models, different interior materials).
[0038] Data annotation: Each image has been meticulously annotated, including key points of people (e.g., head, hand, and leg positions), behavior categories (e.g., normal driving, fatigued driving, using a mobile phone, yawning, abdominal pain, and scratching the head), and the location of items (e.g., the location of lost items).
[0039] It should be noted that after training the generative adversarial network, the performance of the generated model can be evaluated and validated, thereby simultaneously increasing the corresponding validation set and test set. For example, the validation set contains 200 images and the test set contains 400 images. The specific method is existing technology and this application does not make specific limitations on it.
[0040] In some embodiments of this application, multiple labeled images containing different cabin environments and user behaviors include at least: multiple labeled sub-images with different lighting conditions, multiple labeled sub-images with different seat positions in the cabin, multiple labeled sub-images with different cabin interior configurations, and multiple labeled sub-images with different user postures in the cabin.
[0041] This ensures the comprehensiveness and professionalism of the training sample set, thereby indirectly proving the effectiveness of the subsequently generated image generation model and improving its generalization ability.
[0042] The second step involves using the training sample set as training data and fine-tuning the generative adversarial network using the LoRa method to obtain the image generation model.
[0043] Specifically, the LoRa method is used to fine-tune the pre-trained image generation model, employing the Adam optimizer with specific parameters for parameter optimization and a cosine annealing learning rate scheduler to adjust the learning rate.
[0044] In some embodiments of this application, the LoRa method is employed and optimized in conjunction with an advanced deep learning architecture (Generative Adversarial Network). The specific configuration is as follows: 1. Model Architecture: A generative adversarial network (GAN) is used as the base model, fine-tuned using LoRa techniques. LoRa, through low-rank decomposition, improves the model's adaptability to specific tasks without significantly increasing the number of model parameters.
[0045] 2. Optimizer Selection: The Adam optimizer with specific parameters is used because it exhibits good convergence speed and stability when training deep generative models. The parameter settings for the Adam optimizer are as follows: Learning Rate: The initial learning rate is set to 1×10. −4 ; Momentum parameters: First-order momentum (β1) = 0.9, second-order momentum (β2) = 0.999.
[0046] Learning rate scheduler: A cosine annealing scheduler is used to dynamically adjust the learning rate during training. The minimum learning rate is set to 1×10. −6 Each cycle consists of 10 epochs (iterations) of training.
[0047] 3. Training Strategy: To improve the robustness of the image generation model and the quality of the generated images, the following training strategies can be adopted: Training epochs and steps: The model was trained for 50 epochs, with each epoch containing 1,000 steps, for a total of 50,000 training steps.
[0048] Data augmentation: During training, random cropping, rotation, flipping, and color jittering are performed on the input image to improve the generalization ability of the image generation model.
[0049] Adversarial Training: During the training of the Generative Adversarial Network (GAN), the generator and discriminator are updated alternately. The generator aims to minimize the discriminator's loss function, while the discriminator aims to maximize its ability to distinguish between real and generated images. The loss function of the Wasserstein GAN (WGAN) can be used here to avoid mode collapse. Furthermore, to prevent overfitting, Dropout (dropout rate: 0.2) and L2 regularization (weight decay coefficient: 1×10⁻⁶) can be added during training. −5 ).
[0050] This can be used as a reference. Figure 3 The main process shown is to fine-tune the generative adversarial network to obtain the image generation model, which includes: 301: Obtain the dataset, which means obtaining multiple images containing different cockpit environments and user behaviors.
[0051] 302: Use VLM vision technology for labeling. Here, a large visual language model can be used to label each image in the dataset obtained in 301, thereby obtaining multiple labeled images containing different cockpit environments and user behaviors.
[0052] 303: Perform preliminary dataset preprocessing. This may involve data augmentation and / or outlier filtering on the multiple labeled images obtained in 302 that contain different cockpit environments and user behaviors, thereby obtaining the training sample set.
[0053] 304: Training and Round Selection. Here, the training sample set obtained in 303 can be used as training data, and the LoRa method can be used to fine-tune the generative adversarial network to obtain the image generation model; in the fine-tuning stage, relevant rounds and other parameters can be determined.
[0054] 305: Test the output model and compare the quality of the output images. Here, you can evaluate the performance of the image generation model and score the quality of its output images.
[0055] It should be noted that step 305 is the performance evaluation and verification stage after the image generation model is built.
[0056] Thus, by using the LoRa method, combined with an optimized learning rate scheduler, data augmentation, and regularization techniques, the robustness of the image generation model can be successfully improved, thereby further enhancing the quality of generated images in intelligent cockpit image generation tasks.
[0057] Here, in the process of fine-tuning the generative adversarial network to generate an image generation model, the kl_optimall scheduler and dpmpp_3m_sde_gpu sampler can be used to further optimize the training and inference process of the model, thereby significantly improving the system's operating efficiency.
[0058] In addition, to verify the effectiveness of the image generation model provided in this application in generating smart cockpit images, the following experimental control group can be designed: Experimental group: Image generation model 1 was constructed using the method described in this application.
[0059] Control group 1: Image generation model 2 was obtained by fine-tuning the generative adversarial network using LoRa but without data augmentation, and based on multiple labeled images containing different cockpit environments and user behaviors.
[0060] Control group 2: LoRa fine-tuning was used, but without a cosine annealing learning rate scheduler; a fixed learning rate of 1×10⁻⁶ was maintained. −4 Based on multiple labeled images containing different cockpit environments and user behaviors, an image generation model 3 was obtained by fine-tuning the generative adversarial network.
[0061] The corresponding comparison results are as follows: 1. Training time: The training time for image generation model 1 in the experimental group was 48 hours, accelerated using a single NVIDIA A100 Tensor Core GPU based on the Ampere architecture. The training time for image generation model 2 in control group 1 was 45 hours, and the training time for image generation model 3 in control group 2 was 46 hours.
[0062] 2. Generated Image Quality: The quality of the generated images was evaluated by calculating the initial inception score (IS) and the Frechet inception distance (FID). The experimental group had an IS of 8.5 and an FID of 12.3, which were significantly better than those of control group 1 (IS = 7.2, FID = 18.5) and control group 2 (IS = 7.8, FID = 15.2).
[0063] 3. Model Robustness: The robustness of the models (image generation model 1, image generation model 2, and image generation model 3) was evaluated on the validation and test sets. The experimental group achieved an average accuracy of 90.5% under different lighting conditions, person posture, and seat position, while control group 1 achieved 85.2% and control group 2 achieved 87.3%.
[0064] Thus, after comprehensively evaluating the performance of the experimental group, control group 1, and control group 2, the improved Lora training method in the experimental group (i.e., the image generation model 1 constructed in this application) showed the best performance and can be used in intelligent cockpit image generation tasks. In other words, by combining an optimized learning rate scheduler, data augmentation, and regularization techniques, the improved Lora method and fine-tuning of the generative adversarial network can successfully enhance the robustness of the image generation model, thereby further improving the quality of generated images in intelligent cockpit image generation tasks. Furthermore, experimental results show that the method provided in this application outperforms existing methods in terms of detail reproduction, semantic consistency, and generalization ability of generated images. Further exploration of more advanced generative model architectures (such as Diffusion Models) and training strategies can further improve the performance of intelligent cockpit image generation tasks.
[0065] It should be noted that, in the process of using the training sample set as training data and fine-tuning the generative adversarial network using the LoRa method to obtain the image generation model, the training efficiency and quality of the image generation model can be further improved by deploying the parameters of the running hardware. Regarding the computing power system: 1. The graphics processor requires a single-precision floating-point computing power of at least 40 trillion floating-point operations per second (FLOPS) and at least 24GB of video memory.
[0066] 2. The central processing unit (CPU) is required to have at least 12 cores / 24 threads to ensure multitasking and parallel computing capabilities, with a high clock speed of 3.5 GHz or higher.
[0067] 3. The motherboard must support PCIe 4.0 or 5.0 x16 to ensure that the graphics processor can fully utilize the bandwidth.
[0068] 4. Memory requirement is at least 32GB, 128GB is recommended.
[0069] 5. Power supply requirement is at least 1000W.
[0070] 6. Storage requirements: 1TB or more NVMe SSD is recommended.
[0071] In addition, it is recommended to use a high-efficiency liquid cooling radiator for heat dissipation, and the operating system requirement is Ubuntu 20.04 or later, ensuring that the latest NVIDIA drivers are supported.
[0072] This can further improve the performance of the image generation model in intelligent cockpit image generation tasks, while ensuring the quality and generalization ability of the generated images.
[0073] In some embodiments of this application, the following step A1 may also be performed before performing step 102: Step A1: Adjust the model parameters of the image generation model according to the preset scenario requirements to obtain the adjusted image generation model.
[0074] In some embodiments of this application, preset scene requirements are defined as scene requirements corresponding to the cockpit environment. For example, if the vehicle corresponding to the cockpit is in an indoor or aerial environment, the corresponding scene requirements are: the art style and image quality requirements for the indoor or aerial environment are a tranquil observation style, and the image quality requirement level is general. If the vehicle corresponding to the cockpit is in a tense atmosphere, the art style in the corresponding scene requirements may involve information such as alarms or alerts, and the image quality requirement level in the scene requirements is extra high.
[0075] In some embodiments of this application, step A1 can be implemented by the following steps A11 to A13: Step A11: Obtain the artistic style and image quality requirements for the preset scene.
[0076] Step A12: Adjust the generation parameters of the image generation model according to the artistic style and image quality requirements to obtain the adjusted generation parameters.
[0077] Step A13: Based on the adjusted generation parameters, reconfigure the image generation model to obtain the adjusted image generation model.
[0078] In some embodiments of this application, the generation parameters of the image generation model include, but are not limited to: the type of sampler, the number of sampling steps, the cue word relevance guidance coefficient, and the image size.
[0079] In this way, by adjusting the generation parameters of the image generation model according to the artistic style and image quality requirements of the preset scene, it is possible to achieve a high degree of alignment between the image content generated by the adjusted image generation model and the scene requirements, thereby improving the controllability and determinism of the generated image.
[0080] Correspondingly, step 102 above can be achieved through the following step A2: Step A2: Input the text description information into the adjusted image generation model to obtain the initial smart cockpit image set.
[0081] In some embodiments of this application, the number of images included in the initial smart cockpit image set can be determined according to actual needs, and each image in the initial smart cockpit image set embodies "textual description information for describing the scene inside the cockpit", while any two images in the initial smart cockpit image set describe different specific information.
[0082] In this way, by transforming abstract scenario requirements into operable parameters of the model and configuring the image generation model in a targeted manner, the image generation process is transformed from "extensive" to "refined". The final output is a controllable result that is highly matched with the specific application scenario in terms of artistry, quality and consistency, while improving the utilization efficiency of computing resources.
[0083] Step 103: Perform quality assessment on each image in the initial smart cockpit image set to obtain the quality assessment value of each image.
[0084] In some embodiments of this application, the quality of each image in the initial smart cockpit image set is evaluated through the following aspects: sharpness, color, texture, noise, flash, etc., to obtain the corresponding quality score for each aspect, and then the corresponding quality scores for each aspect are integrated according to preset rules to obtain the corresponding quality evaluation value of each image.
[0085] Here, step 103 can be achieved through the following steps 1031 to 1034 ( Figure 1 (not shown): Step 1031: Construct a library of prompt word templates corresponding to image quality concerns in the cockpit scene.
[0086] Step 1032: Input each image in the initial intelligent cockpit image set into the preset visual language model (VLM) in sequence, and conduct multi-round dialogue with the prompt word templates in the prompt word template library to generate multi-round text responses corresponding to each image.
[0087] Step 1033: Quantify the multi-round text responses corresponding to each image to obtain the multi-round quality scores corresponding to each image.
[0088] Step 1034: Fuse the multi-round quality scores corresponding to each image to obtain the quality assessment value of each image.
[0089] In some embodiments of this application, the preset VLM does not require prior training and can rely on a preset prompt template library to guide the preset VLM in making quality judgments about images. Correspondingly, the prompt template library refers to a pre-designed, categorized, stored, and reusable set of prompts, which is designed to obtain high-quality output conforming to a specific style or format from AI models (such as large language models or text-to-image models) more reliably and efficiently.
[0090] In some embodiments of this application, the quality scores of each image are fused to obtain the quality evaluation value of each image. Here, the quality scores of each image are assigned corresponding weights, so that the quality scores of each round are first fused with the corresponding weights to obtain the fused score of each round, and then the fused scores of each image are fused to obtain the quality evaluation value of each image.
[0091] In this way, evaluation criteria are defined by a prompt word template library, deep semantic-level image understanding is achieved by utilizing the multi-turn dialogue capability of the pre-set visual language model (VLM), and finally an interpretable and highly consistent comprehensive quality evaluation value is output through the quantification and fusion of multi-turn responses.
[0092] Step 104: If the quality assessment value of the image to be optimized is less than the preset threshold, the Florence2 visual understanding model with an attention adaptive mechanism is used to extract and analyze the visual content of the image to be optimized, and obtain the semantic information of the image to be optimized.
[0093] The images to be optimized are those within the initial smart cockpit image set.
[0094] In some embodiments of this application, the number of images to be optimized may be one, two or more, and this application does not impose any limitation on this.
[0095] In some embodiments of this application, the Florence2 visual understanding model is an advanced visual foundation model that performs well in image captioning, object detection, and other visual language evaluation tasks. Here, regarding the introduction of an attention-adaptive mechanism into the Florence2 visual understanding model, the corresponding model can be composed of the following two main parts: A multimodal encoder with a self-attention layer is introduced.
[0096] A decoder consisting of a masked sub-attention layer and a cross-attention layer is introduced.
[0097] In this way, the encoder-decoder architecture based on the attention adaptive mechanism can transform visual understanding from a series of scattered, specific classification tasks into a unified, prompt-based "sequence-to-sequence" generation task, thereby enabling the accurate extraction and analysis of the visual content of the image to be optimized, so as to obtain semantic information that accurately expresses the information of the image to be optimized.
[0098] Step 105: Using inpaint technology, based on semantic information, repair and optimize the images to be optimized in the initial intelligent cockpit image set to obtain the target intelligent cockpit image set.
[0099] In some embodiments of this application, inpaint technology is a process of intelligently filling specified areas in an image or video using artificial intelligence and computer vision techniques. Its goal is to learn from known parts (context) of the image to generate visually plausible, content-coherent, and stylistically consistent pixels to replace removed or damaged portions.
[0100] Here, each image in the target intelligent cockpit image set is an image that accurately describes the textual description information of the scene inside the cockpit. At the same time, the information described by any two images in the target intelligent cockpit image set may have some differences or subtle differences.
[0101] In some embodiments of this application, after performing step 105, the application may also perform step B: The target smart cockpit image set is used as training data to iteratively train the preset smart cockpit image recognition model until the output of the target smart cockpit image recognition model meets the preset conditions.
[0102] Here, the specific model structure of the preset intelligent cockpit image recognition model can be determined according to actual needs, and this application does not impose any restrictions on it.
[0103] Here, because the target intelligent cockpit image set contains higher quality images and a larger amount of data, the preset intelligent cockpit image recognition model trained using these data can significantly improve the accuracy and robustness of the trained target intelligent cockpit image recognition model.
[0104] The method for generating smart cockpit images based on AIGC provided in this application has two aspects. First, by using the LoRa method, an image generation model is obtained from multiple labeled images containing different cockpit environments and user behaviors to generate an initial smart cockpit image set. This ensures from the source that the generated images conform to the professional specifications of smart cockpits in terms of style, layout, and basic elements. Second, based on the quality assessment of each image, a Florence2 visual understanding model with an attention-adaptive mechanism is used to extract and analyze the visual content of the images to be optimized, obtaining the semantic information of the images to be optimized. The semantic information is then used to repair and optimize the images to be optimized in the initial smart cockpit image set. In this way, while the quality of the images in the initial smart cockpit image set can be screened to ensure the visual quality of the subsequent output image set, the introduction of the Florence2 visual understanding model with an attention-adaptive mechanism achieves a leap from pixel-level evaluation to semantic-level evaluation of the images to be optimized. This enables the discovery and definition of deep semantic information in the images to be optimized and the realization of closed-loop automatic optimization. Furthermore, it can self-diagnose problems (Florence2 visual understanding model) and perform precise repairs (inpaint), ensuring that the resulting target intelligent cockpit image set is not only visually high-quality but also highly consistent and reasonable in semantics and logic. In other words, this application forms a closed loop through four stages: fine-tuning technology (Lora method), automatic quality screening strategy, semantic-level diagnosis (introducing the Florence2 visual understanding model with attention adaptation mechanism), and semantic-guided repair (inpaint). This evolves the traditional, singular generation process into an iterative, optimizable, and controllable intelligent system, achieving synergistic effects to ensure that the final generated target intelligent cockpit image set reaches extremely high levels in visual quality, domain professionalism, and logical rationality. This significantly improves the reliability and practicality of AIGC technology in vertical applications, providing high-quality synthetic data for the training of subsequent intelligent cockpit image recognition models, thereby promoting the further development of intelligent cockpit technology in ADAS (Advanced Driver Assistance Systems) and HMI (Hardware-Machine Interface) fields.
[0105] The method for generating smart cockpit images based on AIGC will be described below with reference to two specific embodiments. However, it is worth noting that these specific embodiments are only for better illustration of this application and do not constitute an improper limitation of this application.
[0106] Implementation Case 1: Image Recognition and Generation Solution for Car Smart Cockpit. The background is as follows: Car Company A needs to develop a high-precision image recognition system for its smart cockpit to detect the behavior of occupants and lost items. However, the market lacks sufficient high-quality training data, resulting in low recognition accuracy of existing models. The implementation process is based on the AIGC-based smart cockpit image generation method provided in this application embodiment: 1. Determine data requirements analysis: Company A analyzed the requirements for intelligent cockpit image recognition and determined that it needs to generate a large amount of high-quality image data, covering various scenarios in the cockpit (such as different lighting conditions, different postures of people, different placement of objects, etc.).
[0107] 2. Technical Solution Design: The technical solution of this application is adopted. ComfyUI calls the XAI ImageGen model, which is the image generation model generated by this application. A dedicated workflow is designed to be suitable for the interface design of staff to generate images according to their needs, and is used to generate images of people and objects in the cockpit.
[0108] 3. Parameter optimization: Based on the specific needs of Company A, the parameters of the XAI ImageGen model can be adjusted accordingly (such as: adjustment of the generated mask, engineering of prompt words, number of iterations, selection of sampler / scheduler, etc.) to ensure that the quality of the generated images meets the requirements.
[0109] 4. Data generation and correction: A large amount of high-quality image data was generated, and the accuracy and consistency of the image data were ensured through a combination of manual correction and automatic detection.
[0110] 5. Model Training and Testing: Train the intelligent cockpit image recognition model using the generated image data and test it in a real cockpit environment. (The details are as follows:) 5.1. Image Quality: The generated image quality is significantly superior to existing technologies, with substantial reductions in distortion, artifacts, and aspect ratio issues. It also supports specific optimizations for lens distortion, overexposure, and ambient light simulation from different cameras.
[0111] 5.2. Data volume: More than 50,000 high-quality images were successfully generated, meeting the training needs of Company A.
[0112] 5.3. Model Performance: The accuracy of the trained image generation model in practical applications has increased from 70% to 90%, significantly improving the reliability of the system and the user experience.
[0113] Implementation Case 2: Solution for Intelligent Cockpit Recognition Requirements (B Automotive Company). Background: B Automotive Company plans to introduce advanced image recognition capabilities into its next-generation intelligent cockpit to improve driving safety and user experience. However, existing image generation technologies cannot meet its need for high-quality training data. The implementation process is based on the AIGC-based intelligent cockpit image generation method provided in this application embodiment: 1. Needs Survey: Company B conducted a detailed survey on the needs of intelligent cockpit image recognition, and clarified the types of images and scenarios that need to be generated.
[0114] 2. Technology Selection: The technology solution of this application was selected, which utilizes ComfyUI to call the XAI ImageGen model, adds a judgment model to judge the generated image using if-else nodes and designs the generation loop, and develops a customized image generation workflow.
[0115] 3. Parameter Adjustment: Based on the design requirements of B Automotive Company, the parameters of the XAI ImageGen model were optimized to generate high-quality images that conform to the cockpit design style of B Automotive Company.
[0116] 4. Data generation and verification: A large amount of high-quality image data was generated, and the accuracy and usability of the data were ensured through internal verification.
[0117] 5. Model Training and Deployment: Train the intelligent cockpit image recognition model using the generated image data. (The details are as follows:) 5.1. Image Quality: The generated images are of high quality, accurately reproducing the realistic scene inside the cockpit and meeting the stringent standards of Company B. Furthermore, a deep LoRa model was trained to analyze the gender, clothing, behavior, and actions of the characters.
[0118] 5.2. Data richness: More than 50,000 high-quality image data were generated, covering a variety of cockpit scenarios and conditions.
[0119] 5.3. Model Performance: The accuracy of the trained images in the actual application of the recognition model has increased from 65% to 90%, which significantly improves the performance and reliability of the system.
[0120] Based on the above description, the solution provided in this application has the following improvements compared to the prior art: 1. Multi-model integration and collaborative operation: This application integrates multiple advanced deep learning models, including generative adversarial networks, visual language models, and the Florence2 visual understanding model, to achieve collaborative operation of image generation, feature extraction, object detection, and text analysis, significantly improving the efficiency and accuracy of image generation and processing. This multi-model integration method is relatively rare in applications that generate smart cockpit images.
[0121] 2. Lens, Style, and Environment Fine-tuning (Lora): This application employs lens, style, and environment fine-tuning technology to generate images highly consistent with the real cockpit environment using the XAI ImageGen image generation model. This fine-tuning technology can precisely control the style and environmental characteristics of the generated images, enhancing their realism and adaptability.
[0122] 3. VLM Judgment Model and Inpaint Model: The VLM judgment model is introduced to evaluate the quality of the generated image, and the inpaint model is used to repair and optimize the image when necessary. This quality control mechanism is relatively advanced in existing technologies and can ensure the high quality and consistency of the generated image.
[0123] 4. Attention Adaptive Mechanism and Florence2 Language Segmentation Model: An attention adaptive mechanism and the Florence2 language segmentation model are employed to extract and analyze textual information from the generated images. This textual information extraction and analysis method is relatively innovative in existing technologies, and can further enrich the semantic information of images, improving their comprehensibility.
[0124] 5. kl_optimall scheduler and dpmpp_3m_sde_gpu sampler: The kl_optimall scheduler and dpmpp_3m_sde_gpu sampler are used to optimize the model's training and inference processes, significantly improving the system's operating efficiency. This optimization method is relatively advanced among existing technologies and can significantly improve the system's processing speed and stability.
[0125] For reference here. Figure 4 The diagram shown is a test environment diagram corresponding to the method for generating smart cockpit images based on AIGC provided in the embodiments of this application; the following information is shown: 1: The XAI ImageGen image generation model generates images that are highly consistent with the real cockpit environment through its internal lens, style, and environment fine-tuning technology; at the same time, the kl_optimall scheduler and dpmpp_3m_sde_gpu sampler can be used to optimize the training and inference process of the XAI ImageGen image generation model.
[0126] 2: The Florence2 visual understanding model, which incorporates an attention-adaptive mechanism, can extract and analyze the visual content of uploaded vehicle cabin images. It can also generate image masks from uploaded vehicle cabin images, and before generating the image mask, it can pre-create the mask pixels to be generated, and then perform feathering processing on the generated mask pixels (i.e., make the edges of the mask blurry and soft).
[0127] 3. After image recognition using a semantic segmentation model (Florence2 visual understanding model) on the uploaded cabin images, a Ksampler sampler can be used for image generation. This generation process involves not only inpainting techniques (i.e., an internal inpaint model) to repair and optimize the generated images, but also the use of the parameters employed in the internal inpaint model for image repair and optimization. Specifically, this involves using the Florence2 visual understanding model with an attention-adaptive mechanism to extract and analyze the visual content of the images, obtaining semantic information. Furthermore, it involves using a Variational Autoencoder (VAE) to perform corresponding dimensionality reduction and feature extraction operations.
[0128] Furthermore, a dual Clip scheduler can be used to guide the internal inpaint model to perform corresponding operations by providing positive and negative prompts through positive and negative prompt words.
[0129] 4. After generating and outputting the latent space image using the Ksampler, the image can be decoded using a VAE decoder and input into the VLM judgment model for evaluation of the latent space image's image value. Here, in the latent space image output stage, the Florence2 visual understanding model with an attention-adaptive mechanism can also be used to perform pixel-by-pixel expansion of the masking nodes.
[0130] 5. Regarding the output latent space image, a quality assessment value can be used to determine if it meets the preset requirements. If so, super-resolution nodes are output, and the final image is output using a VAE decoder. Here, the final output image can be labeled using the YOLO semantic segmentation model.
[0131] It should be noted that if the latent space image does not meet the preset requirements based on the preset condition judgment, i.e. the quality assessment value, it can correspond to the information of "converting the output text prompt into a real number", and regenerate the corresponding judgment conditions inside the VLM model, and update the Florence2 visual understanding model with attention adaptive mechanism, so as to realize the regeneration and adaptive dynamic adjustment of subsequent images.
[0132] Correspondingly, the AIGC method for generating smart cockpit images provided in this application aims to protect an innovative image generation and processing system. This system achieves efficient and accurate image generation and processing through multi-model integration, fine-tuning technology, quality control mechanisms, text information extraction and analysis, and optimization methods. Furthermore, compared with existing technologies, it has significant improvements in the following aspects: 1. High-quality image generation: The finely tuned XAI ImageGen model significantly improves image quality, better reproducing scenes within the real cockpit and reducing distortion, artifacts, and scale inconsistencies. This makes the generated images more consistent with the objective laws of the natural world, providing higher-quality training data for intelligent cockpit image recognition models.
[0133] 2. Solving the problem of data scarcity: This method can generate a large amount of training data for intelligent cockpit image recognition, effectively solving the problem of data scarcity in the market. The image data generated by the method provided in this application is rich and diverse, covering various possible scenarios and situations inside the cockpit, providing more comprehensive data support for the subsequent training of intelligent cockpit image recognition models.
[0134] 3. Improved Model Training Performance: Due to the higher quality and larger volume of generated images, the intelligent cockpit image recognition model trained using this data has seen significant improvements in accuracy and robustness. For example, the model's accuracy in detecting passenger behavior and lost items within the cockpit may increase from the current 60% to 95%.
[0135] 4. Reduced Development Costs and Time: ComfyUI allows for calling the XAI ImageGen model, making the entire image generation and data preparation process more efficient and convenient. Compared to traditional data collection and processing methods, developers can save significant time and effort, enabling faster model training and deployment.
[0136] 5. Flexibility and Scalability: This solution boasts strong flexibility and scalability. By adjusting the parameters of the XAIImageGen model and the ComfyUI workflow, it can easily adapt to different cockpit scenarios and requirements, generating various types of images of people or objects. Furthermore, this technology can be combined with other image processing and machine learning techniques to further enhance the system's performance and functionality.
[0137] 6. Image correction judgment process: This application has the ability to identify missing conditions or errors in images based on VLM, and the success rate reaches 95% in practical applications. It has the ability to self-check, correct, reject, regenerate, insert and align.
[0138] In other words, the method for generating smart cockpit images based on AIGC provided in this application proposes a solution based on the XAI ImageGen image generation model. This method can generate high-quality images of people or objects within a specific area of the smart cockpit, thereby generating a large amount of training data for smart cockpit image recognition. This application achieves high-quality image data generation by designing a workflow, training and fine-tuning the model, and testing and correcting the generated data, effectively solving the problems of insufficient training data and inadequate accuracy in smart cockpit image recognition in existing technologies.
[0139] Based on the foregoing embodiments, see Figure 5 As shown in the illustration, this application also provides a system for generating smart cockpit images based on AIGC. The system 500 includes: The acquisition module 501 is used to acquire text description information that describes the scene inside the cockpit.
[0140] The image generation module 502 is used to input text description information into the image generation model to obtain an initial intelligent cockpit image set. The image generation model is obtained by fine-tuning the generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using the LoRa method.
[0141] The quality assessment module 503 is used to assess the quality of each image in the initial smart cockpit image set and obtain the quality assessment value of each image.
[0142] The extraction and analysis module 504 is used to extract and analyze the visual content of the image to be optimized when the quality assessment value of the image to be optimized is less than a preset threshold, by using the Florence2 visual understanding model with an attention adaptive mechanism to obtain the semantic information of the image to be optimized; wherein, the image to be optimized is an image in the initial intelligent cockpit image set.
[0143] The repair and optimization module 505 is used to repair and optimize the images to be optimized in the initial smart cockpit image set based on semantic information using inpaint technology, so as to obtain the target smart cockpit image set.
[0144] It should be noted that the description of the above system embodiments is similar to the description of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the system embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0145] It should be noted that, in the embodiments of this application, if the above-described method for generating intelligent cockpit images based on AIGC is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0146] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the steps in any of the methods for generating smart cockpit images based on AIGC described in the above embodiments. Correspondingly, embodiments of this application also provide a computer program product, which, when executed by a processor of an electronic device, is used to implement the steps in any of the methods for generating smart cockpit images based on AIGC described in the above embodiments.
[0147] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0148] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0149] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0150] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0151] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0152] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the device automatic test line to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0153] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0154] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0155] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating an intelligent cockpit image based on AIGC, characterized in that, The method comprises: obtaining text description information for describing a scene in a cockpit; inputting the text description information into an image generation model to obtain an initial intelligent cockpit image set; wherein the image generation model is obtained by fine-tuning a generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using a Lora method; performing quality assessment on each image in the initial intelligent cockpit image set to obtain a quality assessment value of each image; in a case where the quality assessment value of the to-be-optimized image is less than a preset threshold, using a Florence2 visual understanding model with an attention adaptive mechanism to extract and analyze the visual content of the to-be-optimized image to obtain semantic information of the to-be-optimized image; wherein the to-be-optimized image is an image in the initial intelligent cockpit image set; using inpainting technology to repair and optimize the to-be-optimized image in the initial intelligent cockpit image set based on the semantic information to obtain a target intelligent cockpit image set.
2. The method of claim 1, wherein, The construction process of the image generation model comprises: performing data augmentation on the multiple labeled images containing different cockpit environments and user behaviors to obtain a training sample set; using the training sample set as training data, fine-tuning the generative adversarial network using a Lora method to obtain the image generation model; wherein an Adam optimizer with specific parameters is used for parameter optimization during fine-tuning of the pre-trained image generation model using the Lora method, and a cosine annealing learning rate scheduler is used to adjust the learning rate.
3. The method according to claim 1 or 2, characterized in that, The multiple labeled images containing different cockpit environments and user behaviors at least include: multiple labeled sub-images of different lighting conditions, multiple labeled sub-images of different seat positions in the cockpit, multiple labeled sub-images of different interior configurations in the cockpit, and multiple labeled sub-images of different user postures in the cockpit.
4. The method of claim 1, wherein, Before inputting the text description information into the image generation model to obtain the initial intelligent cockpit image set, the method further comprises: adjusting model parameters of the image generation model according to a preset scene requirement to obtain an adjusted image generation model; inputting the text description information into the image generation model to obtain the initial intelligent cockpit image set, comprising: inputting the text description information into the adjusted image generation model to obtain the initial intelligent cockpit image set.
5. The method of claim 4, wherein, Adjusting the model parameters of the image generation model according to the preset scene requirement to obtain the adjusted image generation model, comprising: obtaining an artistic style and image quality requirement of the preset scene requirement; adjusting the generation parameters of the image generation model according to the artistic style and the image quality requirement to obtain adjusted generation parameters; reconfiguring the image generation model according to the adjusted generation parameters to obtain the adjusted image generation model.
6. The method of claim 1, wherein, Performing quality assessment on each image in the initial intelligent cockpit image set to obtain a quality assessment value of each image, comprising: constructing a prompt word template library corresponding to image quality focus points in a cockpit scene; inputting each image in the initial intelligent cockpit image set into a preset visual language model VLM in turn, and performing multi-round dialogue with prompt word templates in the prompt word template library to generate a multi-round text reply corresponding to each image; Quantifying the multi-turn text reply corresponding to each image to obtain a multi-turn quality score corresponding to each image; Fusing the multi-turn quality scores corresponding to each image to obtain a quality evaluation value of each image.
7. The method of claim 1, wherein introducing The Florence2 visual understanding model with attention adaptive mechanism includes: a multi-modal encoder introducing a self-attention layer, and a decoder composed of a mask sub-attention layer and a cross-attention layer.
8. The method of claim 1, wherein, After obtaining the target intelligent cockpit image set by repairing and optimizing the to-be-optimized image in the initial intelligent cockpit image set based on semantic information using the inpaint technology, the method further includes: Taking the target intelligent cockpit image set as training data, iteratively training a preset intelligent cockpit image recognition model until a target intelligent cockpit image recognition model whose output meets a preset condition is obtained.
9. A system for generating an intelligent cockpit image based on AIGC, characterized in that, The system includes: An acquisition module configured to acquire text description information describing a scene in a cockpit; An image generation module configured to input the text description information into an image generation model to obtain an initial intelligent cockpit image set; wherein the image generation model is obtained by fine-tuning a generative adversarial network based on multiple labeled images containing different cockpit environments and user behaviors using a Lora method; A quality evaluation module configured to evaluate the quality of each image in the initial intelligent cockpit image set to obtain a quality evaluation value of each image; An extraction and analysis module configured to, if the quality evaluation value of the to-be-optimized image is less than a preset threshold, extract and analyze the visual content of the to-be-optimized image using the Florence2 visual understanding model with attention adaptive mechanism to obtain semantic information of the to-be-optimized image; wherein the to-be-optimized image is an image in the initial intelligent cockpit image set; A repairing and optimizing module configured to repair and optimize the to-be-optimized image in the initial intelligent cockpit image set based on the semantic information using the inpaint technology to obtain a target intelligent cockpit image set.