Driver driving state image data generation method based on conditional diffusion model
By generating driver driving state image data through a conditional diffusion model, the problem of high cost of driver data acquisition is solved, and efficient synthesis of multi-view images is achieved, supporting low-cost training and testing of intelligent vehicle driver monitoring systems.
Patent Information
- Application Number
- CN202511194259.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-18
AI Technical Summary
Acquiring driver status data is costly and makes it difficult to simulate dangerous driving behaviors, which affects the training and testing of driver monitoring systems.
An image generation method based on a conditional diffusion model is adopted. Through latent space modeling and conditional guidance mechanism, the geometric transformation relationship between the original camera and the target camera is used to generate driver image data under multiple perspectives and camera parameter configurations.
This technology synthesizes driver images under various camera configurations and viewing angles using a limited amount of real-vehicle data, reducing data acquisition costs, improving the realism and stability of image data, and possessing high generalization ability and flexible configuration, thus supporting the training and testing of intelligent vehicle driver monitoring systems.
Smart Images

Figure CN120976047A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent vehicle technology, specifically a method for generating driver driving state image data based on a conditional diffusion model. Background Technology
[0002] In the field of intelligent vehicle technology, the Driver Monitoring System (DMS) is a crucial component of active safety in intelligent vehicles, and its performance directly impacts the ability to prevent traffic accidents and the overall safety level of the vehicle. DMS systems typically rely on technologies such as computer vision and deep learning, using in-vehicle cameras to acquire images of the driver's face and upper body, and then using algorithms to determine the driver's state to provide warnings or safety interventions. To ensure the DMS system accurately identifies the driver's driving state, the training and testing phases of the model must utilize a large amount of high-quality and diverse driver driving state data.
[0003] However, in actual engineering, obtaining sufficient driver status data faces numerous difficulties. A test vehicle often has only one in-cabin camera installed in a single location. To enrich the database, a large number of drivers need to be recruited to collect data under cameras with different installation locations and parameters. This is extremely costly in terms of time, manpower, and funding. Moreover, some dangerous driving behaviors (such as prolonged eye closure or frequent head turning) cannot be frequently simulated in real vehicles for safety reasons, making it difficult to obtain driver status data and affecting model training and testing. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides a method for generating driver driving state image data based on a conditional diffusion model. Using an image diffusion generation network as its core, and combining the geometric transformation relationship between the original camera and the target camera, it achieves generalized generation from a single image to multi-view images through latent spatial modeling and conditional guidance mechanisms. The aim is to synthesize realistic images from other perspectives and camera parameter configurations using a small number of original images.
[0005] The technical solution of this invention is described below in conjunction with the accompanying drawings:
[0006] This invention provides a method for generating driver driving state image data based on a conditional diffusion model, comprising the following steps:
[0007] Step 1: Acquisition of raw driver driving status data and construction of camera parameters;
[0008] A driver driving status data acquisition platform based on a real experimental vehicle was built. Different working conditions were designed, data acquisition schemes were determined, and driving status data was collected. Finally, the collected data was checked, bad data was removed, and data was collected again. At the same time, the intrinsic and extrinsic parameters of the camera used were collected to form the complete perspective information of the original image, and the target camera parameters were set to determine the geometric and optical target conditions of the image data to be generated.
[0009] Step 2: Extraction of latent representation of the original image and construction of geometric control vectors;
[0010] The raw image data collected in step one is encoded into a latent space representation by a trained autoencoder. At the same time, the raw and target camera parameters are embedded into a unified geometric condition vector to provide conditional guidance information for the subsequent generation process.
[0011] Step 3: Generation of latent representation of the target image;
[0012] Using the conditional diffusion model, guided by the latent representation of the original image and the geometric control vector, iterative denoising is performed on random noise to generate a latent representation that meets the requirements of the target image, thus realizing the transfer of the latent representation of the original image to the latent representation of the target image.
[0013] Step 4: Image reconstruction and post-processing;
[0014] The generated latent image representation is mapped back to the image space through a decoder to obtain the image from the target camera's perspective, and post-processing operations are performed as needed.
[0015] Furthermore, the specific method for step one is as follows:
[0016] 11) Select a data acquisition vehicle, and based on the data acquisition vehicle, select a camera device that meets the key parameter requirements of the vehicle DMS camera;
[0017] The camera is installed inside the cabin and is not obstructed by objects inside the vehicle;
[0018] Select data acquisition software and equipment that can support data acquisition personnel in real time monitoring the video recorded by the camera and store video data for different drivers and different working conditions separately.
[0019] 12) Establish a data acquisition plan;
[0020] The data collection scheme includes dangerous driving behaviors and actions, and standardizes these behaviors and actions. It also includes different facial expressions of drivers and designs different lighting conditions.
[0021] 13) Recruit drivers, including those of different genders, ages, heights, and weights;
[0022] Each driver collects data according to a pre-established data collection plan. Before data collection, each driver receives systematic training to clarify the requirements of each action. After data collection is completed, the original driver driving status database is obtained.
[0023] 14) Check the collected driver driving status data, remove data affected by external factors, and re-collect the data;
[0024] 15) Let the original camera be C. s The original image captured is denoted as I. s Completely record the parameters of the cameras used to collect driver's driving status data;
[0025] Collect the following camera intrinsic parameters;
[0026] Focal length: Indicates the focusing ability of a camera, denoted as f. s ;
[0027] Principal point: the coordinates of the image center point (c x c y Set the center point of the image's width and height;
[0028] Resolution: Indicates the number of pixels contained in an image per inch;
[0029] Field of view: The range of spatial angles that a camera can observe, affected by both the focal length and the size of the image sensor, denoted as θ. s ;
[0030] Collect the following camera external parameters:
[0031] Rotation matrix R s : Describes the transformation relationship between the camera coordinate system and the world coordinate system;
[0032] Translation vector t s : Describes the position of the camera center in the world coordinate system;
[0033] 16) Based on the desired target image, design the target camera C. t The parameters include the target focal length f. t Target field of view angle θ t Target resolution (W) t H t ), target rotation matrix R t Target translation vector t t .
[0034] Furthermore, the specific method for step two is as follows:
[0035] 21) A pre-trained image autoencoder is used as the feature extraction network to extract the input image I. s ∈R H×W×3 Transformation into the latent spatial representation z s ∈R h×w×c Where h << H, w << W, H is the height of the original input image in pixel space; W is the width of the original input image in pixel space; h is the height of the latent spatial feature map output by the encoder; w is the width of the latent spatial feature map output by the encoder; R is the camera extrinsic rotation matrix; c is the channel dimension, and the mapping process is expressed as:
[0036] z s =E(I s (1)
[0037] In the formula, E() is the encoder function;
[0038] 22) Normalize each parameter by using min-max scaling or z-score standardization to ensure that all input parameters are within a uniform numerical range;
[0039] 23) Change the original camera parameters P s With the target camera parameter P t Encode them sequentially into two vectors:
[0040] v s =[f s ,θ s ,R s ,t s W s H s (2)
[0041] v t =[f t ,θ t ,R t ,t t W t H t (3)
[0042] Construct the difference vector:
[0043] △v=v t -v s (4)
[0044] In the formula, v s v is the original camera parameter vector; t is the target camera parameter vector; △v is the difference vector between the original camera parameters and the target camera parameters;
[0045] Finally, the vectors are concatenated to form a complete geometric control vector:
[0046] c geom =Concat(v s ,v t ,△v) (5)
[0047] In the formula, c geom This is the complete geometric control vector.
[0048] Furthermore, the specific method for step three is as follows:
[0049] 31) During the training phase, the low-dimensional representation x0 of the original image in the latent space is used as the basis, and Gaussian noise is introduced to construct a noise sequence x1, x2…x T This forms a transition process from data distribution to pure noise, which is modeled as follows:
[0050]
[0051] In the formula, x t The latent representation after adding noise at t time steps; α t is the cumulative retention factor at step t, which controls the degree of retention of the original signal; ∈ is the standard Gaussian noise added at (0,1), i.e., ∈ ~ N(0,1), with a mean of 0 and a covariance matrix that is the identity matrix; t is the current time step; T is the maximum time step.
[0052] The model learning is guided by denoising and reconstructing the target noise ∈, and the following loss function is constructed:
[0053]
[0054] In the formula, L is the expected loss value for the entire training process; ∈ θ The model is based on the current input z t The latent representation z after encoding the original image s The structured conditional vector c, composed of target parameters and camera, geom and the noise value predicted at the current time step t; || || is the L2 norm; This is the expectation over all training samples, noisy samples, and time steps;
[0055] 32) In the actual image generation stage, the model will generate a latent representation from a completely random noisy image. Initially, denoising operations are applied step by step to recover the latent representation of the given constraints, as follows:
[0056]
[0057] In the formula, z is the intermediate latent representation at the k-th iteration; sThis is the latent representation encoded from the original image, used to maintain the consistency of image semantics; c geom The structured conditional vector, composed of target parameters and camera elements, is used to control the geometry of the generated image; t k This represents the current time step; `Denoise()` is the denoising function implemented in a neural network, tasked with predicting the noise contained in the current image. The process iterates T times and outputs the final latent representation. That is, the latent features of the image from the target's perspective;
[0058] The diffusion model's generative network employs a variant of the U-Net structure for feature encoding and reconstruction, supplemented by a conditional control module for embedding the latent vector z from the original image. s With control vector c geom The generation process is guided by a time step embedding module, which maps time step t to a vector and embeds it into each convolutional layer.
[0059] Furthermore, the specific method for step four is as follows:
[0060] 41) Use a pre-trained image autoencoder decoder module D() to generate the latent representation z of the target image. t Mapped to target image I t The process is represented as follows:
[0061] I t =D(z) t (9)
[0062] 42) For the decoded image I t Inverse normalization is performed to convert the pixel values from the [0,1] range output by the model into the actual image values. The process is represented as follows:
[0063]
[0064] The generated image is refined and enhanced using a GAN-based image inpainting module.
[0065] The beneficial effects of this invention are as follows:
[0066] 1) The driver driving state image data generation method based on the conditional diffusion model provided by the present invention utilizes the conditional diffusion mechanism and combines the geometric parameter relationship between the original camera and the target camera to achieve high-fidelity, multi-view image data generation in the latent space.
[0067] 2) This invention can automatically synthesize driver images under various camera configurations and viewing angles with a small amount of real vehicle data collection, which greatly reduces the manpower and time costs of data collection and alleviates the current problem of scarcity of multi-view driver image data.
[0068] 3) This invention performs semantic encoding on the original image, constructs a structured geometric condition vector and inputs it into a diffusion generation network, and then combines it with an image decoder to recover the image spatial data. The final generated image shows good realism and stability in terms of content consistency, geometric correctness and texture quality.
[0069] 4) This invention can not only be used to expand the dataset required for training and testing of DMS system models, but also has high generalization ability and flexible configuration. It can meet the image synthesis requirements under different vehicle and camera layouts, and provides efficient, low-cost and highly controllable image data support for the training optimization and performance evaluation of driver monitoring related algorithms in intelligent vehicles. It has significant practical value and promotion significance. Attached Figure Description
[0070] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0071] Figure 1 This is a flowchart of the present invention;
[0072] Figure 2 This is a structural diagram of the conditional diffusion model. Detailed Implementation
[0073] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0074] Example 1
[0075] See Figure 1 This invention provides a method for generating driver driving state image data based on a conditional diffusion model, comprising the following steps:
[0076] Step 1: Acquisition of raw driver driving status data and construction of camera parameters;
[0077] A driver driving status data acquisition platform based on a real experimental vehicle was built. Different working conditions were designed, and a data acquisition scheme was determined. Drivers of different genders, ages, heights, and weights were recruited to collect driving status data. Finally, the collected data was checked, and data affected by external factors was removed and re-acquired. Simultaneously, the intrinsic parameters (such as focal length, field of view, and resolution) and extrinsic parameters (such as spatial position) of the cameras used for data acquisition were simultaneously set to constitute the complete perspective information of the original image. Target camera parameters were also set to determine the geometric and optical target conditions for the generated image data, as detailed below:
[0078] 11) Select a suitable data acquisition vehicle, and based on the data acquisition vehicle, select a camera device with high resolution, low distortion rate, and wide operating temperature range to meet the key parameter requirements of the vehicle DMS camera.
[0079] Install the camera in a suitable location inside the cabin (such as the inside of the windshield and near the A-pillar) to ensure that the camera can clearly capture the features of the driver's face and hands and other key positions when the driver is in different driving positions. It should also be well concealed, avoiding direct sunlight and being blocked by objects inside the vehicle to avoid blind spots.
[0080] Select appropriate data acquisition software and equipment to enable data acquisition personnel to monitor the video recorded by the camera in real time, and store video data of different drivers and different working conditions separately for later data retrieval.
[0081] 12) Establish a data acquisition scheme with reference to relevant national standards such as GB / T 41797-2022 "Performance Requirements and Test Methods for Driver Attention Monitoring Systems";
[0082] The plan should cover dangerous driving behaviors such as fatigued driving (e.g., closing eyes, yawning) and distracted driving (e.g., using a mobile phone, smoking, or talking to other passengers for extended periods), and standardize these behaviors (specifying the execution standards and duration of each behavior). It should also cover different driver facial states (e.g., bare eyes, wearing glasses, wearing sunglasses, and wearing a mask), and design different lighting conditions (e.g., daytime and nighttime) to enhance the authenticity and comprehensiveness of the raw data.
[0083] 13) Recruit a certain number of drivers with rich driving experience, covering different genders, ages, heights and weights, to fully simulate the individual differences of drivers in real life and ensure the breadth of the collected data;
[0084] Each driver collects data according to a pre-established data collection plan. Before data collection, each driver receives systematic training to clarify the requirements of each action. After data collection is completed, the original driver driving status database is obtained.
[0085] 14) To ensure data quality, carefully examine the collected driver driving status data, remove data that is seriously flawed due to external factors, and re-collect the data;
[0086] 15) Let the original camera be C. s The original image it captured is denoted as I. s To ensure the controllability and consistency of the generated images, the parameters of the cameras used to collect driver's driving status data must be fully recorded.
[0087] Collect the following camera intrinsic parameters;
[0088] Focal length: Indicates the focusing ability of a camera, denoted as f. s ;
[0089] Principal point: the coordinates of the image center point (c x c y In most cameras, this can be approximated as the center point of the image's width and height.
[0090] Resolution: Represents the number of pixels per inch (W). Examples include 640*480, 1280*720, and 1920*1080. It is typically defined as (W). s H s The higher the resolution of an image, the higher the pixel density, and the clearer the image.
[0091] Field of view: The range of spatial angles that a camera can observe, affected by both focal length and sensor size, denoted as θ. s ;
[0092] The above parameters can be obtained through camera calibration or provided directly by the camera manufacturer;
[0093] Collect the following camera external parameters:
[0094] Rotation matrix R s : Describes the transformation relationship between the camera coordinate system and the world coordinate system;
[0095] Translation vector t s : Describes the position of the camera center in the world coordinate system;
[0096] The above parameters can be obtained through calibration using an external positioning system;
[0097] 16) Based on the desired target image, design the target camera C. t The parameters include the target focal length f. t Target field of view angle θ t Target resolution (W) t H t), the target rotation matrix R t , the target translation vector t t etc.;
[0098] Step 2: Extraction of the latent representation of the original image and construction of the geometric control vector;
[0099] Encode the original image data collected in Step 1 into a latent space representation through a trained autoencoder to reduce the dimension and extract high-level semantic features. At the same time, embed the original and target camera parameters into a unified geometric condition vector to provide conditional guidance information for the subsequent generation process, as follows:
[0100] 21) Use a pre-trained image autoencoder as the feature extraction network to convert the input image I s ∈R H×W×3 into a latent space representation z s ∈R h×w×c , where h << H, w << W, H is the height of the original input image in the pixel space; W is the width of the original input image in the pixel space; h is the height of the latent space feature map output by the encoder; w is the width of the latent space feature map output by the encoder; R is the camera extrinsic rotation matrix; c is the channel dimension, and the mapping process is expressed as:
[0101] z s = E(I s ) (1)
[0102] In the formula, E() is the encoder function. The encoder usually uses a convolutional neural network (such as ResNet, Unet-Encoder, etc.). Through the encoder, local structures, edge information, etc. in the image are abstracted into a compact latent vector. This representation is more stable and easier to model than pixels and is suitable for the continuous sampling operation of the subsequent diffusion process;
[0103] 22) Since there are differences in the order of magnitude of different parameters (such as focal length, field of view angle, resolution, etc.), directly inputting them will cause unstable network training. Therefore, normalize each parameter and use min-max scaling or z-score normalization to ensure that all input parameters are within a unified numerical range (such as [0,1]);
[0104] 23) Encode the original camera parameter P s and the target camera parameter P t into two vectors in sequence:
[0105] v s = [f s , θ s , R s , t s , W s , Hs (2)
[0106] v t =[f t ,θ t ,R t ,t t W t H t (3)
[0107] To enhance the model's sensitivity to changes in viewpoint, a difference vector can be further constructed:
[0108] △v=v t -v s (4)
[0109] Finally, the vectors are concatenated to form a complete geometric control vector:
[0110] c geom =Concat(v s ,v t ,△v) (5).
[0111] Step 3: Generation of latent representation of the target image;
[0112] Using a conditional diffusion model, guided by the latent representation of the original image and the geometric control vector, iterative denoising from random noise generates a latent representation that meets the requirements of the target image, achieving the transfer of the latent representation of the original image to the latent representation of the target image, as detailed below:
[0113] 31) During the training phase, the low-dimensional representation x0 of the original image in the latent space is used as the basis, and Gaussian noise is gradually introduced to construct a noise sequence x1, x2…x T This process forms a transition from data distribution to pure noise, which is modeled as follows:
[0114]
[0115] In the formula, x t The latent representation after adding noise at t time steps; α t is the cumulative retention factor at step t, which controls the degree of retention of the original signal; ∈ is the standard Gaussian noise added at (0,1), i.e., ∈ ~ N(0,1), with a mean of 0 and a covariance matrix that is the identity matrix; t is the current time step; T is the maximum time step.
[0116] The model learning is guided by denoising and reconstructing the target noise ∈, and the following loss function is constructed:
[0117]
[0118] In the formula, L is the expected loss value for the entire training process; ∈ θ The model is based on the current input z t The latent representation z after encoding the original image s The structured conditional vector c, composed of target parameters and camera, geom and the noise value predicted at the current time step t; || || is the L2 norm; This is the expectation over all training samples, noisy samples, and time steps;
[0119] This loss function measures the model's prediction accuracy against real noise, thereby enhancing the model's ability to recover reasonable underlying representations from noise under specific conditions;
[0120] 32) In the actual image generation stage, the model will generate a latent representation from a completely random noisy image. Initially, denoising operations are applied step by step to recover the latent representation of the given constraints, as follows:
[0121]
[0122] In the formula, z is the intermediate latent representation at the k-th iteration; s This is the latent representation encoded from the original image, used to maintain the consistency of image semantics; c geom The structured conditional vector, composed of target parameters and camera elements, is used to control the geometry of the generated image; t k This represents the current time step; `Denoise()` is a denoising function implemented in a neural network, whose task is to predict the noise contained in the current image. This process is iterated T times and outputs the final latent representation. That is, the latent features of the image from the target's perspective;
[0123] The diffusion model's generative network employs a variant of the U-Net structure, with the backbone used for feature encoding and reconstruction, supplemented by a conditional control module that embeds the latent vector z from the original image. s With control vector c geom The generation process is guided by a time step embedding module, which maps time step t to a vector and embeds it into each convolutional layer.
[0124] The diffusion model includes:
[0125] The latent space coding module is used to map the original image from pixel space to a low-dimensional latent space;
[0126] The conditional control module is used to embed the latent vector zs of the original graph and the control vector cgeom;
[0127] The time step embedding module is used to map time step t into a vector and embed it into each convolutional layer;
[0128] The denoising generation module is used to iteratively denoise in the latent space and gradually recover the latent representation of the target image.
[0129] Step 4: Image reconstruction and post-processing;
[0130] The generated latent image representation is mapped back to the image space through a decoder to obtain the image from the target camera's perspective. Post-processing operations such as dynamic range normalization, detail enhancement, and artifact suppression can be performed as needed to ensure high realism and consistency in the generated image's structure, texture, and geometry, as detailed below:
[0131] 41) Use a pre-trained image autoencoder decoder module D() to generate the latent representation z of the target image. t Mapped to target image I t This process can be represented as:
[0132] I t =D(z) t (9)
[0133] 42) For the decoded image I t Inverse normalization is performed to convert the pixel values from the [0,1] range output by the model into the actual image values. A common approach is to normalize to the range of [0.255], which can be represented as:
[0134]
[0135] Use GAN-based image inpainting modules, such as ESRGAN, to refine and enhance the generated images, improving their clarity and realism.
[0136] In summary, this invention provides a method for generating driver driving state image data based on a conditional diffusion model, which can synthesize other images based on a small amount of real-world data, providing strong data support for the research and development and testing of DMS systems.
[0137] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for generating driver driving state image data based on a conditional diffusion model, characterized in that, Includes the following steps: Step 1: Acquisition of raw driver driving status data and construction of camera parameters; A driver driving status data acquisition platform based on a real experimental vehicle was built. Different working conditions were designed, data acquisition schemes were determined, and driving status data was collected. Finally, the collected data was checked, bad data was removed, and data was collected again. At the same time, the intrinsic and extrinsic parameters of the camera used were collected to form the complete perspective information of the original image, and the target camera parameters were set to determine the geometric and optical target conditions of the image data to be generated. Step 2: Extraction of latent representation of the original image and construction of geometric control vectors; The raw image data collected in step one is encoded into a latent space representation by a trained autoencoder. At the same time, the raw and target camera parameters are embedded into a unified geometric condition vector to provide conditional guidance information for the subsequent generation process. Step 3: Target image latent representation generation; Using the conditional diffusion model, guided by the latent representation of the original image and the geometric control vector, iterative denoising is performed on random noise to generate a latent representation that meets the requirements of the target image, thus realizing the transfer of the latent representation of the original image to the latent representation of the target image. Step 4: Image reconstruction and post-processing; The generated latent image representation is mapped back to the image space through a decoder to obtain the image from the target camera's perspective, and post-processing operations are performed as needed.
2. The method for generating driver driving state image data based on a conditional diffusion model according to claim 1, characterized in that, The specific method for step one is as follows: 11) Select a data acquisition vehicle, and based on the data acquisition vehicle, select a camera device that meets the key parameter requirements of the vehicle DMS camera; The camera is installed inside the cabin and is not obstructed by objects inside the vehicle; Select data acquisition software and equipment that can support data acquisition personnel in real time monitoring the video recorded by the camera and store video data for different drivers and different working conditions separately. 12) Establish a data acquisition plan; The data collection scheme includes dangerous driving behaviors and actions, and standardizes these behaviors and actions. It also includes different facial expressions of drivers and designs different lighting conditions. 13) Recruit drivers, including those of different genders, ages, heights, and weights; Each driver collects data according to a pre-established data collection plan. Before data collection, each driver receives systematic training to clarify the requirements of each action. After data collection is completed, the original driver driving status database is obtained. 14) Check the collected driver driving status data, remove data affected by external factors, and re-collect the data; 15) Let the original camera be C. s The original image captured is denoted as I. s Completely record the parameters of the cameras used to collect driver's driving status data; Collect the following camera intrinsic parameters; Focal length: Indicates the focusing ability of a camera, denoted as f. s ; Principal point: the coordinates of the image center point (c x c y Set the center point of the image's width and height; Resolution: Indicates the number of pixels contained in an image per inch; Field of view: The range of spatial angles that a camera can observe, affected by both the focal length and the size of the image sensor, denoted as θ. s ; Collect the following camera extrinsic parameters: Rotation matrix R s : Describes the transformation relationship between the camera coordinate system and the world coordinate system; Translation vector t s : Describes the position of the camera center in the world coordinate system; 16) Based on the desired target image, design the target camera C. t The parameters include the target focal length f. t Target field of view angle θ t Target resolution (W) t H t ), target rotation matrix R t Target translation vector t t .
3. The method for generating driver driving state image data based on a conditional diffusion model according to claim 1, characterized in that, The specific method for step two is as follows: 21) A pre-trained image autoencoder is used as the feature extraction network to extract the input image I. s ∈R H×W×3 Transformation into latent spatial representation z s ∈R h×w×c Where h << H, w << W, H is the height of the original input image in pixel space; W is the width of the original input image in pixel space; h is the height of the latent spatial feature map output by the encoder; w is the width of the latent spatial feature map output by the encoder; R is the camera extrinsic rotation matrix; c is the channel dimension, and the mapping process is expressed as: z s =E(I s ) (1) In the formula, E() is the encoder function; 22) Normalize each parameter by using min-max scaling or z-score standardization to ensure that all input parameters are within a uniform numerical range; 23) Change the original camera parameters P s With the target camera parameter P t Encode them sequentially into two vectors: v s =[f s ,θ s ,R s ,t s ,W s ,H s ] (2) v t =[f t ,θ t ,R t ,t t ,W t ,H t ] (3) Construct the difference vector: △v=v t -v s (4) In the formula, v s v is the original camera parameter vector; t is the target camera parameter vector; △v is the difference vector between the original camera parameters and the target camera parameters; Finally, the vectors are concatenated to form a complete geometric control vector: c geom =Concat(in s ,in t ,△v) (5) In the formula, c geom This is the complete geometric control vector.
4. The method for generating driver driving state image data based on a conditional diffusion model according to claim 1, characterized in that, The specific method for step three is as follows: 31) During the training phase, the low-dimensional representation x0 of the original image in the latent space is used as the basis, and Gaussian noise is introduced to construct a noise sequence x1, x2…x T This forms a transition process from data distribution to pure noise, which is modeled as follows: In the formula, x t The latent representation after adding noise at t time steps; α t is the cumulative retention factor at step t, which controls the degree of retention of the original signal; ∈ is the standard Gaussian noise added at (0,1), i.e., ∈ ~ N(0,1), with a mean of 0 and a covariance matrix that is the identity matrix; t is the current time step; T is the maximum time step. The model learning is guided by denoising and reconstructing the target noise ∈, and the following loss function is constructed: In the formula, L is the expected loss value for the entire training process; ∈ θ The model is based on the current input z t The latent representation z after encoding the original image s The structured conditional vector c, composed of target parameters and camera, geom and the predicted noise value at the current time step t; |||| is the L2 norm; This is the expectation over all training samples, noisy samples, and time steps; 32) In the actual image generation stage, the model will generate a latent representation from a completely random noisy image. Initially, denoising operations are applied step by step to recover the latent representation of the given constraints, as follows: In the formula, z is the intermediate latent representation at the k-th iteration; s This is the latent representation encoded from the original image, used to maintain the consistency of image semantics; c geom The structured conditional vector, composed of target parameters and camera elements, is used to control the geometry of the generated image; t k This represents the current time step; `Denoise()` is the denoising function implemented in a neural network, tasked with predicting the noise contained in the current image. The process iterates T times and outputs the final latent representation. That is, the latent features of the image from the target's perspective; The diffusion model's generative network employs a variant of the U-Net structure for feature encoding and reconstruction, supplemented by a conditional control module that embeds the latent vector z from the original image. s With control vector c geom The generation process is guided by a time step embedding module, which maps time step t to a vector and embeds it into each convolutional layer.
5. The method for generating driver driving state image data based on a conditional diffusion model according to claim 1, characterized in that, The specific method for step four is as follows: 41) Use a pre-trained image autoencoder decoder module D() to map the latent representation of the target image to the target image I. t The process is represented as follows: I t =D(z t ) (9) 42) For the decoded image I t Inverse normalization is performed to convert the pixel values from the [0,1] range output by the model into the actual image values. The process is represented as follows: The generated image is refined and enhanced using a GAN-based image inpainting module.
Citation Information
Cited By
Driving video generation method based on space-time factorization architecture and hybrid modulation
CN122093518A