Driving decision analysis method and system based on multi-modal large model
By generating panoramic images from multi-view cameras and combining them with multimodal large model analysis, the specific application problem of multimodal large models in understanding scenarios in autonomous driving is solved, thereby improving the reliability and accuracy of driving decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
In the field of autonomous driving, existing multimodal large models do not specifically describe how to understand scenarios and lack practical application capabilities.
By configuring multi-view cameras to collect continuous image data, generating panoramic images and stitching them together, using a multimodal large model for scene analysis and driving decision-making, and combining text prompts for output.
It enables comprehensive visual feature analysis of targets around the vehicle, improving the reliability and accuracy of driving decisions and outputting key information and driving operation suggestions.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to a driving decision analysis method and system based on a multimodal large model. Background Technology
[0002] In recent years, multimodal large models have made progress in cross-modal understanding, demonstrating unprecedented "world knowledge," powerful contextual reasoning capabilities, and cross-modal understanding abilities, and have been widely applied in the field of autonomous driving. Autonomous driving not only needs to "see" the road, but also needs to "understand" the scenario. For example, understanding whether a vehicle ahead's flashing lights indicates a signal for me to go first or a warning against changing lanes, or predicting that a child playing football on the roadside might run into the road in the next second. This deep semantic understanding and intent prediction is precisely where the strength of large models lies. Deeply integrating multimodal large models with autonomous driving technology is considered a key technological path to achieving fully driverless driving and is one of the most intensely competitive areas in the industry today.
[0003] Patent CN119514635A provides an autonomous driving model, training, and autonomous driving method based on a multimodal large model. The specific implementation scheme includes: acquiring a training corpus dataset, comprising at least visual text alignment corpus and spatial understanding training corpus for autonomous driving scenarios; encoding the visual data in the visual text alignment corpus using a visual encoder to obtain encoded data; mapping the encoded data using a mapping layer; processing the mapped encoded data, text data, and spatial understanding training corpus using a generation layer to obtain a first prediction result and a second prediction result of the autonomous driving model; and adjusting the parameters of the autonomous driving model based at least on the first and second prediction results. The autonomous driving model disclosed in this patent possesses both multimodal information understanding capabilities and reasoning capabilities in autonomous driving scenarios.
[0004] However, the patent only provides a basic framework for applying multimodal large models to the field of autonomous driving, without specifically describing how to understand the scenario or what capabilities the multimodal large models provide to solve specific autonomous driving problems. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention uses continuous frame panoramic images for analysis, which can obtain the visual features of all targets of interest around the vehicle. Driving decisions rely on these surrounding targets and road conditions, making the decisions more reliable. This invention analyzes the visual features of the panoramic image using a multimodal large model, simulating human driving decisions based on current panoramic visual information.
[0006] To achieve the above objectives, the technical solution of the present invention provides a driving decision analysis method based on a multimodal large model, comprising the following steps: S1: acquiring continuous image data during vehicle movement using a multi-view camera configured on a moving vehicle; S2: stitching and fusing the acquired continuous multi-view image data into a panoramic image using camera parameters, thereby obtaining a continuous panoramic image sequence; S3: sampling the continuous panoramic image sequence obtained in step S2 at a fixed frequency to generate a sampled panoramic image sequence; S4: inputting the sampled panoramic image sequence and text prompts into a multimodal large model; S5: the multimodal large model performs scene analysis and driving decision judgment based on the sampled panoramic image sequence, and outputs the required information according to the input text prompts.
[0007] Further, in step S2, the multi-view image data is stitched and fused according to the following steps: S21: Distortion removal processing is performed on each input image; S22: The pixel coordinates of the distortion-removed image are back-projected onto 3D rays in the camera coordinate system, and the 3D rays are normalized to unit vectors to obtain spherical coordinates; S23: Based on the equidistant cylindrical projection method, the spherical coordinates are mapped to equidistant cylindrical image coordinates to complete the reprojection from the multi-view images to the equidistant cylindrical coordinate system; S24: Image registration and alignment are performed on the reprojected multi-view images; S25: The registered images are stitched and fused to generate a seamless panoramic image.
[0008] Further, step S24 specifically includes: S241: estimating the transformation relationship between adjacent images using a feature point detection and matching algorithm; S242: optimizing the camera pose in spherical space to ensure reprojection consistency; S243: performing global optimization of multi-image consistency using a method based on spherical harmonics or graph optimization.
[0009] Further, step S25 specifically includes: S251: For each pixel of the equidistant cylindrical image, reverse mapping back to the original image plane and obtaining the color value using an interpolation algorithm; S252: When a single pixel is covered by multiple original images, pixel fusion is performed; S253: Calculate the illumination difference between the images and perform exposure compensation and color correction on the images; S254: Use multi-band fusion or feathering techniques to perform fusion processing on the image seams.
[0010] Furthermore, the multimodal large model is trained according to the following steps: constructing training data containing prompt words, continuous frame panoramic image sequences, and manually labeled standard answers; using the training data to train the multimodal large model using instruction fine-tuning; and using a validation set to validate the model after each training round, selecting the model with the best performance as the final model.
[0011] Furthermore, after the multimodal large model is trained, the model is iteratively optimized by collecting decision biases from human driver intervention events or simulation replays.
[0012] Furthermore, the output of the multimodal large model includes auxiliary judgment information, key information, and driving operation suggestions, wherein the key information includes important targets and their related attribute information.
[0013] Furthermore, the auxiliary judgment information includes at least one of the following: weather, time, visibility, image quality, current lane of the vehicle, current road conditions, current lane information, and current navigation instructions.
[0014] Furthermore, the important targets include at least one of the following: motor vehicles, non-motor vehicles, pedestrians, animals, obstacles, and road abnormalities.
[0015] The technical solution of this invention also provides a driving decision analysis system based on a multimodal large model, which includes the following modules: a multi-view image acquisition module: acquiring continuous image data during vehicle movement using multi-view cameras configured on the moving vehicle; a panoramic image generation module: stitching and fusing the acquired continuous multi-view image data into a panoramic image using camera parameters to obtain a continuous panoramic image sequence; a temporal sampling module: sampling the continuous panoramic image sequence obtained by the panoramic image generation module at a fixed frequency to generate a sampled panoramic image sequence; a multimodal large model input module: inputting the sampled panoramic image sequence and text prompts into the multimodal large model; and a multimodal large model processing module: the multimodal large model performs scene analysis and driving decision judgment based on the sampled panoramic image sequence, and outputs the required information according to the input text prompts. Attached Figure Description
[0016] none Detailed Implementation
[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] This invention stitches and merges images acquired from multiple perspectives into a panoramic image. For example, it merges and stitches photos taken by cameras at the left rear, left front, front, right front, and right rear of the vehicle. Using image generation technology, it achieves seamless stitching, ensuring the final panoramic image has no duplicate parts. This allows the model to obtain visual information from the entire field of view of the areas of interest around the vehicle, enabling comprehensive judgment. The panoramic image can encompass all key targets during the vehicle's movement, detecting and analyzing non-motorized vehicles, motorized vehicles, pedestrians, and obstacles on the road, ultimately outputting driving decisions.
[0019] The specific implementation of this invention mainly includes the following aspects:
[0020] 1. Configure multi-view cameras on moving vehicles, such as cameras installed in front, left front, right front, left rear, and right rear, to collect image data during vehicle movement;
[0021] 2. Using camera intrinsic parameters, extrinsic parameters, distortion and other parameters, project multi-view images onto a unified coordinate system to generate a panoramic fusion image.
[0022] This invention uses equirectangular projection to stitch and fuse acquired multi-view images into a panoramic image. Equirectangular projection is a method of mapping points on a sphere (or viewing sphere) to a two-dimensional plane.
[0023] Longitude λ is mapped to the horizontal coordinate x of the image.
[0024] Latitude φ is mapped to the vertical coordinate y of the image.
[0025] The image width corresponds to 360° longitude, and the height corresponds to 180° latitude (from -90° to +90°).
[0026] Mapping formula (sphere → plane):
[0027]
[0028] in,
[0029] λ∈[-π,π] (radians)
[0030]
[0031] Image width w, height h
[0032] Inverse mapping (plane → sphere):
[0033]
[0034] The present invention specifically includes the following steps for stitching and merging a panorama:
[0035] (1) Reprojection of input image onto equidistant cylindrical surface
[0036] Assuming the input consists of images taken by multiple ordinary perspective cameras, each image must first be reprojected onto an equidistant cylindrical coordinate system:
[0037] step:
[0038] Obtain camera intrinsic parameters (focal length f, principal point (cx, cy), distortion parameters, etc.)
[0039] Perform distortion correction on each input image (e.g., using OpenCV's undistort function).
[0040] 3D rays that project pixel coordinates backwards into the camera coordinate system:
[0041]
[0042] Normalizing the 3D rays to unit vectors yields spherical coordinates:
[0043]
[0044] Map to the coordinates of the equidistant cylindrical image (using the mapping formula above).
[0045] (2) Image registration and alignment
[0046] Use feature point detection and matching (such as SIFT + FLANN + RANSAC) to estimate homography or relative rotation between adjacent images.
[0047] Optimize the pose in spherical space (e.g., using Bundle Adjustment or spherical optimization) to ensure reprojection consistency.
[0048] Global optimization for multi-graph consistency can be performed using spherical harmonics or graph-based optimization.
[0049] (3) splicing and fusion
[0050] (3.1) Sampling and Interpolation
[0051] For each pixel of the target equidistant cylindrical image, reverse map it back to the original image plane.
[0052] Use bilinear interpolation or higher quality interpolation (such as Lanczos) to obtain color values.
[0053] If a pixel is covered by multiple images, they need to be merged.
[0054] (3.2) Exposure compensation and color correction
[0055] Calculate the lighting differences between images and use gain adjustment (such as OpenCV's exposure compensator).
[0056] Common method: linear correction based on the average brightness difference of the overlapping areas.
[0057] (3.3) Seam Blending
[0058] Use multi-band blending or feathering to reduce seams.
[0059] In overlapping regions, pixels are blended according to a weight map (such as the inverse of the distance from the boundary).
[0060] (4) Boundary treatment and pole problems
[0061] Pole distortion: In equidistant cylindrical projection, the pole (φ = ±90°) is compressed into a single point in the horizontal direction, resulting in severe stretching at the top / bottom.
[0062] Solution: Avoid placing important content near the poles; or use alternative projections such as cube maps.
[0063] Image boundaries: The left and right boundaries of equidistant cylindrical images should be seamlessly connected (λ = -π and λ = π correspond to the same meridian).
[0064] Periodic boundaries need to be handled during feature matching and fusion.
[0065] 3. Sample consecutive frames at a fixed frequency to generate a sequence of consecutive frames;
[0066] Frames are extracted from a continuous sequence of images at fixed time intervals (frequency) to generate multiple new sequences of image frames that are evenly distributed in time.
[0067] When performing sampling, first determine the following values: the sampling frequency is FPS. t The sampling interval is calculated as Δt = 1.0 / FPS. t The frame rate of the original video stream or image sequence is FPS. s Each data point requires an image sequence of length L frames.
[0068] The steps are as follows:
[0069] (1) Calculate the number of frames to be skipped: frame_interval = round(FPS) s / FPS t For example, if the source is 30 FPS and the target is 5 FPS, then frame_interval = 6, meaning one frame is sampled every 6 frames.
[0070] (2) Initialize a global counter count_global = 0, generate a sequence counter count_local = 0, and generate an empty list of image sequences.
[0071] (3) In a loop, continuously read the original video frames or image sequences.
[0072] (4) If count_global % frame_interval equals 0 and count_local is less than L, then add the current frame to the generated image sequence list, increment count_global by 1, and increment count_local by 1.
[0073] (5) When count_local equals L-1, set count_local to 0, generate the image sequence list, add it to the final data list, and clear the generated image sequence.
[0074] (6) Loop until the current original video or image sequence has been read and the final data list is obtained.
[0075] 4. Construct a dedicated multimodal large model architecture, defining the model's capabilities and input / output content, as follows:
[0076] Model capability definition
[0077] The model has the following output capabilities, which can be controlled by input prompts. Based on the visual information provided by the continuous frame panoramic images, the multimodal large model will perform some analysis on the current scene. The auxiliary judgment information obtained from the analysis includes weather (sunny, rainy, snowy, foggy), time (early morning or evening, daytime, nighttime), visibility (clear, limited, severely insufficient), image quality (whether the lens has dirt or water stains that affect decision-making; whether it is overexposed; whether there are decoding anomalies that affect decision-making), the current lane of the vehicle (motor vehicle lane, non-motor vehicle lane, sidewalk, mixed road, etc.), the current road surface condition (good, slippery, uneven, unpaved road, completely undrivable, water accumulation, snow accumulation), the current lane information (unknown, merging lane, diverging lane, straight, left turn, right turn, U-turn, bus lane), and the current navigation command (corresponding icons are drawn on the panoramic image, and the navigation command includes straight, left turn, right turn, U-turn). The key information derived from the analysis includes the current vehicle state. The multimodal large model analyzes the vehicle's current driving state based on continuous frame information, including whether it is driving normally in the current lane, turning left, turning right, changing lanes to the left, changing lanes to the right, avoiding obstacles to the left or right, cutting in from the left or right, waiting in a queue, driving forward on the line, driving against traffic, making a U-turn, or reversing. Simultaneously, the key information also includes important targets. Based on the vehicle's current state, it determines which targets are important and require attention, such as any objects that may overlap with the vehicle's driving path or affect its current driving state. These typically include motor vehicles, non-motor vehicles, pedestrians, animals, obstacles, and road anomalies. The model also analyzes how these targets affect the final decision-making suggestions. Some targets directly affect the final operation suggestions or the current vehicle state, while others do not affect the operation suggestions or the vehicle state, but their dynamics need to be monitored during driving. For different key targets, the model is capable of outputting some related attributes. For pedestrians, non-motorized vehicles, and motorized vehicles, the model determines the target's direction relative to the vehicle (left, right, left front, right front, across left and right front, left rear, right rear), the target's relative speed to the vehicle (approaching the vehicle, moving away from the vehicle, maintaining distance), whether the target is stationary or moving, and whether the target exhibits abnormal behavior, typically including running red lights or not walking in the designated lane. For motorized vehicles, the model also identifies additional abnormal situations, such as open doors, abnormal parking, special vehicles, light status (brake lights on, left turn signal flashing, right turn signal flashing, hazard lights, no lights), and lane-changing status (changing lanes to the center or crossing the line, changing lanes to the left or crossing the line, changing lanes to the right or crossing the line, no lane-changing status). For obstacles, they are divided into non-compact obstacles and compactable obstacles, and the model provides subcategories and specific obstacle names.For road anomalies, the model identifies specific anomaly categories, such as road closures, speed bumps, potholes, slopes, and abnormal manhole covers (severely raised or sunken). Most importantly, the multimodal large model provides driving operation suggestions, including maintaining the current vehicle state, swerving to the left, swerving to the right, swerving to either side, waiting in a queue, stopping and waiting, and continuing straight. All output results are presented in JSON format for easy parsing.
[0078] Model Input: A continuous sequence of panoramic images and text prompts. The text prompts can be specified based on the desired output based on the model's capabilities. For example, if you want the model to output the location of key targets and current driving operation suggestions, the text prompts could be as follows: "You are an experienced driver. The input image is a panoramic view stitched together from images taken by cameras at the front, left front, left rear, right front, and right rear of the vehicle. The image is divided into five parts, showing the vehicle's field of view from left to right: left rear, left front, front, right front, and right rear. The vehicle needs to go straight / turn right / turn left / make a U-turn (selected according to navigation instructions). Please identify the key targets that need attention. Important targets include non-motorized vehicles, motorized vehicles, pedestrians, animals, non-compact obstacles, compactible obstacles, and abnormal road conditions. Output all target types and related attributes, their impact, location coordinates, and current driving operation suggestions in JSON format."
[0079] Model Output: The output format is JSON, containing the model fields to be output. Taking the important object detection task as an example, the output example is {["bbox":[x1,x2,y1,y1],"type":object type,"impact of the object":impact], ["bbox":[x1,x2,y1,y1],"type":object type,"impact of the object":impact]}, where x1 and y1 refer to the x and y coordinates of the upper left corner of the bounding box, and x2 and y2 refer to the x and y coordinates of the lower right corner of the bounding box.
[0080] 5. For each image sequence and model capability definition, develop an annotation scheme and manually annotate the relevant content on a per-image-sequence basis.
[0081] 6. Construct a multimodal large model architecture
[0082] Model architecture:
[0083] (1) Vision Encoder: such as ViT (Vision Transformer), which encodes images into visual tokens.
[0084] (2) Language Model Backbone: Processes text and multimodal tokens, and understands the input text (or the converted multimodal information).
[0085] (3) Modality Alignment: Maps visual tokens to the embedding space of LLM through a learnable projection layer (such as MLP).
[0086] (4) Multimodal input format: Supported Labels can be embedded in image paths or base64 and mixed with text input.
[0087] 7. The trained multimodal large model is used for tasks such as important target detection and driving decision-making. The training steps are as follows:
[0088] (1) Organize the continuous frame panoramic image data and model sequence into a training format. Each training data includes prompt words, continuous frame panoramic image sequence, and manually labeled standard answers to construct high-quality multimodal instruction-response pair data.
[0089] (2) Employing instruction fine-tuning, the multimodal large model is fine-tuned using training data, specifically through full parameter fine-tuning, LoRa fine-tuning, etc. In this stage, the cross-entropy loss is calculated for the generated part (i.e., the answer that the model should output). Through training, the model acquires the following capabilities: multimodal understanding (image + text joint semantics), instruction following ability (such as question answering, description, reasoning, creation, etc.), safe, useful, and consistent generation behavior.
[0090] (3) After each training round, the validation set is used for validation, and the model with the best performance is selected as the final model.
[0091] 8. Collect decision-making biases from human driver intervention events or simulation replays, and perform iterative optimization of the model.
[0092] In an embodiment of the present invention, a driving decision analysis method based on a multimodal large model is provided, comprising the following steps: S1: acquiring continuous image data during vehicle movement using a multi-view camera configured on a moving vehicle; S2: stitching and fusing the acquired continuous multi-view image data into a panoramic image using camera parameters, thereby obtaining a continuous panoramic image sequence; S3: sampling the continuous panoramic image sequence obtained in step S2 at a fixed frequency to generate a sampled panoramic image sequence; S4: inputting the sampled panoramic image sequence and text prompts into a multimodal large model; S5: the multimodal large model performs scene analysis and driving decision judgment based on the sampled panoramic image sequence, and outputs the required information according to the input text prompts.
[0093] Further, in step S2, the multi-view image data is stitched and fused according to the following steps: S21: Distortion removal processing is performed on each input image; S22: The pixel coordinates of the distortion-removed image are back-projected onto 3D rays in the camera coordinate system, and the 3D rays are normalized to unit vectors to obtain spherical coordinates; S23: Based on the equidistant cylindrical projection method, the spherical coordinates are mapped to equidistant cylindrical image coordinates to complete the reprojection from the multi-view images to the equidistant cylindrical coordinate system; S24: Image registration and alignment are performed on the reprojected multi-view images; S25: The registered images are stitched and fused to generate a seamless panoramic image.
[0094] Further, step S24 specifically includes: S241: estimating the transformation relationship between adjacent images using a feature point detection and matching algorithm; S242: optimizing the camera pose in spherical space to ensure reprojection consistency; S243: performing global optimization of multi-image consistency using a method based on spherical harmonics or graph optimization.
[0095] Further, step S25 specifically includes: S251: For each pixel of the equidistant cylindrical image, reverse mapping back to the original image plane and obtaining the color value using an interpolation algorithm; S252: When a single pixel is covered by multiple original images, pixel fusion is performed; S253: Calculate the illumination difference between the images and perform exposure compensation and color correction on the images; S254: Use multi-band fusion or feathering techniques to perform fusion processing on the image seams.
[0096] Furthermore, the multimodal large model is trained according to the following steps: constructing training data containing prompt words, continuous frame panoramic image sequences, and manually labeled standard answers; using the training data to train the multimodal large model using instruction fine-tuning; and using a validation set to validate the model after each training round, selecting the model with the best performance as the final model.
[0097] Furthermore, after the multimodal large model is trained, the model is iteratively optimized by collecting decision biases from human driver intervention events or simulation replays.
[0098] Furthermore, the output of the multimodal large model includes auxiliary judgment information, key information, and driving operation suggestions, wherein the key information includes important targets and their related attribute information.
[0099] Furthermore, the auxiliary judgment information includes at least one of the following: weather, time, visibility, image quality, current lane of the vehicle, current road conditions, current lane information, and current navigation instructions.
[0100] Furthermore, the important targets include at least one of the following: motor vehicles, non-motor vehicles, pedestrians, animals, obstacles, and road abnormalities.
[0101] In another embodiment of the present invention, a driving decision analysis system based on a multimodal large model is also provided, comprising the following modules: a multi-view image acquisition module: acquiring continuous image data during vehicle movement using multi-view cameras configured on a moving vehicle; a panoramic image generation module: stitching and fusing the acquired continuous multi-view image data into a panoramic image using camera parameters to obtain a continuous panoramic image sequence; a temporal sampling module: sampling the continuous panoramic image sequence obtained by the panoramic image generation module at a fixed frequency to generate a sampled panoramic image sequence; a multimodal large model input module: inputting the sampled panoramic image sequence and text prompts into the multimodal large model; and a multimodal large model processing module: the multimodal large model performs scene analysis and driving decision judgment based on the sampled panoramic image sequence, and outputs the required information according to the input text prompts.
[0102] The beneficial technical effects of the technical solution of the present invention are as follows:
[0103] (1) Panoramic perception: Based on the fusion of panoramic images from various perspectives, the model can extract the visual features of all important targets during vehicle movement, making the prediction results more reliable;
[0104] (2) Continuous frame analysis: Analysis based on continuous frames provides necessary temporal clues for some inferences, making the prediction results more accurate;
[0105] (3) Multimodal deep semantic understanding: Utilize multimodal large model to deeply understand visual information and conduct multidimensional analysis on driving operation suggestions.
[0106] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A driving decision analysis method based on a multimodal large model, characterized in that, Includes the following steps: S1: Collect continuous image data during the vehicle's movement by using a multi-view camera configured on the moving vehicle; S2: Using camera parameters, the acquired continuous multi-view image data are stitched and fused into a panoramic image, thereby obtaining a continuous panoramic image sequence; S3: Sample the continuous panoramic image sequence obtained in step S2 at a fixed frequency to generate a sampled panoramic image sequence; S4: Input the sampled panoramic image sequence and text prompts into the multimodal large model; S5: The multimodal large model performs scene analysis and driving decision-making based on the sampled panoramic image sequence, and outputs the required information according to the input text prompts.
2. The method according to claim 1, characterized in that, In step S2, the multi-view image data is stitched and fused according to the following steps: S21: Perform distortion correction processing on each input image; S22: Project the pixel coordinates of the distorted image back onto a 3D ray in the camera coordinate system, and normalize the 3D ray into a unit vector to obtain spherical coordinates. S23: Based on the equidistant cylindrical projection method, the spherical coordinates are mapped to the equidistant cylindrical image coordinates to complete the reprojection from multi-view images to the equidistant cylindrical coordinate system; S24: Perform image registration and alignment on the reprojected multi-view images; S25: Stitch and merge the registered images to generate a seamless panoramic image.
3. The method according to claim 2, characterized in that, Step S24 specifically includes: S241: Use feature point detection and matching algorithms to estimate the transformation relationship between adjacent images; S242: Optimize camera pose in spherical space to ensure reprojection consistency; S243: Use a method based on spherical harmonics or graph optimization to perform global optimization of multi-graph consistency.
4. The method according to claim 3, characterized in that, Step S25 specifically includes: S251: For each pixel of the equidistant cylindrical image, reverse map it back to the original image plane and use an interpolation algorithm to obtain the color value; S252: Pixel fusion is performed when a single pixel is covered by multiple original images; S253: Calculate the lighting differences between images and perform exposure compensation and color correction on the images; S254: Use multi-band fusion or feathering techniques to fuse image seams.
5. The method according to claim 1, characterized in that, The multimodal large model is trained according to the following steps: Construct training data that includes prompt words, a sequence of continuous panoramic images, and manually annotated standard answers; A large multimodal model is trained using training data by employing instruction fine-tuning. After each training epoch, a validation set is used for validation, and the model with the best performance is selected as the final model.
6. The method according to claim 5, characterized in that, After the multimodal large model is trained, the model is iteratively optimized by collecting decision biases from human driver intervention events or simulation replays.
7. The method according to claim 1, characterized in that, The output of the multimodal large model includes auxiliary judgment information, key information, and driving operation suggestions, wherein the key information includes important targets and their related attribute information.
8. The method according to claim 7, characterized in that, The auxiliary judgment information includes at least one of the following: weather, time, visibility, image quality, current lane of the vehicle, current road conditions, current lane information, and current navigation instructions.
9. The method according to claim 8, characterized in that, The key targets include at least one of the following: motor vehicles, non-motor vehicles, pedestrians, animals, obstacles, and road abnormalities.
10. A driving decision analysis system based on a multimodal large model, characterized in that, Includes the following modules: Multi-view image acquisition module: Acquires continuous image data during vehicle movement using multi-view cameras configured on the moving vehicle; Panoramic image generation module: Using camera parameters, the acquired continuous multi-view image data is stitched and fused into a panoramic image, thereby obtaining a continuous panoramic image sequence; The temporal sampling module samples the continuous panoramic image sequence obtained by the panoramic image generation module at a fixed frequency to generate a sampled panoramic image sequence. Multimodal large model input module: Inputs the sampled panoramic image sequence and text prompts into the multimodal large model; Multimodal large model processing module: The multimodal large model performs scene analysis and driving decision judgment based on the sampled panoramic image sequence, and outputs the required information according to the input text prompt words.
Citation Information
Patent Citations
Automatic driving model based on multi-modal large model, training method and automatic driving method
CN119514635A