Zero-sample semantic segmentation map construction method combining road visual language model and dynamic object
By using vector quantization and multimodal data fusion methods, image features in severe weather conditions are restored and semantic descriptions are generated, solving the perception and map construction problems of autonomous driving systems in severe weather conditions and achieving high-precision dynamic object detection and map updates.
Patent Information
- Application Number
- CN202510812721.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Under severe weather conditions, the sensor performance of the autonomous driving system is limited, resulting in a decrease in perception capabilities, affecting the accuracy of map construction and the precision of dynamic object detection. Existing technologies are difficult to adapt to various severe weather changes.
A vector quantized image restoration model is used to recover image features in severe weather conditions, and a road visual language model is combined for semantic description. Multimodal data fusion and BEV encoder processing are then used to generate a high-precision dynamic object segmentation map.
It effectively improves the perception accuracy and map building capabilities under severe weather conditions, updates high-precision vector maps in real time, and improves the perception capability and decision-making efficiency of the autonomous driving system.
Smart Images

Figure CN120747503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a zero-shot semantic segmentation map construction method, and in particular to a zero-shot semantic segmentation map construction method combining a road visual language model with dynamic objects, belonging to the technical field of zero-shot semantic segmentation map construction methods. Background Art
[0002] With the continuous development of autonomous driving technology, building high-precision, dynamically updated maps has become key to ensuring the stable operation of autonomous driving systems. This is especially true in severe weather conditions (such as rain, snow, and haze). The performance of autonomous driving vehicle sensors (such as cameras and lidar) is greatly limited, resulting in a decline in perception capabilities and affecting the accuracy of map construction. Therefore, how to improve the perception accuracy and map construction capabilities of autonomous driving systems in complex weather conditions is one of the challenges currently facing autonomous driving technology.
[0003] Traditional image restoration methods are often optimized for specific severe weather conditions and are difficult to adapt to a variety of weather changes. In autonomous driving systems, map updates and generation not only need to restore image quality in severe weather conditions, but also need to promptly identify and process changes in dynamic objects. To address these issues, zero-shot semantic segmentation and map construction technology for dynamic objects combined with visual language models has emerged. Through multimodal data fusion, image restoration, and semantic segmentation technology, the perception capabilities of autonomous driving systems in dynamic environments can be effectively improved, and accurate vector maps can be constructed.
[0004] This technology proposes a map construction method that combines road visual language model and zero-shot semantic segmentation of dynamic objects. Summary of the Invention
[0005] The main purpose of this invention is to provide a zero-shot semantic segmentation map construction method that combines a road visual language model with dynamic objects, aiming to solve the problems of image restoration, dynamic object detection, and semantic information generation in severe weather conditions, thereby enhancing the perception accuracy and decision-making capabilities of the autonomous driving system.
[0006] The purpose of the present invention can be achieved by adopting the following technical solutions:
[0007] A zero-shot semantic segmentation map construction method combining a road visual language model and dynamic objects includes the following steps:
[0008] Step 1: Input the degraded surround view image into the image restoration model based on vector quantization;
[0009] Step 2: Input the restored image into the road visual language model;
[0010] Step 3: Based on the above-mentioned visual feature map with semantic description, combined with LiDAR and millimeter-wave radar data, multimodal information is integrated through a multimodal fusion strategy;
[0011] Step 4: Use the BEV encoder to solve the feature map alignment problem and extract temporal features;
[0012] Step 5: Fuse the associated features with the BEV features of the current frame and process them through the convolution layer to obtain the moving target segmentation map.
[0013] Preferably, in step 1, by means of vector quantization and optimization means, the image features under severe weather conditions are effectively restored to image features under normal weather conditions;
[0014] The road visual language model in step 2 semantically describes the information about vehicles, sidewalks, and road boundaries on the road.
[0015] Preferably, in step 1, the image feature extraction aims to extract important information that can reflect the core content of the image from the input image;
[0016] Assuming that the input image is set to I, CNN will gradually map the original image into a high-dimensional feature space through several layers of transformation operations. Specifically, it uses convolution operations and pooling operations to extract image features.
[0017] Assume that each image I generates a feature vector after passing through the corresponding processing flow of the network Here, d represents the dimension of the feature space;
[0018] In terms of constructing the normal image codebook, the process is to use a large number of images under normal weather conditions to carry out training to understand the feature distribution of images under normal weather conditions. Assume that there is a training set D normal ={I1,I2,...,I N}, this training set consists of N images under normal weather conditions;
[0019] After each image Ii undergoes CNN feature extraction, a feature vector
[0020] The feature vectors of all normal images Construct a feature matrix Each row of this matrix represents the characteristics of a normal weather image;
[0021] After clustering these feature vectors, a codebook consisting of K cluster centers is obtained, and each cluster center CK They all represent a certain mode of image features under normal weather conditions;
[0022] The codebook is defined as a set Each of these
[0023] Map the input vector into a finite set of vectors, and this finite set of vectors is called the codebook;
[0024] The input severe weather image is defined as I bad , after CNN feature extraction, its feature vector is obtained
[0025] To restore the features of bad weather images, they need to be mapped to the feature space of normal weather images. The specific implementation process is achieved with the help of vector quantization.
[0026] The goal of vector quantization is to find the feature vector F that corresponds to the bad weather image from the codebook. bad The closest cluster center C k , this process can be expressed as the following mathematical formula:
[0027]
[0028] Among these, represents the feature vector obtained after vector quantization, and |F bad -C k |refers to F bad and a cluster center C in the codebook k The Euclidean distance between them. By selecting the cluster center C with the smallest distance k , and achieved the mapping of severe weather image features.
[0029] Preferably, the design of the loss function needs to consider how to reduce the difference between the characteristics of bad weather images and normal weather images;
[0030] Use the Euclidean distance loss function as the distance loss function and define it as:
[0031]
[0032] in, It refers to the feature vector corresponding to the i-th severe weather image, and is the restored feature obtained by means of vector quantization;
[0033] The purpose of this loss function is to minimize the difference between the features of the bad weather image and the restored features;
[0034] Add a regularization term to ensure that the restored features can maintain stability and consistency. Using L2 regularization, the optimized objective function can be expressed as:
[0035]
[0036] L exists as a regularization coefficient, and θ belongs to the category of model parameters. is the L2 norm of the model parameters;
[0037] The image features presented under extreme weather conditions, with the help of image restoration models, will gradually move towards the image feature distribution F close to that under normal weather conditions after a series of vector quantization and optimization processes. normal The features obtained after the recovery operation can be expressed as follows:
[0038]
[0039] With the help of vector quantization and optimization methods, the image features under severe weather conditions can be effectively restored to image features close to those under normal weather conditions, thereby improving the perception performance of autonomous driving under severe weather conditions.
[0040] Preferably, the step 2 includes inputting the restored image into a road visual language model, and the model begins to conduct an in-depth analysis of the road scene in the image and generate a semantic description;
[0041] This process first involves identifying and locating various elements in the image, including vehicles, pedestrians, sidewalks, lane markings, traffic signs, and road boundaries;
[0042] The model extracts image features through convolutional neural network visual processing algorithms and identifies the specific location and category of each object;
[0043] Next, the visual language model generates corresponding semantic descriptions based on these visual features;
[0044] The key to the visual language model is its deep learning architecture, which can understand the relationship between each element in the image and its surrounding environment;
[0045] By training on large-scale datasets, the model continuously learns how to extract effective information from images and convert this information into concise and accurate language descriptions.
[0046] Preferably, the step three includes: by introducing semantic tags, the representation capability of the visual feature map is enhanced;
[0047] Data fusion is achieved through various strategies:
[0048] Feature-level fusion refers to combining features from different sensors. Specifically, visual feature maps and radar point cloud data are often jointly processed in a deep neural network.
[0049] The convolutional neural network model automatically selects the influence of different sensor data through an attention mechanism or weighted averaging strategy, thereby deciding how to process various types of information based on the specific characteristics of the scene;
[0050] Decision-level fusion combines the predictions output by different sensor data after they have been independently processed. In this process, the vision, radar, and millimeter-wave radar perception modules perform target detection, positioning, and tracking tasks respectively, and finally integrate this information into the final decision output.
[0051] The BEV encoder maps the three-dimensional point cloud data, visual image, and distance information of the laser radar into the same coordinate system;
[0052] The BEV encoder plays a key role in the alignment of multimodal data.
[0053] Beneficial technical effects of the present invention:
[0054] The present invention provides a method for constructing a zero-sample semantic segmentation map of dynamic objects by combining a road visual language model with a road visual language model. The method uses a vector quantized image restoration model to perform feature recovery on degraded surround view images under severe weather conditions, thereby restoring the image features under normal weather conditions. Then, the road visual language model is used to analyze the restored image, generate a semantic description, and extract information about key elements on the road (such as vehicles, pedestrians, lane lines, and traffic signs). Subsequently, by combining lidar and millimeter-wave radar data, a multimodal fusion strategy is adopted to integrate visual and radar information, and the feature maps are aligned through a BEV encoder to extract temporal features. Finally, the fused features are combined with the BEV features of the current frame and processed through a convolutional layer to generate a dynamic object segmentation map. This method can effectively deal with image degradation problems under different severe weather conditions, and improve the detection accuracy of dynamic objects through semantic segmentation technology. By combining the visual language model with multimodal data fusion, the high-precision vector map required for the autonomous driving system is updated and accurately constructed in real time, significantly improving the system's perception capability and decision-making efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 The present invention is a flowchart of a preferred embodiment of a method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects. DETAILED DESCRIPTION
[0056] In order to make the technical solution of the present invention more clear and specific to those skilled in the art, the present invention is further described in detail below with reference to embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0057] The method comprises:
[0058] Step S1: Input the degraded surround view image into a vector quantization-based image restoration model. With the help of vector quantization and optimization methods, the image features under severe weather conditions are effectively restored to image features close to those under normal weather conditions.
[0059] Step S2: Subsequently, the restored image is input into a road visual language model, which semantically describes information about vehicles, sidewalks, and road boundaries on the road.
[0060] Step S3: Based on the above-mentioned visual feature extraction map with semantic description, combined with lidar and millimeter-wave radar data, multimodal information is integrated through a multimodal fusion strategy. Then, the BEV encoder is used to solve the feature map alignment problem and extract temporal features. Finally, the associated features are fused with the BEV features of the current frame, and the moving target segmentation map is obtained through convolutional layer processing.
[0061] Step S1 includes,
[0062] Severe weather conditions, such as rain, snow, and fog, not only reduce the perception capabilities of autonomous vehicle sensors (such as cameras and lidar), but also severely impact the accuracy of map construction. Conventional image restoration techniques are often optimized for specific weather conditions and struggle to cope with diverse types of inclement weather. To improve map accuracy, we propose a universal image restoration model that uses vector quantization to effectively restore image features in various adverse weather conditions, enabling map construction to maintain high accuracy in complex environments.
[0063] Image feature extraction aims to extract important information from an input image that reflects its core content. Traditional image processing methods often rely on manually designed algorithms, such as SIFT and HOG. However, with the continuous development of deep learning, convolutional neural networks (CNNs) have gradually become the mainstream tool for image feature extraction. CNNs utilize multi-layer convolution operations, pooling, and fully connected layers to autonomously learn and extract high-level features from images.
[0064] Assuming that the input image is set to I, CNN will gradually map the original image into a high-dimensional feature space through several layers of transformation operations. Specifically, it uses convolution operations and pooling operations to extract image features. Assuming that each image I will generate a feature vector after passing through the corresponding processing flow of the network The d here represents the dimension of the feature space.
[0065] In terms of constructing the normal image codebook, the process is to use a large number of images under normal weather conditions to carry out training and understand the feature distribution of images under normal weather conditions. Assume that there is a training set D normal ={I1,I2,...,I N}, this training set is composed of N images under normal weather conditions. After each image Ii undergoes CNN feature extraction, a feature vector
[0066] In order to construct the codebook of normal images, we first need to convert the feature vectors of all normal images into Construct a feature matrix Each row of this matrix represents the characteristics of a normal weather image.
[0067] After clustering these feature vectors, a codebook consisting of K cluster centers can be obtained. Each cluster center C K They all represent a certain mode of image features under normal weather conditions. The codebook can be defined as a set Each of these
[0068] Vector quantization, or VQ for short, is a process that maps an input vector into a finite set of vectors, known as a codebook. In this paper, we apply VQ to the task of recovering features from severe weather images.
[0069] We define the input severe weather image as I bad , after CNN feature extraction, its feature vector is obtained To restore the features of severe weather images, they need to be mapped into the feature space of normal weather images. The specific implementation process is achieved with the help of vector quantization.
[0070] The goal of vector quantization is to find the feature vector F that corresponds to the bad weather image from the codebook. bad The closest cluster center C k, this process can be expressed as the following mathematical formula:
[0071] Among these, represents the feature vector obtained after vector quantization, and |F bad -C k |refers to F bad and a cluster center C in the codebook k The Euclidean distance between them. By selecting the cluster center C with the smallest distance k , and achieved the mapping of severe weather image features.
[0072] Vector quantization optimization process: To make the characteristics of severe weather images as close as possible to the characteristic distribution of normal weather images, it is necessary to adjust the model parameters by optimizing the loss function in the vector quantization process. Specifically, the design of the loss function needs to consider how to reduce the differences between the characteristics of severe weather images and normal weather images.
[0073] We use the Euclidean distance loss function as the distance loss function and define it as:
[0074] in, Refers to the feature vector corresponding to the i-th severe weather image, and is the restored feature obtained by vector quantization. The purpose of this loss function is to minimize the difference between the features of the severe weather image and the restored features.
[0075] To make the recovery process more effective, in addition to the Euclidean distance loss, a regularization term can be added to ensure that the recovered features maintain stability and consistency. We use L2 regularization, and the optimization objective function can be expressed as:
[0076]
[0077] Here, L exists as a regularization coefficient, and θ belongs to the category of model parameters. is the L2 norm of the model parameters.
[0078] The image features presented under extreme weather conditions, with the help of image restoration models, will gradually move towards the image feature distribution F close to that under normal weather conditions after a series of vector quantization and optimization processes. normal The features obtained after the recovery operation can be expressed as follows:
[0079]
[0080] With the help of vector quantization and optimization methods, we can effectively restore image features in severe weather conditions to image features close to those in normal weather conditions, thereby improving the perception performance of autonomous driving in adverse weather conditions.
[0081] The step S2 includes:
[0082] After the restored image is input into the road visual language model, the model begins to conduct an in-depth analysis of the road scene in the image and generate a semantic description. This process first involves the recognition and positioning of various elements in the image, including vehicles, pedestrians, sidewalks, lane lines, traffic signs, and road boundaries. The model extracts image features through the convolutional neural network (CNN) visual processing algorithm to identify the specific location and category of each object. For example, the model can detect that there is a red car parked in the lane on the road ahead, or there is a clear sidewalk on one side of the road. At the same time, the model can also recognize the presence of traffic signs and traffic lights, and mark their locations in the image.
[0083] Next, the visual language model generates corresponding semantic descriptions based on these visual features. For example, after identifying a vehicle, the model might generate a statement like "50 meters ahead, a red car is parked on the side of the road, and the lane is 3 meters wide." For road boundary recognition, the model might generate a description like "There is a clear lane line on the right side of the road." The model can not only describe static objects but also understand the spatial relationships between them, thereby generating more detailed and accurate semantic information. For example, when the model identifies a car, it may automatically infer its positional relationship, such as "The car is parked on the right side of the road, about 50 meters from the intersection," and reflect the relative positions of the objects in the description.
[0084] The key to the visual language model lies in its deep learning architecture, which understands the relationship between every element in an image and its surroundings. By training on large-scale datasets, the model continuously learns how to extract meaningful information from images and translate this information into concise and accurate language descriptions. In this way, the model can translate complex visual information into natural language, providing the autonomous driving system with clear environmental awareness. These semantic descriptions not only help the system better understand the road scene but also provide fundamental support for path planning, obstacle avoidance, and driving decisions, thereby improving the safety and reliability of autonomous driving.
[0085] Finally, after the restored image is input into the visual language model, the system can generate a detailed semantic description based on various information in the image, providing accurate environmental perception for the autonomous driving system.
[0086] The step S3 comprises:
[0087] The introduction of semantic labels enhances the representational power of visual feature maps. For example, a red car in an image will be labeled as a "vehicle" after semantic segmentation. The model can use this semantic information to determine that the object within the area is a vehicle and infer its motion state and size characteristics. For static objects, the model further identifies their spatial structure, such as lane lines or traffic signs.
[0088] Multimodal data fusion is one of the core technologies for improving autonomous driving perception accuracy. In practical applications, vision, lidar, and millimeter-wave radar each perceive different physical features. Effectively combining this information to generate unified, accurate perception results is a technical challenge. Specifically, data fusion can be achieved through a variety of strategies:
[0089] Feature-level fusion combines features from different sensor data. Specifically, visual feature maps and radar point cloud data are often jointly processed within a deep neural network. Visual data provides the model with information about an object's texture, color, and shape, while lidar and millimeter-wave radar provide spatial features such as distance and velocity. Feature-level fusion effectively integrates multi-source data and improves the accuracy of the perception system.
[0090] To achieve feature-level fusion, convolutional neural networks (CNNs) are often used. These networks can learn the relationships between data from different modalities and integrate them within the same feature space. During this process, the model automatically selects the influence of different sensor data through an attention mechanism or weighted averaging strategy, thus determining how to process each type of information based on the specific characteristics of the scene.
[0091] Decision-level fusion combines the predictions output by different sensor data after they have been independently processed. During this process, the vision, radar, and millimeter-wave radar perception modules each perform target detection, localization, and tracking tasks, ultimately integrating this information into the final decision output. For example, the vision module detects a vehicle in an image, while the radar module provides its speed and relative position. The model combines this information to generate a comprehensive judgment. This strategy is often used to address situations where sensor data is incomplete or of poor quality.
[0092] Based on multimodal data fusion, BEV encoders are widely used in autonomous driving systems to solve the spatial alignment problem of different modal data. The BEV view captures a scene from a bird's-eye view, clearly displaying the spatial layout of the road, lane markings, traffic signs, and the surrounding environment. It converts spatial information from different sensors into a unified planar coordinate system, helping to solve the problem of aligning data from different perspectives and resolutions.
[0093] The BEV encoder's primary task is to map the 3D point cloud data from the lidar, visual images, and distance information from the millimeter-wave radar into a common coordinate system. This process first projects the sensor data into the BEV space through geometric transformations. It then extracts scene features using convolutional layers or other deep learning methods. These features represent scene objects and environmental elements, such as road boundaries, lane markings, and the locations of obstacles.
[0094] The BEV encoder plays a key role in aligning multimodal data. It unifies the spatial relationships between different modal data into a shared coordinate system, enabling the model to process this data from the same viewpoint, thereby reducing errors caused by differences in sensor perspective and resolution. This process not only improves perception accuracy but also enhances model robustness.
[0095] In autonomous driving, extracting temporal features is crucial for detecting and tracking dynamic objects. As the position and state of objects in a scene change over time, temporal features must be extracted from multiple time frames to better understand the objects' motion trajectories and behavior patterns. The BEV encoder processes historical frame data using time series information to extract time-related features, helping the model predict the future position and motion state of dynamic objects.
[0096] Temporal feature extraction is typically achieved through temporal convolutional networks (TCNs). These networks can capture temporal dependencies in time series data, thereby inferring the motion trends of objects. In autonomous driving applications, temporal feature extraction not only helps detect static objects but also identifies the trajectories of dynamic objects and, in turn, predicts their future motion. By effectively processing temporal features, autonomous driving systems can react proactively, improving both responsiveness and safety.
[0097] After completing the fusion of multimodal data and extracting temporal features, the next task is to generate a moving object segmentation map. This map distinguishes moving objects in a scene (such as moving vehicles and pedestrians) from the static background and accurately identifies their outlines and positions. To achieve this, the system uses a U-Net network architecture to perform image segmentation.
[0098] When generating a moving object segmentation map, the system first extracts features from the feature map obtained by the BEV encoder. It then processes the feature map through a convolutional layer to generate a binary segmentation map, where the moving object area is labeled "foreground" and other areas are labeled "background." This allows the model to accurately segment dynamic objects in the scene and generate a final semantic segmentation map containing these objects.
[0099] 1. Application Scenario Overview
[0100] In urban autonomous driving scenarios, autonomous vehicles need to perceive surrounding objects and build high-precision maps in real time within complex and ever-changing road environments. In particular, in adverse weather conditions such as rain, snow, and fog, sensor performance is limited, image quality degrades, and perception accuracy is reduced, impacting map construction accuracy. This embodiment applies a zero-shot semantic segmentation map construction method that combines a road visual language model with dynamic objects to urban autonomous driving scenarios. This approach addresses image degradation in adverse weather conditions, improves dynamic object detection accuracy, and enables real-time updates and precise construction of high-precision vector maps.
[0101] Implementation steps
[0102] Construction and training of image restoration models;
[0103] Collect a large number of urban road images under normal weather conditions to construct a training set. The images should cover urban road scenes in different time periods and weather conditions (sunny days), including various vehicles, pedestrians, lane markings, and traffic signs.
[0104] A convolutional neural network (CNN) is used to extract features from the training set images. Each image is fed into the CNN, where it undergoes convolution and pooling operations to generate a high-dimensional feature vector. Assume the input image is I, and after processing it through the CNN, the feature vector V is obtained.
[0105] Perform clustering on the feature vectors of all normal weather images to construct a codebook for normal images. Determine the number of cluster centers, K, and use a clustering algorithm (such as K-Means) to obtain K cluster centers. Each cluster center represents a mode of normal weather image features, forming the codebook C.
[0106] A vector quantization (VQ) model is constructed. For severe weather images, a feature vector V' is obtained through CNN feature extraction. V' is mapped to the codebook C through VQ, and the cluster center C_i closest to V' is found to achieve the mapping of severe weather image features to the normal weather image feature space.
[0107] Design a loss function, including a Euclidean distance loss function and an L2 regularization term. The Euclidean distance loss function measures the difference between the features of the severe weather image and the restored features, while the L2 regularization term ensures the stability and consistency of the restored features. By optimizing the loss function and adjusting the model parameters, the severe weather image features are closely aligned with the normal weather image feature distribution. Train the image restoration model until it achieves satisfactory image restoration results on the validation set.
[0108] Training and application of road visual language model
[0109] Collect image datasets containing urban road scenes, which include vehicles, pedestrians, sidewalks, lane lines, traffic signs, and road boundary elements, and annotate these elements.
[0110] The system uses a CNN visual processing algorithm to extract features from images and identify the specific location and category of each object. For example, it can detect a red car parked in a lane on the road ahead or a clear sidewalk on the side of the road, while also identifying the locations of traffic signs and traffic lights.
[0111] Build a road visual language model based on a deep learning architecture, such as the Transformer or its variants. Extracted visual features are fed into the model, which then learns how to convert visual information into natural language descriptions. The model is trained on a large-scale dataset, and by adjusting model parameters, it generates accurate and concise semantic descriptions.
[0112] The restored image is fed into a trained road visual language model. The model conducts in-depth analysis of the road scene in the image and generates semantic descriptions. For example, after identifying a vehicle, it generates a statement such as "50 meters ahead, a red car is parked on the roadside, and the lane is 3 meters wide." After identifying the road boundary, it generates a description such as "There is a clear lane marking on the right side of the road." The model not only describes static objects but also understands the spatial relationships between them, such as "The car is parked on the right side of the road, about 50 meters from the intersection," providing clear environmental awareness for the autonomous driving system.
[0113] (3) Application of multimodal data fusion and BEV encoder
[0114] Collect LiDAR and millimeter-wave radar data. LiDAR provides 3D point cloud information about the environment, including object distance, height, and density; millimeter-wave radar provides object distance, velocity, and angle information. This data is synchronized with the restored image data to ensure that multimodal data from the same scene has the same timestamp.
[0115] Perform feature-level fusion. The visual feature map and radar point cloud data are fed into a convolutional neural network (CNN). The network learns the relationship between the different modal data and maps them into the same feature space. Through an attention mechanism or weighted averaging strategy, the network automatically selects the influence of different sensor data and processes various types of information based on the characteristics of the scene. For example, when dealing with congested scenes involving a large number of vehicles, the network may assign a higher weight to radar data to utilize the vehicle speed and distance information it provides; while when dealing with intersections with dense traffic signs, it may place greater emphasis on sign recognition results in the visual data.
[0116] Build a BEV encoder. The 3D point cloud data from the LiDAR, visual images, and distance information from the millimeter-wave radar are projected into the BEV space through geometric transformation. Within the BEV space, convolutional layers or other deep learning methods are used to extract scene features, representing objects and environmental elements in the scene, such as road boundaries, lane markings, and obstacle locations. The BEV encoder solves the spatial alignment problem of different modal data, unifying the spatial relationships of multimodal data into a shared coordinate system, thereby reducing errors caused by differences in sensor perspective and resolution.
[0117] Extracting temporal features. A temporal convolutional network (TCN) processes historical frame data, capturing temporal dependencies in time series data and inferring the motion trends of objects. TCN extracts time-related features from multiple time frames, helping the model predict the future position and motion state of dynamic objects. For example, for a moving vehicle, TCN predicts its next position based on its position changes over the past few frames, enabling the autonomous driving system to react in advance, improving response speed and safety.
[0118] Generation of moving target segmentation map;
[0119] The feature map obtained by the BEV encoder is input into the U-Net network architecture. The U-Net network processes the feature map through convolutional layers, pooling layers, and upsampling layers.
[0120] In a U-Net network, convolution operations extract deeper features from feature maps. Pooling reduces the size and computational complexity of feature maps while retaining important features. Upsampling restores feature maps to a size close to the original image, facilitating pixel-level classification.
[0121] The final result is a binary segmentation map, where the moving target area is marked as "foreground" and the rest of the area is marked as "background." For example, moving vehicles and pedestrians are accurately segmented, while static lane lines and traffic signs are marked as background.
[0122] The generated moving target segmentation map is combined with the semantic description to construct a semantic segmentation map containing dynamic objects. This map provides basic support for the autonomous driving system's path planning, obstacle avoidance, and driving decision-making, improving the system's safety and reliability.
[0123] Implementation effect evaluation
[0124] Field tests were conducted in urban autonomous driving scenarios under diverse weather conditions, including sunny, rainy, snowy, and foggy. Representative road sections, including straight roads, curves, intersections, school zones, and commercial areas, were selected to evaluate the performance of the map-building method in various scenarios.
[0125] Evaluation metrics include the accuracy of the semantic segmentation map (such as the segmentation accuracy of lane lines and traffic signs, and the IOU, i.e., the overlap rate between the predicted value and the true value).
[0126] Comparative experiments with traditional map construction methods verify the method's ability to recover from image degradation in adverse weather conditions and improve the accuracy of dynamic object detection. By comparing test results, we analyze the advantages and disadvantages of the method and provide a basis for further optimization.
[0127] The calculation formula of segmentation accuracy is:
[0128]
[0129] in:
[0130] TP: True Positive, that is, pixels correctly predicted to be in the target area.
[0131] TN: True Negative, that is, pixels correctly predicted as background.
[0132] FP: False Positive, that is, background pixels that are incorrectly predicted to be in the target area.
[0133] FN: False Negative, that is, pixels in the target area that are incorrectly predicted to be background.
[0134] The IOU calculation formula is:
[0135]
[0136] Where A: prediction area.
[0137] B: The actual marked area.
[0138] |A∩B|: the number of pixels in the intersection area of the predicted area and the true annotated area, |A∪B|: the number of pixels in the union area of the predicted area and the true annotated area.
[0139] Based on the above indicators, the baseline method and the patented method were experimentally compared, and the final experimental data were as follows:
[0140] method Segmentation accuracy IOU MoSeg 49.43 26.0 SimpleBEV_Motion 73.79 60.24 This patent model 75.35 62.59
[0141] The above is only a further embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can replace or change the technical solution and concept of the present invention within the scope disclosed by the present invention, which falls within the scope of protection of the present invention.
Claims
1. A zero-shot semantic segmentation map construction method combining a road visual language model and dynamic objects, characterized by: The steps include: Step 1: Input the degraded surround view image into the image restoration model based on vector quantization; Step 2: Input the restored image into the road visual language model; Step 3: Based on the above-mentioned visual feature map with semantic description, combined with LiDAR and millimeter-wave radar data, multimodal information is integrated through a multimodal fusion strategy; Step 4: Use the BEV encoder to solve the feature map alignment problem and extract temporal features; Step 5: Fuse the associated features with the BEV features of the current frame and process them through the convolution layer to obtain the moving target segmentation map.
2. The method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects according to claim 1, characterized in that: In step 1, by using vector quantization and optimization methods, the image features under severe weather conditions are effectively restored to image features close to normal weather conditions; The road visual language model in step 2 semantically describes the information about vehicles, sidewalks, and road boundaries on the road.
3. The method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects according to claim 1, characterized in that: In step 1, the goal of image feature extraction is to extract important information that can reflect the core content of the image from the input image. Assuming that the input image is set to I, CNN will gradually map the original image into a high-dimensional feature space through several layers of transformation operations. Specifically, it uses convolution operations and pooling operations to extract image features. Assume that each image I generates a feature vector after passing through the corresponding processing flow of the network Here, d represents the dimension of the feature space; In terms of constructing the normal image codebook, the process is to use a large number of images under normal weather conditions to carry out training to understand the feature distribution of images under normal weather conditions. Assume that there is a training set D normal ={I1,I2,...,I N }, this training set consists of N images under normal weather conditions; After each image Ii undergoes CNN feature extraction, a feature vector The feature vectors of all normal images Construct a feature matrix Each row of this matrix represents the characteristics of a normal weather image; After clustering these feature vectors, a codebook consisting of K cluster centers is obtained, and each cluster center C K They all represent a certain mode of image features under normal weather conditions; The codebook is defined as a set Each of these Map the input vector into a finite set of vectors, and this finite set of vectors is called the codebook; The input severe weather image is defined as I bad , after CNN feature extraction, its feature vector is obtained To restore the features of bad weather images, they need to be mapped to the feature space of normal weather images. The specific implementation process is achieved with the help of vector quantization. The goal of vector quantization is to find the feature vector F that corresponds to the bad weather image from the codebook. bad The closest cluster center C k , this process can be expressed as the following mathematical formula: Among these, represents the feature vector obtained after vector quantization, and |F bad -C k |refers to F bad and a cluster center C in the codebook k The Euclidean distance between them is used to select the cluster center C with the smallest distance. k , and achieved the mapping of severe weather image features.
4. The method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects according to claim 3, characterized in that: The design of the loss function needs to consider how to reduce the difference between the characteristics of bad weather images and normal weather images; Use the Euclidean distance loss function as the distance loss function and define it as: in, It refers to the feature vector corresponding to the i-th severe weather image, and is the restored feature obtained by means of vector quantization; The purpose of this loss function is to minimize the difference between the features of the bad weather image and the restored features; Add a regularization term to ensure that the restored features can maintain stability and consistency. Using L2 regularization, the optimized objective function can be expressed as: L exists as a regularization coefficient, and θ belongs to the category of model parameters. is the L2 norm of the model parameters; The image features presented under extreme weather conditions, with the help of image restoration models, will gradually move towards the image feature distribution F close to that under normal weather conditions after a series of vector quantization and optimization processes. normal The characteristics obtained after the recovery operation can be expressed in the following form: With the help of vector quantization and optimization methods, the image features under severe weather conditions can be effectively restored to image features close to those under normal weather conditions, thereby improving the perception performance of autonomous driving under severe weather conditions.
5. The method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects according to claim 1, characterized in that: The second step includes inputting the restored image into the road visual language model, and the model begins to conduct an in-depth analysis of the road scene in the image and generate a semantic description; This process first involves identifying and locating various elements in the image, including vehicles, pedestrians, sidewalks, lane markings, traffic signs, and road boundaries; The model extracts image features through convolutional neural network visual processing algorithms and identifies the specific location and category of each object; Next, the visual language model generates corresponding semantic descriptions based on these visual features; The key to the visual language model is its deep learning architecture, which can understand the relationship between each element in the image and its surrounding environment; By training on large-scale datasets, the model continuously learns how to extract effective information from images and convert this information into concise and accurate language descriptions.
6. The method for constructing a zero-shot semantic segmentation map combining a road visual language model and dynamic objects according to claim 1, characterized in that: The step three includes: by introducing semantic labels, the representation capability of the visual feature map is enhanced; Data fusion is achieved through various strategies: Feature-level fusion refers to combining data features from different sensors. Specifically, visual feature maps and radar point cloud data are usually jointly processed in a deep neural network. The convolutional neural network model automatically selects the influence of different sensor data through an attention mechanism or weighted averaging strategy, thereby deciding how to process various types of information based on the specific characteristics of the scene; Decision-level fusion combines the predictions output by different sensor data after they have been independently processed. During this process, the vision, radar, and millimeter-wave radar perception modules perform target detection, positioning, and tracking tasks respectively, and finally integrate this information into the final decision output. The BEV encoder maps the three-dimensional point cloud data, visual image, and distance information of the laser radar into the same coordinate system; The BEV encoder plays a key role in the alignment of multimodal data.
Citation Information
Patent Citations
Real scene severe weather image restoration method based on visual language model
CN118537264A
Wireless vehicular systems and methods for detecting roadway conditions
US20210097311A1
Cited By
Self-supervised multi-modal fusion and collaborative optimization method suitable for curve and ramp scenes
CN121861611A