Safe driving environment sensing method and system combined with artificial intelligence
By constructing multi-scene data sets, transfer learning and introduction of attention mechanisms, and combining the generation of adversarial networks to simulate extreme scene data, the problem of insufficient adaptability of existing environmental perception methods in complex and extreme driving scenarios is solved, and the accuracy and safety of environmental perception are improved.
Patent Information
- Application Number
- CN202511053837.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing environmental perception methods are difficult to adapt to in complex and changeable driving scenarios, especially in extreme conditions, which lacks effective training, resulting in misjudgment and misjudgment, affecting driving safety.
A driving scenario data set covering a variety of weather conditions, different road conditions and behavioral patterns of traffic participants is constructed, and a driving scenario adaptive fine-tuning is used to use transfer learning for pre-training and scene adaptive fine-tuning, an attention mechanism module is introduced, and an intensive training is carried out in combination with the generation of adversarial networks to simulate extreme driving scenario data to generate an enhanced perception model.
It significantly improves the adaptability and generalization ability of the model in complex and variable driving scenarios, reduces the probability of misjudgment and misjudgment, and improves the accuracy of environmental perception and driving safety.
Smart Images

Figure CN120564154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a safe driving environment perception method and system combined with artificial intelligence. Background Art
[0002] Amid the rapid development of intelligent vehicles and autonomous driving technology, safe driving environment perception, as a core component for ensuring driving safety and enabling autonomous driving, has garnered widespread attention. Currently, traditional environmental perception methods primarily rely on single sensor data or simple rule-based models to identify environmental elements. For example, perception methods based on camera image processing can only capture two-dimensional visual information and are sensitive to light variations. In inclement weather (such as heavy rain or dense fog) or low-light conditions, image quality degrades, significantly reducing recognition accuracy. LiDAR-based perception methods, while capable of acquiring highly accurate three-dimensional spatial information, perform poorly for detecting transparent objects (such as glass) and objects with low reflectivity (such as black vehicles).
[0003] Furthermore, existing environmental perception models often employ fixed architectures and lack the ability to adapt to diverse driving scenarios. Faced with complex and changing real-world road conditions, such as congested urban areas and rugged rural roads, as well as the diverse behaviors of various traffic participants (pedestrians, non-motorized vehicles, and other vehicles), these models struggle to accurately perceive key environmental elements, leading to misjudgments and omissions, which compromise driving safety. Furthermore, data for extreme driving scenarios (such as sudden accidents and emergency avoidance in inclement weather) is scarce, and existing models lack effective training for these scenarios. This results in inability to make accurate decisions in these extreme situations, seriously threatening driving safety. Summary of the Invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a safe driving environment perception method combined with artificial intelligence, the method comprising: Constructing a driving scenario dataset, wherein the driving scenario dataset covers driving scenario data under various weather conditions, different road conditions, and various traffic participant behavior patterns; Based on the driving scenario dataset, pre-training the initial environment perception model and fine-tuning the scene adaptability are performed using transfer learning to obtain a basic environment perception model; Introducing an attention mechanism module into the basic environmental perception model to enable the basic environmental perception model to focus on environmental elements critical to driving safety, thereby generating an enhanced perception model that enhances the perception capability of key elements; Generate extreme driving scenario data by simulating using a generative adversarial network, input the extreme driving scenario data into the enhanced perception model for enhanced training, and obtain a target environment perception model; The target environment perception model is used to perceive the driving scene data acquired by different sensors and output a safe driving environment perception result.
[0005] On the other hand, an embodiment of the present invention also provides a safe driving environment perception system combined with artificial intelligence, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0006] Based on the above aspects, the embodiment of the present invention constructs a driving scenario dataset covering various weather conditions, different road conditions and various traffic participant behavior patterns, enabling the model to learn environmental characteristics in different scenarios, thereby significantly improving the model's adaptability and generalization ability to complex and changeable driving scenarios. The initial environmental perception model is pre-trained and fine-tuned for scenario adaptability using transfer learning, which can fully utilize the knowledge of existing models and accelerate the training process of new models. At the same time, the model can better adapt to specific driving scenarios, improve training efficiency and model performance, introduce an attention mechanism module, enable the model to automatically focus on key environmental elements for driving safety, effectively filter out irrelevant information, enhance the model's perception of key elements, reduce the probability of misjudgment and missed judgment, and improve the accuracy of environmental perception. Generative adversarial networks are used to simulate and generate extreme driving scenario data, and the data is input into the enhanced perception model for enhanced training, which makes up for the scarcity of extreme scenario data in actual data, enabling the model to make correct decisions when facing extreme situations, and further improve driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 1 is a schematic diagram of the execution flow of the safe driving environment perception method combined with artificial intelligence provided by an embodiment of the present invention.
[0008] Figure 2 Schematic diagram of exemplary hardware and software components of a safe driving environment perception system combined with artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0009] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a safe driving environment perception method combined with artificial intelligence provided by an embodiment of the present invention. The safe driving environment perception method combined with artificial intelligence is introduced in detail below.
[0010] Step S110: Constructing a driving scenario dataset, which covers driving scenario data of various weather conditions, different road conditions, and various traffic participant behavior patterns.
[0011] In this embodiment, in order to construct a data set that can fully reflect the actual driving situation, a variety of sensor devices are required to obtain information. Visual sensors, such as high-definition cameras, are placed in multiple key locations of the test vehicle, such as the front, rear, and sides of the vehicle. The front camera is mainly used to capture the scene in the direction of the vehicle's travel, including road conditions, vehicles in front, and pedestrians; the rear camera is used to monitor the dynamics of the rear vehicle; and the cameras on both sides can assist in detecting the situation around the vehicle, such as vehicles and pedestrians in adjacent lanes. These cameras have different parameter settings, such as different focal lengths and resolutions, to meet different acquisition requirements. Radar sensors, such as millimeter-wave radars and lidars, are also reasonably installed on the vehicle. Millimeter-wave radars are usually installed at the front and rear of the vehicle and can measure the distance, speed, and angle information of the target object in real time. Lidars are generally installed on the top of the vehicle and can generate high-precision three-dimensional point cloud data.
[0012] Furthermore, data collection operations must be conducted at different times and seasons to cover a variety of weather conditions. On sunny days, ample light provides clear environmental image information, facilitating accurate recognition of road signs, traffic lights, and other vehicle features. Under these conditions, the images captured by visual sensors are vibrant and high-contrast, allowing radar sensors to accurately detect target objects. On cloudy days, the light is relatively dim, and image brightness and contrast vary, placing higher demands on the model's recognition capabilities. Rainy and snowy days pose additional driving challenges. Road surfaces become slippery, affecting vehicle performance and interfering with sensor operation. On rainy days, raindrops form on camera lenses, blurring images. On snowy days, snowflakes scatter the lidar's light, adding noise to the point cloud data. Therefore, data collected under these adverse weather conditions is crucial for improving the model's adaptability in complex environments.
[0013] Furthermore, data is collected on a variety of road types, including urban roads, highways, and rural roads, tailored to different road conditions. Urban roads are characterized by high traffic volume, numerous intersections, and dense pedestrian traffic, requiring vehicles to frequently start, stop, change lanes, and make turns. Highways offer faster speeds and relatively wide spacing between vehicles, but require attention to the detection and early warning of distant targets. Rural roads can present narrow roads, complex road conditions, and a lack of traffic signs. By collecting data under these diverse road conditions, the model can be adapted to various driving scenarios.
[0014] At the same time, we focus on the behavioral patterns of various traffic participants. Vehicle behavior patterns include normal driving, acceleration, deceleration, lane changes, and overtaking. Pedestrian behavior patterns include normal walking, sudden crossing of the road, and stopping by the roadside. Cyclists may arbitrarily change lanes or ride against traffic. For each behavior pattern, sufficient data samples must be collected so that the model can accurately identify and predict it.
[0015] Collected raw data may contain noise, errors, and duplication, necessitating preprocessing. Data cleaning is a crucial preprocessing step. It checks data integrity, accuracy, and consistency to remove noise, errors, and duplication. For example, image data is checked for blur and occlusion; point cloud data is checked for uniform density and the presence of outliers. Data labeling is another key step, labeling the collected driving scene data with labels such as weather conditions, road conditions, and traffic participant behavior patterns. This labeling requires professional labelers to ensure accuracy and consistency. A combination of manual and semi-automatic labeling can be used to improve labeling efficiency. After preprocessing, the data is classified and stored according to predefined rules. Classification can be performed by weather conditions, road condition types, and traffic participant behavior patterns to facilitate subsequent use and management. Ultimately, a driving scene dataset encompassing a variety of weather conditions, road conditions, and traffic participant behavior patterns is generated.
[0016] Step S120: Based on the driving scene dataset, the initial environment perception model is pre-trained and fine-tuned for scene adaptability using transfer learning to obtain a basic environment perception model.
[0017] In this embodiment, transfer learning can utilize existing knowledge and experience to accelerate model training and improve model performance. Before using transfer learning, it is necessary to select a suitable initial environment perception model.
[0018] Step S121: selecting an initial environment perception model that includes general image feature extraction capabilities, wherein the initial environment perception model includes a feature extraction network and a perception task output layer.
[0019] In this embodiment, multiple factors should be considered when selecting the initial environmental perception model. A convolutional neural network model pre-trained on large-scale image datasets is selected. These models have achieved good results in common image classification, object detection, and other tasks, and have strong capabilities for extracting general image features. For example, some classic convolutional neural network models can automatically learn image features such as edges, textures, and shapes when processing image data. These features are universal and applicable to a variety of image-related tasks.
[0020] The selected initial environmental perception model includes a feature extraction network and a perception task output layer. The feature extraction network is the core of the model, responsible for extracting environmental features from the input driving scene data. The feature extraction network typically consists of multiple convolutional layers, pooling layers, and activation functions. Through continuous convolution and pooling operations, it converts the input image data into feature representations with different levels of abstraction. The perception task output layer outputs the corresponding perception results based on the extracted features, such as the identified environmental element category and location. The perception task output layer can be a fully connected layer, a classifier, etc., and is designed according to the specific task requirements.
[0021] Step S122: pre-training the feature extraction network of the initial environment perception model on a general image dataset, so that the feature extraction network has the ability to recognize basic environmental elements.
[0022] In this embodiment, the feature extraction network of the selected initial environmental perception model is pre-trained on a general image dataset. This dataset contains a large number of different types of images, covering a wide range of categories, including natural scenery, animals, people, and vehicles. During the pre-training process, the feature extraction network continuously learns the features of these images, gradually acquiring the ability to recognize basic environmental elements.
[0023] The pre-training process is an iterative optimization process. First, images from a general image dataset are input into the feature extraction network, which extracts features from the images through convolutional layers. The convolution kernels in the convolutional layers slide across the image, extracting local features at different locations. The pooling layer then performs dimensionality reduction on the extracted features, reducing the number of features while retaining important information. Activation functions, such as the ReLU function, perform nonlinear transformations on the features, increasing the model's expressive power.
[0024] At each iteration, the difference between the model's output and the true label is calculated, known as the loss. This value reflects the degree to which the model's prediction deviates from the true situation. Using the backpropagation algorithm, the gradient of each parameter is calculated based on the loss value. The parameters are then updated in the opposite direction of the gradient, and this process continues until the model's performance reaches a stable state.
[0025] After pre-training, the feature extraction network learns common image features that can be used to identify various environmental elements. For example, it can identify basic environmental elements such as vehicles, pedestrians, and road signs in an image.
[0026] Step S123: dividing the driving scene data set into a scene adaptability training subset and a verification subset, wherein the scene adaptability training subset includes driving scene data under different weather conditions, and the verification subset includes driving scene data under different road conditions.
[0027] In this embodiment, to effectively fine-tune the initial environment perception model for scene adaptation, it is necessary to partition the driving scene dataset into a scene adaptation training subset and a validation subset. The purpose of this partitioning is to train and validate the model separately to ensure the model's performance in different scenarios.
[0028] During the segmentation process, data is categorized based on its characteristics and task requirements. The scenario-adaptive training subset primarily contains driving scene data under various weather conditions. This data helps the model learn the characteristic variations of environmental elements in different weather conditions. For example, the appearance of a vehicle, the color, and texture of the road will differ on sunny and rainy days. By learning from this data, the model can better adapt to various weather conditions.
[0029] The validation subset contains driving scenario data under different road conditions. Different road conditions, such as urban roads, highways, and rural roads, have different road structures, traffic rules, and driving characteristics. On urban roads, vehicles frequently start, stop, and change lanes; on highways, vehicles must maintain high speeds and safe distances; and on rural roads, the roads may be narrow and complex. By validating the validation subset, we can evaluate the model's performance under different road conditions, ensuring that the model can accurately perceive and respond to various road conditions.
[0030] The partitioning method can be adjusted based on actual circumstances. A random partitioning method can be used to randomly divide the driving scenario dataset into a scenario-adaptive training subset and a validation subset. Alternatively, a stratified partitioning method can be used based on data distribution to ensure that each subset contains a variety of data types. After the partitioning is completed, the scenario-adaptive training subset and validation subset can be managed and used separately.
[0031] Step S124: Input the scene adaptability training subset into the pre-trained feature extraction network, freeze the bottom parameters of the feature extraction network, and adjust the top parameters of the feature extraction network and the parameters of the perception task output layer to perform the first stage of fine-tuning.
[0032] Step S1241: Determine the dividing boundary between the bottom-layer parameters and the top-layer parameters of the feature extraction network, where the bottom-layer parameters include the parameters of the first half of the convolution layer close to the input layer, and the top-layer parameters include the parameters of the second half of the convolution layer close to the output layer.
[0033] In this embodiment, when performing the first stage of fine-tuning, it is necessary to clarify the dividing boundary between the bottom-layer parameters and the top-layer parameters of the feature extraction network. The bottom-layer parameters are usually the parameters of the first half of the convolution layer close to the input layer. These parameters learn general, bottom-layer image features, such as edges, textures, etc. These features are highly universal in different tasks and scenarios, so they need to be retained during the fine-tuning process. The top-layer parameters are the parameters of the second half of the convolution layer close to the output layer. These parameters learn high-level features related to specific tasks. These features can better reflect information in specific scenarios and need to be adjusted according to driving scene data. The method for determining the dividing boundary can be based on the structure and experience of the model. For example, it can be divided equally into the first half and the second half according to the number of convolution layers.
[0034] Step S1242: Freeze the underlying parameters of the feature extraction network so that the underlying parameters remain unchanged during the first stage fine-tuning process.
[0035] In this embodiment, the underlying parameters of the feature extraction network are frozen to prevent the destruction of common features learned during pre-training during fine-tuning. This can be achieved by setting the trainable flag of the parameter to non-trainable. This prevents the underlying parameters from changing with gradient updates during subsequent training, thereby preserving the useful information learned during pre-training.
[0036] Step S1243: Input the driving scene data in the scene adaptability training subset into the pre-trained feature extraction network in preset batches, extract basic environmental features through the bottom-level parameters of the feature extraction network, and then perform scene adaptability conversion on the basic environmental features through the top-level parameters of the feature extraction network to generate scene adaptation features.
[0037] In this embodiment, the driving scene data in the scene adaptability training subset are input into the pre-trained feature extraction network in preset batches. During the input process, the bottom-level parameters of the feature extraction network first extract basic environmental features from the input driving scene data. These basic environmental features are universal features related to a variety of scenes, such as the edges of the road, the outlines of the vehicle, etc. Then, the basic environmental features are subjected to scene adaptability conversion through the top-level parameters of the feature extraction network to generate scene adaptation features. The scene adaptation features can better reflect the characteristics of driving scenes under different weather conditions. For example, on rainy days, the scene adaptation features may highlight the reflective effect of the vehicle on a slippery road surface. Inputting data in preset batches can improve the efficiency of training, and also help the model better learn the distribution characteristics of the data.
[0038] Step S1244: inputting the scene adaptation feature into the perception task output layer, and outputting the driving scene perception prediction result through the perception task output layer.
[0039] In this embodiment, the generated scene adaptation features are input into the perception task output layer. This layer, which can be a fully connected layer, performs a linear transformation and activation function on the scene adaptation features to output information such as the predicted category and location of environmental elements. Based on the input scene adaptation features and its own parameter settings, the perception task output layer analyzes and determines the driving scene, ultimately outputting a perception prediction result for the driving scene.
[0040] Step S1245: Calculate the loss value between the driving scene perception prediction result and the actual perception result corresponding to the driving scene data in the scene adaptability training subset.
[0041] In this example, the loss value is calculated between the driving scene perception prediction results and the actual perception results corresponding to the driving scene data in the scene adaptability training subset. The loss value is used to measure the difference between the model's prediction results and the actual results. Common loss functions include cross-entropy loss function and mean squared error loss function. The appropriate loss function is selected based on the specific task requirements. By comparing the predicted results with the actual results, the model's performance at the current stage can be quantified.
[0042] Step S1246: Based on the loss value, adjust the top-level parameters of the feature extraction network and the parameters of the perception task output layer through the back propagation algorithm to gradually reduce the loss value.
[0043] In this embodiment, the backpropagation algorithm is used to adjust the top-level parameters of the feature extraction network and the parameters of the perception task output layer based on the calculated loss value. The backpropagation algorithm is an optimization algorithm based on gradient descent. It calculates the gradient of each parameter based on the loss value and then updates the parameter in the opposite direction of the gradient. During the parameter update process, an optimizer is used to control the parameter update step size. Common optimizers include stochastic gradient descent (SGD) and adaptive moment estimation (Adam). By continuously adjusting the parameters, the loss value is gradually reduced, thereby improving the performance of the model.
[0044] Step S1247: Repeat the step of inputting the driving scene data in the scene adaptability training subset into the feature extraction network in preset batches to adjust the parameters until the preset number of first-stage fine-tuning iterations is reached, thereby completing the first-stage fine-tuning.
[0045] In this embodiment, the driving scene data from the scene adaptability training subset is fed into the feature extraction network in preset batches. A series of steps, including feature extraction, scene adaptability feature generation, perception task output, loss calculation, and parameter adjustment, are repeated until a preset number of first-stage fine-tuning iterations are reached. This number of first-stage fine-tuning iterations is determined based on experience and experimentation to ensure that the model fully learns on the scene adaptability training subset and achieves optimal performance. Once the preset number of iterations is reached, the first-stage fine-tuning is complete, and the model's scene adaptability under various weather conditions has been significantly improved.
[0046] Step S125: After the first stage of fine-tuning is completed, the underlying parameters of the feature extraction network are unfrozen, the verification subset is input into the feature extraction network, and all parameters of the feature extraction network and the parameters of the perception task output layer are adjusted to perform the second stage of fine-tuning.
[0047] In this embodiment, after the first stage of fine-tuning is completed, the underlying parameters of the feature extraction network are unfrozen. Because in the first stage, in order to retain the pre-trained common features, the underlying parameters are frozen. In the second stage, the underlying parameters need to be involved in the adjustment to further optimize the performance of the model. The verification subset is input into the feature extraction network. The verification subset contains driving scene data under different road conditions, which allows the model to learn the characteristic changes of environmental elements under different road conditions. At the same time, all the parameters of the feature extraction network and the parameters of the perception task output layer are adjusted. Through continuous iterative optimization, the model can also perform well under different road conditions. During the second stage of fine-tuning, the loss value between the output result of the model and the true label of the verification subset is also calculated, and then the backpropagation algorithm is used to update the parameters until the performance of the model reaches a satisfactory state, and finally the basic environmental perception model is obtained.
[0048] Step S130: Introducing an attention mechanism module into the basic environmental perception model so that the basic environmental perception model focuses on key environmental elements for driving safety, and generates an enhanced perception model that enhances the perception capability of key elements.
[0049] In this embodiment, an attention mechanism module is introduced to enable the basic environmental perception model to pay more attention to environmental factors critical to driving safety. The attention mechanism module can help the model automatically focus on important information and improve its perception of key elements.
[0050] Step S131: inserting an attention mechanism module between the feature extraction network and the perception task output layer of the basic environment perception model, wherein the attention mechanism module includes a spatial attention submodule and a channel attention submodule.
[0051] In this embodiment, an attention mechanism module is inserted between the feature extraction network and the perception task output layer of the basic environmental perception model. This attention mechanism module consists of a spatial attention submodule and a channel attention submodule. The spatial attention submodule focuses on the importance of different spatial locations in the feature map, while the channel attention submodule focuses on the importance of different channels. Through the collaborative work of these two submodules, the characteristics of key environmental elements for driving safety can be more comprehensively captured.
[0052] Step S132: defining a set of key driving safety environment elements, wherein the set of key driving safety environment elements includes a vehicle ahead, pedestrians, traffic lights, and road signs.
[0053] In this embodiment, a set of key environmental elements for driving safety is clearly defined. This set includes elements such as vehicles ahead, pedestrians, traffic lights, and road signs. These elements are crucial for driving safety, and the model needs to focus on the status and location of these elements.
[0054] Step S133: Input the environmental feature map output by the feature extraction network of the basic environmental perception model into the spatial attention sub-module, and generate a spatial attention weight map by calculating the semantic similarity between each spatial position in the environmental feature map and the elements in the set of driving safety key environmental elements.
[0055] In this embodiment, the environmental feature map output by the feature extraction network of the basic environmental perception model is input into the spatial attention submodule. In the spatial attention submodule, the semantic similarity between each spatial position in the environmental feature map and the elements in the set of key environmental elements for driving safety is calculated. The semantic similarity reflects the degree of match between the features of each spatial position and the features of the key environmental elements. Based on the calculated semantic similarity, a spatial attention weight map is generated. Each position in the spatial attention weight map corresponds to a spatial position in the environmental feature map, and the weight value indicates the importance of the position. The higher the weight value, the stronger the correlation between the position and the key environmental elements, and the model should pay more attention to the position.
[0056] Step S134: Input the environmental feature map into the channel attention submodule, and generate a channel attention weight vector by analyzing the response strength of different channels in the environmental feature map to key environmental elements for driving safety.
[0057] Step S1341: decomposing the environment feature map into multiple single-channel feature maps according to the channel dimension, each single-channel feature map corresponds to a feature channel of the environment feature map.
[0058] In this embodiment, the environmental feature map is decomposed by channel dimension to obtain multiple single-channel feature maps. Each single-channel feature map corresponds to a feature channel in the environmental feature map, and these single-channel feature maps contain characteristic information from different aspects. By decomposing the environmental feature map, it is easier to analyze the response of each channel to environmental factors that are critical to driving safety.
[0059] Step S1342: Selecting sample elements from the set of driving safety critical environmental elements, and extracting the marked areas corresponding to the sample elements in the driving scene dataset.
[0060] In this example, sample elements, such as vehicles ahead and pedestrians, are selected from a set of environmental elements critical to driving safety. The annotated regions corresponding to these sample elements in the driving scene dataset are then extracted. The annotated regions, created during the data collection and preprocessing stages, clearly define the specific locations of the sample elements within the image.
[0061] Step S1343: Perform regional feature matching on each single-channel feature map and the annotated area of the sample element, and calculate the average response value of each single-channel feature map in the annotated area. The average response value is used to represent the response intensity of the feature channel to the sample element.
[0062] In this embodiment, regional feature matching is performed between each single-channel feature map and the annotated region of the sample element. The average response value is calculated by comparing the feature values of the single-channel feature map within the annotated region. The average response value reflects the response strength of the feature channel to the sample element. A higher average response value indicates a stronger ability of the channel to express the characteristics of the sample element.
[0063] Step S1344: normalize the average response values of all single-channel feature maps to obtain the normalized response strength of each feature channel, where the value range of the normalized response strength is within a preset interval.
[0064] In this embodiment, the average response values of all single-channel feature maps are normalized. Normalization can unify the average response values of different channels into a preset range, facilitating subsequent comparison and calculation. Through normalization, the scale differences of the response values between different channels are eliminated, allowing the importance of each channel to be evaluated on the same scale.
[0065] Step S1345: constructing a channel importance scoring function based on the normalized response strength, and calculating the importance score of each feature channel through the channel importance scoring function to form an initial channel weight vector.
[0066] In this embodiment, a channel importance scoring function is constructed based on the normalized response strength. This channel importance scoring function calculates the importance score of each feature channel based on the normalized response strength. Using this channel importance scoring function, each channel is assigned an importance score, which forms an initial channel weight vector. This initial channel weight vector reflects the importance of each channel to the environmental elements critical to driving safety.
[0067] Step S1346: Smoothing the initial channel weight vector to generate a final channel attention weight vector.
[0068] In this embodiment, the initial channel weight vector is smoothed. Smoothing can reduce noise and fluctuations in the weight vector, making the weight vector more stable and reliable. Through smoothing, the final channel attention weight vector is generated, which will be used for subsequent feature weighting processing.
[0069] Step S135: Perform spatial dimension weighted multiplication processing on the spatial attention weight map and the environmental feature map to obtain a spatial enhancement feature map.
[0070] In this embodiment, a weighted multiplication of the spatial attention weight map and the environmental feature map is performed in the spatial dimension. Each position in the spatial attention weight map has a weight value. This weight value is multiplied by the feature value of the corresponding position in the environmental feature map to generate a spatial enhancement feature map. This can highlight the spatial locations in the environmental feature map that are related to environmental elements critical to driving safety and enhance the feature expression of these locations.
[0071] Step S136: Perform channel dimension weighted multiplication processing on the channel attention weight vector and the spatial enhancement feature map to obtain an enhanced feature map that integrates spatial and channel attention.
[0072] In this embodiment, the channel attention weight vector and the spatial enhancement feature map are weighted and multiplied in the channel dimension. Each element in the channel attention weight vector corresponds to a channel of the spatial enhancement feature map. The elements of the channel attention weight vector are multiplied by the eigenvalues of the corresponding channel of the spatial enhancement feature map to obtain an enhanced feature map that integrates spatial and channel attention. In this way, the channel features related to key environmental elements of driving safety are further highlighted, making the enhanced feature map more focused on key elements.
[0073] Step S137: Input the enhanced feature map into the perception task output layer, adjust the connection parameters between the attention mechanism module and the perception task output layer, so that the basic environmental perception model has a higher perception response intensity to driving safety-critical environmental elements than to non-critical environmental elements, and generate an enhanced perception model.
[0074] In this embodiment, the enhanced feature map is input into the perception task output layer. Simultaneously, the connection parameters between the attention mechanism module and the perception task output layer are adjusted. Through continuous training and optimization, the basic environmental perception model is made to perceive environmental elements critical to driving safety more responsively than non-critical environmental elements. This enables the model to more accurately identify and focus on critical environmental elements, generating an enhanced perception model with enhanced perception capabilities for these elements.
[0075] Step S140: Generate extreme driving scenario data by using a generative adversarial network simulation, input the extreme driving scenario data into the enhanced perception model for enhanced training, and obtain a target environment perception model.
[0076] In this embodiment, to further improve the performance of the enhanced perception model in extreme driving scenarios, a generative adversarial network (GAN) was used to simulate and generate data for these scenarios. While data for these scenarios can be difficult to collect in practice, GANs can simulate these scenarios, providing more training data for the model.
[0077] Step S141: constructing a generative adversarial network, wherein the generative adversarial network includes a generator and a discriminator, wherein the generator is used to generate simulated extreme driving scene data, and the discriminator is used to distinguish between real driving scene data and generated simulated extreme driving scene data.
[0078] In this embodiment, a generative adversarial network (GAN) is constructed, consisting of a generator and a discriminator. The generator generates simulated extreme driving scenario data based on random noise input. This data is as close to real-world extreme driving scenarios as possible. The discriminator distinguishes whether the input data is real driving scenario data or simulated data generated by the generator. Through continuous adversarial training, the generator and discriminator gradually improve the quality of the generated data.
[0079] Step S142: extracting an extreme scene feature set from the driving scene dataset, wherein the extreme scene feature set includes environmental visual features under extreme weather conditions, road surface structure features under complex road conditions, and sudden traffic participant behavior features.
[0080] In this embodiment, an extreme scene feature set is extracted from the driving scene dataset. This extreme scene feature set includes environmental visual features under extreme weather conditions, such as image features in weather conditions such as heavy rain and dense fog; road surface structural features under complex road conditions, such as rugged mountain roads and flooded roads; and unexpected traffic participant behavior features, such as sudden vehicle braking and pedestrians jaywalking.
[0081] Step S143: Input the extreme scene feature set as a conditional constraint into the generator, perform up-sampling processing on the random noise vector through the multi-layer transposed convolution operation of the generator, and generate preliminary simulated extreme driving scene data.
[0082] In this embodiment, the extreme scenario feature set is fed into the generator as a conditional constraint. The generator uses multi-layer transposed convolution operations to upsample the random noise vector. This transposed convolution operation converts the low-resolution random noise vector into high-resolution image data. Combined with the constraints of the extreme scenario feature set, preliminary simulated extreme driving scenario data is generated. This data reflects, to a certain extent, the characteristics of extreme driving scenarios.
[0083] Step S144: The preliminary simulated extreme driving scene data is mixed with the real extreme scene data in the driving scene dataset and input into the discriminator, and features are extracted through the convolution operation of the discriminator and the authenticity judgment probability is output.
[0084] In this embodiment, preliminary simulated extreme driving scenario data is mixed with real extreme scenario data from the driving scenario dataset and then input into the discriminator. The discriminator extracts features from the input data through a convolution operation and then outputs a probability of authenticity based on these features. This probability indicates the likelihood that the input data is real. By continuously training the discriminator, it can accurately distinguish between real and simulated data.
[0085] Step S145: Calculate the generator loss function and the discriminator loss function based on the authenticity judgment probability, and alternately optimize the parameters of the generator and the discriminator so that the authenticity judgment probability of the simulated extreme driving scene data generated by the generator is close to a preset threshold.
[0086] In this embodiment, the generator loss function and the discriminator loss function are calculated based on the authenticity probabilities output by the discriminator. The goal of the generator is to generate simulated data that can deceive the discriminator, so the generator loss function measures the degree to which the generated data can be identified as real data. The goal of the discriminator is to accurately distinguish between real data and simulated data, and the discriminator loss function measures the discriminator's classification accuracy. By alternately optimizing the parameters of the generator and discriminator, the quality of the generator's generated data is continuously improved, bringing the authenticity probability of the generated data for simulated extreme driving scenarios close to a preset threshold.
[0087] Step S146: Perform quality screening on the simulated extreme driving scene data generated by the trained generator, and retain the simulated extreme driving scene data whose feature similarity with the real extreme scene data is higher than the set standard to form an extreme driving scene enhanced dataset.
[0088] Step S1461: selecting real extreme scene data from the driving scene data set as a reference data set, where the reference data set contains real data samples of different types of extreme scenes.
[0089] In this embodiment, real extreme scene data is selected from the driving scene dataset as a reference dataset. The reference dataset contains real data samples of different types of extreme scenes, such as scenes in extreme weather conditions such as heavy rain and snow, as well as scenes under complex mountain roads and accident scenes.
[0090] Step S1462: extracting a scene feature vector of each real data sample in the reference data set, wherein the scene feature vector includes a combination of environmental visual features, road surface structure features, and traffic participant behavior features.
[0091] In this embodiment, a scene feature vector is extracted for each real-world data sample in the reference dataset. This scene feature vector comprises a combination of environmental visual features, road surface structural features, and traffic participant behavioral characteristics. By extracting these feature vectors, the characteristics of the real-world data sample can be quantified, facilitating subsequent comparison with simulated data.
[0092] Step S1463: Extract the scene feature vector of each simulated extreme driving scene data generated by the trained generator, where the dimension of the scene feature vector is consistent with the dimension of the scene feature vector of the real data sample in the reference data set.
[0093] In this example, the scene feature vectors for each simulated extreme driving scenario data set generated by the trained generator are extracted. To enable effective comparison, the dimensions of the scene feature vectors for the simulated data are kept consistent with those of the real data samples in the reference dataset. This ensures that subsequent feature similarity calculations are performed in the same dimensional space.
[0094] Step S1464: Calculate the feature similarity between the scene feature vector of each simulated extreme driving scene data and the scene feature vectors of all real data samples in the reference data set.
[0095] In this embodiment, feature similarity is calculated between the scene feature vectors of each simulated extreme driving scenario data and the scene feature vectors of all real data samples in the reference dataset. Feature similarity can be measured using a similarity calculation method, such as cosine similarity. By calculating feature similarity, the degree of similarity between each simulated data set and the real data can be understood.
[0096] Step S1465: The maximum feature similarity corresponding to each simulated extreme driving scene data is used as the feature similarity between the simulated extreme driving scene data and the real extreme scene data.
[0097] In this embodiment, for each simulated extreme driving scenario data, the maximum value of its feature similarity with all real data samples in the reference data set is selected as the feature similarity between the simulated data and the real extreme scenario data, thereby ensuring that the quality of the simulated data is evaluated with the similarity closest to the real data.
[0098] Step S1466: Mark the simulated extreme driving scenario data with feature similarity higher than the set standard as qualified simulated data.
[0099] In this embodiment, simulated extreme driving scenario data with feature similarity exceeding the specified criteria is marked as qualified simulated data. This standard is determined based on practical needs and experience, ensuring that the retained simulated data is of high quality and can be effectively used for enhanced model training.
[0100] Step S1467: Collect all qualified simulation data to form an extreme driving scenario enhanced dataset, where the number of simulation data in the extreme driving scenario enhanced dataset maintains a preset proportional relationship with the number of real data samples in the reference dataset.
[0101] In this example, all qualified simulated data are collected to form an enhanced dataset for extreme driving scenarios. To ensure the rationality and effectiveness of the dataset, the amount of simulated data in the enhanced dataset maintains a preset ratio with the number of real data samples in the reference dataset. This ensures that the enhanced dataset contains sufficient simulated data to enrich the training samples, while not affecting the model's learning performance due to excessive simulated data.
[0102] Step S147: The extreme driving scene enhanced dataset is mixed with the driving scene dataset, and the mixed dataset is input into the enhanced perception model for enhanced training. The parameters of the enhanced perception model are adjusted through the back propagation algorithm so that the perception error rate of the enhanced perception model on the extreme driving scene data is lower than the preset error threshold, thereby obtaining the target environment perception model.
[0103] In this example, an enhanced dataset of extreme driving scenarios is mixed with a driving scenario dataset and then fed into an enhanced perception model for enhanced training. During training, the loss between the model's output and the true labels of the mixed data is calculated, and then the parameters of the enhanced perception model are updated using a backpropagation algorithm. Training is iterated until the enhanced perception model's perception error rate for extreme driving scenario data falls below a preset error threshold. When this threshold is reached, the model's performance in extreme driving scenarios has been significantly improved, ultimately resulting in the target environment perception model.
[0104] Step S150: Perceiving the driving scene data acquired by different sensors through the target environment perception model, and outputting a safe driving environment perception result.
[0105] In this embodiment, the target environment perception model is used to perceive the driving scene data acquired by different sensors, and finally outputs the safe driving environment perception result.
[0106] Step S151: Acquire driving scene data collected by different types of sensors, where the different types of sensors include visual sensors and radar sensors, and the driving scene data includes environmental image information collected by the visual sensors and environmental point cloud information collected by the radar sensors.
[0107] In this embodiment, driving scene data collected by different types of sensors is acquired. Visual sensors, such as high-definition cameras, capture environmental image information, which includes the appearance characteristics of roads, vehicles, pedestrians, etc. Radar sensors, such as millimeter-wave radars and lidars, collect environmental point cloud information. This point cloud information provides information such as the distance, speed, and three-dimensional spatial position of target objects. By combining data from these two types of sensors, more comprehensive driving scene information can be obtained.
[0108] Step S152: Input the environmental image information into the image feature extraction subnetwork of the target environmental perception model, and extract environmental image features through convolution operations and pooling operations. The environmental image features include visual appearance features and two-dimensional spatial position features of environmental elements.
[0109] In this embodiment, environmental image information is input into the image feature extraction subnetwork of the target environment perception model. This subnetwork processes the environmental image information through convolution and pooling operations. Convolution extracts local features from the image, such as edges and textures. Pooling reduces the dimensionality of the extracted features, reducing the number of features while retaining important information. After processing, environmental image features are extracted. These features include the visual appearance characteristics and two-dimensional spatial position characteristics of environmental elements.
[0110] Step S153: Input the environmental point cloud information into the point cloud feature extraction subnetwork of the target environment perception model, and extract environmental point cloud features through voxel processing and attention pooling operations. The environmental point cloud features include three-dimensional spatial position features and distance depth features of environmental elements.
[0111] Step S1531: performing coordinate system normalization processing on the environmental point cloud information, and converting the point cloud coordinates collected by different radar sensors into a unified vehicle coordinate system.
[0112] In this embodiment, the environmental point cloud information is normalized to a uniform coordinate system. Because different radar sensors may have different installation locations and measurement methods, the collected point cloud coordinates may differ. Therefore, these point cloud coordinates must be converted to a unified vehicle coordinate system for subsequent processing and analysis. Coordinate normalization can be achieved using methods such as coordinate transformation matrices to ensure that all point cloud data is represented in the same coordinate system.
[0113] Step S1532: setting three-dimensional voxel grid parameters, dividing the environmental point cloud information in the unified coordinate system into a plurality of three-dimensional voxel units, each of which contains point cloud data within a preset spatial range.
[0114] In this embodiment, 3D voxel grid parameters are set to divide the environmental point cloud information within a unified coordinate system into multiple 3D voxel units. Each 3D voxel unit contains point cloud data within a preset spatial range. Voxelization discretizes continuous point cloud data, facilitating subsequent feature extraction and analysis. By properly setting the 3D voxel grid parameters, the size and distribution of the voxel units can be tailored to meet actual needs.
[0115] Step S1533: extracting features from the point cloud data within each three-dimensional voxel unit to generate a voxel feature vector, which includes the number of point clouds within the voxel, average coordinates, and reflection intensity statistical features.
[0116] In this embodiment, feature extraction is performed on the point cloud data within each 3D voxel unit to generate a voxel feature vector. The voxel feature vector contains information such as the number of point clouds within the voxel, average coordinates, and statistical characteristics of reflection intensity. These voxel feature vectors can reflect the distribution and characteristics of the point cloud data within the voxel unit.
[0117] Step S1534: Arrange the voxel feature vectors of all three-dimensional voxel units according to their spatial positions to construct a three-dimensional voxel feature matrix.
[0118] In this embodiment, the voxel feature vectors of all 3D voxel units are arranged according to their spatial positions to construct a 3D voxel feature matrix. The 3D voxel feature matrix can more intuitively represent the feature distribution of the entire environment point cloud data, providing suitable input for subsequent attention pooling operations.
[0119] Step S1535: Input the three-dimensional voxel feature matrix into the attention pooling layer of the point cloud feature extraction sub-network to calculate the contribution weight of each three-dimensional voxel unit to the key environmental elements of driving safety.
[0120] In this embodiment, the 3D voxel feature matrix is input into the attention pooling layer of the point cloud feature extraction subnetwork. The attention pooling layer calculates the contribution weight of each 3D voxel unit to the key environmental elements of driving safety. By analyzing the correlation between each voxel unit's features and the key environmental elements of driving safety, its contribution weight is determined. The higher the contribution weight, the more important the voxel unit is in perceiving the key environmental elements.
[0121] Step S1536: performing weighted aggregation processing on the three-dimensional voxel feature matrix based on the contribution weight, and performing dimensionality reduction processing on the weighted aggregated voxel features through a convolution operation to generate an environmental point cloud feature containing the three-dimensional spatial position features and distance depth features of the environmental elements.
[0122] In this embodiment, a weighted aggregation process is performed on the 3D voxel feature matrix based on the calculated contribution weights. The feature vector of each voxel unit is multiplied by its contribution weight and then aggregated. Next, a dimensionality reduction process is performed on the weighted aggregated voxel features through a convolution operation, reducing the number of features while retaining important information. Ultimately, an environmental point cloud feature is generated, which contains the 3D spatial position characteristics and distance and depth characteristics of the environmental elements.
[0123] Step S154: inputting the environmental image features and the environmental point cloud features into the feature fusion layer of the target environmental perception model, and merging the environmental image features and the environmental point cloud features into a target fusion feature vector through a feature splicing operation.
[0124] In this embodiment, the environmental image features and environmental point cloud features are input into the feature fusion layer of the target environment perception model. In this layer, these two features are combined into a target fused feature vector through a feature concatenation operation. This concatenation operation combines different types of features, fully utilizing information from both the visual and radar sensors.
[0125] Step S155: performing dimension normalization processing on the target fusion feature vector to eliminate the dimensional differences between the features of different sensors and obtain a standardized fusion feature.
[0126] In this embodiment, the target fused feature vector is dimensionally normalized. Because the environmental image features and the environmental point cloud features come from different sensors, they may have different dimensions. Dimensional normalization can unify the different sensor features to the same scale, eliminating these differences and enabling the model to more effectively process the fused features. After normalization, standardized fused features are obtained.
[0127] Step S156: Input the standardized fusion features into the perception result output layer of the target environment perception model, and generate a safe driving environment perception result including environmental element categories, position coordinates and motion state parameters through a fully connected network and activation function processing.
[0128] In this embodiment, the standardized fused features are input into the perception result output layer of the target environment perception model. This layer performs a linear transformation on the standardized fused features using a fully connected network, followed by nonlinear processing using an activation function. Ultimately, a safe driving environment perception result is generated, including environmental element categories, location coordinates, and motion state parameters. This safe driving environment perception result can provide drivers with important information, helping them make safe driving decisions.
[0129] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a safe driving environment perception system 100 incorporating artificial intelligence, which can implement the concepts of the present application, as provided in some embodiments of the present application. For example, a processor 120 can be used in the safe driving environment perception system 100 incorporating artificial intelligence to perform the functions described in the present application.
[0130] The safe driving environment perception system 100 incorporating artificial intelligence can be a general-purpose server or a special-purpose server, both of which can be used to implement the safe driving environment perception method incorporating artificial intelligence in this application. Although this application only shows a single server, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0131] For example, the safe driving environment perception system 100 incorporating artificial intelligence may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the safe driving environment perception system 100 incorporating artificial intelligence may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The safe driving environment perception system 100 incorporating artificial intelligence also includes an I / O interface 150 between the computer and other input and output devices.
[0132] For ease of explanation, only one processor is described in the safe driving environment perception system 100 combined with artificial intelligence. However, it should be noted that the safe driving environment perception system 100 combined with artificial intelligence in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the safe driving environment perception system 100 combined with artificial intelligence executes step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0133] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned safe driving environment perception method combined with artificial intelligence is implemented.
[0134] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A safe driving environment perception method combined with artificial intelligence, characterized in that: The method comprises: Constructing a driving scenario dataset, wherein the driving scenario dataset covers driving scenario data under various weather conditions, different road conditions, and various traffic participant behavior patterns; Based on the driving scenario dataset, pre-training the initial environment perception model and fine-tuning the scene adaptability are performed using transfer learning to obtain a basic environment perception model; Introducing an attention mechanism module into the basic environmental perception model to enable the basic environmental perception model to focus on environmental elements critical to driving safety, thereby generating an enhanced perception model that enhances the perception capability of key elements; Generate extreme driving scenario data by using a generative adversarial network simulation, input the extreme driving scenario data into the enhanced perception model for enhanced training, and obtain a target environment perception model; The target environment perception model is used to perceive the driving scene data acquired by different sensors and output a safe driving environment perception result.
2. The method for safe driving environment perception combined with artificial intelligence according to claim 1, characterized in that: Based on the driving scene dataset, the initial environment perception model is pre-trained and fine-tuned for scene adaptability using transfer learning to obtain a basic environment perception model, including: Selecting an initial environment perception model that includes general image feature extraction capabilities, wherein the initial environment perception model includes a feature extraction network and a perception task output layer; Pre-training the feature extraction network of the initial environment perception model on a general image dataset so that the feature extraction network has basic environmental element recognition capabilities; Dividing the driving scene dataset into a scene adaptability training subset and a validation subset, wherein the scene adaptability training subset includes driving scene data under different weather conditions, and the validation subset includes driving scene data under different road conditions; Inputting the scene adaptation training subset into the pre-trained feature extraction network, performing first-stage fine-tuning by freezing the bottom-layer parameters of the feature extraction network and adjusting the top-layer parameters of the feature extraction network and the parameters of the perception task output layer; After the first stage of fine-tuning is completed, the underlying parameters of the feature extraction network are unfrozen, the verification subset is input into the feature extraction network, and all parameters of the feature extraction network and the parameters of the perception task output layer are adjusted to perform the second stage of fine-tuning; The verification subset is used to monitor the perception accuracy change trend of the initial environmental perception model during the scene adaptation fine-tuning process. When the perception accuracy change trend remains stable for multiple consecutive iterations, the fine-tuning is stopped and the current model parameters are saved to obtain the basic environmental perception model.
3. The method for safe driving environment perception combined with artificial intelligence according to claim 2, characterized in that: Inputting the scene adaptability training subset into the pre-trained feature extraction network, freezing the bottom parameters of the feature extraction network, and adjusting the top parameters of the feature extraction network and the parameters of the perception task output layer to perform first-stage fine-tuning, including: Determining a dividing boundary between bottom-layer parameters and top-layer parameters of the feature extraction network, wherein the bottom-layer parameters include parameters of the first half of the convolutional layer close to the input layer, and the top-layer parameters include parameters of the second half of the convolutional layer close to the output layer; Freezing the underlying parameters of the feature extraction network so that the underlying parameters remain unchanged during the first stage fine-tuning process; Inputting the driving scene data in the scene adaptation training subset into the pre-trained feature extraction network in preset batches, extracting basic environmental features using the bottom-level parameters of the feature extraction network, and then performing scene adaptation conversion on the basic environmental features using the top-level parameters of the feature extraction network to generate scene adaptation features; Inputting the scene adaptation feature into the perception task output layer, and outputting the driving scene perception prediction result through the perception task output layer; Calculating a loss value between the driving scene perception prediction result and a true perception result corresponding to the driving scene data in the scene adaptability training subset; Based on the loss value, adjusting the top-level parameters of the feature extraction network and the parameters of the perception task output layer through a back-propagation algorithm so that the loss value gradually decreases; Repeat the steps of inputting the driving scene data in the scene adaptability training subset into the feature extraction network in preset batches to adjust the parameters until the preset number of first-stage fine-tuning iterations is reached, thus completing the first-stage fine-tuning.
4. The method for safe driving environment perception combined with artificial intelligence according to claim 1, characterized in that: The introduction of an attention mechanism module into the basic environmental perception model enables the basic environmental perception model to focus on key environmental elements for driving safety and generate an enhanced perception model that enhances the perception capability of key elements, including: Inserting an attention mechanism module between the feature extraction network and the perception task output layer of the basic environment perception model, wherein the attention mechanism module includes a spatial attention submodule and a channel attention submodule; Defining a set of key environmental elements for driving safety, wherein the set of key environmental elements for driving safety includes vehicles ahead, pedestrians, traffic lights, and road signs; Inputting the environmental feature map output by the feature extraction network of the basic environmental perception model into the spatial attention submodule, and generating a spatial attention weight map by calculating the semantic similarity between each spatial position in the environmental feature map and the elements in the set of driving safety key environmental elements; Inputting the environmental feature map into the channel attention submodule, and generating a channel attention weight vector by analyzing the response strength of different channels in the environmental feature map to key environmental elements for driving safety; Performing a spatial dimension weighted multiplication process on the spatial attention weight map and the environmental feature map to obtain a spatial enhancement feature map; Performing a channel dimension weighted multiplication process on the channel attention weight vector and the spatial enhancement feature map to obtain an enhanced feature map that integrates spatial and channel attention; The enhanced feature map is input into the perception task output layer, and the connection parameters between the attention mechanism module and the perception task output layer are adjusted so that the basic environmental perception model has a higher perception response intensity to environmental elements critical to driving safety than to non-critical environmental elements, thereby generating an enhanced perception model.
5. The method for safe driving environment perception combined with artificial intelligence according to claim 4 is characterized in that: The step of inputting the environmental feature map into the channel attention submodule and generating a channel attention weight vector by analyzing the response strength of different channels in the environmental feature map to key environmental elements for driving safety comprises: Decomposing the environmental feature map into multiple single-channel feature maps according to the channel dimension, each single-channel feature map corresponds to a feature channel of the environmental feature map; Selecting sample elements from the set of driving safety key environmental elements, and extracting the marked areas corresponding to the sample elements in the driving scene dataset; Performing regional feature matching on each single-channel feature map and the marked area of the sample element, and calculating the average response value of each single-channel feature map in the marked area, wherein the average response value is used to represent the response intensity of the feature channel to the sample element; Normalizing the average response values of all single-channel feature maps to obtain a standardized response intensity for each feature channel, where the value range of the standardized response intensity is within a preset interval; Constructing a channel importance scoring function based on the normalized response strength, and calculating the importance score of each feature channel using the channel importance scoring function to form an initial channel weight vector; The initial channel weight vector is smoothed to generate a final channel attention weight vector.
6. The method for safe driving environment perception combined with artificial intelligence according to claim 1, characterized in that: The method of using a generative adversarial network to simulate and generate extreme driving scenario data, inputting the extreme driving scenario data into the enhanced perception model for enhanced training, and obtaining a target environment perception model includes: Constructing a generative adversarial network, the generative adversarial network comprising a generator and a discriminator, the generator being used to generate simulated extreme driving scenario data, and the discriminator being used to distinguish between real driving scenario data and the generated simulated extreme driving scenario data; Extracting an extreme scene feature set from the driving scene dataset, wherein the extreme scene feature set includes environmental visual features under extreme weather conditions, road surface structure features under complex road conditions, and unexpected traffic participant behavior features; Inputting the extreme scene feature set as a conditional constraint into the generator, performing upsampling processing on the random noise vector through the multi-layer transposed convolution operation of the generator, and generating preliminary simulated extreme driving scene data; Mixing the preliminary simulated extreme driving scene data with the real extreme scene data in the driving scene dataset and inputting the mixed data into the discriminator, extracting features through the convolution operation of the discriminator and outputting the probability of true or false judgment; Calculating a generator loss function and a discriminator loss function based on the authenticity judgment probability, and alternately optimizing the parameters of the generator and the discriminator so that the authenticity judgment probability of the simulated extreme driving scene data generated by the generator is close to a preset threshold; The simulated extreme driving scene data generated by the trained generator is quality-screened, and the simulated extreme driving scene data whose feature similarity with the real extreme scene data is higher than the set standard is retained to form the extreme driving scene enhanced dataset; The extreme driving scene enhanced dataset is mixed with the driving scene dataset and input into the enhanced perception model for enhanced training. The parameters of the enhanced perception model are adjusted through the back propagation algorithm so that the perception error rate of the enhanced perception model on the extreme driving scene data is lower than the preset error threshold, thereby obtaining the target environment perception model.
7. The method for safe driving environment perception combined with artificial intelligence according to claim 6, characterized in that: The simulated extreme driving scene data generated by the trained generator is quality screened, and the simulated extreme driving scene data with feature similarity to the real extreme scene data higher than the set standard is retained to form an extreme driving scene enhanced dataset, including: Selecting real extreme scene data from the driving scene dataset as a reference dataset, wherein the reference dataset includes real data samples of different types of extreme scenes; Extracting a scene feature vector for each real data sample in the reference data set, wherein the scene feature vector comprises a combination of environmental visual features, road surface structural features, and traffic participant behavioral features; Extracting a scene feature vector for each simulated extreme driving scene data generated by the trained generator, where the dimension of the scene feature vector is consistent with the dimension of the scene feature vector of the real data sample in the reference dataset; Calculate the feature similarity between the scene feature vector of each simulated extreme driving scene data and the scene feature vectors of all real data samples in the reference dataset; The maximum feature similarity corresponding to each simulated extreme driving scene data is used as the feature similarity between the simulated extreme driving scene data and the real extreme scene data; Mark the simulated extreme driving scenario data whose feature similarity is higher than the set standard as qualified simulated data; All qualified simulation data are collected to form an extreme driving scenario enhanced dataset, where the amount of simulation data in the extreme driving scenario enhanced dataset maintains a preset proportional relationship with the amount of real data samples in the reference dataset.
8. The method for safe driving environment perception combined with artificial intelligence according to claim 1, characterized in that: The target environment perception model is used to perceive driving scene data acquired by different sensors and output a safe driving environment perception result, including: Acquiring driving scene data collected by different types of sensors, the different types of sensors including visual sensors and radar sensors, the driving scene data including environmental image information collected by the visual sensors and environmental point cloud information collected by the radar sensors; Inputting the environmental image information into the image feature extraction subnetwork of the target environment perception model, extracting environmental image features through convolution and pooling operations, wherein the environmental image features include visual appearance features and two-dimensional spatial position features of environmental elements; Inputting the environmental point cloud information into the point cloud feature extraction subnetwork of the target environment perception model, extracting environmental point cloud features through voxelization processing and attention pooling operations, wherein the environmental point cloud features include three-dimensional spatial position features and distance depth features of environmental elements; Inputting the environmental image features and the environmental point cloud features into a feature fusion layer of the target environment perception model, and merging the environmental image features and the environmental point cloud features into a target fusion feature vector through a feature splicing operation; Performing dimension normalization on the target fusion feature vector to eliminate the dimensional differences between the features of different sensors and obtain a standardized fusion feature; The standardized fusion features are input into the perception result output layer of the target environment perception model, and processed through a fully connected network and activation function to generate a safe driving environment perception result including environmental element categories, position coordinates and motion state parameters.
9. The method for safe driving environment perception combined with artificial intelligence according to claim 8, characterized in that: The environmental point cloud information is input into the point cloud feature extraction subnetwork of the target environment perception model, and environmental point cloud features are extracted through voxelization processing and attention pooling operations. The environmental point cloud features include three-dimensional spatial position features and distance depth features of environmental elements, including: Performing coordinate system unification processing on the environmental point cloud information, converting the point cloud coordinates collected by different radar sensors into a unified vehicle coordinate system; Set the 3D voxel grid parameters to divide the environmental point cloud information in the unified coordinate system into multiple 3D voxel units, each of which contains point cloud data within a preset spatial range; Extracting features from the point cloud data within each three-dimensional voxel unit to generate a voxel feature vector, wherein the voxel feature vector includes the number of point clouds within the voxel, average coordinates, and reflection intensity statistical features; Arrange the voxel feature vectors of all three-dimensional voxel units according to spatial positions to construct a three-dimensional voxel feature matrix; Inputting the three-dimensional voxel feature matrix into the attention pooling layer of the point cloud feature extraction subnetwork to calculate the contribution weight of each three-dimensional voxel unit to the key environmental elements of driving safety; The three-dimensional voxel feature matrix is weightedly aggregated based on the contribution weights, and the voxel features after weighted aggregation are reduced in dimension through a convolution operation to generate environmental point cloud features containing three-dimensional spatial position features and distance depth features of environmental elements.
10. A safe driving environment perception system combined with artificial intelligence, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the safe driving environment perception method combined with artificial intelligence as described in any one of claims 1 to 9.
Citation Information
Cited By
Intelligent driving perception data processing method and device, electronic equipment and vehicle
CN121071438A