A drivable area identification method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202310504926.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-05-04
AI Technical Summary
[0004]然而,采用上述方法时,由于复杂的编码器的内存访问代价较大,因此图形处理器(graphics processing unit,GPU)的计算效率较低,限制了识别速度
[0051]本申请实施例中,在驾驶对象触发可行驶区域识别请求时,服务器响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像,再按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图,然后,对各道路特征图在空间位置维度和通道维度分别进行注意力加权处理,获得各注意力特征图,最后,基于目标道路图像的图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。
Smart Images

Figure CN116543369B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and in particular to a method, device, electronic device and storage medium for identifying drivable areas. Background Technology
[0002] Currently, detecting scene information during vehicle operation is crucial in driver assistance systems and autonomous driving systems. Scene information detection includes lane lines, traffic signs, road signs, and drivable areas. Among these, drivable area detection plays a vital role in driver assistance systems and autonomous driving systems. During vehicle operation, it provides drivable routes and avoids drivable areas. Drivable areas mainly include areas with static obstacles on the road, such as cones and guardrails, as well as areas with dynamic obstacles, such as vehicles and pedestrians.
[0003] In related technologies, when identifying drivable areas, a complex encoder is typically used to extract features from the acquired images, resulting in a four-fold downsampled feature map. For example, a high-resolution network (High-Resolution Net, HRNet) is used. The four-fold downsampled feature map is then enlarged to restore the output image to the size of the input image, thereby obtaining the drivable area.
[0004] However, when using the above method, the computational efficiency of the graphics processing unit (GPU) is low due to the high memory access cost of the complex encoder, which limits the recognition speed.
[0005] In addition, in order to improve the recognition speed, the image size is directly restored by magnification. However, since the pixels of the four-fold downsampled feature map are small, the object cannot be clearly observed. Therefore, if the four-fold downsampled feature map is directly magnified, it will cause inaccurate recognition.
[0006] For example, when a 4x downsampled feature map contains several pixels in a drivable area that are not recognized as drivable areas, after magnification, the pixels become larger, and the object can observe that a small area is not recognized as a drivable area, resulting in a hole problem.
[0007] Therefore, the accuracy and efficiency of drivable area identification in related technologies need to be improved. Summary of the Invention
[0008] This application provides a method, apparatus, electronic device, and storage medium for identifying drivable areas, in order to improve the accuracy and efficiency of drivable area identification results.
[0009] The specific technical solutions provided in this application are as follows:
[0010] Firstly, a method for identifying drivable areas is provided, including:
[0011] In response to a drivable area recognition request triggered by a driving object, the target road image to be identified is acquired, and the target road image has the target image scale;
[0012] According to the preset candidate image scales, feature extraction is performed on the target road image to obtain the corresponding road feature map;
[0013] For each road feature map, the following operations are performed: Based on the image position information of each pixel in a road feature map, and combined with the image channel information in a road feature map, attention weighting processing is performed on the road feature map to obtain the corresponding attention feature map;
[0014] Based on the target image scale, the obtained attention feature maps are scaled to obtain at least one target feature map, and the drivable area of the target vehicle is obtained based on at least one target feature map.
[0015] Secondly, a drivable area identification device is provided, comprising:
[0016] The acquisition module is used to acquire the target road image to be identified in response to the drivable area recognition request triggered by the driving object. The target road image has the target image scale.
[0017] The extraction module is used to extract features from the target road image according to the preset candidate image scales to obtain the corresponding road feature map;
[0018] The first processing module is used to perform the following operations on each road feature map: based on the image position information of each pixel contained in a road feature map and combined with the image channel information contained in a road feature map, perform attention weighting processing on the road feature map to obtain the corresponding attention feature map;
[0019] The second processing module is used to adjust the image scale of each attention feature map obtained based on the target image scale to obtain at least one target feature map, and to obtain the drivable area of the target vehicle based on at least one target feature map.
[0020] Optionally, when extracting features from the target road image according to preset candidate image scales to obtain the corresponding road feature maps, the extraction module is used for:
[0021] For each preset candidate image scale, perform the following operations respectively:
[0022] Based on a candidate image scale, a first feature is extracted from the target road image to obtain a first feature map, which has a first number of image channels.
[0023] According to the preset receptive field of each candidate image, the second feature is extracted from the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, the receptive field of each candidate image represents: the size of the area on the target road image mapped by each pixel contained in the corresponding second feature map;
[0024] Based on the first number of image channels, the superimposed feature map is adjusted to obtain a road feature map corresponding to a candidate image scale.
[0025] Optionally, based on the image position information of each pixel in a road feature map and combined with the image channel information in the road feature map, when performing attention weighting processing on the road feature map to obtain the corresponding attention feature map, the first processing module is used to:
[0026] Based on the image location information of each pixel in a road feature map, the position attention weight of each pixel in the road feature map is determined. The position attention weight represents the importance of the image location of the corresponding pixel when identifying drivable areas.
[0027] Based on the image channel information contained in a road feature map, the channel attention weight of each image channel contained in the road feature map is determined. The channel attention weight represents the importance of the corresponding image channel when identifying drivable areas.
[0028] Based on the obtained attention weights for each location and channel, a road feature map is weighted to obtain the attention feature map corresponding to the road feature map.
[0029] Optionally, based on the target image scale, the obtained attention feature maps are scaled to obtain at least one target feature map. The second processing module is used for:
[0030] Based on the image scale of each attention feature map, the attention feature maps are fused and updated to obtain at least one fused feature map.
[0031] Based on the preset number of target image channels, image channel adjustment is performed on at least one fused feature map to obtain the corresponding intermediate feature map.
[0032] Based on the target image scale, at least one intermediate feature map is scaled to obtain the corresponding target feature map.
[0033] Optionally, based on the image scale of each obtained attention feature map, the second processing module further performs fusion and updating on each attention feature map to obtain at least one fused feature map, and then further performs the following:
[0034] The attention feature maps are sorted according to their respective image scales to obtain the sorting results.
[0035] Based on the ranking results, each pair of adjacent attention feature maps is sequentially read and fused until all maps are read. Each fusion process includes:
[0036] Read two adjacent attention feature maps and fuse them to obtain a fused feature map;
[0037] Save the fused feature map and use it as a new attention feature map to replace the two attention feature maps in the ranking results.
[0038] Optionally, when obtaining the drivable area of the target vehicle based on at least one target feature map, the second processing module is further configured to:
[0039] If at least one target feature map includes a target feature map, then the target feature map is used as the result feature map;
[0040] If at least one target feature map includes multiple target feature maps, then image fusion is performed on the multiple target feature maps to obtain the result feature map;
[0041] The values of each pixel in the resulting feature map are compared with the preset drivable area label values to determine the target pixels in the resulting feature map that belong to the drivable area.
[0042] In the target road image, each pixel corresponding to its respective image position is marked to obtain an image containing the drivable area.
[0043] Optionally, the drivable area is obtained by inputting a target road image into a target recognition model. The device also includes a training module, which is used for:
[0044] The target recognition model is iteratively trained based on the training sample set to obtain the target recognition model; each training sample includes: image data of the sample road, wherein each iteration process performs the following operations:
[0045] According to the preset sample image scale, feature extraction is performed on the selected training samples to obtain the corresponding sample feature maps. The training samples have the sample image scale.
[0046] For each sample feature map, the following operations are performed: Based on the sample image position information of each pixel contained in a sample feature map, and combined with the sample image channel information contained in a sample feature map, attention weighting processing is performed on the sample feature map to obtain the corresponding sample attention feature map.
[0047] Based on the sample image scale, the image scale of each obtained sample attention feature map is adjusted to obtain at least one target sample feature map.
[0048] Based on the feature map of at least one target sample, the drivable area prediction result of the training sample is obtained, and the parameters are tuned based on the loss value corresponding to the drivable area prediction result.
[0049] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in any of the first aspects above.
[0050] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the first aspects above.
[0051] In this embodiment, when a driver triggers a drivable area recognition request, the server responds to the request by acquiring the target road image to be recognized, then extracts features from the target road image according to preset candidate image scales to obtain corresponding road feature maps. Next, attention weighting is applied to each road feature map in both the spatial location and channel dimensions to obtain attention feature maps. Finally, based on the image scale of the target road image, the obtained attention feature maps are scaled to obtain at least one target feature map, and the drivable area of the target vehicle is obtained based on at least one target feature map.
[0052] In this way, by obtaining multi-scale road feature maps, weighting each road feature map in terms of spatial location and channel dimensions, and then fusing the weighted road feature maps, the spatial and semantic information of the multi-scale road feature maps can be better balanced, avoiding the hole problem, improving the accuracy and efficiency of drivable area identification, and ensuring that the target feature map and the target road image have the same image scale, which can more accurately obtain the drivable area of the target vehicle. Attached Figure Description
[0053] Figure 1 This is a schematic diagram illustrating possible application scenarios in the embodiments of this application;
[0054] Figure 2This is a schematic diagram of the training process of the target recognition model in one round in an embodiment of this application;
[0055] Figure 3 This is a schematic diagram of the encoder in an embodiment of this application;
[0056] Figure 4 This is a schematic diagram of the high-performance GPU module layer in an embodiment of this application;
[0057] Figure 5 This is a schematic diagram of the spatial attention module and the channel attention module in the embodiments of this application;
[0058] Figure 6 This is a flowchart illustrating the drivable area identification method in the embodiments of this application;
[0059] Figure 7 This is a schematic diagram illustrating the process of obtaining the road feature map in an embodiment of this application;
[0060] Figure 8 This is a schematic diagram illustrating the obtained superimposed feature map in an embodiment of this application;
[0061] Figure 9 This is a schematic diagram of the process for obtaining an attention feature map corresponding to a road feature map in an embodiment of this application;
[0062] Figure 10 This is a schematic diagram of obtaining the attention feature map in an embodiment of this application;
[0063] Figure 11 This is a schematic diagram illustrating the process of obtaining at least one target feature map according to an embodiment of this application;
[0064] Figure 12 This is a schematic diagram illustrating the process of obtaining at least one fused feature map according to an embodiment of this application;
[0065] Figure 13 This is a schematic diagram illustrating the acquisition of at least one fused feature map in an embodiment of this application;
[0066] Figure 14 This is a schematic diagram illustrating the acquisition of at least one target feature map in an embodiment of this application;
[0067] Figure 15 This is a schematic diagram illustrating the process of obtaining the drivable area of the target vehicle in an embodiment of this application;
[0068] Figure 16 This is a schematic diagram of an image containing a drivable area in an embodiment of this application;
[0069] Figure 17 This is a schematic diagram of the drivable area identification device in the embodiments of this application;
[0070] Figure 18 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation
[0071] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0072] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:
[0073] Driving area: Without considering traffic rules, all road areas where vehicles are allowed to drive normally, excluding all obstacles (including guardrails, vehicles, pedestrians and other obstacles) on the road where vehicles travel.
[0074] Loss function: In machine learning, a function used to measure and optimize the distance between model predictions and actual sample labeled values.
[0075] The design concept of the embodiments of this application is briefly introduced below:
[0076] Currently, detecting scene information during vehicle operation is crucial in driver assistance systems and autonomous driving systems. Scene information detection includes lane lines, traffic signs, road signs, and drivable areas. Among these, drivable area detection plays a vital role in driver assistance systems and autonomous driving systems. During vehicle operation, it provides drivable routes and avoids drivable areas. Drivable areas mainly include areas with static obstacles on the road, such as cones and guardrails, as well as areas with dynamic obstacles, such as vehicles and pedestrians.
[0077] Currently, commonly used methods for identifying drivable areas include:
[0078] Method 1: Manually design features to obtain drivable areas, such as traditional thresholding, clustering, support vector machines, decision trees, etc.
[0079] When using Method 1, due to the need for manual design of features, it is limited by professional theories and prior knowledge, making it difficult to represent the variability of complex traffic environments and the diversity of road structures. It can only be applied in limited specific environments, which has limitations and low efficiency.
[0080] Method 2: For the acquired images, use a complex encoder to extract features and obtain a four-fold downsampled feature map. For example, a high-resolution network (High-Resolution Net, HRNet) can be used. Then, the four-fold downsampled feature map is enlarged so that the output image is restored to the size of the input image, and the drivable area is obtained.
[0081] When using method two, the computational efficiency of the graphics processing unit (GPU) is low due to the high memory access cost of the complex encoder, which limits the recognition speed.
[0082] In addition, in order to improve the recognition speed, the image size is directly restored by magnification. However, since the pixels of the four-fold downsampled feature map are small, the object cannot be clearly observed. Therefore, if the four-fold downsampled feature map is directly magnified, it will cause inaccurate recognition.
[0083] For example, when a 4x downsampled feature map contains several pixels in a drivable area that are not recognized as drivable areas, after magnification, the pixels become larger, and the object can observe that a small area is not recognized as a drivable area, resulting in a hole problem.
[0084] In view of this, this application proposes a drivable area identification method, apparatus, electronic device, and storage medium. When a driving object triggers a drivable area identification request, a target road image to be identified can be acquired. Then, according to preset candidate image scales, features are extracted from the target road image to obtain corresponding road feature maps. For each road feature map, the following operations are performed: based on the image position information of each pixel in a road feature map, combined with the image channel information in another road feature map, attention-weighted processing is applied to the road feature map to obtain a corresponding attention feature map. Finally, based on the target image scale, the obtained attention feature maps are scaled to obtain at least one target feature map. Based on at least one target feature map, the drivable area of the target vehicle is obtained. This weighted processing and fusion operation on road feature maps of different scales ensures that the target feature map and the target road image have consistent sizes, better balancing the spatial and semantic information of multi-scale road feature maps, avoiding the hole problem, and improving the accuracy and efficiency of drivable area identification.
[0085] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.
[0086] See Figure 1 The diagram shown is a possible application scenario illustration in the embodiments of this application. The application scenario diagram includes a server 110 and terminal devices 120 (including terminal devices 1201, 1202, ..., 120n).
[0087] Server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal device 120 and server 110 can be connected directly or indirectly via wired or wireless communication; this application does not impose any restrictions on this.
[0088] Terminal device 120 can be a mobile phone, portable computer, or other device carried by the driver during the driving of the target vehicle. Terminal device can also be a computer device with certain computing capabilities, such as an in-vehicle terminal installed in the target vehicle.
[0089] In this embodiment of the application, either or both of the server 110 and the terminal device 120 may be configured with a trained target recognition model, so that the trained target recognition model can be used to identify the drivable area of the target road image to be identified.
[0090] It should be noted that, in the embodiments of this application, the recognition model used on the device may be trained by itself or by other devices.
[0091] Taking the example of server 110 using a trained recognition model to perform drivable area recognition operation on the target road image to be recognized, the recognition model used by server 110 can be trained by itself, or it can be a recognition model trained by other devices and then directly sent to server 110.
[0092] The following explanation will use the training of the target recognition model implemented on the server as an example to illustrate the relevant training process in detail.
[0093] In addition, in this embodiment of the application, the server training of the target recognition model can be a periodic process according to the actual processing needs. Training samples can be regenerated periodically to train the target recognition model.
[0094] Based on the training sample set, the recognition model to be trained is iteratively trained to obtain the target recognition model; each training sample includes: image data of the sample road, see [link / reference] Figure 2 The diagram shown illustrates the process of one round of training for the target recognition model in this embodiment of the application, specifically including:
[0095] Step 20: Extract features from the selected training samples according to the preset sample image scale to obtain the corresponding sample feature maps.
[0096] The training samples have a sample image scale.
[0097] In this embodiment, an encoder is used to extract features from the training samples. According to the preset image scale of each sample, features are extracted from the selected training samples to obtain the corresponding sample feature maps.
[0098] The preset sample image scales can be 1 / 2, 1 / 4, 1 / 8, or 1 / 16, and this embodiment does not impose any restrictions on them.
[0099] For example, see Figure 3 The diagram shown is a schematic of the encoder in an embodiment of this application. The encoder mainly consists of a network stem layer and four stage layers. The network stem layer consists of a 3*3 convolutional layer, a batch normalization layer, and a FReLU layer to obtain a first sample feature map with an image size of 1 / 2. The stage layers consist of a high-performance GPU module (HG block) layer and a downsampling (LDS layer) layer. The downsampling layer is a 2*2 max pooling layer with a stride of 2 to obtain first sample feature maps with image sizes of 1 / 4, 1 / 8, and 1 / 16, respectively. Based on each first sample feature map, the HG block layer is used to obtain sample feature maps with image sizes of 1 / 2, 1 / 4, 1 / 8, and 1 / 16, respectively.
[0100] Specifically, in the HG Block layer, for the first sample feature map corresponding to each sample image scale, the following operations are performed respectively: according to the preset receptive field of each sample image, the second feature is extracted from the first sample feature map corresponding to the sample image scale to obtain the corresponding second sample feature map, and the obtained second sample feature maps are superimposed to obtain the superimposed sample feature map. Based on the number of image channels of the first sample feature map, the image channels of the superimposed feature map are adjusted to obtain the sample feature map corresponding to the sample image scale.
[0101] The receptive field of each sample image represents the size of the region on the training sample that each pixel in the corresponding second sample feature map is mapped to. The receptive field of each sample image can be 1x1, 3x3, 5x5, 7x7, 9x9, or 11x11. This embodiment does not impose any restrictions on this.
[0102] For example, see Figure 4 The diagram shown is a schematic of the high-performance GPU module layer in this embodiment. The first sample feature map is first passed through a series of 3x3 convolutional layers to obtain second sample feature maps with different receptive fields, including: second sample feature maps with a receptive field of 3x3, second sample feature maps with a receptive field of 5x5, second sample feature maps with a receptive field of 7x7, second sample feature maps with a receptive field of 9x9, and second sample feature maps with a receptive field of 11x11. The obtained second sample feature maps are then superimposed to obtain a superimposed sample feature map. A 1x1 convolution is used to adjust the image channels of the superimposed sample feature map to obtain a sample feature map, so that the number of image channels of the sample feature map is the same as the number of image channels of the first sample feature map.
[0103] Thus, since the impact of memory access cost on computational efficiency is far greater than the computational cost and parameter count of the model, especially for the convolutional layers used extensively in the model, the memory access cost is minimized only when the number of input and output image channels is equal. Therefore, by using high-performance GPU module layers and designing that the number of input and output image channels is the same, the memory access cost can be minimized. This improves computational efficiency while enhancing the network's multi-feature extraction capability through layered convolutional overlap.
[0104] Optionally, residual connections and squeeze stimuli can be added to the high-performance GPU module layer included in the final stage layer, such as... Figure 4 As shown, the dashed lines represent residual connections. While residual connections excel at accelerating model convergence and preventing gradient vanishing, the element-wise addition operations they introduce are computationally intensive on GPUs, impacting model speed. Squeezing excitation, a channel attention mechanism, effectively improves model accuracy but also introduces significant latency. To balance accuracy and speed, the encoder uses residual connections and squeezing excitation only in the final stage layer.
[0105] Step 21: For each sample feature map, perform the following operations: Based on the sample image position information of each pixel contained in a sample feature map, and combined with the sample image channel information contained in a sample feature map, perform attention weighting processing on the sample feature map to obtain the corresponding sample attention feature map.
[0106] In this embodiment of the application, the following operations are performed for each sample feature map: using a spatial attention module, based on the sample image position information of each pixel in a sample feature map, the sample position attention weight of each pixel in the sample feature map is determined; using a channel attention module, based on the sample image channel information in the sample feature map, the sample channel attention weight of each image channel in the sample feature map is determined; and based on the obtained sample position attention weight and sample channel attention weight, the sample feature map is weighted to obtain the sample attention feature map corresponding to the sample feature map.
[0107] Among them, the sample position attention weight represents the importance of the sample image position of the corresponding pixel when identifying the drivable area, and the sample channel attention weight represents the importance of the corresponding sample image channel when identifying the drivable area.
[0108] For details, please refer to Figure 5 The diagram illustrates the spatial attention module and channel attention module in this embodiment. The spatial attention module performs mean pooling and max pooling operations on the sample feature map along the spatial position dimension to obtain a spatial mean pooling sample feature map and a spatial max pooling sample feature map. Then, the spatial mean pooling sample feature map and the spatial max pooling sample feature map are concatenated to obtain a spatial pooling concatenated sample feature map. Finally, a convolutional layer and a sigmoid function are applied to activate the concatenated feature map to obtain the sample position attention weights for each sample pixel in the sample feature map. The channel attention module performs mean pooling and max pooling operations on the sample feature map along the channel dimension to obtain a channel mean pooling sample feature map and a channel max pooling sample feature map. Then, the channel mean pooling sample feature map and the channel max pooling sample feature map are concatenated to obtain a channel pooling concatenated feature map. Finally, a convolutional layer and a sigmoid function are applied to activate the concatenated feature map to obtain the sample channel attention weights for each sample image channel in the sample feature map.
[0109] By employing spatial attention and channel attention modules to obtain attention feature maps for each sample, the recognition model can focus more on important information in the sample feature maps, suppress useless information, and improve the accuracy of the recognition model.
[0110] Step 22: Based on the sample image scale, adjust the image scale of each obtained sample attention feature map to obtain at least one target sample feature map.
[0111] In this embodiment, based on the image scale of each obtained sample attention feature map, the sample attention feature maps are fused and updated to obtain at least one sample fusion feature map. Based on the preset number of sample image channels, the image channels of the at least one sample fusion feature map are adjusted to obtain the corresponding sample intermediate feature map. Based on the sample image scale, the image scale of the at least one sample intermediate feature map is adjusted to obtain the corresponding sample target feature map.
[0112] For example, using a convolutional network, based on a preset number of sample image channels, image channel adjustment is performed on at least one sample fusion feature map to obtain the corresponding sample intermediate feature map. Using bilinear interpolation, based on the sample image scale, image scale adjustment is performed on the obtained at least one sample intermediate feature map to obtain the corresponding sample target feature map.
[0113] In this way, at least one sample target feature map with the same number of sample image channels and sample image scale is obtained, which integrates the differences in channel and scale features of different sample attention feature maps and improves the accuracy of the target recognition model.
[0114] Specifically, there are multiple methods for fusing and updating the attention feature maps of each sample.
[0115] The first fusion update method is as follows: according to the image scale of each sample attention feature map, the sample attention feature maps are sorted to obtain the sample sorting result. Based on the sample sorting result, every two adjacent sample attention feature maps are read and fused in turn until all are read. One fusion process includes: reading two adjacent sample attention feature maps, fusing the two sample attention feature maps to obtain a sample fusion feature map, saving the sample fusion feature map, and adding the sample fusion feature map as a new sample attention feature map to replace the two sample attention feature maps in the sample sorting result.
[0116] For example, suppose each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, and sample attention feature map A4. The image scale of sample attention feature map A1 is 1 / 2, the image scale of sample attention feature map A2 is 1 / 4, the image scale of sample attention feature map A3 is 1 / 8, and the image scale of sample attention feature map A4 is 1 / 16. Then, sorting the sample attention feature maps according to their image scale from smallest to largest, the resulting sample sorting result is A1, A2, A3, and A4. Reading A1 and A2, and fusing A1 and A2, we obtain the sample attention feature map. This fusion feature map B1 is saved, and B1 is used as a new sample attention feature map, replacing A1 and A2, and added to the sample ranking result. At this time, the sample ranking result is B1, A3, and A4. B1 and A3 are read and fused to obtain the sample fusion feature map B2. B2 is saved, and B2 is used as a new sample attention feature map, replacing B1 and A3, and added to the sample ranking result. At this time, the sample ranking result is B2 and A4. B2 and A4 are read and fused to obtain the sample fusion feature map B3. B3 is saved. Finally, at least one sample fusion feature map is obtained, including B1, B2, and B3.
[0117] Furthermore, before fusing and updating the sample attention feature maps corresponding to different sample image scales, the sample feature maps corresponding to the small sample image scale are upsampled to the same sample image scale as the sample feature maps corresponding to the large sample image scale, and then the sample attention feature maps corresponding to the small sample image scale are obtained through the spatial attention module and the channel attention module.
[0118] The second fusion update method is as follows: according to the image scale of each sample attention feature map, sort the sample attention feature maps to obtain the sample sorting result, and based on the sample sorting result, sequentially read every two adjacent sample attention feature maps and fuse them until all are read. One fusion process includes: reading two adjacent sample attention feature maps, fusing the two sample attention feature maps to obtain the sample fusion feature map, and saving the sample fusion feature map.
[0119] For example, suppose each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, and sample attention feature map A4. The image scale of sample attention feature map A1 is 1 / 2, the image scale of sample attention feature map A2 is 1 / 4, the image scale of sample attention feature map A3 is 1 / 8, and the image scale of sample attention feature map A4 is 1 / 16. Then, sort the sample attention feature maps according to the image scale from smallest to largest, and the resulting sample sorting result is A1, A2, A3, and A4. Read A1 and A2, and fuse A1 and A2 to obtain sample fused feature map B4. Save B4. Read A2 and A3, and fuse A2 and A3 to obtain sample fused feature map B5. Save B5. Read A3 and A4, and fuse A3 and A4 to obtain sample fused feature map B6. Save B6. Finally, at least one sample fused feature map is obtained, including B4, B5, and B6.
[0120] The third fusion update method is to fuse the attention feature maps of each sample to obtain a sample fusion feature map.
[0121] For example, assuming that each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, and sample attention feature map A4, then A1, A2, A3, and A4 are fused to obtain a sample fusion feature map B7.
[0122] Step 23: Based on the feature map of at least one target sample, obtain the drivable area prediction result of the training sample, and perform parameter tuning based on the loss value corresponding to the drivable area prediction result.
[0123] In this embodiment, a sample result feature map is obtained based on at least one target sample feature map. The values of each sample pixel in the sample result feature map are compared with preset sample drivable area label values to determine each sample target pixel in the sample result feature map that belongs to the drivable area. Each sample pixel in the training sample corresponding to the image position of each sample target pixel is labeled to obtain the drivable area prediction result of the training sample. Based on the drivable area prediction result, the real label corresponding to the training sample and the weight of easy and difficult samples, the loss value is calculated, and the network parameters of the recognition model are adjusted based on the loss value.
[0124] In this embodiment of the application, obtaining the feature map of the sample result includes, but is not limited to, the following two cases:
[0125] Case 1: If at least one target sample feature map is included, use that target sample feature map as the sample result feature map.
[0126] Case 2: At least one target sample feature map includes multiple target sample feature maps. Image fusion is performed on the multiple target sample feature maps to obtain the sample result feature map.
[0127] The preset drivable area label value can be 1, and this embodiment does not impose any limitation on this. The objective function for model training is loss = 0.9 * L. FocalLoss +0.1*L LovaszLoss L FocalLoss To apply corresponding weighted loss based on the difficulty of identification—that is, assigning smaller weights to easily identifiable samples and larger weights to difficult-to-identify samples—L... LovaszLoss This is a convex Lovasz extension based on submodulus loss, which optimizes the mean intersection over union (Mean IoU) ratio for semantic segmentation tasks.
[0128] For example, the sigmoid function is used to activate the feature map of the sample result. The values of each sample pixel in the feature map of the sample result are compared with the preset sample drivable area label value 1 to determine each sample target pixel in the feature map of the sample result that belongs to the drivable area. The sample pixels in the training sample that correspond to the image positions of each sample target pixel are marked to obtain the drivable area prediction results of the training sample.
[0129] The following description, with reference to the accompanying diagram, illustrates the process of applying the trained target recognition model:
[0130] See Figure 6 As shown, it is a flowchart illustrating the drivable area identification method in an embodiment of this application. The following is a detailed explanation in conjunction with the attached diagram. Figure 6 The specific operation will be explained in detail:
[0131] Step 60: In response to the drivable area recognition request triggered by the driving object, acquire the current target road image to be identified.
[0132] The target road image has the target image scale.
[0133] In this embodiment of the application, after the driver performs certain operations on the terminal device, the terminal device can be triggered to send a drivable area identification request to the server. In response to the drivable area identification request triggered by the driver, the server calls the camera to obtain the image of the target road to be identified.
[0134] Step 61: Extract features from the target road image according to the preset candidate image scales to obtain the corresponding road feature maps.
[0135] In this embodiment of the application, after obtaining the target road image, features are extracted from the target road image according to each preset candidate image scale to obtain the road feature map corresponding to each candidate image scale.
[0136] The preset candidate image scales can be 1 / 2, 1 / 4, 1 / 8, or 1 / 16, and this embodiment does not impose any restrictions on them.
[0137] Specifically, when obtaining a road feature map corresponding to a candidate image scale, the server performs the following operations. See also... Figure 7 As shown, it is a schematic diagram of the process of obtaining the road feature map in an embodiment of this application. The following is in conjunction with the attached diagram. Figure 7 The specific operations to be performed will be explained in detail:
[0138] Step 610: Based on a candidate image scale, perform first feature extraction on the target road image to obtain a first feature map.
[0139] The first feature map has a first number of image channels.
[0140] In this embodiment of the application, based on a candidate image scale, a corresponding downsampling operation is used to extract the first feature from the target road image to obtain a first feature map, and the image scale of the first feature map is the candidate image scale.
[0141] For example, assuming a candidate image has a scale of 1 / 2, the first feature extraction is performed on the target road image to obtain a first feature map with an image scale of 1 / 2.
[0142] Step 611: According to the preset receptive field of each candidate image, perform second feature extraction on the first feature map to obtain the corresponding second feature map, and superimpose the obtained second feature maps to obtain the superimposed feature map.
[0143] The receptive field of each candidate image represents the size of the region on the target road image mapped by each pixel in the corresponding second feature map. The receptive field of each candidate image can be 1x1, 3x3, 5x5, 7x7, 9x9, or 11x11. This embodiment does not impose any restrictions on this.
[0144] In this embodiment, after obtaining the first feature map, the following operations are performed for each preset receptive field of a candidate image: based on the receptive field of a candidate image, a second feature is extracted from the first feature map to obtain a corresponding second feature map. The obtained second feature maps are then superimposed to obtain a superimposed feature map.
[0145] For example, see Figure 8The diagram shown is a schematic of obtaining the superimposed feature map in an embodiment of this application. Assuming that the receptive fields of each candidate image are 1x1, 3x3, 5x5, 7x7, 9x9, and 11x11, a series of 3x3 convolutional layers are used to obtain second feature maps with different receptive fields, including: a second feature map with a receptive field of 3x3, a second feature map with a receptive field of 5x5, a second feature map with a receptive field of 7x7, a second feature map with a receptive field of 9x9, and a second feature map with a receptive field of 11x11. The obtained second sample feature maps are then superimposed to obtain the superimposed feature map.
[0146] Step 612: Based on the first image channel number, adjust the image channels of the superimposed feature map to obtain a road feature map corresponding to a candidate image scale.
[0147] In this embodiment of the application, after obtaining the superimposed feature map, the superimposed feature map is adjusted by 1*1 convolution based on the first number of image channels to obtain the road feature map, where the number of image channels of the road feature map is the first number of image channels.
[0148] Since the impact of memory access cost on computational efficiency is much greater than the computational cost and parameter count of the model, especially for convolution operations, the memory access cost is minimized only when the number of input and output image channels is equal. Therefore, adopting a design where the number of image channels of the road feature map and the first feature map is the same can minimize the memory access cost. While improving computational efficiency, multiple features of the target road image can be obtained through layered convolution stacking.
[0149] Optionally, based on the first image channel number, when adjusting the superimposed feature map corresponding to the first feature map with the smallest image scale to obtain the road feature map corresponding to the smallest candidate image scale, residual connections and squeezing excitations are added, such as... Figure 8 As shown, the dashed lines represent residual connections. While residual connections excel at avoiding gradient vanishing, the element-wise addition operations they introduce are computationally intensive on GPUs, impacting the speed of drivable region recognition. Squeeze excitation, a channel attention mechanism, effectively improves the accuracy of drivable region recognition, but also introduces significant latency. To balance accuracy and speed in drivable region recognition, residual connections and squeeze excitation are used only when obtaining the road feature map corresponding to the smallest candidate image scale.
[0150] Furthermore, it is worth noting that in this embodiment, a first feature map at a scale of 1 / 2 can be obtained firstly through a 3*3 convolutional layer, a batch normalization layer, and a FReLU layer, and a corresponding road feature map can be obtained. Then, a 2*2 max pooling layer with a stride of 2 is used to extract features from the road feature map corresponding to the first feature map at a scale of 1 / 2, obtaining a first feature map at a scale of 1 / 4, and a corresponding road feature map. Then, a 2*2 max pooling layer with a stride of 2 is used to extract features from the road feature map corresponding to the first feature map at a scale of 1 / 4, obtaining a first feature map at a scale of 1 / 8, and a corresponding road feature map. Finally, a 2*2 max pooling layer with a stride of 2 is used to extract features from the road feature map corresponding to the first feature map at a scale of 1 / 8, obtaining a first feature map at a scale of 1 / 16, and a corresponding road feature map.
[0151] Step 62: For each road feature map, perform the following operations: Based on the image position information of each pixel in a road feature map, and combined with the image channel information in a road feature map, perform attention weighting processing on the road feature map to obtain the corresponding attention feature map.
[0152] Specifically, during step 62, the server performs the following operations. See also... Figure 9 As shown, this is a schematic diagram of the process for obtaining an attention feature map corresponding to a road feature map in an embodiment of this application. The following is a detailed explanation in conjunction with the attached diagram. Figure 9 The specific operations to be performed will be explained in detail:
[0153] Step 620: Based on the image location information of each pixel in a road feature map, determine the positional attention weight of each pixel in the road feature map.
[0154] Among them, the position attention weight represents the importance of the image position of the corresponding pixel when identifying drivable areas.
[0155] In this embodiment, based on the image location information of each pixel in a road feature map, average pooling and max pooling operations are performed on the road feature map in the spatial location dimension to obtain a spatial average pooling feature map and a spatial max pooling feature map. Then, the spatial average pooling feature map and the spatial max pooling feature map are concatenated to obtain a spatial pooling concatenated feature map. Finally, the location attention weights of each pixel in the road feature map are obtained through convolutional layers and sigmoid function activation.
[0156] Step 621: Based on the image channel information contained in a road feature map, determine the channel attention weights of each image channel contained in the road feature map.
[0157] Among them, the channel attention weight represents the importance of the corresponding image channel when identifying drivable areas.
[0158] In this embodiment, based on the image channel information contained in a road feature map, average pooling and max pooling operations are performed on the road feature map in the channel dimension to obtain channel average pooling feature maps and channel max pooling feature maps. Then, the channel average pooling feature maps and channel max pooling feature maps are concatenated to obtain channel pooling concatenated feature maps. Finally, the channel attention weights of each image channel contained in the road feature map are obtained through convolutional layers and sigmoid function activation.
[0159] Step 622: Based on the obtained attention weights at each location and the attention weights at each channel, a road feature map is weighted to obtain the attention feature map corresponding to the road feature map.
[0160] In this embodiment of the application, after obtaining the attention weights of each location and the attention weights of each channel, a road feature map is multiplied by its attention weights of each location in the spatial location dimension to obtain a location attention feature map. In the channel dimension, the road feature map is multiplied by its attention weights of each channel to obtain a channel attention feature map. Finally, the location attention feature map and the channel attention feature map are added together to obtain the attention feature map corresponding to the road feature map.
[0161] For example, see Figure 10 The diagram shown is a schematic of obtaining the attention feature map in an embodiment of this application, assuming road features. Figure 1 If the location attention weight matrix is A and the channel attention weight matrix is B, then the road features will be... Figure 1 Multiplying the position attention weight matrix A by the position attention weight matrix A yields the position attention features. Figure 1 Road features Figure 1 Multiplying the channel attention weight matrix B yields the channel attention features. Figure 1 Finally, the location attention features Figure 1 and channel attention features Figure 1 Add them together to obtain road features. Figure 1 Corresponding attention features Figure 1 .
[0162] In this way, based on the image position information of each pixel in a road feature map and the image channel information contained in a road feature map, attention weighting processing is performed on the road feature map to obtain the corresponding attention feature map. This can focus on important information in the road feature map, suppress useless information, and improve the accuracy of drivable area identification.
[0163] Step 63: Based on the target image scale, adjust the image scale of each attention feature map to obtain at least one target feature map, and obtain the drivable area of the target vehicle based on at least one target feature map.
[0164] In this embodiment of the application, when adjusting the image scale of each attention feature map based on the target image scale to obtain at least one target feature map, the server specifically performs the following operations. See also... Figure 11 As shown, it is a schematic diagram of the process of obtaining at least one target feature map according to an embodiment of this application. The following is in conjunction with the attached diagram. Figure 11 The specific operations to be performed will be explained in detail:
[0165] Step 630: Based on the image scale of each attention feature map, merge and update each attention feature map to obtain at least one fused feature map.
[0166] Specifically, there are several methods for fusing and updating the attention feature maps.
[0167] First method: See reference Figure 12 As shown, this is a schematic diagram of the process for obtaining at least one fused feature map according to an embodiment of this application. The following is a detailed explanation in conjunction with the attached diagram. Figure 12 The specific operations to be performed will be explained in detail:
[0168] Step 6300: Sort the attention feature maps according to their respective image scales to obtain the sorting results.
[0169] In this embodiment of the application, the attention feature maps are sorted in ascending order of image scale according to their respective image scales to obtain the sorting result.
[0170] For example, see Figure 13The diagram shown is a schematic of obtaining at least one fusion feature map in an embodiment of this application. Assuming that each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, and attention feature map P4, and the image scale size of attention feature map P1 is 1 / 2, the image scale size of attention feature map P2 is 1 / 4, the image scale size of attention feature map P3 is 1 / 8, and the image scale size of attention feature map P4 is 1 / 16, then the attention feature maps are sorted according to the image scale from smallest to largest, and the sorting result is P1, P2, P3, and P4.
[0171] Step 6301: Based on the ranking results, sequentially read every two adjacent attention feature maps and fuse them until all are read.
[0172] One fusion process includes: reading two adjacent attention feature maps and fusing them to obtain a fused feature map; saving the fused feature map; and adding the fused feature map as a new attention feature map, replacing the two existing attention feature maps, to the ranking result.
[0173] For example, such as Figure 13 As shown, the sample ranking results are P1, P2, P3, and P4. Read P1 and P2, and fuse them to obtain a fused feature map R1. Save R1, and use R1 as a new attention feature map, replacing P1 and P2, and add it to the ranking results. The ranking results are now R1, P3, and P4. Read R1 and P3, and fuse them to obtain a fused feature map R2. Save R2, and use R2 as a new attention feature map, replacing R1 and P3, and add it to the ranking results. The ranking results are now R2 and P4. Read R2 and P4, and fuse them to obtain a fused feature map R3. Save R3. Finally, at least one fused feature map is obtained, including R1, R2, and R3.
[0174] Furthermore, before fusing and updating the attention feature maps corresponding to different image scales, the road feature map corresponding to the small image scale is upsampled to the same image scale as the road feature map corresponding to the large image scale. Then, based on the image position information of each pixel contained in the upsampled road feature map, combined with the image channel information contained in the upsampled road feature map, attention weighting processing is performed on the upsampled road feature map to obtain the attention feature map corresponding to the small image scale.
[0175] The second method is to sort the attention feature maps according to their respective image scales, obtain the sorting results, and then, based on the sorting results, sequentially read and fuse every two adjacent attention feature maps until all have been read.
[0176] One fusion process includes: reading the attention feature maps of two adjacent samples, fusing the two sample attention feature maps to obtain a sample fusion feature map, and saving the sample fusion feature map.
[0177] For example, suppose each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, and attention feature map P4. Sort each attention feature map according to the image scale from smallest to largest, and the sorted result is P1, P2, P3, and P4. Read P1 and P2, and fuse P1 and P2 to obtain fused feature map R4. Save R4. Read P2 and P3, and fuse P2 and P3 to obtain fused feature map R5. Save R5. Read P3 and P4, and fuse P3 and P4 to obtain fused feature map R6. Save R6. Finally, at least one fused feature map is obtained, including R4, R5, and R6.
[0178] The third method is to fuse the attention feature maps to obtain a fused feature map.
[0179] For example, assuming that each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, and attention feature map P4, then P1, P2, P3, and P4 are fused to obtain a fused feature map R7.
[0180] Step 631: Based on the preset number of target image channels, adjust the image channels of at least one fused feature map to obtain the corresponding intermediate feature map.
[0181] In this embodiment of the application, after obtaining at least one fused feature map, the following operations are performed on each of the at least one fused feature map: using a convolutional layer, based on a preset number of target image channels, the image channels of a fused feature map are adjusted to obtain the intermediate feature map corresponding to the fused feature map.
[0182] The preset number of target image channels can be the number of channels of the target road image, and this application embodiment does not limit this.
[0183] For example, see Figure 14 The diagram shown is a schematic of obtaining at least one target feature map in an embodiment of this application. Assuming that at least one fused feature map includes fused feature map R1, fused feature map R2 and fused feature map R3, and the preset number of target image channels is 3, a convolutional layer is used to adjust the image channels of fused feature map R1, fused feature map R2 and fused feature map R3 respectively to obtain intermediate feature map Z1, intermediate feature map Z2 and intermediate feature map Z3 with 3 image channels.
[0184] Step 632: Based on the target image scale, adjust the image scale of at least one intermediate feature map to obtain the corresponding target feature map.
[0185] In this embodiment of the application, after obtaining at least one intermediate feature map, the following operations are performed on each intermediate feature map: using bilinear interpolation, based on the target image scale, to adjust the image scale of an intermediate feature map to obtain the target feature map corresponding to the intermediate feature map.
[0186] For example, such as Figure 14 As shown, at least one intermediate feature map includes intermediate feature map Z1, intermediate feature map Z2 and intermediate feature map Z3. The target image scale is 640x480. Bilinear interpolation is used to adjust the image scale of intermediate feature map Z1, intermediate feature map Z2 and intermediate feature map Z3 respectively to obtain target feature map M1, target feature map M2 and target feature map M3 with image scale of 640x480.
[0187] In this way, at least one sample target feature map with the same number of image channels and image scale is obtained, which integrates the differences in channel and scale features of different attention feature maps and improves the accuracy of drivable area recognition.
[0188] In this embodiment of the application, when obtaining the drivable area of the target vehicle based on at least one target feature map, the server specifically performs the following operations. See also... Figure 15 As shown, this is a schematic diagram of the process for obtaining the drivable area of the target vehicle according to an embodiment of this application. The following is a description of the process in conjunction with the attached diagram. Figure 15 The specific operations to be performed will be explained in detail:
[0189] Step 633: Obtain the result feature map based on at least one target feature map.
[0190] In this embodiment of the application, obtaining the feature map of the sample result includes, but is not limited to, the following two cases:
[0191] Case 1: If at least one target feature map includes a target feature map, then the target feature map is used as the result feature map.
[0192] For example, if at least one target feature map includes only target feature map M7, then target feature map M7 is used as the result feature map.
[0193] Case 2: If at least one target feature map includes multiple target feature maps, then perform image fusion on the multiple target feature maps to obtain the result feature map.
[0194] For example, assuming that at least one target feature map includes target feature map M1, target feature map M2 and target feature map M3, then M1, M2 and M3 are fused to obtain the result feature map.
[0195] Step 634: Compare the values of each pixel in the result feature map with the preset drivable area label values to determine the target pixels in the result feature map that belong to the drivable area.
[0196] In this embodiment of the application, for each pixel in the result feature map, the following operations are performed respectively: the value of a pixel is compared with the preset drivable area label value. If the value of the pixel is the same as the drivable area label value, then the pixel is determined to be the target pixel.
[0197] The preset drivable area label value can be 1, and this application embodiment does not limit this.
[0198] For example, assuming the label value of the drivable area is 1, the values of each pixel in the resulting feature map are compared with the label value 1. If the pixels in the resulting feature map that are the same as the label value 1 are pixel 1, pixel 2, and pixel 3, then the target pixels in the resulting feature map that belong to the drivable area are determined to be pixel 1, pixel 2, and pixel 3.
[0199] Step 635: Mark the pixels in the target road image that correspond to the respective image positions of each target pixel to obtain an image containing the drivable area.
[0200] For example, see Figure 16 As shown, this is a schematic diagram of an image containing a drivable area in an embodiment of this application. After obtaining each target pixel, the pixels in the target road image corresponding to the respective image positions of each target pixel are marked to segment the drivable area, thereby obtaining an image containing the drivable area.
[0201] Based on the same inventive concept, this application also provides a drivable area identification device, see reference. Figure 17 The diagram shown is a structural schematic of the drivable area identification device in this embodiment of the application, specifically including:
[0202] The acquisition module 1701 is used to acquire the target road image to be identified in response to the drivable area recognition request triggered by the driving object. The target road image has the target image scale.
[0203] The extraction module 1702 is used to extract features from the target road image according to each preset candidate image scale to obtain the corresponding road feature map;
[0204] The first processing module 1703 is used to perform the following operations on each road feature map: based on the image position information of each pixel contained in a road feature map and combined with the image channel information contained in a road feature map, perform attention weighting processing on the road feature map to obtain the corresponding attention feature map.
[0205] The second processing module 1704 is used to adjust the image scale of each attention feature map obtained based on the target image scale to obtain at least one target feature map, and to obtain the drivable area of the target vehicle based on at least one target feature map.
[0206] Optionally, when extracting features from the target road image according to preset candidate image scales to obtain the corresponding road feature maps, the extraction module 1702 is used for:
[0207] For each preset candidate image scale, perform the following operations respectively:
[0208] Based on a candidate image scale, a first feature is extracted from the target road image to obtain a first feature map, which has a first number of image channels.
[0209] According to the preset receptive field of each candidate image, the second feature is extracted from the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, the receptive field of each candidate image represents: the size of the area on the target road image mapped by each pixel contained in the corresponding second feature map;
[0210] Based on the first number of image channels, the superimposed feature map is adjusted to obtain a road feature map corresponding to a candidate image scale.
[0211] Optionally, when performing attention-weighted processing on the road feature map based on the image position information of each pixel contained in the road feature map and combining the image channel information contained in the road feature map to obtain the corresponding attention feature map, the first processing module 1703 is used to:
[0212] Based on the image location information of each pixel in a road feature map, the position attention weight of each pixel in the road feature map is determined. The position attention weight represents the importance of the image location of the corresponding pixel when identifying drivable areas.
[0213] Based on the image channel information contained in a road feature map, the channel attention weight of each image channel contained in the road feature map is determined. The channel attention weight represents the importance of the corresponding image channel when identifying drivable areas.
[0214] Based on the obtained attention weights for each location and channel, a road feature map is weighted to obtain the attention feature map corresponding to the road feature map.
[0215] Optionally, based on the target image scale, the obtained attention feature maps are scaled to obtain at least one target feature map. The second processing module 1704 is used for:
[0216] Based on the image scale of each attention feature map, the attention feature maps are fused and updated to obtain at least one fused feature map.
[0217] Based on the preset number of target image channels, image channel adjustment is performed on at least one fused feature map to obtain the corresponding intermediate feature map.
[0218] Based on the target image scale, at least one intermediate feature map is scaled to obtain the corresponding target feature map.
[0219] Optionally, when fusing and updating the attention feature maps based on their respective image scales to obtain at least one fused feature map, the second processing module 1704 is further configured to:
[0220] The attention feature maps are sorted according to their respective image scales to obtain the sorting results.
[0221] Based on the ranking results, each pair of adjacent attention feature maps is sequentially read and fused until all maps are read. Each fusion process includes:
[0222] Read two adjacent attention feature maps and fuse them to obtain a fused feature map;
[0223] Save the fused feature map and use it as a new attention feature map to replace the two attention feature maps in the ranking results.
[0224] Optionally, when obtaining the drivable area of the target vehicle based on at least one target feature map, the second processing module 1704 is further configured to:
[0225] If at least one target feature map includes a target feature map, then the target feature map is used as the result feature map;
[0226] If at least one target feature map includes multiple target feature maps, then image fusion is performed on the multiple target feature maps to obtain the result feature map;
[0227] The values of each pixel in the resulting feature map are compared with the preset drivable area label values to determine the target pixels in the resulting feature map that belong to the drivable area.
[0228] In the target road image, each pixel corresponding to its respective image position is marked to obtain an image containing the drivable area.
[0229] Optionally, the drivable area is obtained by inputting a target road image into a target recognition model. The device also includes a training module 1705, which is used for:
[0230] The target recognition model is iteratively trained based on the training sample set to obtain the target recognition model; each training sample includes: image data of the sample road, wherein each iteration process performs the following operations:
[0231] According to the preset sample image scale, feature extraction is performed on the selected training samples to obtain the corresponding sample feature maps. The training samples have the sample image scale.
[0232] For each sample feature map, the following operations are performed: Based on the sample image position information of each pixel contained in a sample feature map, and combined with the sample image channel information contained in a sample feature map, attention weighting processing is performed on the sample feature map to obtain the corresponding sample attention feature map.
[0233] Based on the sample image scale, the image scale of each obtained sample attention feature map is adjusted to obtain at least one target sample feature map.
[0234] Based on the feature map of at least one target sample, the drivable area prediction result of the training sample is obtained, and the parameters are tuned based on the loss value corresponding to the drivable area prediction result.
[0235] Based on the above embodiments, see Figure 18 The diagram shown is a structural schematic of the electronic device in an embodiment of this application.
[0236] This application provides an electronic device that may include a processor 1810 (Center Processing Unit, CPU), a memory 1820, an input device 1830, and an output device 1840. The input device 1830 may include a keyboard, a mouse, a touch screen, etc., and the output device 1840 may include a display device, such as a liquid crystal display (LCD) or a cathode ray tube (CRT).
[0237] The memory 1820 may include a read-only memory (ROM) and a random access memory (RAM), and provides the processor 1810 with program instructions and data stored in the memory 1820. In this embodiment, the memory 1820 may be used to store the program of any drivable area identification method in this embodiment.
[0238] The processor 1810 executes any of the drivable area identification methods in this application embodiment according to the program instructions stored in the memory 1820.
[0239] Based on the above embodiments, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the drivable area identification method in any of the above method embodiments.
[0240] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0241] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0242] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0243] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0244] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A travelable area recognition method characterized by comprising: Applied to target vehicles, including: In response to a drivable area recognition request triggered by a driving object, a target road image to be identified is acquired, wherein the target road image has a target image scale; For each preset candidate image scale, perform the following operations respectively: Based on a candidate image scale, the target road image is subjected to a first feature extraction to obtain a first feature map, wherein the first feature map has a first number of image channels; According to the preset receptive field of each candidate image, the second feature is extracted from the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, the receptive field of each candidate image represents: the size of the area on the target road image mapped by each pixel contained in the corresponding second feature map; Based on the first number of image channels, the superimposed feature map is adjusted to obtain the road feature map corresponding to the scale of the candidate image; For each road feature map, the following operations are performed: Based on the image position information of each pixel in a road feature map, and combined with the image channel information in the road feature map, attention weighting processing is performed on the road feature map to obtain the corresponding attention feature map; Based on the target image scale, the obtained attention feature maps are scaled to obtain at least one target feature map, and based on the at least one target feature map, the drivable area of the target vehicle is obtained.
2. The method as described in claim 1, characterized in that, The process involves applying attention-weighted processing to a road feature map based on the image position information of each pixel within the map, combined with the image channel information contained in the road feature map, to obtain a corresponding attention feature map. This includes: Based on the image position information of each pixel in a road feature map, the position attention weight of each pixel in the road feature map is determined, wherein the position attention weight represents the importance of the image position of the corresponding pixel when identifying a drivable area. Based on the image channel information contained in the road feature map, the channel attention weight of each image channel contained in the road feature map is determined, wherein the channel attention weight represents the importance of the corresponding image channel when identifying drivable areas; Based on the obtained attention weights for each location and channel, the road feature map is weighted to obtain the attention feature map corresponding to the road feature map.
3. The method according to any one of claims 1-2, characterized in that, The step of adjusting the image scale of each attention feature map based on the target image scale to obtain at least one target feature map includes: Based on the image scale of each attention feature map, the attention feature maps are fused and updated to obtain at least one fused feature map; Based on the preset target image channel number, the image channels of the at least one fused feature map are adjusted respectively to obtain the corresponding intermediate feature map; Based on the target image scale, at least one intermediate feature map is scaled to obtain the corresponding target feature map.
4. The method as described in claim 3, characterized in that, The step of fusing and updating the attention feature maps based on their respective image scales to obtain at least one fused feature map includes: The attention feature maps are sorted according to their respective image scales to obtain a sorting result. Based on the sorting result, each pair of adjacent attention feature maps is sequentially read and fused until all attention feature maps have been read. One fusion process includes: Read two adjacent attention feature maps and fuse them to obtain a fused feature map; Save the fused feature map, and use the fused feature map as a new attention feature map to replace the two attention feature maps, and add it to the ranking result.
5. The method according to any one of claims 1-2, characterized in that, Obtaining the drivable area of the target vehicle based on the at least one target feature map includes: If the at least one target feature map includes a target feature map, then the target feature map is taken as the result feature map; If the at least one target feature map includes multiple target feature maps, then image fusion is performed on the multiple target feature maps to obtain a result feature map; The values of each pixel in the resulting feature map are compared with the preset drivable area label values to determine the target pixels in the resulting feature map that belong to the drivable area. In the target road image, each pixel corresponding to its respective image position is marked to obtain an image containing the drivable area.
6. The method according to any one of claims 1-2, characterized in that, The drivable area is obtained by inputting the target road image into a target recognition model, wherein the target recognition model is trained in the following manner: The target recognition model is iteratively trained based on the training sample set to obtain the target recognition model; each training sample includes: image data of the sample road, wherein each iteration process performs the following operations: According to the preset sample image scale, feature extraction is performed on the selected training samples to obtain the corresponding sample feature maps. The training samples have sample image scale. For each sample feature map, the following operations are performed: Based on the sample image position information of each pixel contained in a sample feature map, and combined with the sample image channel information contained in the sample feature map, attention weighting processing is performed on the sample feature map to obtain the corresponding sample attention feature map. Based on the sample image scale, the image scale of each obtained sample attention feature map is adjusted to obtain at least one target sample feature map. Based on the feature map of the at least one target sample, the drivable area prediction result of the training sample is obtained, and the parameters are tuned based on the loss value corresponding to the drivable area prediction result.
7. A drivable area identification device, characterized in that, include The acquisition module is used to acquire the target road image to be identified in response to a drivable area recognition request triggered by a driving object, wherein the target road image has a target image scale; The extraction module performs the following operations for each preset candidate image scale: Based on a candidate image scale, the target road image is subjected to a first feature extraction to obtain a first feature map, wherein the first feature map has a first number of image channels; According to the preset receptive field of each candidate image, the second feature is extracted from the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, the receptive field of each candidate image represents: the size of the area on the target road image mapped by each pixel contained in the corresponding second feature map; Based on the first number of image channels, the superimposed feature map is adjusted to obtain the road feature map corresponding to the scale of the candidate image; The first processing module is used to perform the following operations on each road feature map: based on the image position information of each pixel contained in a road feature map, and combined with the image channel information contained in the road feature map, perform attention weighting processing on the road feature map to obtain the corresponding attention feature map; The second processing module is used to adjust the image scale of each attention feature map obtained based on the target image scale to obtain at least one target feature map, and to obtain the drivable area of the target vehicle based on the at least one target feature map.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Road drivable area fine recommendation method based on layered input and output and double-attention jumper connection
CN114674338A
Image feature extraction method and device, equipment and storage medium
CN115620017A