Negative Obstacle Semantic Segmentation Detection Method and Device Based on Attention Feature Fusion
The proposed method uses dual-camera systems and attention-based feature fusion to enhance negative obstacle detection in mobile robots, addressing inefficiencies in existing methods by improving accuracy and real-time performance.
Patent Information
- Application Number
- CN202310184347.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-01
AI Technical Summary
In the prior art, the negative obstacle detection method has low detection accuracy and insufficient efficiency, and the edge segmentation effect of the multimodal fusion method is poor, which cannot meet the real-time requirements of mobile robots.
Using an attention feature fusion method, a binocular camera is used to obtain the road surface image and process it into a parallax map. Combining RGB images and parallax maps, the features are extracted and fused through the attention feature fusion module, and the negative obstacle semantic segmentation is performed using the encoder-decoder framework.
The edge segmentation accuracy and overall segmentation results of negative obstacles are improved, which meets the real-time requirements of mobile robots and reduces the hardware computing power requirements.
Smart Images

Figure CN116188782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile robot environment perception, and in particular to a negative obstacle semantic segmentation detection method and device based on attention feature fusion. Background Art
[0002] The safe driving of a mobile robot is an important measure of its reliability. During the driving process, the robot will inevitably encounter various obstacles. According to the relationship with the ground, the obstacles can be divided into positive obstacles and negative obstacles. Positive obstacles refer to obstacles higher than the ground horizontal plane, such as vehicles, pedestrians, and stones; negative obstacles refer to obstacles lower than the horizontal plane, usually ditches or terrains with steep negative slopes, such as potholes and ground cracks. Crossing these obstacles will pose a danger to the mobile robot and may cause rollover, jamming of the motion structure, etc. Therefore, the detection of negative obstacles is of great significance for the safe driving of mobile robots.
[0003] Currently, the detection methods for positive obstacles have been widely studied, and a perfect algorithm framework and detection process have been implemented based on various mainstream semantic segmentation algorithms. However, the research on negative obstacle detection is relatively less, and there are still great deficiencies in detection accuracy and efficiency.
[0004] Currently, there are some means to detect negative obstacles based on single-modal semantic segmentation methods, mainly by reading RGB images or disparity maps for segmentation. However, these two methods have the following deficiencies: the visual features of negative obstacles are relatively close to the ground, and it is difficult to accurately segment their features only through RGB images; using disparity maps is easily affected by complex environments, such as water accumulation in pits and water stains around negative obstacles. At the same time, it is necessary to analyze each pixel of the image, which greatly improves the requirement for computing power and cannot meet the real-time performance of mobile robots. Currently, there are also some multi-modal fusion methods for negative obstacle detection, but in terms of the fusion method, only the corresponding channels of the images are added, which results in relatively low detection accuracy of the final result, especially for small negative obstacles, and the edge segmentation effect is poor. Summary of the Invention
[0005] Aiming at the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a negative obstacle semantic segmentation detection method and device based on attention feature fusion to solve the technical problems mentioned in the above background art part. The binocular camera of the mobile robot is used to capture road surface images and process them into disparity maps. Then, the road surface RGB images and disparity maps are input into the negative obstacle semantic segmentation detection model, and the attention feature fusion module is used to extract the features of the two images and further convert them to obtain the negative obstacle semantic segmentation result.
[0006] In a first aspect, the present invention provides a method for negative obstacle semantic segmentation detection based on attention feature fusion, comprising the following steps:
[0007] S1. Obtain a left-eye image and a right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the perspective-transformed left-eye image and right-eye image, and process the perspective-transformed left-eye image and right-eye image through a binocular stereo matching algorithm to obtain a disparity map;
[0008] S2. Construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module, and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined using a nested structure. The attention module is connected to the feature fusion module for reallocating weights;
[0009] S3. Input the perspective-transformed left-eye image and the disparity map into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the perspective-transformed left-eye image and the features of the disparity map, and perform feature fusion through the attention feature fusion module to obtain a final fused feature;
[0010] S4. Input the final fused feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
[0011] Preferably, the feature fusion module includes a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, and a fifth feature fusion module. The attention module includes a first channel attention module, a second channel attention module, a third channel attention module, a first dual attention module, and a second dual attention module. The features of the perspective-transformed left-eye image and the features of the disparity map sequentially pass through the first feature fusion module and the first channel attention module, and the first fused feature is output; the first fused feature and the features of the disparity map sequentially pass through the second feature fusion module and the second channel attention module, and the second fused feature is output; the second fused feature and the features of the disparity map sequentially pass through the third feature fusion module and the third channel attention module, and the third fused feature is output; the third fused feature and the features of the disparity map sequentially pass through the fourth feature fusion module and the first dual attention module, and the fourth fused feature is output; the fourth fused feature and the features of the disparity map sequentially pass through the fifth feature fusion module and the second dual attention module, and the final fused feature is output.
[0012] Preferably, the feature fusion module performs feature fusion using the following steps:
[0013] S31. Pixel - level add the input feature and the feature of the disparity map to obtain the added feature data;
[0014] S32. Input the added feature data into the fusion unit. The fusion unit includes a global feature extraction unit and a local feature extraction unit. Respectively extract features from the added feature data through the global feature extraction unit and the local feature extraction unit to obtain two groups of first weights, and unify the two groups of first weights through a normalization function to obtain two groups of first weights after normalization and unification;
[0015] S33. Corresponding assign the two groups of first weights after normalization and unification to pixel - level multiply with the input feature and the feature of the disparity map respectively to obtain two groups of first multiplication results, and then pixel - level add the two groups of first multiplication results to obtain the preliminary fusion feature;
[0016] S34. Input the preliminary fusion feature into the fusion unit again. Respectively extract features from the preliminary fusion feature through the global feature extraction unit and the local feature extraction unit to obtain two groups of second weights, and unify the two groups of second weights through a normalization function to obtain two groups of second weights after normalization and unification;
[0017] S35. Corresponding assign the two groups of second weights after normalization and unification to pixel - level multiply with the input feature and the feature of the disparity map to obtain two groups of second multiplication results, and then pixel - level add the two groups of second multiplication results to obtain the secondary fusion feature;
[0018] S36. Input the secondary fusion feature into the attention module to obtain the fusion feature.
[0019] Preferably, the calculation process in the feature fusion module is expressed as the following formula:
[0020]
[0021] Among them, R represents the input feature, D represents the feature of the disparity map, A is the global feature extraction unit or the local feature extraction unit, O is the output feature, represents pixel - level multiplication, and + represents pixel - level addition.
[0022] Preferably, the first - channel attention module, the second - channel attention module, and the third - channel attention module are all calculated using the following formula:
[0023]
[0024] Among them, X represents the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, X ′represents the first fusion feature, the second fusion feature, or the third fusion feature, S represents the Sigmoid function, T1 and T2 represent two linear transformations, H and W respectively represent the height and width of the output features of the first feature fusion module, the second feature fusion module, or the third feature fusion module, x(i,j) represents the pixel coordinates of the output features of the first feature fusion module, the second feature fusion module, or the third feature fusion module, i represents the abscissa of the pixel, and j represents the ordinate of the pixel.
[0025] Preferably, both the first dual attention module and the second dual attention module include a channel attention unit and a spatial attention unit. The fourth fusion feature is the sum of the output of the channel attention unit and the output of the spatial attention unit in the first dual attention module, and the final fusion feature is the sum of the output of the channel attention unit and the output of the spatial attention unit in the second dual attention module;
[0026] The calculation process of the spatial attention unit is as follows:
[0027]
[0028] Among them, A j is the output feature of the fourth feature fusion module or the fifth feature fusion module, B i , C j , D i are three feature maps generated by linearly operating on A j respectively. N is the product of the height and width of the output feature of the fourth feature fusion module or the fifth feature fusion module, α is the first scale parameter, P is the output of the spatial attention unit, i represents the abscissa of the pixel of the output feature of the fourth feature fusion module or the fifth feature fusion module, and j represents the ordinate of the pixel of the output feature of the fourth feature fusion module or the fifth feature fusion module;
[0029] The calculation process of the channel attention unit is as follows:
[0030]
[0031] Among them, A j is the output feature of the fourth feature fusion module or the fifth feature fusion module, X is the number of channels of the output feature of the fourth feature fusion module or the fifth feature fusion module, β is the second scale parameter, and Q is the output of the channel attention unit.
[0032] Preferably, the left-eye image and the right-eye image after perspective transformation are top views perpendicular to the ground angle.
[0033] In a second aspect, the present invention provides a negative obstacle semantic segmentation detection device based on attention feature fusion, including:
[0034] The image processing module is configured to obtain a left-eye image and a right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the left-eye image and the right-eye image after perspective transformation, and process the left-eye image and the right-eye image after perspective transformation through a binocular stereo matching algorithm to obtain a disparity map;
[0035] The model construction module is configured to construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined in a nested structure. The attention module is connected to the feature fusion module for reallocating weights;
[0036] The extraction and fusion module is configured to input the left-eye image and the disparity map after perspective transformation into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the left-eye image and the features of the disparity map after perspective transformation, and perform feature fusion through the attention feature fusion module to obtain the final fused features;
[0037] The decoding and output module is configured to input the final fused features into the decoder to obtain the semantic segmentation result of the negative obstacle.
[0038] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] (1) The present invention uses an RGB image and a disparity map as the input of the negative obstacle semantic segmentation detection model, thus solving the problems that the RGB image has a low degree of discrimination for negative obstacles and the disparity map is easily affected by the surrounding complex environment.
[0042] (2) The negative obstacle semantic segmentation detection model proposed by the present invention uses an attention feature fusion module to extract and fuse the features of the RGB image and the disparity map. Among them, the attention module increases the weight for the negative obstacle area; the feature fusion module extracts and fuses the features of the RGB image and the disparity map to improve the accuracy of region division. Compared with the existing multi-modal fusion methods, the edge segmentation accuracy of negative obstacles is greatly increased, thereby improving the overall level of the segmentation result.
[0043] (3) The negative obstacle semantic segmentation detection model proposed by the present invention is based on an encoder-decoder framework, with high overall detection efficiency and low requirements for hardware computing power, and can meet the real-time requirements of mobile robots. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 is an exemplary device architecture diagram to which an embodiment of the present application can be applied;
[0046] Figure 2 is a schematic flowchart of the negative obstacle semantic segmentation detection method based on attention feature fusion according to the embodiment of the present application;
[0047] Figure 3 is a flowchart of the negative obstacle semantic segmentation detection method based on attention feature fusion according to the embodiment of the present application;
[0048] Figure 4 is a schematic diagram of the input negative obstacle RGB image of the negative obstacle semantic segmentation detection method based on attention feature fusion according to the embodiment of the present application;
[0049] Figure 5 is a schematic diagram of the input negative obstacle disparity map of the negative obstacle semantic segmentation detection method based on attention feature fusion according to the embodiment of the present application;
[0050] Figure 6 is a schematic diagram of the negative obstacle semantic segmentation detection model according to the embodiment of the present application;
[0051] Figure 7 is a schematic flowchart of the attention feature fusion module of the negative obstacle semantic segmentation detection method based on attention feature fusion according to the embodiment of the present application;
[0052] Figure 8 The semantic segmentation result of the negative obstacle, which is the output of the negative obstacle semantic segmentation detection method based on attention feature fusion in the embodiment of the present application;
[0053] Figure 9 The schematic diagram of the negative obstacle semantic segmentation detection device based on attention feature fusion in the embodiment of the present application;
[0054] Figure 10 It is the schematic diagram of the structure of the computer device of the electronic device suitable for implementing the embodiment of the present application. Detailed implementation manners
[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Figure 1 The exemplary device architecture 100 to which the negative obstacle semantic segmentation detection method based on attention feature fusion or the negative obstacle semantic segmentation detection device based on attention feature fusion in the embodiment of the present application can be applied is shown.
[0057] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0058] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications, such as data processing applications, file processing applications, etc., may be installed on the terminal devices 101, 102, 103.
[0059] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0060] Server 105 may be a server that provides various services. For example, it may be a background data processing server that processes files or data uploaded by terminal devices 101, 102, and 103. The background data processing server can process the acquired files or data to generate a processing result.
[0061] It should be noted that the method for negative obstacle semantic segmentation detection based on attention feature fusion provided in the embodiments of the present application can be executed by server 105, or can be executed by terminal devices 101, 102, and 103. Correspondingly, the device for negative obstacle semantic segmentation detection based on attention feature fusion can be set in server 105, or can be set in terminal devices 101, 102, and 103.
[0062] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. In the case where the data to be processed does not need to be obtained remotely, the above device architecture may not include a network, but only a server or a terminal device.
[0063] Figure 2 The embodiments of the present application provide a method for negative obstacle semantic segmentation detection based on attention feature fusion, including the following steps:
[0064] S1, obtain a left-eye image and a right-eye image, respectively perform perspective transformation on the left-eye image and the right-eye image to obtain the left-eye image and the right-eye image after perspective transformation, and process the left-eye image and the right-eye image after perspective transformation through a binocular stereo matching algorithm to obtain a disparity map.
[0065] In a specific embodiment, the left-eye image and the right-eye image after perspective transformation are top views perpendicular to the ground angle.
[0066] Specifically, referring to Figure 3 , the left-eye image and the right-eye image are a pair of RBG images, and this pair of RGB images should be acquired by the binocular camera at the same time. The mobile robot captures road information through the binocular camera system to obtain a pair of RGB images, and transmits the pair of RGB images to the computer of the mobile robot. Among them, the binocular camera system consists of two industrial cameras and a camera bracket, which is installed at the advancing direction end of the mobile robot. The industrial cameras are vertically installed on the camera bracket. The camera bracket should form a fixed angle with the ground, and the baseline of the binocular camera and the distance between the two industrial cameras should also be fixed.
[0067] After the binocular camera system is installed, the internal and external parameters should be calibrated and the angle formed with the ground should be measured. The corresponding calibration parameters should be input into the computer of the mobile robot for subsequent image processing operations. If the angle between the binocular camera system and the ground changes or the baseline of the binocular camera changes, the above operations need to be performed again.
[0068] Reference Figure 4 , first, perform perspective transformation on a pair of RGB images. First, perform perspective transformation on them to obtain a top view perpendicular to the ground angle. Then, use the binocular stereo matching algorithm to obtain the disparity map of the current RGB image. The purpose of perspective transformation is to process the RGB image into a top view that is ninety degrees with the ground, so that the overall shape of the negative obstacles is more concrete. The perspective transformation processing angle should be related to the fixed angle between the camera and the ground, that is, the angle selection of perspective transformation should be strictly selected according to the angle between the binocular camera system and the ground. Then, use the binocular stereo matching algorithm to process the RGB image after perspective transformation to obtain the corresponding disparity map, as Figure 5 shown.
[0069] The calculation process of perspective transformation is as follows:
[0070]
[0071] Among them, (u, v) are the coordinates of the RGB image, (x = x' / w', y = y' / w') are the coordinates of the RGB image after perspective transformation, and the perspective transformation matrix is as follows:
[0072]
[0073] Among them, represents the linear transformation of the image;
[0074] T2 = [a 13 a 23 is used to generate the perspective transformation of the image;
[0075] T3 = [a 31 a 32 represents the translation of the image.
[0076] The disparity map obtained in this step should be the same size as the RGB image after perspective transformation.
[0077] S2, construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module, and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined using a nested structure. The attention module is connected to the feature fusion module and is used to reallocate weights.
[0078] In a specific embodiment, referring to Figure 6 , the feature fusion module includes a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, and a fifth feature fusion module, and the attention module includes a first channel attention module, a second channel attention module, a third channel attention module, a first dual attention module, and a second dual attention module; the features of the left-eye image after perspective transformation and the features of the disparity map sequentially pass through the first feature fusion module and the first channel attention module, and the first fused feature is output; the first fused feature and the features of the disparity map sequentially pass through the second feature fusion module and the second channel attention module, and the second fused feature is output; the second fused feature and the features of the disparity map sequentially pass through the third feature fusion module and the third channel attention module, and the third fused feature is output; the third fused feature and the features of the disparity map sequentially pass through the fourth feature fusion module and the first dual attention module, and the fourth fused feature is output; the fourth fused feature and the features of the disparity map sequentially pass through the fifth feature fusion module and the second dual attention module, and the final fused feature is output.
[0079] S3. Input the left-eye image after perspective transformation and the disparity map into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the left-eye image after perspective transformation and the features of the disparity map, and perform feature fusion through the attention feature fusion module to obtain the final fused feature.
[0080] Specifically, input the left-eye image after perspective transformation and the disparity map into the negative obstacle semantic segmentation detection model, extract the features of the two through the encoder and the attention feature fusion module, and fuse them to realize the segmentation of the negative obstacle. First, the left-eye image after perspective transformation and the disparity map are input into the encoder to respectively extract the features of the left-eye image after perspective transformation and the features of the disparity map. In a specific embodiment, the structure of the encoder in the negative obstacle semantic segmentation detection model uses ResNet30 as the BackBone, changes the input channel number from 3 to 512, and changes the image size from 512*512 to 16*16; through Figure 7 The shown attention feature fusion module extracts and fuses the features of the two to obtain relevant information about the negative obstacle. The attention feature fusion module includes a feature fusion module and an attention module. Among them, the features of the left-eye image and the features of the disparity map are input into the feature fusion module. First, pixel-level addition is performed on them, and then the added feature data is input into the fusion unit. The fusion unit includes a global feature extraction unit and a local feature extraction unit, and the two respectively extract features from the data to obtain two sets of weights, and finally unify them through a normalization function. In a preferred embodiment, the activation function is Sigmoid, and the weights are mapped to the interval of (0,1):
[0081]
[0082] After the weights after normalization and unification are correspondingly assigned, they are multiplied pixel by pixel with the features of the left-eye image and the features of the disparity map to increase the weights of negative obstacles in the image. Then, the two are added pixel by pixel to obtain the preliminary fusion features. The feature fusion module adopts a nested structure. The feature results of the previous layer are input into the fusion unit of the next layer, and the above operations are repeated to obtain the secondary fusion features. The secondary fusion features are multiplied pixel by pixel with the features of the left-eye image and the features of the disparity map and then added to obtain the final result. At this time, the weights of negative obstacles and the background in the image are further separated. The subsequent feature fusion modules all adopt the same calculation process.
[0083] In a specific embodiment, the feature fusion module performs feature fusion using the following steps:
[0084] S31, add the input features and the features of the disparity map pixel by pixel to obtain the added feature data;
[0085] S32, input the added feature data into the fusion unit. The fusion unit includes a global feature extraction unit and a local feature extraction unit. The global feature extraction unit and the local feature extraction unit respectively extract features from the added feature data to obtain two sets of first weights, and the two sets of first weights are unified through a normalization function to obtain two sets of first weights after normalization and unification;
[0086] S33, correspondingly assign the two sets of first weights after normalization and unification to multiply with the input features and the features of the disparity map pixel by pixel to obtain two sets of first multiplication results, and then add the two sets of first multiplication results pixel by pixel to obtain the preliminary fusion features;
[0087] S34, input the preliminary fusion features into the fusion unit again. The global feature extraction unit and the local feature extraction unit respectively extract features from the preliminary fusion features to obtain two sets of second weights, and the two sets of second weights are unified through a normalization function to obtain two sets of second weights after normalization and unification;
[0088] S35, correspondingly assign the two sets of second weights after normalization and unification to multiply with the input features and the features of the disparity map pixel by pixel to obtain two sets of second multiplication results, and then add the two sets of second multiplication results pixel by pixel to obtain the secondary fusion features;
[0089] S36, input the secondary fusion features into the attention module to obtain the fusion features.
[0090] In a specific embodiment, the calculation process in the feature fusion module is expressed as the following formula:
[0091]
[0092] Among them, R represents the input feature, D represents the feature of the disparity map, A is the global feature extraction unit or the local feature extraction unit, and O is the output feature. represents pixel-level multiplication, and + represents pixel-level addition.
[0093] Further, an attention module is connected behind each feature fusion module. The attention module includes a channel attention module and a dual attention module. The channel attention module calculates the importance of each channel of the input image. If the channel contains information related to negative obstacles, then this channel will receive more attention, that is, the corresponding weight increases, so as to achieve the purpose of improving the feature representation ability. The dual attention module includes a channel attention unit and a spatial attention unit. Among them, the spatial attention unit is a supplement to the channel attention unit. Based on the channel direction on the basis of the channel where the negative obstacle is located, the position where the most information related to the negative obstacle gathers is found and a higher weight is given.
[0094] In a specific embodiment, the first channel attention module, the second channel attention module, and the third channel attention module are all calculated using the following formula:
[0095]
[0096] Among them, X represents the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, and X ′ represents the first fusion feature, the second fusion feature, or the third fusion feature, S represents the Sigmoid function, T1 and T2 represent two linear transformations, H and W respectively represent the height and width of the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, x(i,j) represents the pixel point of the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, i represents the abscissa of the pixel point, and j represents the ordinate of the pixel point.
[0097] In a specific embodiment, the first dual attention module and the second dual attention module both include a channel attention unit and a spatial attention unit. The fourth fusion feature is the sum of the output of the channel attention unit and the output of the spatial attention unit in the first dual attention module. The final fusion feature is the sum of the output of the channel attention unit and the output of the spatial attention unit in the second dual attention module.
[0098] The calculation process of the spatial attention unit is as follows:
[0099]
[0100] Among them, Aj is the output feature of the fourth or fifth feature fusion module, B i , C j , D i are three feature maps generated by performing linear operations on A j respectively. N is the product of the height and width of the output feature of the fourth or fifth feature fusion module, α is the first scale parameter, P is the output of the spatial attention unit, i represents the abscissa of the pixel point of the output feature of the fourth or fifth feature fusion module, and j represents the ordinate of the pixel point of the output feature of the fourth or fifth feature fusion module;
[0101] The calculation process of the channel attention unit is as follows:
[0102]
[0103] Among them, A j is the output feature of the fourth or fifth feature fusion module, X is the number of channels of the output feature of the fourth or fifth feature fusion module, β is the second scale parameter, and Q is the output of the channel attention unit.
[0104] Specifically, in step S31, the feature input to the first feature fusion module is the feature of the left-eye image after perspective transformation, the feature input to the second feature fusion module is the first fusion feature, the feature input to the third feature fusion module is the second fusion feature, the feature input to the fourth feature fusion module is the third fusion feature, and the feature input to the fifth feature fusion module is the fourth fusion feature. In step S36, the fusion feature corresponding to the output of the first channel attention module is the first fusion feature, the fusion feature corresponding to the output of the second channel attention module is the second fusion feature, the fusion feature corresponding to the output of the third channel attention module is the third fusion feature, the fusion feature corresponding to the output of the first dual attention module is the fourth fusion feature, and the fusion feature corresponding to the output of the second dual attention module is the final fusion feature. The attention module redistributes the weights of the ROI region of the target image, making its weight higher than that of other irrelevant regions, thereby improving the segmentation accuracy of the model.
[0105] S4. Input the final fusion feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
[0106] Specifically, the final fused feature is restored to the image size through a decoder to obtain the semantic segmentation result of the negative obstacle. The decoder consists of five layers of networks, which perform a scale restoration operation on the final fused feature and convert it into an image for output. In one embodiment, the decoder changes the input channels from 512 to 2, and the image size from 16*16 to 512*512. Finally, the semantic segmentation result of the negative obstacle is obtained as Figure 8 shown.
[0107] For further reference Figure 9 , as an implementation of the methods shown in the above figures, an embodiment of a negative obstacle semantic segmentation detection device based on attention feature fusion is provided in the present application. This device embodiment corresponds to the Figure 2 method embodiment shown, and this device can be specifically applied to various electronic devices.
[0108] An embodiment of a negative obstacle semantic segmentation detection device based on attention feature fusion is provided in the embodiments of the present application, including:
[0109] An image processing module 1, configured to obtain a left-eye image and a right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the perspective-transformed left-eye image and right-eye image, and process the perspective-transformed left-eye image and right-eye image through a binocular stereo matching algorithm to obtain a disparity map;
[0110] A model construction module 2, configured to construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module, and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined using a nested structure. The attention module is connected to the feature fusion module for reassigning weights;
[0111] An extraction and fusion module 3, configured to input the perspective-transformed left-eye image and the disparity map into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the perspective-transformed left-eye image and the features of the disparity map, and perform feature fusion through the attention feature fusion module to obtain the final fused feature;
[0112] A decoding and output module 4, configured to input the final fused feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
[0113] Next, refer to Figure 10 , which shows a schematic structural diagram of a computer device 1000 suitable for implementing the embodiments of the present application (such as Figure 1 the server or terminal device shown). Figure 10The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0114] As Figure 10 shown, the computer device 1000 includes a central processing unit (CPU) 1001 and a graphics processing unit (GPU) 1002, which can perform various appropriate actions and processes according to programs stored in the read-only memory (ROM) 1003 or programs loaded from the storage section 1009 into the random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the device 1000 are also stored. The CPU 1001, GPU 1002, ROM 1003, and RAM 1004 are connected to each other via a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus 1005.
[0115] The following components are connected to the I / O interface 1006: an input section 1007 including a keyboard, a mouse, etc.; an output section 1008 including, for example, a liquid crystal display (LCD), etc. and speakers, etc.; a storage section 1009 including a hard disk, etc.; and a communication section 1010 including a network interface card such as a LAN card, a modem, etc. The communication section 1010 performs communication processing via a network such as the Internet. A drive 1011 may also be connected to the I / O interface 1006 as needed. A removable medium 1012, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1011 as needed so that a computer program read therefrom can be installed into the storage section 1009 as needed.
[0116] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1010, and / or installed from the removable medium 1012. When the computer program is executed by the central processing unit (CPU) 1001 and the graphics processing unit (GPU) 1002, the above functions defined in the methods of the present application are executed.
[0117] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution device, apparatus, or component. In this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0118] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0120] The modules described in the embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0121] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a left-eye image and a right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the left-eye image and the right-eye image after perspective transformation, process the left-eye image and the right-eye image after perspective transformation through a binocular stereo matching algorithm to obtain a disparity map; construct and train a negative obstacle semantic segmentation detection model, where the negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module, and a decoder, the attention feature fusion module includes a feature fusion module and an attention module, the feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined in a nested structure, and the attention module is connected to the feature fusion module for reallocating weights; input the left-eye image after perspective transformation and the disparity map into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the left-eye image after perspective transformation and the features of the disparity map, perform feature fusion through the attention feature fusion module to obtain a final fused feature; input the final fused feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
[0122] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. A negative obstacle semantic segmentation detection method based on attention feature fusion, characterized in that, It includes the following steps: S1. Obtain the left-eye image and the right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the left-eye image and the right-eye image after perspective transformation, and process the left-eye image and the right-eye image after perspective transformation through a binocular stereo matching algorithm to obtain a disparity map; S2. Construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined in a nested structure. The attention module is connected to the feature fusion module and is used to reallocate weights; S3. Input the left-eye image and the disparity map after perspective transformation into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the left-eye image and the features of the disparity map after perspective transformation. Feature fusion is performed through the attention feature fusion module. The number of the feature fusion modules is five. The attention module includes three channel attention modules and two dual attention modules; The features of the left-eye image and the features of the disparity map after perspective transformation pass through the first feature fusion module and the first channel attention module in sequence, and the first fusion feature is output; The first fusion feature and the features of the disparity map pass through the second feature fusion module and the second channel attention module in sequence, and the second fusion feature is output; The second fusion feature and the features of the disparity map pass through the third feature fusion module and the third channel attention module in sequence, and the third fusion feature is output; The third fusion feature and the features of the disparity map pass through the fourth feature fusion module and the first dual attention module in sequence, and the fourth fusion feature is output; The fourth fusion feature and the features of the disparity map pass through the fifth feature fusion module and the second dual attention module in sequence, and the final fusion feature is output; The feature fusion module performs feature fusion through the following steps: S31. Add the input features and the features of the disparity map pixel by pixel to obtain the added feature data; S32. Input the added feature data into the fusion unit. The fusion unit includes the global feature extraction unit and the local feature extraction unit. The global feature extraction unit and the local feature extraction unit respectively extract features from the added feature data to obtain two sets of first weights, and unify the two sets of first weights through a normalization function to obtain two sets of first weights after normalization and unification; S33. Corresponding assign the two sets of first weights after normalization and unification to multiply the input features and the features of the disparity map pixel by pixel respectively to obtain two sets of first multiplication results, and then add the two sets of first multiplication results pixel by pixel to obtain a preliminary fusion feature; S34. Input the preliminary fusion feature into the fusion unit again, extract features from the preliminary fusion feature through the global feature extraction unit and the local feature extraction unit respectively to obtain two sets of second weights, and unify the two sets of second weights through a normalization function to obtain two sets of second weights after normalization and unification; S35. Corresponding to the two sets of second weights after normalization and unification, perform pixel-wise multiplication with the features of the input feature and the disparity map to obtain two sets of second multiplication results, and then perform pixel-wise addition on the two sets of second multiplication results to obtain a secondary fusion feature; S36. Input the secondary fusion feature into the attention module to obtain a fusion feature; S4. Input the final fusion feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
2. The negative obstacle semantic segmentation detection method based on attention feature fusion according to claim 1, wherein The calculation process in the feature fusion module is expressed as follows: Among them, R represents the input feature, D represents the feature of the disparity map, A is a global feature extraction unit or a local feature extraction unit, and O is the output feature. denotes pixel-level multiplication, and + denotes pixel-level addition.
3. The negative obstacle semantic segmentation detection method based on attention feature fusion according to claim 1, wherein, The first channel attention module, the second channel attention module, and the third channel attention module are all calculated using the following formula: Among them, X represents the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, and X ′ represents the first fusion feature, the second fusion feature, or the third fusion feature, S represents the Sigmoid function, T1 and T2 represent two linear transformations, H and W respectively represent the height and width of the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, x(i,j) represents the pixel point coordinates of the output feature of the first feature fusion module, the second feature fusion module, or the third feature fusion module, i represents the abscissa of the pixel point, and j represents the ordinate of the pixel point.
4. The negative obstacle semantic segmentation detection method based on attention feature fusion according to claim 1, wherein The first dual attention module and the second dual attention module both include a channel attention unit and a spatial attention unit. The fourth fusion feature is the sum of the outputs of the channel attention unit and the spatial attention unit in the first dual attention module, and the final fusion feature is the sum of the outputs of the channel attention unit and the spatial attention unit in the second dual attention module; The calculation process of the spatial attention unit is as follows: Among them, A j is the output feature of the fourth feature fusion module or the fifth feature fusion module, B i , C j , D i are three feature maps generated by performing linear operations on A j respectively. N is the product of the height and width of the output feature of the fourth feature fusion module or the fifth feature fusion module. α is the first scale parameter. P is the output of the spatial attention unit. i represents the abscissa of the pixel point of the output feature of the fourth feature fusion module or the fifth feature fusion module, and j represents the ordinate of the pixel point of the output feature of the fourth feature fusion module or the fifth feature fusion module; The calculation process of the channel attention unit is as follows: Among them, A j is the output feature of the fourth feature fusion module or the fifth feature fusion module, X is the number of channels of the output feature of the fourth feature fusion module or the fifth feature fusion module, β is the second scale parameter, and Q is the output of the channel attention unit.
5. The negative obstacle semantic segmentation detection method based on attention feature fusion according to claim 1, characterized in that, The left-eye image and the right-eye image after perspective transformation are top views perpendicular to the ground angle.
6. A negative obstacle semantic segmentation detection device based on attention feature fusion, characterized in that, Using the negative obstacle semantic segmentation detection method based on attention feature fusion according to any one of claims 1-5, including: An image processing module configured to obtain a left-eye image and a right-eye image, perform perspective transformation on the left-eye image and the right-eye image respectively to obtain the left-eye image and the right-eye image after perspective transformation, and process the left-eye image and the right-eye image after perspective transformation through a binocular stereo matching algorithm to obtain a disparity map; A model construction module configured to construct and train a negative obstacle semantic segmentation detection model. The negative obstacle semantic segmentation detection model includes an encoder, an attention feature fusion module, and a decoder. The attention feature fusion module includes a feature fusion module and an attention module. The feature fusion module includes a global feature extraction unit and a local feature extraction unit, and is combined using a nested structure. The attention module is connected to the feature fusion module and is used to reallocate weights; An extraction fusion module configured to input the left-eye image and the disparity map after perspective transformation into the encoder in the negative obstacle semantic segmentation detection model to extract the features of the left-eye image and the disparity map after perspective transformation, and perform feature fusion through the attention feature fusion module to obtain a final fusion feature; A decoding output module configured to input the final fusion feature into the decoder to obtain the semantic segmentation result of the negative obstacle.
7. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1-5 is implemented.