Road segmentation method based on local-global double-flow information collaborative awareness

Through the road segmentation method of local-global dual-stream information collaborative perception, dynamic shadow synthesis and dual-path fusion modules are used to achieve efficient feature extraction and fusion, solving the problem of insufficient fusion of local details and global semantic information in the existing technology, and improving the accuracy and robustness of road segmentation.

CN120375320AActive Publication Date: 2025-07-25MO NI XUEDI (JIANGXI) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510879077.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing road segmentation method is difficult to take into account the effective fusion of local details and global semantic information at the same time, resulting in poor segmentation consistency in complex scenarios, especially in extreme lighting, occlusion or long-distance road scenarios.

Method used

The road segmentation method based on local-global dual-stream information collaborative perception is adopted, and dynamic shadow synthesis data enhancement module, local collaborative perception branch and global perception branch are combined with dynamic dual-path fusion module to achieve dynamic balance and adaptive feature calibration of high-resolution local features and global context.

Benefits of technology

It significantly improves the road feature recognition accuracy and robustness of the model in complex environments, enhances the detection ability of small targets, improves the accuracy and detailed integrity of segmentation, and overcomes the problems of information loss and low fusion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375320A_ABST
    Figure CN120375320A_ABST
Patent Text Reader

Abstract

The invention discloses a road segmentation method based on local-global double-flow information collaborative awareness. The method comprises the following steps: collecting a high-resolution road image; constructing a road segmentation model; inputting the image into a road segmentation model, firstly introducing a shadow into the high-resolution road image in the road segmentation model for enhancement, and then reducing the enhanced image; obtaining a global semantic feature map of the reduced image through a global perception branch; acquiring a local detail feature map of the enhanced image through a local collaborative perception branch; performing feature alignment on the global semantic feature map, performing element-by-element addition on the global semantic feature map and the local detail feature map to obtain a preliminary fusion feature map, and inputting the preliminary fusion feature map into a dynamic double-path fusion module to obtain a road segmentation map; according to the method, full-link upgrading from data generation, feature extraction and feature fusion to training optimization is realized, and the recognition precision and robustness of the model on road features and the detection capability on small targets in a complex environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically provides a road segmentation method based on collaborative perception of local-global dual-stream information. Background Art

[0002] In the road segmentation task, significant progress has been made in image segmentation methods based on deep learning. However, existing methods usually adopt a single feature extraction method, making it difficult to effectively integrate local details and global semantic information simultaneously. In the process of feature extraction by traditional methods, the receptive field is often enlarged at the expense of spatial resolution, resulting in a decline in the perception ability of small-scale structures (such as road markings and cracks), or ignoring the long-distance context dependence when emphasizing local details, affecting the segmentation consistency in complex scenarios.

[0003] In the prior art, although multi-scale feature fusion (such as FPN, U-Net, etc.) can alleviate this problem to a certain extent, there are still key defects. Local-global feature fragmentation: Most methods adopt a serial stacked convolutional structure, and local fine-grained features are gradually diluted during the propagation in the deep network, and key microscopic structure information cannot be retained in the final prediction; Insufficient feature interaction mechanism: Existing dual-stream architectures (such as high-low resolution parallel networks) usually rely on simple feature concatenation or addition fusion, and fail to establish an effective cross-level dynamic interaction mechanism, resulting in insufficient collaborative optimization of local details and global semantics; Poor adaptability to complex scenarios: In challenging scenarios such as extreme lighting, occlusion, or long-distance roads, existing methods are prone to generate fragmented mis-segmentations (such as blurred road edges and missed detection of small targets) due to local feature ambiguity or global context loss. Notably, although some studies attempt to introduce attention mechanisms (such as Non-local modules) to enhance long-range dependence modeling, their computational complexity is high, and it is difficult to achieve precise calibration of global semantics while maintaining high-resolution local features. In addition, the existing methods have insufficient explicit modeling ability for road topology structures and cannot adaptively adjust the contribution weights of local and global features, resulting in limited generalization performance of the algorithm in complex urban scenarios.

[0004] Therefore, there is an urgent need for a new local-global dual-stream information collaborative perception mechanism that can achieve a dynamic balance between high-resolution local feature extraction and efficient global context modeling, and improve the accuracy and robustness of road segmentation through structured cross-level interaction and adaptive feature calibration. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a road segmentation method based on collaborative perception of local-global dual-stream information, aiming to solve the problems in the background art.

[0006] To achieve the above object, the present invention provides the following technical solution: A road segmentation method based on local-global dual-stream information collaborative perception, comprising the following steps: Step S1: Construct a data set, which includes a number of high-resolution road images, and the road areas in the high-resolution road images are all labeled with tags; Step S2: Construct a road segmentation model; the road segmentation model consists of a dynamic shadow synthesis data enhancement module, a local collaborative perception branch, a global perception branch, and a dynamic dual-path fusion module; Step S3: Input the high-resolution road image into the dynamic shadow synthesis data enhancement module. First, introduce an ellipse or tree-shaped shadow into the high-resolution road image to obtain a high-resolution road enhanced image, and then reduce the high-resolution road enhanced image to obtain a low-resolution road enhanced image; Step S4: Obtain the global semantic feature map of the low-resolution road enhanced image through the global perception branch, and obtain the local detail feature map of the high-resolution road enhanced image through the local collaborative perception branch; Step S5: After aligning the global semantic feature map, add it element-wise to the local detail feature map to obtain a preliminary fusion feature map, and input the preliminary fusion feature map into the dynamic dual-path fusion module to obtain a road segmentation map.

[0007] Further, the dynamic dual-path fusion module consists of a spatial attention module, a channel attention module, an edge assistance branch, a pixel attention module, and a segmentation head; input the preliminary fusion feature map into the dynamic dual-path fusion module to obtain the specific process of the road segmentation map is as follows: Input the preliminary fusion feature map into the spatial attention module to obtain a spatial attention map , which is expressed as: ; In the formula, represents convolutional layer; represents the concatenation operation; represents the average pooling layer; represents the max pooling layer; Input the spatial attention map and the preliminary fusion feature map into the edge assistance branch to obtain an edge-enhanced spatial attention map , which is expressed as: ; In the formula, represents function; represents Convolution layer; Input the preliminary fusion feature map into the channel attention module to obtain the channel attention map ; Add the edge-enhanced spatial attention map and the channel attention map through broadcast operation to obtain the spatial-channel attention map ; Input the spatial-channel attention map and the preliminary fusion feature map into the pixel attention module simultaneously to obtain the spatial-channel-pixel attention map , denoted as: ; In the formula, denotes grouped convolution layer; denotes batch normalization operation; Through the spatial-channel-pixel attention map dynamically adaptively fuse the global semantic feature map and the local detail feature map , add the output after dynamic adaptive fusion to the preliminary fusion feature map and then input it into the segmentation head to obtain the road segmentation map , denoted as: ; In the formula, denotes the segmentation head.

[0008] Furthermore, the specific process of introducing elliptical or tree-shaped shadows into the high-resolution road image to obtain the high-resolution road enhanced image is as follows: Step S3.11: Randomly select different types of shadows, where different types of shadows include elliptical shadows and tree shadows; Step S3.12: After determining the shadow type, based on the label information of the road area in the high-resolution road image, dynamically generate the enhancement parameters of the corresponding type of shadow; Step S3.13: Based on the generated enhancement parameters of the corresponding type of shadow, call the elliptical shadow mask generation function or the tree shadow mask generation function to generate the shadow mask of the corresponding shape , denoted as: ; In the formula, denotes the elliptical shadow mask generation function; denotes the tree shadow mask generation function; Represents the coordinates of any pixel point in the high - resolution road image; Represents the type of shadow; Represents the enhancement parameter of the generated elliptical shadow; Represents the enhancement parameter of the generated tree shadow; Represents the label information of the road area in the high - resolution road image; Represents an elliptical shadow; Step S3.14: Through the weighted fusion formula, overlay the shadow mask and the high - resolution road image according to the transparency parameter, and output the enhanced high - resolution road image .

[0009] Furthermore, the specific process of step S3.12 is as follows: For the elliptical shadow, first randomly sample the center coordinates of the elliptical shadow within the road area in the high - resolution road image ; Generate the random size parameters of the elliptical shadow , represents the horizontal semi - axis length of the elliptical shadow, represents the vertical semi - axis length of the elliptical shadow; Generate the transparency parameter for controlling the depth of the elliptical shadow ; For the tree shadow, first set a fixed template of a preset size in the road area of the high - resolution road image, stack three predefined elliptical shadows on the fixed template in a staggered manner to generate the tree shadow; Then calculate the scaling factor to control the size of the fixed template, and finally generate the transparency parameter for controlling the depth of the tree shadow and the horizontal flip flag for controlling the horizontal flip of the tree shadow ; Among them, the three predefined elliptical shadows include the central elliptical shadow, the left - hand elliptical shadow, and the right - hand elliptical shadow; The process of dynamically generating the enhancement parameters of the corresponding type of shadow is expressed as: ; ; ; In the formula, represents the enhancement parameter of the generated corresponding type of shadow; represents the fixed template in the position coordinates within the road area of the high - resolution road image; The elliptical shadow mask generation function is expressed as: ; The tree shadow mask generation function Expressed as: ; In the formula, represents a resizing operation for resizing the fixed template according to ; represents a flipping operation for flipping the fixed template according to ; represents the fixed template; represents the coordinates of any pixel point in the fixed template.

[0010] Furthermore, the local collaborative perception branch has a three-layer structure. The first layer of the local collaborative perception branch uses the initialization layer of the ResNet18 network, the second layer of the local collaborative perception branch uses a residual layer, and the third layer of the local collaborative perception branch uses a multi-directional local feature enhancement module; the global perception branch has a five-layer structure. The first layer of the global perception branch uses a PatchEmbed module, the second, third, and fourth layers of the global perception branch respectively use the first, second, and third layers of the MambaVision network, and the fifth layer of the global perception branch uses an adaptive channel interaction attention module.

[0011] Furthermore, the specific process for the global perception branch to obtain the global semantic feature map of the low-resolution road enhancement image is as follows: Step S4.11: Input the low-resolution road enhancement image into the PatchEmbed module to obtain the embedded feature ; among them, the PatchEmbed module consists of two convolutional layers with a stride of 2; input sequentially through two convolutional layers with a stride of 2 to obtain the embedded feature ; Step S4.12: Input the embedded feature into the first layer of the MambaVision network to obtain the first feature map ; among them, the first layer of the MambaVision network consists of a convolutional block, and the convolutional block contains two convolutional layers with a stride of 1; input the embedded feature sequentially through two convolutional layers with a stride of 1 to obtain the intermediate feature map , perform regularization on the intermediate feature map using the random path dropout operation, and add the regularized output to the embedded feature to obtain the first feature map ; Step S4.13: Perform a downsampling operation on the first feature map to obtain the first downsampled feature map ; Step S4.14: Input the first downsampled feature map into the second layer of the MambaVision network to obtain a second feature map ; wherein, the second layer of the MambaVision network consists of three sequentially connected convolutional blocks, and the convolutional blocks within the second layer of the MambaVision network have the same structure as the convolutional blocks within the first layer of the MambaVision network; the first downsampled feature map passes through the three convolutional blocks within the second layer of the MambaVision network in sequence to obtain a second feature map ; Step S4.14: Perform a downsampling operation on the second feature map to obtain a second downsampled feature map ; Step S4.15: Input the second downsampled feature map into the third layer of the MambaVision network to obtain a third feature map ; wherein, the third layer of the MambaVision network consists of a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence; first, flatten the second downsampled feature map , and the flattened passes through the first hybrid module, the second hybrid module, and the third hybrid module in sequence to obtain a third feature map ; Step S4.16: Input the third feature map into the adaptive channel interaction attention module to obtain a global semantic feature map .

[0012] Furthermore, the specific process of step S4.15 is as follows: Flatten the second downsampled feature map in the spatial dimension to obtain a feature sequence ; Input the feature sequence into the first hybrid module to obtain the output of the first hybrid module ; wherein, the first hybrid module consists of a first MLP branch and a Mamba branch connected in parallel; The first MLP branch consists of two fully connected layers; the feature sequence passes through the two fully connected layers in sequence to obtain a first feature enhancement sequence , that is, the output of the first MLP branch; In the Mamba branch, the feature sequence is upsampled through a linear transformation layer; the upsampled feature sequence After passing through a one-dimensional convolutional layer, a second feature enhancement sequence is obtained ; The second feature enhancement sequence generates a gating signal through a linear transformation layer ; Through the gating signal to control the information flow of the second feature enhancement sequence and obtain a third feature enhancement sequence ; Perform a selective scanning operation on the third feature enhancement sequence through a learnable time step parameter ranging from 0 to 1 to dynamically adjust the fusion ratio between the third feature enhancement sequence input at the current moment and the third feature enhancement sequence input at the previous moment; Pass the third feature enhancement sequence with the adjusted fusion ratio through a linear transformation layer to restore the original dimension and then concatenate it with the second feature enhancement sequence to obtain a fourth enhanced feature sequence , which is the output of the Mamba branch; Pass the fourth enhanced feature sequence through a linear transformation layer and then concatenate it with the first feature enhancement sequence in the dimension to obtain a fifth enhanced feature sequence . Pass the fifth enhanced feature sequence through the first MLP branch and perform a random path dropout operation, and then add it element-wise to the feature sequence to obtain the output of the first mixing module ; Input the output of the first mixing module into the second mixing module; The second mixing module consists of a parallel self-attention branch and a second MLP branch; The second MLP branch of the second mixing module has the same structure as the first MLP branch of the first mixing module; Pass the output of the first mixing module through the self-attention branch and the second MLP branch respectively to obtain the output of the self-attention branch and the output of the second MLP branch. Concatenate the output of the self-attention branch and the output of the second MLP branch to obtain a sixth feature enhancement sequence . Pass the sixth feature enhancement sequence through the second MLP branch and perform a random path dropout operation, and then add it element-wise to the output of the first mixing module to obtain the output of the second mixing module ; The structure of the third mixing module is the same as that of the second mixing module; The output of the second mixing module is input into the third mixing module to obtain the output of the third mixing module ; The output of the third mixing module is converted from the feature sequence into a feature map to obtain the third feature map .

[0013] Furthermore, a high-resolution road enhancement image is obtained through the local collaborative perception branch of the local detail feature map The specific process is as follows: Step S4.21: Input the high-resolution road enhancement image into the initialization layer of the ResNet18 network to obtain the fourth feature map ; where the initialization layer of the ResNet18 network consists of a convolutional layer with a stride of 2 and a padding of 3 and a max pooling layer; the high-resolution road enhancement image passes through the convolutional layer and the max pooling layer in the initialization layer of the ResNet18 network in sequence to obtain the fourth feature map ; Step S4.22: Input the fourth feature map into the residual layer to obtain the output of the residual layer ; where the residual layer consists of a first residual block and a second residual block, and the first residual block and the second residual block have the same structure, both consisting of two convolutional layers with a stride of 1; the fourth feature map passes through the two convolutional layers with a stride of 1 in the first residual block to obtain the output of the first residual block, and the output of the first residual block passes through the two convolutional layers with a stride of 1 in the second residual block and is added to the output of the first residual block to obtain the output of the residual layer ; Step S4.23: Input the output of the residual layer into the multi-directional local feature enhancement module to obtain the local detail feature map .

[0014] Furthermore, the multi-directional local feature enhancement module includes a depthwise separable convolutional layer, a convolutional layer, and a local attention mechanism; input the output of the residual layer into the multi-directional local feature enhancement module to obtain the local detail feature map The specific process is as follows: Input the output of the residual layer sequentially through the depthwise separable convolutional layer and a convolutional layer in the horizontal direction to obtain the horizontal feature ; input the output of the residual layer sequentially through the depthwise separable convolutional layer and a convolutional layer in the vertical direction to obtain the vertical feature ; add the horizontal feature , the vertical feature and the output of the residual layer to obtain the multi-directional fusion feature ; Generate multi-directional fusion features through a local attention mechanism of the channel weights , and multiply the channel weights with the output of the residual layer and then add it to the multi-directional fusion feature to obtain a local detail feature map .

[0015] Compared with the existing technologies, the present invention has the following beneficial effects:

[0016] (1) The present invention organically combines dynamic shadow synthesis data augmentation, parallel global and local collaborative perception branches, and dynamic dual-path fusion to form a complete and efficient technical system. By generating diverse difficult samples through data augmentation, it provides rich training materials for the global and local collaborative perception branches, enabling them to fully play the complementary feature extraction advantage under complex lighting and occlusion conditions; by extracting comprehensive features through the global and local collaborative perception branches and then performing adaptive interaction based on the triple attention mechanism by the dynamic dual-path fusion module, it realizes the deep fusion and optimization of features; the present invention realizes the full-link upgrade from data generation, feature extraction, feature fusion to training optimization, significantly improving the model's recognition accuracy, robustness of road features, and detection ability for small targets in complex environments, showing obvious advantages in the accuracy, detail integrity of road segmentation, and handling of unbalanced data, effectively overcoming problems such as information loss, low fusion efficiency, and training bias existing in the prior art.

[0017] (2) The dynamic shadow synthesis data augmentation module designed by the present invention can simulate complex lighting and occlusion scenarios, enabling the model to encounter diverse difficult samples during the training phase and enhancing the model's robust perception ability of road features in actual complex environments such as shadow coverage and partial occlusion.

[0018] (3) Through the design of parallel global perception branch and local collaborative perception branch, the present invention realizes the complementary extraction of low-resolution global semantic information and high-resolution local details; the global perception branch captures semantic context through the hybrid layer and adaptive channel interaction attention module of the MambaVision network, and the local collaborative perception branch utilizes the ResNet18 network and multi-directional local feature enhancement module. The combination of the two improves the comprehensiveness of feature expression. Compared with traditional neural networks, it can more accurately balance global scene understanding and local detail characterization, avoiding problems such as semantic ambiguity or detail loss caused by relying only on single-scale features.

[0019] (4)The dynamic dual-path fusion module designed in the present invention performs dynamic weight adjustment on the outputs of the global perception branch and the local collaborative perception branch through a triple attention mechanism of space, channel, and pixel, achieving adaptive interaction of features. Compared with fixed-weight fusion, this module can intelligently allocate fusion strategies according to the complexity of different regional features, enhancing the segmentation accuracy of the model and avoiding the problem that key information is diluted or disturbed by secondary information caused by the traditional fixed fusion method's inability to perceive the difference in feature importance, significantly improving the accuracy and detail richness of the segmentation results.

[0020] (5)The present invention introduces a loss function based on hard sample mining. By dynamically screening hard samples with predicted probabilities lower than the threshold to calculate the loss and combining class weight to balance the differences, it avoids simple samples dominating the training, effectively alleviates the segmentation bias caused by imbalance, and improves the recognition ability of small target roads. Description of the Drawings

[0021] Figure 1 It is a processing flow chart of the road segmentation model of the present invention. Detailed Embodiment

[0022] The present invention provides a technical solution: a road segmentation method based on local-global dual-stream information collaborative perception, including:

[0023] Step S1: Construct a data set, which includes several high-resolution road images, and the road areas in the high-resolution road images are all labeled.

[0024] Select the KITTI data set as the core data source. The KITTI data set is collected from rich and diverse real-world scenarios, covering typical traffic environments such as urban streets, rural roads, and highways. It not only completely records real road images but also synchronously collects corresponding radar information, fully restoring the details of the road scene. The road types in the KITTI data set can be classified into three categories: urban unmarked roads UU (urban unmarked), urban marked roads UM (urban marked) with clear single-lane markings, and urban multiply marked roads UMM (urban multiply marked) with multiple lane markings. This detailed classification provides accurate data support for the training and testing of the subsequent road segmentation model in different road scenarios.

[0025] Step S2: Construct a road segmentation model, as Figure 1 shown; the road segmentation model consists of a dynamic shadow synthesis data augmentation module, a local collaborative perception branch, a global perception branch, and a dynamic dual-path fusion module.

[0026] Step S3: Input the high-resolution road image into the dynamic shadow synthesis data augmentation module. First, introduce elliptical or tree-shaped shadows to the high-resolution road image to obtain a high-resolution road enhanced image, and then downscale the high-resolution road enhanced image to obtain a low-resolution road enhanced image.

[0027] Among them, the dynamic shadow synthesis data augmentation module aims to simulate the complex situation of roads covered by shadows in real scenes. Through a probability-driven dual-mode shadow addition strategy, it enhances the adaptability and robustness of the road segmentation model to illumination changes.

[0028] The specific process of introducing elliptical or tree-shaped shadows to the high-resolution road image to obtain a high-resolution road enhanced image is as follows: Step S3.11: Randomly select different types of shadows (elliptical shadow or tree shadow) with a 50% probability, expressed as: ; where, represents the type of shadow; represents an elliptical shadow; represents a tree shadow; represents a random number between 0 and 1.

[0029] After determining the shadow type, based on the label information of the road area in the high-resolution road image, dynamically generate the enhancement parameters of the corresponding type of shadow; specifically: For an elliptical shadow, first randomly sample the center coordinates of the elliptical shadow within the road area in the high-resolution road image, and generate the random size parameters of the elliptical shadow, represents the horizontal semi-axis length of the elliptical shadow, represents the vertical semi-axis length of the elliptical shadow; These two parameters jointly determine the size and shape of the elliptical shadow; in addition, a transparency parameter for controlling the depth of the elliptical shadow will also be generated.

[0030] For a tree shadow, first set a fixed template with a size of 400×200 in the road area of the high-resolution road image, stack three predefined elliptical shadows on the fixed template with misalignment to generate a tree shadow; then calculate the scaling factor according to the road direction in the high-resolution road image to control the size of the fixed template, and finally generate a transparency parameter for controlling the depth of the tree shadow and a horizontal flip flag for controlling the horizontal flip of the tree shadow.

[0031] Among them, the three predefined elliptical shadows include the central elliptical shadow, the left elliptical shadow, and the right elliptical shadow; the central coordinates of the central elliptical shadow are (200, 67), its horizontal semi-axis length is 133, and its vertical semi-axis length is 50; the central coordinates of the left elliptical shadow are (100, 100), its horizontal semi-axis length is 80, and its vertical semi-axis length is 67; the central coordinates of the right elliptical shadow are (300, 100), its horizontal semi-axis length is 80, and its vertical semi-axis length is 67.

[0032] The process of dynamically generating the enhancement parameters of the corresponding type of shadow can be expressed as: ; ; ; In the formula, represents the enhancement parameter of the generated corresponding type of shadow; represents the enhancement parameter of the generated elliptical shadow; represents the enhancement parameter of the generated tree shadow; represents the label information of the road area in the high-resolution road image; represents the position coordinates of the fixed template within the road area in the high-resolution road image.

[0033] Step S3.13: Based on the enhancement parameter of the generated corresponding type of shadow, call the elliptical shadow mask generation function or the tree shadow mask generation function to generate a shadow mask of the corresponding shape , which can be expressed as: ; In the formula, represents the elliptical shadow mask generation function; represents the tree shadow mask generation function; represents the coordinates of any pixel point in the high-resolution road image.

[0034] Among them, the elliptical shadow mask generation function can be expressed as: ; Among them, the tree shadow mask generation function can be expressed as: ; In the formula, represents the resize operation, which is used to resize the fixed template according to ; represents the flip operation, which is used to flip the fixed template according to ; represents the fixed template; represents the coordinates of any pixel point in the fixed template.

[0035] Among them, is expressed as: .

[0036] Step S3.14: Through the weighted fusion formula, overlay the shadow mask and the high-resolution road image according to the transparency parameter to output the high-resolution road enhanced image .

[0037] This process not only simulates the morphological changes of real shadows, but also ensures the diversity of each enhancement through dynamic parameter adjustment, thereby effectively improving the segmentation performance of the model in complex lighting scenarios.

[0038] Among them, shrinking the high-resolution road enhanced image to obtain the low-resolution road enhanced image The specific process is as follows: The size of the high-resolution road enhanced image is 1240×375. The size of the high-resolution road enhanced image is adjusted to 310×188 through the bilinear interpolation method to obtain the low-resolution road enhanced image .

[0039] Step S4: Obtain the global semantic feature map of the low-resolution road enhanced image through the global perception branch, and obtain the local detail feature map of the high-resolution road enhanced image through the local collaborative perception branch.

[0040] Among them, the local collaborative perception branch is a three-layer structure. The first layer of the local collaborative perception branch uses the initialization layer of the ResNet18 network, the second layer of the local collaborative perception branch uses the residual layer, and the third layer of the local collaborative perception branch uses the multi-directional local feature enhancement module.

[0041] Among them, the global perception branch is a five-layer structure. The first layer of the global perception branch uses the PatchEmbed module, the second, third, and fourth layers of the global perception branch respectively use the first, second, and third layers of the MambaVision network, and the fifth layer of the global perception branch uses the adaptive channel interaction attention module.

[0042] Among them, the specific process of the global perception branch obtaining the global semantic feature map of the low-resolution road enhanced image is as follows: Step S4.11: Input the low-resolution road enhanced image into the PatchEmbed module to obtain the embedded feature ; among them, the PatchEmbed module consists of two convolutional layers (3×3) with a stride of 2; Pass through two convolutional layers (3×3) with a stride of 2 in sequence to obtain the embedded features , which is expressed as: ; ; In the formula, represents the output after the first convolutional layer in the PatchEmbed module; represents the activation function; represents the batch normalization operation; represents the 3×3 convolutional layer.

[0043] Step S4.12: Input the embedded features into the first layer of the MambaVision network to obtain the first feature map ; Among them, the first layer of the MambaVision network consists of a convolutional block, and the convolutional block contains two convolutional layers (3×3) with a stride of 1; Input the embedded features into two convolutional layers (3×3) with a stride of 1 in sequence to obtain the intermediate feature map , perform regularization on the intermediate feature map using the stochastic depth operation, add the output after regularization to the embedded features to obtain the first feature map , which is expressed as: ; ; ; In the formula, represents the embedded features after the first convolutional layer (3×3) with a stride of 1 in the convolutional block of the MambaVision network; represents the activation function; represents the stochastic depth operation.

[0044] Step S4.13: Perform downsampling on the first feature map to obtain the first downsampled feature map , which is expressed as: . Step S4.14: Input the first downsampled feature map into the second layer of the MambaVision network to obtain the second feature map ; Among them, the second layer of the MambaVision network consists of three sequentially connected convolutional blocks, and the convolutional blocks within the second layer of the MambaVision network have the same structure as the convolutional blocks within the first layer of the MambaVision network; the first downsampled feature map After passing through the three convolutional blocks within the second layer of the MambaVision network in sequence, a second feature map is obtained .

[0045] Step S4.14: Perform a downsampling operation on the second feature map to obtain a second downsampled feature map , which is expressed as: ; Step S4.15: Input the second downsampled feature map into the third layer of the MambaVision network to obtain a third feature map ; Among them, the third layer of the MambaVision network consists of a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence; first, flatten the second downsampled feature map , and the flattened pass through the first hybrid module, the second hybrid module, and the third hybrid module in sequence to obtain a third feature map ; Specifically: Flatten the second downsampled feature map in the spatial dimension to obtain a feature sequence , which is expressed as: ; In the formula, represents the flattening operation.

[0046] Input the feature sequence into the first hybrid module to obtain the output of the first hybrid module ; Among them, the first hybrid module consists of a first MLP branch and a Mamba branch connected in parallel.

[0047] The first MLP branch consists of two fully connected layers; input the feature sequence through the two fully connected layers in sequence to obtain a first feature enhancement sequence , that is, the output of the first MLP branch, which is expressed as: ; In the formula, represents the first fully connected layer in the first MLP branch; represents the second fully connected layer in the first MLP branch.

[0048] In the Mamba branch, the feature sequence Dimensionality is increased through a linear transformation layer; the feature sequence after dimensionality increase passes through a one-dimensional convolutional layer to obtain a second feature enhancement sequence ; the second feature enhancement sequence generates a gating signal through a linear transformation layer ( taking values between 0 and 1, which is equivalent to assigning a learnable "switch" to each feature dimension, dynamically determining which information needs to be retained and which information can be ignored, thereby enhancing the model's attention to important features and suppressing noise); through the gating signal to control the information flow of the second feature enhancement sequence to obtain a third feature enhancement sequence , expressed as: ; ; ; In the formula, represents the activation function; represents the one-dimensional convolutional layer; represents the linear transformation layer; represents the element-wise multiplication operation.

[0049] Performs a selective scanning operation on the third feature enhancement sequence . In this process, the Mamba branch combines the third feature enhancement sequence input at the current moment with the historical state (the third feature enhancement sequence input at the previous moment), through a learnable time step parameter ranging from 0 to 1, to dynamically adjust the fusion ratio of new and old information (when is close to 1, it will pay more attention to the input at the current moment, and when is close to 0, it will rely more on the third feature enhancement sequence input at the previous moment); the third feature enhancement sequence with the adjusted fusion ratio is restored to the original dimension through a linear transformation layer and concatenated with the second feature enhancement sequence to obtain a fourth enhanced feature sequence , which is the output of the Mamba branch, expressed as: ; ; In the formula, represents the concatenation operation.

[0050] After passing the fourth enhanced feature sequence through a linear transformation layer and concatenating it with the first feature enhancement sequence in the dimension, a fifth enhanced feature sequence is obtained . The fifth enhanced feature sequence is passed through the first MLP branch and undergoes a random path dropout operation, and then added element-wise to the feature sequence to obtain the output of the first mixing module , which is expressed as: ; ; In the formula, represents the first MLP branch.

[0051] The output of the first mixing module is input into the second mixing module; the second mixing module consists of a parallel self-attention branch (Attention) and a second MLP branch; the second MLP branch of the second mixing module has the same structure as the first MLP branch of the first mixing module, and its processing flow will not be elaborated here; the output of the first mixing module is respectively passed through the self-attention branch and the second MLP branch to obtain the output of the self-attention branch and the output of the second MLP branch, and the output of the self-attention branch and the output of the second MLP branch are concatenated to obtain a sixth feature enhancement sequence . The sixth feature enhancement sequence is passed through the second MLP branch and undergoes a random path dropout operation, and then added element-wise to the output of the first mixing module to obtain the output of the second mixing module .

[0052] The structure of the third mixing module is the same as that of the second mixing module, and its processing flow will not be elaborated here.

[0053] The output of the third mixing module is converted from the feature sequence to a feature map to obtain the third feature map .

[0054] Step S4.16: Input the third feature map into the adaptive channel interaction attention module (ECA) to obtain the global semantic feature map .

[0055] The role of the adaptive channel interaction attention module (ECA) is to adaptively adjust the importance of each channel of the feature map, effectively enhance the distinction between the road area and the background, and improve the segmentation accuracy.

[0056] In the adaptive channel interaction attention module, the spatial dimension of the third feature map is compressed and transformed through global average pooling to obtain a feature vector ; After passing the feature vector through a one-dimensional convolutional layer, its dimension is expanded. The output after dimension expansion is multiplied element-wise with the third feature map to obtain the global semantic feature map . .

[0057] Among them, the specific process of obtaining the local detail feature map of the high-resolution road enhancement image through the local collaborative perception branch is as follows: Step S4.21: Input the high-resolution road enhancement image into the initialization layer of the ResNet18 network to obtain the fourth feature map ; Among them, the initialization layer of the ResNet18 network consists of a convolutional layer (7×7) with a stride of 2 and a padding of 3 and a max pooling layer; The high-resolution road enhancement image passes through the convolutional layer (7×7) and the max pooling layer in the initialization layer of the ResNet18 network in sequence to obtain the fourth feature map , which is expressed as: ; ; In the formula, represents the max pooling layer; represents the 7×7 convolutional layer.

[0058] Step S4.22: Input the fourth feature map into the residual layer to obtain the output of the residual layer ; Among them, the residual layer consists of a first residual block and a second residual block, and the structures of the first residual block and the second residual block are the same, both consisting of two convolutional layers (3×3) with a stride of 1; The fourth feature map passes through two convolutional layers (3×3) with a stride of 1 in the first residual block to obtain the output of the first residual block. The output of the first residual block passes through two convolutional layers (3×3) with a stride of 1 in the second residual block and is added to the output of the first residual block to obtain the output of the residual layer .

[0059] Step S4.23: Input the output of the residual layer into the multi-directional local feature enhancement module to obtain the local detail feature map ; The function of the multi-directional local feature enhancement module is to enhance the local structural information related to the road in the extracted feature map through multi-directional feature extraction, local attention mechanism and residual connection, while retaining spatial details.

[0060] Among them, the multi-directional local feature enhancement module includes a depthwise separable convolutional layer, a convolutional layer, and a local attention mechanism; the output of the residual layer is successively passed through the depthwise separable convolutional layer in the horizontal direction and a convolutional layer to obtain horizontal features ; the output of the residual layer is successively passed through the depthwise separable convolutional layer in the vertical direction and a convolutional layer to obtain vertical features ; the horizontal features , vertical features and the output of the residual layer are added together to obtain multi-directional fusion features , expressed as: ; ; ; In the formula, represents the depthwise separable convolutional layer with a convolution kernel of , that is, the depthwise separable convolutional layer in the horizontal direction; represents the depthwise separable convolutional layer with a convolution kernel of , that is, the depthwise separable convolutional layer in the vertical direction.

[0061] Generate the channel weights of the multi-directional fusion features through the local attention mechanism, multiply the channel weights by the output of the residual layer and add it to the multi-directional fusion features to obtain the local detail feature map .

[0062] Step S5: After aligning the global semantic feature map , add it element-wise to the local detail feature map to obtain the preliminary fusion feature map , and input the preliminary fusion feature map into the dynamic dual-path fusion module to obtain the road segmentation map .

[0063] Among them, the dynamic dual-path fusion module consists of a spatial attention module, a channel attention module, an edge auxiliary branch, a pixel attention module, and a segmentation head; the process of inputting the preliminary fusion feature map into the dynamic dual-path fusion module to obtain the road segmentation map is as follows: Input the preliminary fusion feature map into the spatial attention module to obtain the spatial attention map , expressed as: ; In the formula, represents the average pooling layer; represents the max pooling layer.

[0064] To further enhance the sensitivity to road edges, an edge assistance branch is introduced to enhance the spatial attention map ; specifically, the spatial attention map and the preliminary fusion feature map are input into the edge assistance branch to obtain the edge-enhanced spatial attention map , which is expressed as: ; In the formula, represents function.

[0065] The preliminary fusion feature map is input into the channel attention module to obtain the channel attention map .

[0066] The edge-enhanced spatial attention map and the channel attention map are added through a broadcast operation to obtain the spatial-channel attention map , which is expressed as: ; In the formula, represents the broadcast operation.

[0067] The spatial-channel attention map and the preliminary fusion feature map are simultaneously input into the pixel attention module to obtain the spatial-channel-pixel attention map , which is expressed as: ; In the formula, represents grouped convolutional layer.

[0068] Through the spatial-channel-pixel attention map the global semantic feature map and the local detail feature map are dynamically adaptively fused, and the output after dynamic adaptive fusion is added to the preliminary fusion feature map and then input into the segmentation head to obtain the road segmentation map , which is expressed as: ; In the formula, Indicates a splitting head, which consists of two consecutively connected convolutional layers (1×1).

[0069] In the road segmentation task, the cross-entropy loss function is a commonly used optimization method, but it has limitations in dealing with the problem of sample imbalance. Taking road segmentation as an example, the number of pixels in the road area is often significantly less than that in the non-road area. This makes it possible that in the model training process, the non-road area may dominate the model learning direction, resulting in the model being more likely to misclassify pixels as non-road areas during pixel classification, thereby affecting the segmentation accuracy. To solve this problem, a loss function based on hard sample mining is introduced As the loss function of the road segmentation model, this loss function effectively improves the ability to handle the problem of sample imbalance by dynamically screening hard samples and calculating the cross-entropy loss; DHM-CELoss is expressed as: ; In the formula, represents the predicted probability distribution of the road segmentation model; represents the true category; represents the set of selected hard samples (high-resolution road images); represents the corresponding class weight; represents the negative logarithm function; represents the probability distribution of the road segmentation model predicting the th sample; represents the th sample belonging to probability; represents the th sample's true category; represents the th sample belonging to category probability; represents the total number of categories.

[0070] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A road segmentation method based on collaborative perception of local-global dual-stream information, characterized in that, It includes the following steps: Step S1: Construct a dataset, which includes several high-resolution road images, and the road areas in the high-resolution road images are all labeled; Step S2: Construct a road segmentation model; the road segmentation model is composed of a dynamic shadow synthesis data augmentation module, a local collaborative perception branch, a global perception branch, and a dynamic dual-path fusion module; Step S3: Input the high-resolution road image into the dynamic shadow synthesis data augmentation module. First, introduce an ellipse or tree-shaped shadow to the high-resolution road image to obtain a high-resolution road enhanced image, and then downscale the high-resolution road enhanced image to obtain a low-resolution road enhanced image; Step S4: Obtain the global semantic feature map of the low-resolution road enhanced image through the global perception branch, and obtain the local detail feature map of the high-resolution road enhanced image through the local collaborative perception branch; Step S5: After aligning the global semantic feature map, add it element-wise to the local detail feature map to obtain a preliminary fusion feature map, and input the preliminary fusion feature map into the dynamic dual-path fusion module to obtain a road segmentation map.

2. The road segmentation method based on local-global two-stream information collaborative perception according to claim 1, wherein: The dynamic dual-path fusion module consists of a spatial attention module, a channel attention module, an edge assistance branch, a pixel attention module, and a segmentation head; the preliminary fusion feature map is input into the dynamic dual-path fusion module to obtain a road segmentation map The specific process is as follows: Input the preliminary fusion feature map into the spatial attention module to obtain the spatial attention map , which is expressed as: ; In the formula, represents the convolutional layer; represents the concatenation operation; represents the average pooling layer; represents the max pooling layer; Input the spatial attention map and the preliminary fusion feature map into the edge assistance branch to obtain the edge-enhanced spatial attention map , which is expressed as: ; In the formula, represents a function; represents a convolutional layer; Input the preliminary fusion feature map into the channel attention module to obtain the channel attention map ; Enhanced spatial attention map of edges and channel attention map are added together through a broadcast operation to obtain a spatial-channel attention map ; Input the spatial-channel attention map and the preliminary fusion feature map into the pixel attention module simultaneously to obtain the spatial-channel-pixel attention map , which is expressed as: ; In the formula, represents a grouped convolution layer; represents a batch normalization operation; Through the spatial-channel-pixel attention map for the global semantic feature map and the local detail feature map perform dynamic adaptive fusion, add the output after dynamic adaptive fusion to the preliminary fusion feature map and then input it into the segmentation head to obtain the road segmentation map , which is expressed as: ; In the formula, represents a splitting head.

3. The road segmentation method based on local-global dual-stream information collaborative perception according to claim 2, wherein: The specific process of introducing an ellipse or tree-shaped shadow to the high-resolution road image to obtain a high-resolution road enhanced image is as follows: Step S3.11: Randomly select different types of shadows, and different types of shadows include ellipse shadows and tree shadows; Step S3.12: After determining the shadow type, based on the label information of the road area in the high-resolution road image, dynamically generate the enhancement parameters of the corresponding type of shadow; Step S3.13: Based on the enhanced parameters of the generated corresponding type of shadow, call the elliptical shadow mask generation function or the tree shadow mask generation function to generate a shadow mask of the corresponding shape , expressed as: ; In the formula, represents the elliptical shadow mask generation function; represents the tree shadow mask generation function; represents the coordinates of any pixel point in the high-resolution road image; represents the type of shadow; represents the enhancement parameter of the generated elliptical shadow; represents the enhancement parameter of the generated tree shadow; represents the label information of the road area in the high-resolution road image; represents the elliptical shadow; Step S3.14: Superimpose the shadow mask and the high-resolution road image according to the transparency parameter through the weighted fusion formula to output a high-resolution road enhanced image .​ 4. The road segmentation method based on local-global dual-stream information collaborative perception according to claim 3, wherein: The specific process of Step S3.12 is as follows: For the elliptical shadow, first randomly sample the central coordinates of the elliptical shadow within the road area in the high-resolution road image ; Random size parameters for generating an elliptical shadow , represents the horizontal semi-axis length of the elliptical shadow, represents the vertical semi-axis length of the elliptical shadow; generate a transparency parameter for controlling the depth of the elliptical shadow ; For the tree shadow, first, set a fixed template of a preset size in the road area of the high-resolution road image, and stack three predefined elliptical shadows on the fixed template with misalignment to generate the tree shadow; then calculate the scaling factor to control the size of the fixed template, and finally generate the transparency parameter for controlling the depth of the tree shadow and the horizontal flip flag for controlling the horizontal flip of the tree shadow ; Among them, the three predefined ellipse shadows include a central ellipse shadow, a left ellipse shadow, and a right ellipse shadow; The process of dynamically generating the enhancement parameters of the corresponding type of shadow is expressed as: ; ; ; In the formula, represents the enhancement parameter of the generated shadow of the corresponding type; represents the fixed template The position coordinates within the road area in the high-resolution road image; The ellipse shadow mask generation function is expressed as: ; Tree shadow mask generation function Expressed as: ; In the formula, represents a resizing operation for resizing a fixed template according to ; represents a flipping operation for flipping the fixed template according to ; represents the fixed template; represents the coordinates of any pixel point in the fixed template.

5. A road segmentation method based on local-global dual-stream information collaborative perception according to claim 4, characterized in that: The local collaborative perception branch is a three-layer structure. The first layer of the local collaborative perception branch uses the initialization layer of the ResNet18 network, the second layer of the local collaborative perception branch uses a residual layer, and the third layer of the local collaborative perception branch uses a multi-directional local feature enhancement module; The global perception branch is a five-layer structure. The first layer of the global perception branch uses a PatchEmbed module, the second, third, and fourth layers of the global perception branch respectively use the first, second, and third layers of the MambaVision network, and the fifth layer of the global perception branch uses an adaptive channel interaction attention module.

6. The road segmentation method based on collaborative perception of local-global dual-stream information according to claim 5, characterized in that: The specific process of the global perception branch obtaining the global semantic feature map of the low-resolution road enhanced image is as follows: Step S4.11: Input the low-resolution road enhanced image into the PatchEmbed module to obtain the embedded features ; where the PatchEmbed module consists of two convolutional layers with a stride of 2; Input sequentially through two convolutional layers with a stride of 2 to obtain the embedded features ; Step S4.12: Input the embedded feature into the first layer of the MambaVision network to obtain the first feature map ; where the first layer of the MambaVision network consists of a convolutional block, and the convolutional block contains two convolutional layers with a stride of 1; the embedded feature passes through two convolutional layers with a stride of 1 in sequence to obtain an intermediate feature map , and perform regularization on the intermediate feature map using the random path dropout operation, and add the output after regularization to the embedded feature to obtain the first feature map ; Step S4.13: Perform downsampling on the first feature map to obtain the first downsampled feature map ; Step S4.14: Input the first downsampled feature map into the second layer of the MambaVision network to obtain the second feature map ; where the second layer of the MambaVision network consists of three convolutional blocks connected in sequence, and the convolutional blocks in the second layer of the MambaVision network have the same structure as the convolutional blocks in the first layer of the MambaVision network; the first downsampled feature map passes through the three convolutional blocks in the second layer of the MambaVision network in sequence to obtain the second feature map ; Step S4.14: Perform downsampling on the second feature map to obtain a second downsampled feature map ; Step S4.15: Input the second downsampled feature map into the third layer of the MambaVision network to obtain the third feature map ; wherein, the third layer of the MambaVision network consists of a first hybrid module, a second hybrid module, and a third hybrid module connected in sequence; first, flatten the second downsampled feature map , and pass the flattened through the first hybrid module, the second hybrid module, and the third hybrid module in sequence to obtain the third feature map ; Step S4.16: Input the third feature map into the adaptive channel interaction attention module to obtain the global semantic feature map .

7. The road segmentation method based on local-global two-stream information collaborative perception according to claim 6, characterized in that: The specific process of Step S4.15 is as follows: For the second downsampled feature map Flatten it in the spatial dimension to obtain a feature sequence ; Input the feature sequence into the first mixing module to obtain the output of the first mixing module ; wherein, the first mixing module is composed of a first MLP branch and a Mamba branch in parallel; The first MLP branch consists of two fully connected layers; the feature sequence passes through the two fully connected layers in sequence to obtain the first feature enhancement sequence , which is the output of the first MLP branch; In the Mamba branch, the feature sequence is dimensionally upsampled through a linear transformation layer; the dimensionally upsampled feature sequence passes through a one-dimensional convolutional layer to obtain a second feature enhancement sequence ; the second feature enhancement sequence generates a gating signal through a linear transformation layer; the information flow of the second feature enhancement sequence is controlled by the gating signal to obtain a third feature enhancement sequence ; Perform a selective scanning operation on the third feature enhancement sequence through a learnable time step parameter ranging from 0 to 1 to dynamically adjust the fusion ratio between the third feature enhancement sequence input at the current moment and the third feature enhancement sequence input at the previous moment; pass the third feature enhancement sequence with the adjusted fusion ratio through a linear transformation layer to restore the original dimension and then splice it with the second feature enhancement sequence to obtain the fourth enhanced feature sequence , which is the output of the Mamba branch; The fourth enhanced feature sequence After passing through a linear transformation layer, it is concatenated with the first feature enhancement sequence in the dimension to obtain the fifth enhanced feature sequence . The fifth enhanced feature sequence After passing through the first MLP branch and performing a random path dropout operation, it is element-wise added to the feature sequence to obtain the output of the first mixing module ; Feed the output of the first mixing module into the second mixing module; the second mixing module consists of a self-attention branch and a second MLP branch in parallel; the second MLP branch of the second mixing module has the same structure as the first MLP branch of the first mixing module; feed the output of the first mixing module through the self-attention branch and the second MLP branch respectively to obtain the output of the self-attention branch and the output of the second MLP branch, concatenate the output of the self-attention branch and the output of the second MLP branch to obtain the sixth feature enhancement sequence , feed the sixth feature enhancement sequence through the second MLP branch and perform a random path dropout operation, then add it element-wise to the output of the first mixing module to obtain the output of the second mixing module ; The structure of the third mixing module is the same as that of the second mixing module; the output of the second mixing module is input into the third mixing module to obtain the output of the third mixing module ; The output of the third mixing module is converted from the feature sequence into a feature map to obtain the third feature map .

8. A road segmentation method based on local-global dual-stream information collaborative perception according to claim 7, characterized in that: Obtain a high-resolution road enhanced image through the local collaborative perception branch of the local detail feature map The specific process is as follows: Step S4.21: Input the high-resolution road enhanced image into the initialization layer of the ResNet18 network to obtain the fourth feature map ; where the initialization layer of the ResNet18 network consists of a convolutional layer with a stride of 2 and a padding of 3 and a max pooling layer; pass the high-resolution road enhanced image through the convolutional layer and the max pooling layer in the initialization layer of the ResNet18 network in sequence to obtain the fourth feature map ; Step S4.22: Input the fourth feature map into the residual layer to obtain the output of the residual layer ; where the residual layer consists of a first residual block and a second residual block, and the first residual block and the second residual block have the same structure, both consisting of two convolutional layers with a stride of 1; input the fourth feature map through the two convolutional layers with a stride of 1 in the first residual block to obtain the output of the first residual block, and add the output of the first residual block after passing through the two convolutional layers with a stride of 1 in the second residual block to the output of the first residual block to obtain the output of the residual layer ; Step S4.23: Output of the residual layer is input into the multi-directional local feature enhancement module to obtain a local detail feature map .

9. A road segmentation method based on local-global dual-stream information collaborative perception according to claim 8, characterized in that: The multi-directional local feature enhancement module includes a depthwise separable convolutional layer, a convolutional layer, and a local attention mechanism; the output of the residual layer is input into the multi-directional local feature enhancement module to obtain a local detail feature map The specific process is as follows: The output of the residual layer is sequentially passed through a depthwise separable convolutional layer in the horizontal direction and a convolutional layer to obtain horizontal features ; The output of the residual layer is sequentially passed through a depthwise separable convolutional layer in the vertical direction and a convolutional layer to obtain vertical features ; The horizontal features , vertical features and the output of the residual layer are added together to obtain multi-directional fusion features ; Generate multi-directional fusion features through local attention mechanism Channel weights Multiply the channel weights with the output of the residual layer and then add the result to the multi-directional fusion features to obtain the local detail feature map .

Citation Information

Patent Citations

  • Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism

    CN117078943A

  • Double-resolution real-time semantic segmentation method based on detail enhancement

    CN117409412A

  • Multi-scale double-flow fusion real-time semantic segmentation method for road scene

    CN119888222A

  • Luminous road surface marker and method for its execution

    JP1998037141A

  • Image segmentation and segmentation network training method and apparatus, device, medium, and product

    WO2019238126A1

Cited By

  • Power transmission line external damage hidden danger early warning method based on visual device

    CN121053547A

  • A power transmission line external damage hidden danger early warning method based on a visualization device

    CN121053547B

  • Single composite image shadow generation method based on 3D perception

    CN121353508A