Semantic map road segmentation method and map construction method and system in unmanned aerial vehicle scene

By building a lightweight road segmentation network model based on the fusion of boundary constraints and multi-scale features, the problem of real-time construction of semantic maps under unmanned airport scenes is solved, efficient and precise road segmentation is achieved, and drone combat efficiency is improved.

CN120259649APending Publication Date: 2025-07-04AEROSPACE TIMES FEIHONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510257551.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The real-time construction of semantic maps under unmanned airport scenes faces the problems of high image resolution, diversified road structure scales and unclear boundaries, and existing algorithms are difficult to take into account both real-time and accuracy.

Method used

A lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion is constructed, including basic convolution modules, short-dense duct modules and bilateral guided fusion modules, optimize the network using boundary constraint loss function, extract and fuse low-dimensional detail features and high-dimensional semantic features.

Benefits of technology

Without increasing the amount of network computing parameters, the semantic map road boundary segmentation performance in unmanned airport scenes is improved, and efficient real-time map construction and precise road segmentation are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259649A_ABST
    Figure CN120259649A_ABST
Patent Text Reader

Abstract

The invention relates to the field of real-time map segmentation, and provides a semantic map road segmentation method and a map construction method and system in an unmanned aerial vehicle scene, and the semantic map road segmentation method comprises the steps: S1, constructing a lightweight road segmentation network model based on boundary constraint and multi-scale feature fusion; s2, training the lightweight road segmentation network model to obtain an optimized model; and S3, carrying out road segmentation on the unmanned aerial vehicle map by using the optimized model. The lightweight road segmentation network model comprises a basic convolution module, a short-circuit dense splicing module and a bilateral guide fusion module. The accuracy and prediction efficiency of the algorithm can be considered at the same time; the calculation parameter quantity of the network structure is small, and the low-dimensional detail features and the high-dimensional semantic features of the image can be efficiently fused; the boundary constraint loss function participates in network optimization, and the road boundary segmentation performance of the network can be improved on the premise that the network calculation parameter quantity is not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of real-time map segmentation, and particularly to a method for semantic map road segmentation, a map construction method and a system in the scenario of an unmanned aerial vehicle (UAV). Background Art

[0002] Military UAVs are high-performance information-based weapons and an important and indispensable part of modern information-based warfare. The development of UAVs has attracted great attention from various countries. UAVs can perform tasks such as reconnaissance, surveillance, attack, and interception in combat scenarios. Among them, the real-time construction of semantic maps in the captured scenarios plays a crucial role in the judgment of commanders. However, the real-time construction of semantic maps in the UAV scenario still faces severe challenges. The resolution of UAV-captured images is high, and general intelligent algorithms are difficult to meet the real-time requirements; due to the large variation range of the flight altitude, angle, etc. of UAVs, the road structures in the captured images have scale diversity; at the same time, some road structures may be blocked by surrounding trees, resulting in unclear boundaries.

[0003] Aiming at the difficulties existing in the current road semantic segmentation in the UAV scenario, how to accurately and real-time construct a semantic map during the flight of the UAV is a key technology for obtaining combat situation perception information in a complex battlefield environment. The construction of a semantic map can assist the commander in planning the flight route of the UAV and greatly improve the combat efficiency. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and provide a method for semantic map road segmentation, a map construction method and a system in the UAV scenario, which can take into account both the accuracy and prediction efficiency of the algorithm, and improve the segmentation performance of the network for road boundaries without increasing the number of network calculation parameters.

[0005] The present invention adopts the following technical solutions:

[0006] On the one hand, the present invention provides a method for semantic map road segmentation in the UAV scenario, including:

[0007] S1. Construct a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion;

[0008] S2. Train the lightweight road segmentation network model in step S1 to obtain an optimized model;

[0009] S3. Use the optimized model to segment the roads in the UAV map.

[0010] In any of the above possible implementation manners, a further implementation manner is provided. In step S1, the lightweight road segmentation network model includes a basic convolution module, a short-circuit dense splicing module STDC, and a bilateral guidance fusion module;

[0011] The basic convolution module is used to extract low-dimensional detailed features of UAV images;

[0012] There are multiple short-circuit dense splicing modules STDC connected in sequence, which are used to extract image features without increasing the network width. The image features include low-dimensional detailed features and high-dimensional semantic features; as the network depth increases, the extracted high-dimensional semantic features become more obvious;

[0013] The bilateral guidance fusion module fuses the low-dimensional detailed features and the high-dimensional semantic features, and uses the fused features to predict the segmentation result of the UAV image.

[0014] In any of the possible implementation manners described above, a further implementation manner is provided. Each STDC includes 3 independent convolution modules. The output of the previous level of the STDC first passes through a convolution module with a convolution kernel size of 1×1, and the convolution stride is set to 1, which is used to reduce the feature dimension; then, the dimension-reduced features are input into a convolution module with a convolution kernel size of 3×3, and the convolution stride is 1 or 2; the convolution kernel size of the third convolution module is 3×3, and the convolution stride is 1;

[0015] The outputs of each convolution module are jump-connected as the output of the STDC. The fused feature is expressed as:

[0016] x out = Cat(x1, x2, x3);

[0017] Among them, x1, x2, and x3 respectively represent the features extracted by the three convolution modules of the STDC, and Cat(g) represents the feature fusion operator.

[0018] In any of the possible implementation manners described above, a further implementation manner is provided. The specific method of step S2 includes:

[0019] S21. The UAV image passes through 2 basic convolution modules to extract low-dimensional detailed features;

[0020] S22. The output after passing through 2 basic convolution modules is used as the input of the short-circuit dense splicing module, and passes through 4 STDCs in sequence. The features extracted by the first STDC are directly input into the bilateral guidance fusion module as the low-dimensional detailed feature map, and the features extracted by the fourth STDC are input into the bilateral guidance fusion module as the high-dimensional semantic feature map;

[0021] S23. The bilateral guidance module fuses the low-dimensional detailed feature map and the high-dimensional semantic feature map to obtain the final fused feature map:

[0022] S24. Obtain the predicted segmentation result of the UAV image based on the final fused feature map;

[0023] S25. Calculate the loss function value between the predicted segmentation result and the actual segmentation result, and adjust the network parameters according to the loss function value;

[0024] S26. Repeat steps S21 - S25 until the loss function value is within the set threshold, and the model training ends. At this time, the network parameters are the optimized model parameters.

[0025] In any of the above - mentioned possible implementation manners, a further implementation manner is provided. In step S22, in the 1st, 3rd, and 4th STDCs, each STDC includes 3 independent convolutional modules. The output of the previous level first passes through a convolutional module with a kernel size of 1×1, the convolutional stride is set to 1 to reduce the feature dimension; then, the dimension - reduced features are input into a convolutional module with a kernel size of 3×3, and the convolutional stride is 1; the kernel size of the third convolutional module is 3×3, and the convolutional stride is 1;

[0026] In the 2nd STDC module, it includes 3 independent convolutional modules. The output of the previous level first passes through a convolutional module with a kernel size of 1×1, the convolutional stride is set to 1 to reduce the feature dimension; then, the dimension - reduced features are input into a convolutional module with a kernel size of 3×3, and the convolutional stride is 2; the kernel size of the third convolutional module is 3×3, and the convolutional stride is 1; the output of the first convolutional module is downsampled, and the feature scale of the output of the first convolutional module is changed to half of the original by using a global pooling operation with a kernel size of 3×3 and fused with other features.

[0027] In any of the above - mentioned possible implementation manners, a further implementation manner is provided. In step S23, the high - dimensional feature map to be fused performs downsampling operations 2 times on the basis of the low - dimensional feature map to be fused. The low - dimensional feature map sequentially passes through a convolutional module with a stride of 2 and a kernel size of 3×3 and an average pooling operation with a pooling kernel size of 3×3. The high - dimensional feature map sequentially passes through a depth - separable convolutional module with a kernel size of 3×3 and a convolutional module with a kernel size of 1×1, and then the operated low - dimensional feature map and the high - dimensional feature map are fused through a dot - product operation to obtain the first fused feature;

[0028] The high - dimensional feature map is upsampled and then fused with the low - dimensional feature map through dot - product to obtain the second fused feature;

[0029] Add the first fused feature and the second fused feature to obtain the final fused feature map.

[0030] In any of the above - mentioned possible implementation manners, a further implementation manner is provided. In step S25, the loss function is:

[0031]

[0032] Wherein: L BIOU is the boundary IOU loss function, and L cls is the cross-entropy loss function; N represents the total number of predicted pixels, and x i represents a pixel, and Y seg (x i ) and represent the ground truth and the predicted value respectively, and M represents the range of the road boundary concerned by the network.

[0033] On the other hand, the present invention also provides a method for constructing a semantic map in a drone scenario, and the method for constructing a semantic map in a drone scenario uses the above-mentioned method for segmenting roads in a semantic map of a drone scenario for road segmentation.

[0034] On the other hand, the present invention also provides a system for segmenting roads in a semantic map of a drone scenario, and the system is used to implement the above method, and the system includes:

[0035] A lightweight road segmentation network model construction unit that constructs a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion;

[0036] A model training unit that trains the lightweight road segmentation network model using multi-scale drone images to obtain an optimized model;

[0037] A map road segmentation unit that segments roads in a drone map using the optimized model.

[0038] On the other hand, the present invention also provides an information processing terminal for implementing the above method for constructing a semantic map in a drone scenario.

[0039] The beneficial effects of the present invention are:

[0040] 1. It is applicable to the real-time map construction technology in a drone scenario and can balance the accuracy and prediction efficiency of the algorithm at the same time.

[0041] 2. A lightweight road segmentation network based on boundary constraints and multi-scale feature fusion, with a relatively small number of network structure calculation parameters, can efficiently fuse the low-dimensional detail features and high-dimensional semantic features of the image.

[0042] 3. The present invention uses the boundary constraint loss function to participate in optimizing the network, and can improve the road boundary segmentation performance of the network without increasing the number of network calculation parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 The network structure diagram of the lightweight road segmentation network model in the embodiment of the present invention is shown.

[0044] Figure 2 Shown is a schematic diagram of the structure of the short-circuit dense splicing module (STDC) in the embodiment.

[0045] Figure 3 Shown is a schematic diagram of the structure of the bilateral guidance fusion (GAL) module in the embodiment. Detailed implementation manners

[0046] The specific embodiments of the present invention will be described in detail below with reference to the specific drawings. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered in isolation, and they can be combined with each other to achieve better technical effects.

[0047] An embodiment of the present invention provides a method for semantic map road segmentation in a drone scenario, which realizes the real-time construction task of the semantic map in the drone scenario based on a lightweight road segmentation network model with boundary constraints and multi-scale feature fusion, assists the commander to plan the combat flight route of the drone in real time, and provides strong technical support for combat.

[0048] An embodiment of the present invention provides a method for semantic map road segmentation in a drone scenario, including:

[0049] S1. Construct a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion;

[0050] S2. Train the lightweight road segmentation network model in step S1 to obtain an optimized model;

[0051] S3. Use the optimized model to perform road segmentation on the drone map.

[0052] In a specific embodiment, in step S1, the lightweight road segmentation network model includes a basic convolution module, a short-circuit dense splicing module STDC, and a bilateral guidance fusion module; the overall network is as Figure 1 shown;

[0053] The basic convolution module is used to extract low-dimensional detailed features of the drone image;

[0054] There are multiple short-circuit dense splicing modules STDC connected in sequence, which are used to extract image features without increasing the network width. The image features include low-dimensional detailed features and high-dimensional semantic features; as the network depth increases, the extracted high-dimensional semantic features become more obvious;

[0055] The bilateral guidance fusion module fuses the low-dimensional detailed features and the high-dimensional semantic features, and uses the fused features to predict the segmentation result of the drone image.

[0056] In a specific embodiment, there are 2 basic convolution modules. The low-dimensional feature space contains rich detailed features, and the first convolution module uses a small scale and a small stride.

[0057] In the segmentation task, high-dimensional semantic features have a greater impact on the segmentation performance. The acquisition of high-dimensional features often depends on the depth of the network and the size of the receptive field. Due to the diverse flight heights and angles of drones, the road scales in the captured images vary greatly. When designing the network, the problem of multi-scale feature fusion needs to be considered. Therefore, a Short-Term Dense Concatenation module (STDC) is designed, and the structure diagram is as Figure 2 shown.

[0058] In a specific embodiment, each of the STDCs includes 3 independent convolution modules. The output of the previous stage first passes through a convolution module with a convolution kernel size of 1×1, and the convolution stride is set to 1 to reduce the feature dimension and can reduce the number of calculation parameters of the subsequent convolution. Then, the dimension-reduced features are input into a convolution module with a convolution kernel size of 3×3, and the convolution stride is 1 or 2. The convolution kernel size of the third convolution module is 3×3, and the convolution stride is 1;

[0059] Since more attention needs to be paid to the multi-scale problem in road segmentation, the convolution modules of the STDC extract features of different scales. The outputs of each convolution module are skip-connected as the output of the STDC, and the fused feature representation is:

[0060] x out = Cat(x1, x2, x3);

[0061] where x1, x2, and x3 respectively represent the features extracted by the three convolution modules of the STDC, and Cat(g) represents the feature fusion operator.

[0062] Multiple STDCs extract high-dimensional semantic features of drone images. Compared with the classic lightweight network BiSeNetv2, the detailed features and semantic features are extracted in one branch, reducing the number of parameters of the network. At the same time, the low-dimensional features and high-dimensional features are complementary. The STDC fuses context information and extracts multi-scale features of the image.

[0063] In a specific embodiment, the specific method of step S2 includes:

[0064] S21. The drone image passes through 2 basic convolution modules to extract low-dimensional detailed features;

[0065] S22. The output after passing through 2 basic convolution modules is used as the input and sequentially passes through 4 STDCs. The features extracted by the 1st STDC are directly input into the bilateral guidance fusion module as the low-dimensional detailed feature map, and the features extracted by the 4th STDC are input into the bilateral guidance fusion module as the high-dimensional semantic feature map;

[0066] S23. The bilateral guidance module fuses the low-dimensional detail feature map and the high-dimensional semantic feature map:

[0067] The low-dimensional features encode rich detail information, including boundaries, textures, etc. After deep convolution operations, the high-dimensional features become smooth and the semantic features are more explicit. Both forms of feature expressions play crucial roles in the road segmentation task. Therefore, the low-dimensional features and high-dimensional features are fused by the bilateral guidance fusion module before predicting the segmentation result. The low-dimensional features fused in the present invention come from the output of the first STDC. The structural diagram of the bilateral guidance fusion module is as Figure 3 shown;

[0068] The high-dimensional feature map to be fused is downsampled twice based on the low-dimensional feature map to be fused. The low-dimensional feature map sequentially passes through a convolution module with a stride of 2 and a convolution kernel size of 3×3 and an average pooling operation with a pooling kernel size of 3×3. The high-dimensional feature map sequentially passes through a depthwise separable convolution module with a convolution kernel size of 3×3 and a convolution module with a convolution kernel size of 1×1. Then, the low-dimensional feature map after the operation and the high-dimensional feature map are fused through a dot product operation to obtain the first fused feature;

[0069] The high-dimensional feature map is upsampled and then fused with the low-dimensional feature map through a dot product to obtain the second fused feature;

[0070] The first fused feature and the second fused feature are added to obtain the final fused feature map; as Figure 3 shown;

[0071] S24. Obtain the predicted segmentation result of the UAV image according to the final fused feature map;

[0072] S25. Calculate the loss function value between the predicted segmentation result and the actual segmentation result, and adjust the network parameters according to the loss function value;

[0073] S26. Repeat steps S21 - S25 until the loss function value is within the set threshold, and the model training ends. The network parameters at this time are the optimized model parameters.

[0074] In a specific embodiment, in step S22, in the 1st, 3rd, and 4th STDCs, each STDC includes 3 independent convolution modules. The output of the previous level first passes through a convolution module with a convolution kernel size of 1×1 and a convolution stride set to 1 to reduce the feature dimension. Then, the dimension-reduced features are input into a convolution module with a convolution kernel size of 3×3 and a convolution stride of 1. The convolution kernel size of the third convolution module is 3×3 and the convolution stride is 1; as Figure 2 (a) shown;

[0075] In the second STDC, there are three independent convolutional modules. The output of the previous stage first passes through a convolutional module with a kernel size of 1×1, and the convolutional stride is set to 1 to reduce the feature dimension. Then, the feature with reduced dimension is input into a convolutional module with a kernel size of 3×3 and a convolutional stride of 2. The kernel size of the third convolutional module is 3×3 and the convolutional stride is 1. The output of the first convolutional module is downsampled, and the feature scale of the output of the first convolutional module is changed to half of the original by using a global pooling operation with a kernel size of 3×3 and fused with other features. As Figure 2 shown in (b).

[0076] The present invention uses a boundary constraint loss function to make the network pay more attention to the learning of boundaries in the road structure. In the field of segmentation, the segmentation accuracy of object boundaries has always faced challenges. The IoU loss function is widely used, and positive sample pixels are treated equally, and the segmentation performance of boundaries cannot be significantly distinguished. The boundary IoU loss function only considers the prediction accuracy of pixels near the road boundary, which can improve the network's attention to these regions. The cross-entropy loss function is widely used for optimizing image segmentation networks. The present invention combines the boundary IoU loss function and the cross-entropy loss function to constrain the parameter learning of the model. The network is trained in an end-to-end manner.

[0077] In a specific embodiment, in step S25, the loss function is:

[0078] L total = L BIoU + L cls ;

[0079]

[0080] where: L BIOU is the boundary IOU loss function, and L cls is the cross-entropy loss function; N represents the total number of predicted pixels, x i represents a pixel, and Y seg (x i ) and represent the true value and the predicted value respectively, and M represents the range of the road boundary that the network focuses on.

[0081] The method for constructing a semantic map in a UAV scenario according to an embodiment of the present invention uses the above method for segmenting roads in a semantic map of a UAV scenario to perform road segmentation.

[0082] A system for segmenting roads in a semantic map of a UAV scenario according to an embodiment of the present invention is characterized in that the system includes:

[0083] A lightweight road segmentation network model construction unit that constructs a lightweight road segmentation network model based on boundary constraint and multi-scale feature fusion;

[0084] A model training unit that trains a lightweight road segmentation network model using multi-scale drone images to obtain an optimized model;

[0085] A map road segmentation unit that uses the optimized model to segment roads in the drone map.

[0086] An information processing terminal for implementing the method for constructing a semantic map in the drone scenario described above in an embodiment of the present invention.

[0087] In actual tests, four indicators, namely accuracy (Acc), Dice coefficient (Dice), intersection over union (IoU), and frames per second of prediction (FPS), are used to evaluate the performance of the segmentation algorithm. The comparison of the effects of the method of the present invention with other methods is as follows in the table:

[0088]

[0089] It can be seen that the method proposed by the present invention segments more precisely and completely, and has a higher segmentation efficiency.

[0090] The present invention proposes a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion to achieve the real-time construction task of a semantic map in the drone scenario, assist the commander in real-time planning of the combat flight route of the drone, and provide strong technical support for combat. First, the drone image is input into the network, and two basic convolutional modules extract the low-dimensional features of the image. Since the low-dimensional feature space contains rich detailed features, the convolutional modules use small scales and small strides. Secondly, four STDC modules are used to extract the high-dimensional semantic features of the drone image. Compared with the classic lightweight network BiSeNet v2, the detailed features and semantic features are extracted in one branch, reducing the number of network parameters. At the same time, the low-dimensional features and high-dimensional features are complementary. The STDC module fuses context information and extracts multi-scale features of the image. Then, since the detailed features and semantic features play a decisive role in the final segmentation effect, a bilateral guidance fusion module is used to effectively fuse the low-dimensional features and high-dimensional features. Then, the fused features are used to predict the segmentation result. In network training, the joint cross-entropy loss function and the boundary constraint loss function are used to optimize the parameter learning of the network. By adding the boundary constraint loss function, the segmentation performance of the algorithm for road boundaries is improved without additionally increasing the number of network parameters.

[0091] Although several embodiments of the present invention have been given in this article, those skilled in the art should understand that the embodiments in this article can be changed without departing from the spirit of the present invention. The above embodiments are only exemplary and should not be used as the limitation of the scope of the rights of the present invention.

Claims

1. A semantic map road segmentation method in the scenario of an unmanned aerial vehicle, characterized in that The method includes: S1. Construct a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion; S2. Train the lightweight road segmentation network model in step S1 to obtain an optimized model; S3. Use the optimized model to perform road segmentation on the drone map.

2. The semantic map road segmentation method in the UAV scenario according to claim 1, wherein, In step S1, the lightweight road segmentation network model includes a basic convolution module, a short-circuit dense splicing module STDC, and a bilateral guidance fusion module; The basic convolution module is used to extract low-dimensional detailed features of the drone image; There are multiple STDC modules connected in sequence, which are used to extract image features without increasing the network width. The image features include low-dimensional detailed features and high-dimensional semantic features; The bilateral guidance fusion module fuses the low-dimensional detailed features and the high-dimensional semantic features, and uses the fused features to predict the segmentation result of the drone image.

3. The semantic map road segmentation method in the UAV scenario according to claim 2, wherein, Each STDC includes 3 independent convolution modules. The output of the previous stage of the STDC first passes through a convolution module with a convolution kernel size of 1×1, and the convolution stride is set to 1 for reducing the feature dimension. Then, the dimension-reduced features are input into a convolution module with a convolution kernel size of 3×3, and the convolution stride is 1 or 2. The convolution kernel size of the third convolution module is 3×3, and the convolution stride is 1; The outputs of each convolution module are skip-connected as the output of the STDC. The fused feature is expressed as: x out = Cat(x1, x2, x3); where x1, x2, and x3 respectively represent the features extracted by the three convolution modules of the STDC, and Cat(g) represents the feature fusion operator.

4. The semantic map road segmentation method in the UAV scenario according to claim 1, characterized in that The specific method of step S2 includes: S21. The drone image passes through 2 basic convolution modules to extract low-dimensional detailed features; S22. The output after passing through 2 basic convolution modules is used as the input and passes through 4 STDCs in sequence. The features extracted by the 1st STDC are directly input into the bilateral guidance fusion module as the low-dimensional detailed feature map, and the features extracted by the 4th STDC are input into the bilateral guidance fusion module as the high-dimensional semantic feature map; S23. The bilateral guidance module fuses the low-dimensional detailed feature map and the high-dimensional semantic feature map to obtain the final fused feature map: S24. Obtain the predicted segmentation result of the drone image according to the final fused feature map; S25. Calculate the loss function value between the predicted segmentation result and the actual segmentation result, and adjust the network parameters according to the loss function value; S26. Repeat steps S21 - S25 until the loss function value is within the set threshold, and the model training ends. The network parameters at this time are the optimized model parameters.

5. The semantic map road segmentation method in the UAV scenario according to claim 4, characterized in that, In step S22, in the 1st, 3rd, and 4th STDCs, each STDC includes 3 independent convolution modules. The output of the previous stage first passes through a convolution module with a convolution kernel size of 1×1, and the convolution stride is set to 1 for reducing the feature dimension. Then, the dimension-reduced features are input into a convolution module with a convolution kernel size of 3×3, and the convolution stride is 1. The convolution kernel size of the third convolution module is 3×3, and the convolution stride is 1; In the second STDC, there are three independent convolutional modules. The output of the previous stage first passes through a convolutional module with a kernel size of 1×1, and the convolutional stride is set to 1 to reduce the feature dimension. Then, the feature with reduced dimension is input into a convolutional module with a kernel size of 3×3 and a convolutional stride of 2. The kernel size of the third convolutional module is 3×3, and the convolutional stride is 1. The output of the first convolutional module is downsampled, and the global pooling operation with a kernel size of 3×3 is used to change the feature scale of the output of the first convolutional module to half of the original and fuse it with other features.

6. The semantic map road segmentation method in the UAV scenario according to claim 4, characterized in that In step S23, the high-dimensional feature map to be fused performs downsampling operations twice based on the low-dimensional feature map to be fused. The low-dimensional feature map sequentially passes through a convolutional module with a stride of 2 and a kernel size of 3×3 and an average pooling operation with a pooling kernel size of 3×3. The high-dimensional feature map sequentially passes through a depthwise separable convolutional module with a kernel size of 3×3 and a convolutional module with a kernel size of 1×1. Then, the low-dimensional feature map after the operation and the high-dimensional feature map are fused through a dot product operation to obtain the first fused feature. The high-dimensional feature map is upsampled and then fused with the low-dimensional feature map through a dot product to obtain the second fused feature. The first fused feature and the second fused feature are added to obtain the final fused feature map.

7. The semantic map road segmentation method in the UAV scenario according to claim 4, wherein In step S25, the loss function is: L total = L BIoU + L cls ; Where: L BIOU is the boundary IOU loss function, and L cls is the cross-entropy loss function; N represents the total number of predicted pixels, x i represents a pixel, and Y seg (x i ) and represent the ground truth and the predicted value respectively, and M represents the road boundary range that the network focuses on.

8. A method for constructing a semantic map in an unmanned aerial vehicle scenario, characterized in that, The method for constructing a semantic map in the unmanned aerial vehicle (UAV) scenario uses the method for segmenting roads in the semantic map of the UAV scenario according to any one of claims 1-7 for road segmentation.

9. A semantic map road segmentation system in an unmanned aerial vehicle scenario, characterized in that, The system is used to implement the method according to any one of claims 1-7. The system includes: A lightweight road segmentation network model construction unit that constructs a lightweight road segmentation network model based on boundary constraints and multi-scale feature fusion. A model training unit that trains the lightweight road segmentation network model using multi-scale UAV images to obtain an optimized model. A map road segmentation unit that uses the optimized model to segment roads in the UAV map.

10. An information processing terminal for implementing the method for constructing a semantic map in the UAV scenario according to any one of claims 1-7.