Skyline segmentation method, device, equipment and medium

Through the training of LineFormer model and multi-stage learning strategy, the problem of poor skyline segmentation performance in mountainous environments is solved, and high-precision skyline segmentation is achieved.

CN120495318APending Publication Date: 2025-08-15XIANGJIANG LAB
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510567040.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing skyline segmentation method has poor segmentation performance in mountainous environments, especially in cloudy or foggy conditions, making it difficult to achieve high-precision positioning.

Method used

The LineFormer model is used for training, including embedded modules, multiple Transformer modules, multi-scale feature fusion modules and recursive hollow self-attention modules, combining multi-stage learning strategies and specific loss functions to improve the accuracy and adaptability of skyline segmentation.

Benefits of technology

It significantly improves the accuracy of skyline segmentation and the adaptability of the model to complex environments, and improves the performance of skyline segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495318A_ABST
    Figure CN120495318A_ABST
Patent Text Reader

Abstract

The invention provides a skyline segmentation method, device and equipment and a medium, and the method comprises the steps: training a constructed LineFormer model through employing mountainous area image data collected at a fixed point in a mountainous area, and obtaining a skyline segmentation model; inputting the image data of the target mountainous area into a skyline segmentation model for segmentation to obtain a skyline binarization segmentation result; the LineFormer model comprises an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer sensing module, a second multi-layer sensing module and a recursive cavity self-attention module; and the skyline segmentation performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a skyline segmentation method, device, equipment and medium. Background Art

[0002] With the development of social economy, people's demand for positioning is increasing. The Global Navigation Satellite System (GNSS) is applied to all walks of life in society because it can provide all-weather, high-precision positioning information. However, since the satellites in the GNSS positioning system are far away from the nodes to receive the positioning signal, the positioning signal received by the node is low in power and easily interfered with. Especially in mountainous areas with complex terrain with a lot of obstruction or Global Positioning System (GPS) interference, the satellite positioning system will experience GNSS denial, resulting in the GNSS being unable to provide high-precision positioning information. At this time, positioning in mountainous areas requires a positioning technology that can be used when the GNSS signal is poor or unavailable.

[0003] Many researchers have proposed localization methods based on image retrieval. This method achieves localization by capturing images of the current environment and performing matching searches against an image database containing location information. By associating the location signal with the search image, image retrieval can be used to find the image most similar to the search image, thereby achieving regional localization. Current image retrieval and localization methods are mostly focused on the retrieval of landmark images and street view images, which have fewer interferences and more features. Research on image retrieval and localization in mountainous and outdoor areas is relatively limited.

[0004] Because the skyline—the boundary between mountains and the sky—is typically relatively stable in mountainous environments, existing technologies have also utilized the skyline as a positioning feature. As the boundary between the sky and the ground, its accurate segmentation provides essential preparatory steps for subsequent image retrieval and localization. Existing skyline segmentation methods are primarily divided into two categories: the first utilizes machine learning techniques to segment the sky and ground, accurately extracting the boundaries that constitute the skyline; the second directly identifies and extracts the skyline through edge detection. Past research on machine learning-driven skyline segmentation methods has often relied on analyzing features such as color, texture, and pixel gradients in both sky and non-sky areas to identify and extract the skyline. This approach is robust for irregular skyline extraction but is time-consuming. The second type of skyline segmentation method directly extracts the skyline based on its features. The skyline is typically assumed to be a curve with distinct edge features, but in cloudy or foggy conditions, the skyline is less distinct. Consequently, existing skyline segmentation methods suffer from poor segmentation performance. Summary of the Invention

[0005] The present invention provides a skyline segmentation method, apparatus, device and medium, the purpose of which is to improve the skyline segmentation performance.

[0006] To achieve the above object, the present invention provides a skyline segmentation method, comprising:

[0007] Step 1: Collect mountain image data at a fixed point in the mountainous area;

[0008] Step 2: Use mountain image data to train the constructed LineFormer model to obtain a skyline segmentation model;

[0009] Step 3: Input the image data of the target mountain area into the skyline segmentation model for segmentation to obtain the skyline binary segmentation result;

[0010] The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module;

[0011] The input of the embedding module is the input of the LineFormer model, and the output of the second multi-layer perception module is the output of the Lineformer model;

[0012] The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively;

[0013] The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively;

[0014] The output of the third Transformer module is connected to the input of the fourth Transformer module;

[0015] The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively;

[0016] The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively;

[0017] The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module;

[0018] The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module;

[0019] The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

[0020] Furthermore, the first multi-scale feature fusion module, the second multi-scale feature fusion module, the third multi-scale feature fusion module, and the fourth multi-scale feature fusion module all include:

[0021] The first convolutional layer, the second convolutional layer, the upsampling layer, and the multiplier are connected in sequence;

[0022] The first input end of the multiplier is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the first input end of the fourth multi-scale feature fusion module;

[0023] The input end of the first convolutional layer is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the second input end of the fourth multi-scale feature fusion module;

[0024] The output end of the multiplier is the output end of the first multi-scale feature fusion module, the output end of the second multi-scale feature fusion module, the output end of the third multi-scale feature fusion module, and the output end of the fourth multi-scale feature fusion module;

[0025] The high-level feature map is reduced in dimension through the first convolutional layer to obtain a feature map with the same number of channels as the low-level feature map;

[0026] The feature map with the same number of channels as the low-level feature map is convolved through the second convolutional layer and then input into the upsampling layer for bilinear sampling to obtain a feature map with the same size as the low-level feature map.

[0027] The feature map with the same size as the low-level feature map is multiplied pixel by pixel with the low-dimensional feature map to obtain the fused feature map.

[0028] Specifically, the recursive atrous self-attention module includes:

[0029] First layer normalization unit, second layer normalization unit, first void self-attention unit, second void self-attention unit, third void self-attention unit, first adder, second adder, multi-layer perceptron;

[0030] The input end of the first layer normalization unit and the first input end of the first adder are connected to the output end of the first multi-layer perception module;

[0031] The output of the first normalization unit is connected to the first input of the first dilated self-attention unit;

[0032] The second input terminal of the first void self-attention unit and the input terminal of the second void self-attention unit are both connected to the input terminal of the embedding module;

[0033] The output end of the first void self-attention unit is connected to the first input end of the third void self-attention unit, and the output end of the second void self-attention unit is connected to the second input end of the third void self-attention unit;

[0034] The output end of the third hole self-attention unit is connected to the second input end of the first adder, and the output end of the first adder is connected to the input end of the second layer normalization unit and the first input end of the second adder respectively;

[0035] The output end of the second layer normalization unit is connected to the input end of the multilayer perceptron, and the output end of the multilayer perceptron is connected to the second input end of the second adder;

[0036] The output end of the second adder is connected to the input end of the second multi-layer perception module.

[0037] Furthermore, before training the constructed LineFormer model using mountain image data, the following steps are also included:

[0038] Obtaining the pan / tilt parameters of a device used to collect mountain image data, the pan / tilt parameters including pitch angle and roll angle;

[0039] The mountain image data is angle-corrected using the PTZ parameters to obtain the corrected image sequence.

[0040] Performing panoramic stitching on the rectified image sequence to obtain a stitched image sequence;

[0041] Preprocessing the spliced image sequence to obtain a preprocessed image sequence;

[0042] The skyline in the preprocessed image sequence is labeled to obtain a labeled label image.

[0043] Furthermore, a multi-stage learning adjustment strategy is used to train the constructed LineFormer model.

[0044] Furthermore, a multi-stage learning adjustment strategy is used to train the constructed LineFormer model, including:

[0045] In the initial training stage, the constructed LineFormer model is trained using a linear learning rate strategy;

[0046] In the mid-term training stage, the Poly LR Ratio strategy is used to train the constructed LineFormer model;

[0047] In the final training stage, a constant learning rate is used to train the constructed LineFormer model.

[0048] Furthermore, the loss function of the LineFormer model during training is:

[0049] Loss=λ1FocalLoss+λ2DiceLoss

[0050] FocalLoss=-a t (1-pt)log p t

[0051]

[0052] Among them, Loss represents the total loss value, λ1 and λ2 represent the loss weights, FocalLoss represents the imbalance loss value, DiceLoss represents the prediction loss value, and a t Indicates the importance of balancing positive and negative samples or different categories, p t Represents the model's predicted probability of the target category, X represents the set or tensor corresponding to the predicted segmentation result, and Y represents the set or tensor corresponding to the actual segmentation label.

[0053] The present invention also provides a skyline segmentation device, comprising:

[0054] An acquisition module is used to acquire mountain image data at fixed points in mountainous areas;

[0055] The training module is used to train the constructed LineFormer model using mountain image data to obtain a skyline segmentation model;

[0056] The segmentation module is used to input the image data of the target mountain area into the skyline segmentation model for segmentation to obtain the skyline binary segmentation result;

[0057] The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module;

[0058] The input of the embedding module is the input of the LineFormer model, and the output of the second multi-layer perception module is the output of the Lineformer model;

[0059] The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively;

[0060] The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively;

[0061] The output of the third Transformer module is connected to the input of the fourth Transformer module;

[0062] The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively;

[0063] The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively;

[0064] The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module;

[0065] The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module;

[0066] The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

[0067] The present invention also provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the skyline segmentation method when executing the computer program.

[0068] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the skyline segmentation method is implemented.

[0069] The above solution of the present invention has the following beneficial effects:

[0070] The present invention uses mountain image data collected at a fixed point in a mountainous area to train a constructed LineFormer model to obtain a skyline segmentation model. The image data of the target mountainous area is input into the skyline segmentation model for segmentation to obtain a skyline binary segmentation result. The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive void self-attention module. Compared with the prior art, the present invention introduces multiple multi-scale fusion modules into the LineFormer model to integrate microscopic detail information of the segmented object and contextual information of the macroscopic environment, significantly improving segmentation accuracy. The present invention also introduces a recursive void self-attention module to enhance the model's adaptability to various environments and target recognition accuracy. By adding multiple Transformer modules to increase feature maps with deeper and smaller scales, the present invention not only optimizes the expression of spatial details and boundary information, but also further enhances the model's adaptability to complex environments, thereby achieving the purpose of improving skyline segmentation performance.

[0071] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1A schematic diagram of a flow chart of an embodiment of the present invention;

[0073] Figure 2 Schematic diagram of the structure of the LineFormer model in an embodiment of the present invention;

[0074] Figure 3 2 is a schematic diagram of the structure of a multi-scale feature fusion module in an embodiment of the present invention;

[0075] Figure 4 Schematic diagram of the structure of the recursive hole self-attention module in an embodiment of the present invention;

[0076] Figure 5 A graph showing changes in learning rate during training in an embodiment of the present invention;

[0077] Figure 6 2 is a schematic structural diagram of a skyline segmentation device according to an embodiment of the present invention;

[0078] Figure 7 Schematic diagram of the structure of the terminal device in an embodiment of the present invention. DETAILED DESCRIPTION

[0079] To make the technical problems, technical solutions, and advantages to be solved by the present invention more clear, the following is a detailed description with reference to the accompanying drawings and specific embodiments. It is obvious that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0080] In the description of the present invention, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and should not be understood as indicating or implying relative importance.

[0081] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to a locking connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0082] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0083] In view of the existing problems, the present invention provides a skyline segmentation method, device, equipment and medium.

[0084] like Figure 1 As shown, an embodiment of the present invention provides a skyline segmentation method, including:

[0085] Step 1: Collect mountain image data at a fixed point in the mountainous area;

[0086] Step 2: Use mountain image data to train the constructed LineFormer model to obtain a skyline segmentation model;

[0087] Step 3: Input the target mountain area image data into the skyline segmentation model for segmentation to obtain the skyline binary segmentation result. The specific performance of the skyline binary segmentation result is: the background is completely black, and the skyline area presents a continuous white curve.

[0088] Specifically, before training the constructed LineFormer model using mountain image data, the following steps are also included:

[0089] Obtaining the pan / tilt parameters of a device used to collect mountain image data, the pan / tilt parameters including pitch angle and roll angle;

[0090] The mountain image data is angle-corrected using the PTZ parameters to obtain the corrected image sequence.

[0091] Performing panoramic stitching on the rectified image sequence to obtain a stitched image sequence;

[0092] Preprocessing the spliced image sequence to obtain a preprocessed image sequence;

[0093] The skyline in the preprocessed image sequence is labeled to obtain a labeled label image.

[0094] Since image matching and positioning of specific features in mountainous areas requires clear and unchanging features, the skyline can be extracted as a good feature for positioning information. Therefore, the embodiment of the present invention requires the use of a vehicle-mounted pan-tilt image acquisition device and a posture sensor device to obtain mountain image data and pan-tilt data during navigation in the mountainous area.

[0095] Since the fixed point for acquisition cannot be completely flat, the acquired mountain images have different pitch and roll angles from the images taken when the camera is horizontal. Therefore, the embodiment of the present invention uses all the mountain image data collected by the vehicle-mounted pan-tilt image acquisition device and the pitch and roll angles of the attitude sensor to construct an original data image library, and corrects the polar angle of the mountain image data to ensure the consistency of the shooting angle; then the corrected image sequence is panoramically stitched, and the two adjacent stitched images have a 30-40% overlapping area. The stitching steps include key point feature detection, local invariant features, key feature point matching, random sampling consistency, and perspective transformation. The resulting stitched image sequence is preprocessed, including cropping the images to a resolution of 512×512, random scaling with a scale of 0.5-2.0, left-right flipping, Gaussian blurring, random rotation, color gamut conversion, and other image enhancement transformations. This addresses objective factors such as environmental factors and the device's pixel quality. Through a series of data augmentation techniques, a subset of new images is generated, enabling the network to achieve high generalization capabilities and accurate predictions across multiple scenarios. The skylines in the image data are annotated, resulting in a total of 2050 labeled images. These 2050 images are divided into a training set and a validation set with a ratio of 9:1 (i.e., 1845 and 205 images, respectively), and a dataset is created in the PASCAL VOC 2012 format.

[0096] Specifically, if Figure 2 As shown, the LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module;

[0097] The input of the embedding module is the input of the LineFormer model, and the output of the second multi-layer perception module is the output of the Lineformer model;

[0098] The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively;

[0099] The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively;

[0100] The output of the third Transformer module is connected to the input of the fourth Transformer module;

[0101] The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively;

[0102] The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively;

[0103] The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module;

[0104] The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module;

[0105] The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

[0106] Specifically, the first multi-scale feature fusion module, the second multi-scale feature fusion module, the third multi-scale feature fusion module, and the fourth multi-scale feature fusion module all include:

[0107] The first convolutional layer, the second convolutional layer, the upsampling layer, and the multiplier are connected in sequence, such as Figure 3 As shown;

[0108] The first input end of the multiplier is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the first input end of the fourth multi-scale feature fusion module;

[0109] The input end of the first convolutional layer is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the second input end of the fourth multi-scale feature fusion module;

[0110] The output end of the multiplier is the output end of the first multi-scale feature fusion module, the output end of the second multi-scale feature fusion module, the output end of the third multi-scale feature fusion module, and the output end of the fourth multi-scale feature fusion module;

[0111] The high-level feature map is reduced in dimension through the first convolutional layer to obtain a feature map with the same number of channels as the low-level feature map;

[0112] The feature map with the same number of channels as the low-level feature map is convolved through the second convolutional layer and then input into the upsampling layer for bilinear sampling to obtain a feature map with the same size as the low-level feature map.

[0113] The feature map with the same size as the low-level feature map is multiplied pixel by pixel with the low-dimensional feature map to obtain the fused feature map.

[0114] In an embodiment of the present invention, the first convolution layer is a 1×1 convolution kernel, and the second convolution layer is a 3×3 convolution kernel. Through the above-mentioned first multi-scale feature fusion module, second multi-scale feature fusion module, third multi-scale feature fusion module, and fourth multi-scale feature fusion module, the local features extracted by the smaller receptive field and the wide-area features extracted by the larger receptive field can be comprehensively considered. Then, through the fusion of multi-scale features, the continuous changes from macroscopic mountain undulations to microscopic texture details can be adaptively captured, thereby improving the adaptability and robustness of the LineFormer model to different geographical environments.

[0115] Specifically, if Figure 4 As shown in Figure 2, the recursive hole self-attention module includes:

[0116] First layer normalization unit, second layer normalization unit, first void self-attention unit, second void self-attention unit, third void self-attention unit, first adder, second adder, multi-layer perceptron;

[0117] The input end of the first layer normalization unit and the first input end of the first adder are connected to the output end of the first multi-layer perception module;

[0118] The output of the first normalization unit is connected to the first input of the first dilated self-attention unit;

[0119] The second input terminal of the first void self-attention unit and the input terminal of the second void self-attention unit are both connected to the input terminal of the embedding module;

[0120] The output end of the first void self-attention unit is connected to the first input end of the third void self-attention unit, and the output end of the second void self-attention unit is connected to the second input end of the third void self-attention unit;

[0121] The output end of the third hole self-attention unit is connected to the second input end of the first adder, and the output end of the first adder is connected to the input end of the second layer normalization unit and the first input end of the second adder respectively;

[0122] The output end of the second layer normalization unit is connected to the input end of the multilayer perceptron, and the output end of the multilayer perceptron is connected to the second input end of the second adder;

[0123] The output end of the second adder is connected to the input end of the second multi-layer perception module.

[0124] The recursive hole self-attention module provided in the embodiment of the present invention can combine multi-scale contextual information through a recursive mechanism. Based on a comprehensive consideration of skyline detection performance and computational parameter cost, a two-layer recursive depth is set. This can fully exploit high-level skyline features through multiple hole self-attention units, thereby improving the ability to represent high-level image skyline features without introducing too many parameters.

[0125] Specifically, the LineFormer model is trained by mainly using a multi-stage learning adjustment strategy to train the constructed LineFormer model.

[0126] Furthermore, a multi-stage learning adjustment strategy is used to train the constructed LineFormer model, including:

[0127] In the initial training stage, the constructed LineFormer model is trained using a linear learning rate strategy;

[0128] In the mid-term training stage, the Poly LR Ratio strategy is used to train the constructed Line Former model;

[0129] In the final training stage, a constant learning rate is used to train the constructed Line Former model.

[0130] Specifically, a multi-stage learning adjustment strategy is used to train the constructed Line Former model. The specific process is as follows:

[0131] The training parameters are configured as follows: Batch Size is set to 12, MaxEpoch is set to 2500, and differentiated learning rate adjustment strategies are carefully designed for different training stages to improve model convergence and segmentation accuracy;

[0132] The Adam W optimizer sets the initial learning rate to 0.001 and the weight decay coefficient to 0.1 to ensure the stability of parameter updates;

[0133] In the initial training phase, i.e., the first 0-1200 iterations, a linear learning rate (LR) strategy is used to linearly increase the learning rate from the initial value to 3e-2 to accelerate the model's initial learning of skyline features.

[0134] In the mid-term training phase, i.e., iterations 1200-2400, the Poly LR Ratio strategy is introduced to gradually reduce the learning rate to 3e-2 times the initial value, with the power value set to 0.9, effectively balancing the model's optimization of microscopic details and global context;

[0135] In the final training phase, a constant learning rate (Constant LR) is used to maintain the learning rate at the final value of the previous phase to ensure the fine adjustment and robustness of the model in complex mountainous scenes. The change of learning rate is shown in the figure below. Figure 5 shown.

[0136] Specifically, the loss function of the Line Former model during training is:

[0137] Loss=λ1FocalLoss+λ2DiceLoss

[0138] FocalLoss=-a t (1-pt)logp t

[0139]

[0140] Among them, Loss represents the total loss value, λ1 and λ2 represent the loss weights, λ1,λ2∈[0,1], λ1+λ2=1, FocalLoss represents the imbalance loss value, which is mainly used to deal with the imbalance between foreground and background. By introducing the modulation factor (1-p t ), which makes the model pay more attention to samples that are difficult to classify, thereby improving the overall detection or segmentation accuracy. DiceLoss represents the prediction loss value, which is used to evaluate the overlap between the two sets to measure the overlap between the segmentation area predicted by the model and the true segmentation area. The higher the value, the more consistent the prediction is with the true result. t Indicates the importance of balancing positive and negative samples or different categories, p t Represents the model's predicted probability of the target category, X represents the set or tensor corresponding to the predicted segmentation result, which can usually be regarded as the binary / probability map output by the model, and Y represents the set or tensor corresponding to the actual segmentation annotation.

[0141] In an embodiment of the present invention, FocalLoss loss is a cross-entropy loss with dynamic scaling characteristics. As the prediction confidence of the correct classification increases, its scaling factor gradually decays to zero, thereby effectively focusing on skyline pixels that are difficult to classify and enhancing the model's recognition ability for slender curves. Dice Loss loss takes the similarity of two samples as its core, and by optimizing the overlap between the predicted area and the true skyline, it further alleviates the interference of an excessively high proportion of background pixels on model training. The combined use of these two loss functions not only improves the learning efficiency of LineFormer for sparse skyline features, but also significantly improves the accuracy of the segmentation results.

[0142] In order to verify the effectiveness of the multi-scale feature fusion module, the recursive hole self-attention module, and the multiple Transformer modules in the LineFormer model, the embodiment of the present invention adopts the mean intersection over union (mloU), the maximum skyline F measurement value max (SF β ) and category mean pixel accuracy (MPA) are used as the main evaluation indicators to comprehensively measure the performance of the improved algorithm in the skyline segmentation task. mIoU (mean intersection over union) is calculated by averaging the intersection over union of the skyline and background categories to comprehensively evaluate the segmentation consistency of the model across different categories. The calculation formula for the average intersection over union is:

[0143]

[0144] The calculation formula for the category average pixel accuracy is:

[0145]

[0146] Since there is a conflict between precision and recall, the maximum skyline F-measure is considered as the evaluation indicator, and its calculation formula is:

[0147]

[0148] Among them, TP represents the number of samples correctly predicted as positive by the model, FP represents the number of samples incorrectly predicted as positive but actually negative by the model, FN represents the number of samples incorrectly predicted as negative but actually positive by the model, TN represents the number of samples correctly predicted as negative by the model, and β 2 The role of is to adjust the importance weight of Recall in the comprehensive index, β 2 =0.3, N represents the total number of categories, S represents the Precision and Recall on a subset of objects of a certain type, Precision represents the accuracy, Recall represents the recall rate.

[0149] For the skyline segmentation task, relevant ablation experiments were conducted, and the experimental results are shown in Table 1:

[0150] Table 1 Ablation experiment results of LineFormer model

[0151]

[0152] From the ablation experiment results shown in Table 1, we can see that the model structure adjustment methods such as multi-scale feature fusion module, recursive hole self-attention module, and improved feature map size have a significant impact on mloU, max(SF β ), MPA and other indicators have helped to improve. Among them, after adding the recursive void self-attention module, compared with the existing Segformer algorithm, the model's mloU, max(SF β ), MPA increased by 4.61%, 1.84%, and 0.95% respectively; in addition, the combination of multi-scale feature fusion module + recursive void self-attention module has the greatest improvement on the indicators; compared with the existing Segformer algorithm, after adding these two modules, the model's mloU, max(SF β ), MPA increased by 7.55%, 1.31%, and 1.46% respectively; if the three model adjustment methods introduced in the embodiment of the present invention are combined, compared with the existing Segformer algorithm, the LineFormer model’s mloU, max(SF β ), MPA increased by 7.93%, 3.56% and 3.61% respectively.

[0153] It can be seen that the model structure adjustment methods proposed in this invention, such as the MFMM module, RASA module, and improved feature map size, can effectively improve the performance of image skyline segmentation. Among them, the introduction of the MFMM module and RASA module has the greatest impact on segmentation performance. This is mainly because the MFMM module effectively integrates microscopic object details and large-area context information by fusing features of different scales, thereby improving skyline segmentation accuracy. The RASA module increases the receptive field through the self-attention mechanism and obtains context information of different scales through recursive void convolution, enhancing the algorithm's ability to judge spatial changes. Improving the feature map size can optimize the expression of spatial details and boundary information, thereby further improving the performance of the segmentation algorithm.

[0154] In this embodiment of the present invention, the proposed LineFormer model is compared with FCN, Segformer, PSPNet, DeepLabv3+, SFNet, etc. The input image size is 512×512. The model is trained on the skyline dataset of this embodiment of the present invention. The results on the validation set are shown in Table 2:

[0155] Table 2 Experimental comparison results of LineFormer model on the skyline dataset

[0156] Model mIoU <![CDATA[max(SF β )]]> MPA SIoU FCN 70.24 72.61 75.10 77.35 PSPNet 76.82 80.17 81.22 85.37 DeepLabv3+ 79.58 82.52 84.36 86.73 SFNet 78.23 82.16 83.42 86.06 SegFormer 81.21 90.01 86.12 87.23 LineFormer model 88.65 93.21 89.23 93.17

[0157] As can be seen from Table 2 above, the LineFormer model proposed in the embodiment of the present invention has achieved significant improvements in different indicators compared with other models. In terms of mIoU indicators, the model of the embodiment of the present invention is 26.21%, 15.40%, 11.40%, 13.32% and 9.16% higher than FCN, PSPNet, DeepLabv3+, SFNet and SegFormer respectively. β ), the model of the embodiment of the present invention is 28.37%, 16.27%, 12.95%, 13.45% and 3.56% higher than FCN, PSPNet, DeepLabv3+, SFNet and SegFormer. In terms of MPA indicators, the model of the embodiment of the present invention is 18.81%, 9.86%, 5.77%, 6.96% and 3.61% higher than FCN, PSPNet, DeepLabv3+, SETR and SegFormer. In terms of SIOU indicators, the model of the embodiment of the present invention is 20.45%, 9.14%, 7.43%, 8.26% and 6.81% higher than FCN, PSPNet, DeepLabv3+, SFNet and SegFormer. These indicators show the advanced nature of the improved model proposed in the embodiment of the present invention in processing skyline segmentation tasks, and fully confirm the correctness and effectiveness of the model design of the embodiment of the present invention.

[0158] In the embodiment of the present invention, a LineFormer model is trained using mountain image data collected at a fixed point in a mountainous area to obtain a skyline segmentation model. Image data of a target mountainous area is input into the skyline segmentation model for segmentation, thereby obtaining a skyline binary segmentation result. The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive void self-attention module. Compared with the prior art, the embodiment of the present invention introduces multiple multi-scale fusion modules into the LineFormer model to integrate microscopic detail information of the segmented object and macroscopic contextual information of the environment, significantly improving segmentation accuracy. The recursive void self-attention module is also introduced to enhance the model's adaptability to various environments and target recognition accuracy. By adding multiple Transformer modules to increase feature maps with deeper and smaller scales, the expression of spatial details and boundary information can be optimized and the model's adaptability to complex environments can be further enhanced, thereby achieving the purpose of improving skyline segmentation performance.

[0159] Corresponding to the skyline segmentation method described in the above embodiment, Figure 6 As shown, an embodiment of the present invention further provides a skyline segmentation device 100, which includes:

[0160] The acquisition module 101 is used to acquire mountain image data at a fixed point in the mountain area;

[0161] A training module 102 is used to train the constructed LineFormer model using mountain image data to obtain a skyline segmentation model;

[0162] The segmentation module 103 is used to input the image data of the target mountain area into the skyline segmentation model for segmentation to obtain a skyline binary segmentation result;

[0163] The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module;

[0164] The input of the embedding module is the input of the LineFormer model, and the output of the second multi-layer perception module is the output of the Lineformer model;

[0165] The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively;

[0166] The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively;

[0167] The output of the third Transformer module is connected to the input of the fourth Transformer module;

[0168] The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively;

[0169] The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively;

[0170] The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module;

[0171] The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module;

[0172] The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

[0173] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0175] The embodiment of the present invention further provides a terminal device, such as Figure 7 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the above-mentioned skyline segmentation method when executing the computer program D102.

[0176] The terminal device D10 can be a computing device such as a desktop computer, a notebook, a PDA, a server, a server cluster, a cloud server, etc. The terminal device may include, but is not limited to, a processor D100 and a memory D101. It will be understood by those skilled in the art that Figure 7 This is merely an example of the terminal device D10 and does not constitute a limitation on the terminal device D10 . The terminal device D10 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device D10 may also include input and output devices, network access devices, etc.

[0177] The processor D100 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0178] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card, etc. equipped on the terminal device D10. Furthermore, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or is to be output.

[0179] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0180] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0181] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the skyline segmentation method is implemented.

[0182] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying the computer program code to a construction device / terminal device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.

[0183] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A skyline segmentation method, characterized in that: include: Step 1: Collect mountain image data at a fixed point in the mountainous area; Step 2: Using the mountain image data to train the constructed LineFormer model to obtain a skyline segmentation model; Step 3: Input the image data of the target mountain area into the skyline segmentation model for segmentation to obtain a skyline binary segmentation result; The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module; The input end of the embedding module is the input end of the LineFormer model, and the output end of the second multi-layer perception module is the output end of the Lineformer model; The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively; The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively; The output end of the third Transformer module is connected to the input end of the fourth Transformer module; The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively; The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively; The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module; The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module; The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

2. The skyline segmentation method according to claim 1, wherein: The first multi-scale feature fusion module, the second multi-scale feature fusion module, the third multi-scale feature fusion module, and the fourth multi-scale feature fusion module all include: The first convolutional layer, the second convolutional layer, the upsampling layer, and the multiplier are connected in sequence; The first input end of the multiplier is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the first input end of the fourth multi-scale feature fusion module; The input end of the first convolutional layer is the input end of the first multi-scale feature fusion module, the input end of the second multi-scale feature fusion module, the input end of the third multi-scale feature fusion module, and the second input end of the fourth multi-scale feature fusion module; The output end of the multiplier is the output end of the first multi-scale feature fusion module, the output end of the second multi-scale feature fusion module, the output end of the third multi-scale feature fusion module, and the output end of the fourth multi-scale feature fusion module; The high-level feature map is reduced in dimension by the first convolutional layer to obtain a feature map with the same number of channels as the low-level feature map; The feature map with the same number of channels as the low-level feature map is convolved through the second convolutional layer and then input into the upsampling layer for bilinear sampling to obtain a feature map with the same size as the low-level feature map; The feature map with the same size as the low-level feature map is multiplied pixel by pixel with the low-dimensional feature map to obtain the fused feature map.

3. The skyline segmentation method according to claim 1, wherein: The recursive hole self-attention module includes: First layer normalization unit, second layer normalization unit, first void self-attention unit, second void self-attention unit, third void self-attention unit, first adder, second adder, multi-layer perceptron; The input end of the first layer normalization unit and the first input end of the first adder are both connected to the output end of the first multi-layer perception module; An output end of the first layer normalization unit is connected to a first input end of the first hole self-attention unit; The second input end of the first void self-attention unit and the input end of the second void self-attention unit are both connected to the input end of the embedding module; The output end of the first void self-attention unit is connected to the first input end of the third void self-attention unit, and the output end of the second void self-attention unit is connected to the second input end of the third void self-attention unit; The output end of the third void self-attention unit is connected to the second input end of the first adder, and the output end of the first adder is connected to the input end of the second layer normalization unit and the first input end of the second adder respectively; The output end of the second layer normalization unit is connected to the input end of the multilayer perceptron, and the output end of the multilayer perceptron is connected to the second input end of the second adder; The output end of the second adder is connected to the input end of the second multi-layer perception module.

4. The skyline segmentation method according to claim 1, wherein: Before training the constructed LineFormer model using the mountain image data, the following steps are also included: Acquiring pan / tilt parameters of a device used to collect the mountain image data, wherein the pan / tilt parameters include a pitch angle and a roll angle; Performing angle correction on the mountain image data using the pan / tilt parameters to obtain a corrected image sequence; Performing panoramic stitching on the rectified image sequence to obtain a stitched image sequence; Preprocessing the spliced image sequence to obtain a preprocessed image sequence; The skyline in the preprocessed image sequence is labeled to obtain a labeled label image.

5. The skyline segmentation method according to claim 4, characterized in that: A multi-stage learning adjustment strategy is used to train the constructed LineFormer model.

6. The skyline segmentation method according to claim 5, characterized in that: The constructed LineFormer model is trained using a multi-stage learning adjustment strategy, including: In the initial training stage, the constructed LineFormer model is trained using a linear learning rate strategy; In the mid-term training stage, the Poly LR Ratio strategy is used to train the constructed LineFormer model; In the final training stage, a constant learning rate is used to train the constructed LineFormer model.

7. The skyline segmentation method according to claim 6, characterized in that: The loss function of the LineFormer model during training is: Loss=λ1FocalLoss+λ2DiceLoss WordLoss=-a t (1-p t )log p t Among them, Loss represents the total loss value, λ1 and λ2 represent the loss weights, FocalLoss represents the imbalance loss value, DiceLoss represents the prediction loss value, and a t Indicates the importance of balancing positive and negative samples or different categories, p t Represents the model's predicted probability of the target category, X represents the set or tensor corresponding to the predicted segmentation result, and Y represents the set or tensor corresponding to the actual segmentation label.

8. A skyline segmentation device, characterized in that: include: An acquisition module is used to acquire mountain image data at fixed points in mountainous areas; A training module, configured to train the constructed LineFormer model using the mountain image data to obtain a skyline segmentation model; a segmentation module, configured to input the image data of the target mountain area into the skyline segmentation model for segmentation, and obtain a skyline binary segmentation result; The LineFormer model includes: an embedding module, a first Transformer module, a second Transformer module, a third Transformer module, a fourth Transformer module, a fifth Transformer module, a first multi-scale feature fusion module, a second multi-scale feature fusion module, a third multi-scale feature fusion module, a fourth multi-scale feature fusion module, a first multi-layer perception module, a second multi-layer perception module, and a recursive hole self-attention module; The input end of the embedding module is the input end of the LineFormer model, and the output end of the second multi-layer perception module is the output end of the Lineformer model; The output end of the embedding module is connected to the input end of the first Transformer module, and the output end of the first Transformer module is connected to the input end of the second Transformer module and the first input end of the fourth multi-scale feature fusion module respectively; The output end of the second Transformer module is connected to the input end of the third Transformer module and the first input end of the third multi-scale feature fusion module respectively; The output end of the third Transformer module is connected to the input end of the fourth Transformer module; The output end of the fourth Transformer module is connected to the input end of the fifth Transformer module and the first input end of the first multi-scale feature fusion module respectively; The output end of the fifth Transformer module is connected to the second input end of the first multi-scale feature fusion module and the first input end of the first multi-layer perception module respectively; The first output end of the first multi-scale feature fusion module is connected to the input end of the second multi-scale feature fusion module, the first output end of the second multi-scale feature fusion module is connected to the second input end of the third multi-scale feature fusion module, and the first output end of the third multi-scale feature fusion module is connected to the second input end of the fourth multi-scale feature fusion module; The second output end of the first multi-scale feature fusion module, the second output end of the second multi-scale feature fusion module, the second output end of the third multi-scale feature fusion module, and the second output end of the fourth multi-scale feature fusion module are all connected to the second input end of the first multi-layer perception module; The output end of the first multi-layer perception module is connected to the input end of the recursive void self-attention module, and the output end of the recursive void self-attention module is connected to the input end of the second multi-layer perception module.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the skyline segmentation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the skyline segmentation method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Integrated navigation method based on skyline positioning and IMU prior fusion

    CN121430600A

  • Combined navigation method based on horizon positioning and imu prior fusion

    CN121430600B