Plateau mountain passable area identification method based on hybrid encoder network
By adopting a hybrid encoder network in a plateau mountain photovoltaic power station, combined with UNet and Swin Transformer encoder, the problem of passable area identification in complex terrain is solved, and the accuracy and robustness of detection are improved.
Patent Information
- Application Number
- CN202311652274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively identify accessible areas of complex terrain in plateau mountain photovoltaic power plants, especially in cases where image downsampling leads to semantic ambiguity caused by loss of detail information and occlusion problems.
A hybrid encoder network structure is adopted, combined with UNet's encoder part and Swin Transformer encoder, and a parallel dual encoder structure is formed to use the spatial characteristics and context information of the image to construct a data set of passable areas of the plateau photovoltaic power stations, and the passable area detection results are obtained through training models.
The accuracy, adaptability and robustness of the model to detect passable areas in complex environments can be improved, and can more effectively adapt to the detection needs of complex terrain.
Smart Images

Figure CN120107640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more specifically, to a method for identifying traversable areas in plateau mountains using a hybrid encoder network. Background Art
[0002] In recent years, the rapid development of high-altitude photovoltaic power stations has put forward new requirements for automated autonomous mobile robot equipment. Autonomous mobile robots can participate in the inspection of photovoltaic power stations. If equipped with corresponding agricultural tools, they can also complete power station maintenance work, such as intelligent weeding. However, high-altitude photovoltaic power stations are usually located in high-altitude mountainous and hilly areas with complex terrain. Autonomous mobile robots are required to detect and identify the environment through onboard sensors, which poses a great challenge to visual recognition technology.
[0003] In the prior art, when processing images, convolutional neural networks (CNNs) usually perform downsampling to reduce the amount of computation and the number of parameters. However, this may result in the loss of some important detail information, especially small-scale features. Objects of different semantic categories may have similar sizes, materials, and spectral features, making them difficult to distinguish. In addition, occlusion problems often lead to semantic ambiguity. Therefore, more global context information and fine spatial features are needed as clues for semantic reasoning. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the present invention provides a method for identifying traversable areas in plateau mountains based on a hybrid encoder network, which can simultaneously utilize the spatial features and contextual information of the image, so that the model can adapt to the detection of traversable areas in complex terrain, and improve the accuracy and stability of traversable area detection. The method of the present invention has higher adaptability, flexibility and robustness.
[0005] The technical solution of the present invention is as follows: A method for identifying traversable areas in plateau mountains using a hybrid encoder network comprises the following steps:
[0006] Step 1: Construct a dataset of accessible areas for highland photovoltaic power stations;
[0007] Step 2: Use the encoder part in UNet as the main encoder and the ST (Swin Transformer) encoder as the auxiliary encoder to form a parallel dual encoder structure to form a hybrid encoder network structure based on the convolutional neural network structure block and the transformer structure block;
[0008] Step 3: Use the traversable area dataset of the plateau photovoltaic power station to train the designed hybrid encoder network to obtain the traversable area detection model;
[0009] Step 4: Use the model to detect the plateau photovoltaic power station image and output the passable area detection results classified pixel by pixel.
[0010] Furthermore, Step 1 is specifically as follows:
[0011] Step 1.1: Use the inspection robot for high-altitude photovoltaic power stations to collect images of high-altitude photovoltaic power stations, and perform uniform light and denoising on the images;
[0012] Step 1.2: Label the images after uniform light and denoising, and construct a dataset of accessible areas for photovoltaic power stations on the plateau;
[0013] Step 1.3: Divide the traversable area and obstacle datasets into training set, validation set and test set respectively.
[0014] Furthermore, the denoising process in Step 1.1 is as follows:
[0015] (1) Use a median filter to filter the image and remove small noise points in the image;
[0016] (2) Remove small burrs in the image through opening operation.
[0017] Furthermore, in the denoising operation (1), specifically: when finding the median of the pixel point f(x, y) in the neighborhood window, first find the average of the pixels in the window; if the pixel point f(x, y) is not greater than the average, no filtering is performed; if the pixel point f(x, y) is greater than the average, a recursive method is used to find the median g(x, y), and the median g(x, y) is used to replace the original pixel point f(x, y).
[0018] Furthermore, in Step 1.3, 70% of the traversable area and obstacle data sets are training sets, 20% are validation sets, and 10% are test sets.
[0019] Furthermore, Step 2 is specifically as follows:
[0020] Step 2.1: The ST encoder consists of a linear embedding block and multiple cascaded ST encoding blocks. Each ST encoding block consists of a spatial interaction module SIM and a feature compression module FCM.
[0021] Step 2.2: Use the spatial interaction module SIM to establish pixel-level correlation to encode the spatial information in the ST coding block to enhance the feature representation capability of the occluded object;
[0022] Step 2.3: The global dependencies from the ST encoder are hierarchically integrated into the features of the CNN through the relationship aggregation module RAM;
[0023] Furthermore, in step 2.3, specifically:
[0024] (1) Use CNN to extract features from the input image and obtain low-level and high-level features of the image;
[0025] (2) Using the ST encoder to extract global dependencies from the input image to obtain the global context information of the image;
[0026] (3) Taking the features extracted by CNN and the global dependencies extracted by ST as input, the RAM is used to perform relationship aggregation;
[0027] (4) The aggregated features are fused with the original CNN features to obtain enhanced image features.
[0028] Step 2.4: The output of the encoder is passed through a 2×2 deconvolution layer to expand the resolution; a skip layer is used to connect the encoder and decoder, while reducing the number of channels passing through the 3×3 convolution layer; each convolution layer is accompanied by a BN layer and a ReLU layer; the above process is performed 4 times, so that the features are gradually restored to 1 / 4 of the original image, and then the segmentation result is obtained through a 3×3 convolution layer and linear interpolation upsampling.
[0029] Furthermore, Step 3 is specifically as follows:
[0030] Step 3.1: Set the training parameters, including at least batch size, learning rate, number of detection categories and maximum number of iterations;
[0031] Step 3.2: Input the images of the training set into the hybrid encoder network, calculate the loss of the prediction results, use the loss for back propagation, and calculate the gradient; update the gradient and model parameters in real time, and use the validation set to verify the network accuracy;
[0032] Furthermore, in step 3.2, specifically:
[0033] (1) Load the images of the training set into memory and forward propagate each training image through the hybrid encoder network;
[0034] (2) Calculate the network loss value according to the loss function to quantify the difference between the network prediction and the actual label;
[0035] (3) Using the chain rule, differentiate each weight, calculate the gradient through the loss value, and use the gradient to update the weight of the network;
[0036] (4) Using the calculated gradient, update the model parameters through the Adam optimizer;
[0037] (5) At the end of each round of training, the images of the validation set are input into the network to obtain the prediction results on the validation set, and the accuracy and recall on the validation set are calculated;
[0038] (6) Repeat the above steps (1) to (5) until the preset maximum number of iterations is reached.
[0039] Step 3.3: Repeat Step 3.1 to Step 3.2. After reaching the maximum number of iterations, the hybrid encoder network completes learning and obtains the weights of the drivable area prediction model.
[0040] Furthermore, the loss function is a mean square error function or a cross entropy loss function.
[0041] The present invention according to the above scheme has the beneficial effect that: the present invention provides a plateau mountain passable area recognition method of a hybrid encoder network, firstly constructs a plateau photovoltaic power station passable area data set, and constructs a hybrid encoder network structure based on a convolutional neural network structure block and a Swin transformer structure block; then uses the plateau photovoltaic power station passable area data set to train the designed hybrid encoder network to obtain a passable area detection model; finally, the model is used to detect the plateau photovoltaic power station image, and the passable area detection result of pixel-by-pixel classification is output. This method effectively utilizes the spatial features and contextual information of the image, and improves the accuracy, adaptability and robustness of the model in detecting passable areas in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 It is a flow chart of a method for identifying a plateau mountainous traversable area using a hybrid encoder network in an embodiment of the present invention;
[0044] Figure 2 is a structural block diagram of an ST encoder in an embodiment of the present invention;
[0045] Figure 3 is a structural block diagram of a hybrid encoder network structure in an embodiment of the present invention;
[0046] Figure 44 is a structural block diagram of an ST encoding block in an embodiment of the present invention;
[0047] Figure 5 It is a structural block diagram of a space interaction module SIM in an embodiment of the present invention;
[0048] Figure 6 It is a structural block diagram of a feature compression module FCM in an embodiment of the present invention;
[0049] Figure 7 4 is a structural block diagram of a relationship aggregation module RAM in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following detailed description of the embodiments of the present invention is provided in conjunction with the accompanying drawings and examples. The following detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of the present invention, but cannot be used to limit the scope of the present invention, that is, the present invention is not limited to the described embodiments.
[0051] In order to better understand the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments:
[0052] See also Figures 1 to 7 As shown, an embodiment of the present invention provides a method for identifying a plateau mountainous traversable area using a hybrid encoder network, comprising the following steps:
[0053] Step 1: Construct a dataset of accessible areas for highland photovoltaic power stations;
[0054] Step 2: Use the encoder part in UNet as the main encoder and the ST (Swin Transformer) encoder as the auxiliary encoder. The two are parallelized to form a dual encoder structure, which is used to form a hybrid encoder network based on the convolutional neural network building block and the transformer building block.
[0055] Step 3: Use the traversable area dataset of the plateau photovoltaic power station to train the designed hybrid encoder network to obtain the traversable area detection model;
[0056] Step 4: Use the model to detect the plateau photovoltaic power station image and output the passable area detection results classified pixel by pixel.
[0057] Specifically, the present invention provides a plateau mountain passable area recognition method of a hybrid encoder network. First, a plateau photovoltaic power station passable area dataset is constructed, and a hybrid encoder network structure based on a convolutional neural network structure block and a Swintransformer structure block is constructed; then, the plateau photovoltaic power station passable area dataset is used to train the designed hybrid encoder network to obtain a passable area detection model; finally, the model is used to detect the plateau photovoltaic power station image, and the passable area detection result of pixel-by-pixel classification is output. This method effectively utilizes the spatial features and contextual information of the image, and improves the accuracy, adaptability and robustness of the model in detecting passable areas in complex environments.
[0058] In this embodiment, Step 1 is specifically:
[0059] Step 1.1: Use the plateau photovoltaic power station inspection robot to collect images of the plateau photovoltaic power station, and perform light uniformity and denoising on the images to eliminate uneven lighting and noise interference in the images and improve the image quality and clarity.
[0060] Specifically, the camera carried by the inspection robot of the plateau photovoltaic power station is used for image acquisition. By controlling the exposure time and aperture size of the camera and adjusting the camera's parameters such as sensitivity and focal length, all-round and multi-angle shooting of the plateau photovoltaic power station can be achieved.
[0061] After shooting, the collected images are processed for light uniformity and noise reduction. The light uniformity processing can adopt methods such as histogram equalization and linear stretching to make the brightness distribution of the image more uniform and improve the overall visual effect of the image.
[0062] In this embodiment, the denoising process in Step 1.1 is specifically as follows:
[0063] (1) Use a median filter to filter the image and remove small noise points in the image;
[0064] (2) Remove small burrs in the image through opening operation.
[0065] In this embodiment, in operation (1), specifically: when finding the median of the pixel point f(x, y) in the neighborhood window, first find the average of the pixels in the window; if the pixel point f(x, y) is not greater than the average, no filtering is performed; if the pixel point f(x, y) is greater than the average, a recursive method is used to find the median g(x, y), and the median g(x, y) is used to replace the original pixel point f(x, y).
[0066] Step 1.2: Label the images after uniform illumination and denoising, and construct a dataset of accessible areas for highland photovoltaic power stations.
[0067] Specifically, the image after uniform light and denoising is segmented to distinguish the passable area and the obstacle area. Binarization, edge detection and other methods can be used for image segmentation, so that the passable area and the obstacle area can be accurately identified and marked.
[0068] After the labeling is completed, the datasets of the labeled traversable areas and obstacle areas are stored and managed for subsequent use.
[0069] Step 1.3: Divide the traversable area and obstacle datasets into training set, validation set and test set respectively.
[0070] In this embodiment, in Step 1.3, 70% of the traversable area and obstacle data sets are training sets, 20% are validation sets, and 10% are test sets.
[0071] Specifically, the labeled passable area and obstacle datasets are divided into training set, validation set and test set in a ratio of 70%:20%:10%. The training set is used to train the robot's autonomous navigation and obstacle avoidance control model, the validation set is used to verify the accuracy and generalization ability of the model, and the test set is used to test the final effect and performance of the model.
[0072] In this embodiment, Step 2 is specifically:
[0073] Step 2.1: The ST encoder consists of a linear embedding block and multiple cascaded ST encoding blocks. Each ST encoding block consists of a spatial interaction module SIM and a feature compression module FCM.
[0074] In this embodiment, the spatial interaction module SIM is used to establish pixel-level correlation to enhance the feature representation capability of the occluded object. Specifically, the spatial interaction module SIM generates a feature vector for each pixel that represents its spatial position and the relationship between the surrounding pixels by analyzing the correlation between adjacent pixels. This feature vector contains not only the information of the pixel itself, but also the information of the surrounding pixels, thereby enhancing the feature representation capability of the occluded object in the image.
[0075] The feature compression module FCM is used to compress the output of each ST coding block to reduce the amount of computation and storage overhead. Specifically, the feature compression module FCM compresses the output of the ST coding block through a convolutional neural network, converting the high-dimensional feature vector into a low-dimensional representation while retaining important feature information.
[0076] Step 2.2: Use the spatial interaction module SIM to establish pixel-level correlation to encode the spatial information in the ST coding block to enhance the feature representation capability of the occluded object; more specifically, the spatial interaction module SIM extracts the feature vector of each pixel through a convolutional neural network; the convolutional neural network uses a lightweight convolution kernel and a smaller batch normalization layer to reduce the amount of computation and storage overhead while ensuring the feature extraction effect; after extracting the feature vector of each pixel, the spatial interaction module SIM inputs it into the subsequent ST coding block, and performs feature compression and dimensionality reduction operations together with the feature compression module FCM.
[0077] Step 2.3: The global dependencies from the ST encoder are hierarchically integrated into the features of the CNN through the relation aggregation module RAM.
[0078] Among them, the relationship aggregation module RAM analyzes and processes the output of the ST encoder and converts it into a feature representation suitable for the target segmentation task. Specifically, the relationship aggregation module RAM first globally aggregates the output of the ST encoder and integrates the feature vectors of each pixel into a global feature vector. This global feature vector contains the feature information of the entire image and can more comprehensively reflect the content and structure of the image. The relationship aggregation module RAM inputs the global feature vector into a hierarchical decoder and converts it into a feature representation suitable for the target segmentation task. The decoder is implemented using a convolutional neural network. During the decoding process, the relationship aggregation module RAM can gradually restore the resolution and spatial structure information of the features through layer-by-layer deconvolution, and gradually extract more advanced feature representations.
[0079] In this embodiment, in step 2.3, specifically:
[0080] (1) Use CNN to extract features from the input image and obtain low-level and high-level features of the image;
[0081] (2) Using the ST encoder to extract global dependencies from the input image to obtain the global context information of the image;
[0082] (3) Taking the features extracted by CNN and the global dependencies extracted by ST as input, the RAM is used to perform relationship aggregation;
[0083] (4) The aggregated features are fused with the original CNN features to obtain enhanced image features.
[0084] Step 2.4: The output of the encoder is passed through a 2×2 deconvolution layer to expand the resolution; a skip layer is used to connect the encoder and decoder, while reducing the number of channels passing through the 3×3 convolution layer; each convolution layer is accompanied by a BN layer and a ReLU layer; the above process is performed 4 times, so that the features are gradually restored to 1 / 4 of the original image, and then the segmentation result is obtained through a 3×3 convolution layer and linear interpolation upsampling.
[0085] In this embodiment, Step 3 is specifically:
[0086] Step 3.1: Set the training parameters, including at least batch size, learning rate, number of detection categories and maximum number of iterations;
[0087] Step 3.2: Input the images of the training set into the hybrid encoder network, calculate the loss of the prediction results, use the loss for back propagation, and calculate the gradient; update the gradient and model parameters in real time, and use the validation set to verify the network accuracy;
[0088] In this embodiment, in step 3.2, specifically:
[0089] (1) Load the images of the training set into memory and forward propagate each training image through the hybrid encoder network;
[0090] (2) Calculate the network loss value according to the loss function to quantify the difference between the network prediction and the actual label;
[0091] (3) Using the chain rule, differentiate each weight, calculate the gradient through the loss value, and use the gradient to update the weight of the network;
[0092] (4) Using the calculated gradient, update the model parameters through the Adam optimizer;
[0093] (5) At the end of each round of training, the images of the validation set are input into the network to obtain the prediction results on the validation set, and the accuracy and recall on the validation set are calculated;
[0094] (6) Repeat the above steps (1) to (5) until the preset maximum number of iterations is reached.
[0095] In this embodiment, the loss function is a mean square error function or a cross entropy loss function.
[0096] Step 3.3: Repeat Step 3.1 to Step 3.2. After reaching the maximum number of iterations, the hybrid encoder network completes learning and obtains the weights of the drivable area prediction model.
[0097] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
[0098] The above is an exemplary description of the present invention in conjunction with the accompanying drawings. It is obvious that the implementation of the present invention is not limited to the above-mentioned method. As long as various improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A hybrid encoder network method for identifying traversable areas in plateau mountains. It is characterized in that The following steps are involved: Step 1: Construct a dataset of accessible areas for highland photovoltaic power stations; Step 2: Use the encoder part in UNet as the main encoder and the ST encoder as the auxiliary encoder. The two are parallelized to form a dual encoder structure to form a hybrid encoder network based on the convolutional neural network building block and the transformer building block. Step 3: Use the traversable area dataset of the plateau photovoltaic power station to train the designed hybrid encoder network to obtain the traversable area detection model; Step 4: Use the model to detect the plateau photovoltaic power station image and output the passable area detection results classified pixel by pixel.
2. A method for identifying traversable areas in plateaus and mountains using a hybrid encoder network as claimed in claim 1, Features: Step 1 is as follows: Step 1.1: Use the inspection robot for high-altitude photovoltaic power stations to collect images of high-altitude photovoltaic power stations, and perform light homogenization and denoising on the images; Step 1.2: Label the images after uniform light and denoising, and construct a dataset of accessible areas for photovoltaic power stations on the plateau; Step 1.3: Divide the traversable area and obstacle datasets into training set, validation set and test set respectively.
3. The plateau mountain passable area identification method of a hybrid encoder network as claimed in claim 2, Features: In Step 1.3, 70% of the traversable area and obstacle data sets are training sets, 20% are validation sets, and 10% are test sets.
4. The method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 1, Features: Step 2 is as follows: Step 2.1: The ST encoder consists of a linear embedding block and multiple cascaded ST encoding blocks. Each ST encoding block consists of a spatial interaction module SIM and a feature compression module FCM. Step 2.2: Use the spatial interaction module SIM to establish pixel-level correlation to encode the spatial information in the ST coding block to enhance the feature representation capability of the occluded object; Step 2.3: The global dependencies from the ST encoder are hierarchically integrated into the features of the CNN through the relationship aggregation module RAM; Step 2.4: The output of the encoder is passed through a 2×2 deconvolution layer to expand the resolution; a skip layer is used to connect the encoder and decoder, while reducing the number of channels passing through the 3×3 convolution layer; each convolution layer is accompanied by a BN layer and a ReLU layer; the above process is performed 4 times, so that the features are gradually restored to 1 / 4 of the original image, and then the segmentation result is obtained through a 3×3 convolution layer and linear interpolation upsampling.
5. The method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 1, Features: Step 3 is as follows: Step 3.1: Set the training parameters, including at least batch size, learning rate, number of detection categories and maximum number of iterations; Step 3.2: Input the images of the training set into the hybrid encoder network, calculate the loss of the prediction results, use the loss for back propagation, and calculate the gradient; update the gradient and model parameters in real time, and use the validation set to verify the network accuracy; Step 3.3: Repeat Step 3.1 to Step 3.
2. After reaching the maximum number of iterations, the hybrid encoder network completes learning and obtains the weights of the drivable area prediction model.
6. A method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 2, Features: The denoising process in Step 1.1 is as follows: (1) Use a median filter to filter the image and remove small noise points in the image; (2) Remove small burrs in the image through opening operation.
7. A method for identifying traversable areas in plateaus and mountains using a hybrid encoder network as claimed in claim 6, Features: In operation (1), specifically: when finding the median of a pixel point f(x, y) in a neighborhood window, first find the average of the pixels in the window. If the pixel point f(x, y) is not greater than the average, no filtering is performed. If the pixel point f(x, y) is greater than the average, a recursive method is used to find the median g(x, y), and the median g(x, y) is used to replace the original pixel point f(x, y).
8. The method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 4, Features: In step 2.3, specifically: (1) Use CNN to extract features from the input image and obtain low-level and high-level features of the image; (2) Using the ST encoder to extract global dependencies from the input image to obtain the global context information of the image; (3) Taking the features extracted by CNN and the global dependencies extracted by ST as input, the RAM is used to perform relationship aggregation; (4) The aggregated features are fused with the original CNN features to obtain enhanced image features.
9. The method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 5, Features: In step 3.2, specifically: (1) Load the images of the training set into memory and forward propagate each training image through the hybrid encoder network; (2) Calculate the network loss value according to the loss function to quantify the difference between the network prediction and the actual label; (3) Using the chain rule, differentiate each weight, calculate the gradient through the loss value, and use the gradient to update the weight of the network; (4) Using the calculated gradient, update the model parameters through the Adam optimizer; (5) At the end of each round of training, the images in the validation set are input into the network to obtain the prediction results on the validation set. And calculate its precision and recall on the validation set; (6) Repeat the above steps (1) to (5) until the preset maximum number of iterations is reached.
10. The method for identifying traversable areas in plateau and mountainous areas using a hybrid encoder network as claimed in claim 9, Features: The loss function is the mean square error function or the cross entropy loss function.