Seismic exploration work area road automatic identification method and device
By using deep learning technology in the seismic exploration area, a deep neural network model for road target segmentation is built, the problem of uneven distribution of detection points is solved, efficient automatic road recognition is achieved, and the accuracy and work efficiency of data interpretation are improved.
Patent Information
- Application Number
- CN202311646987.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
In the seismic exploration area, the field environment is complex, and the existence of obstacles such as roads and buildings leads to uneven distribution of detection points, affecting the uniformity of the number of coverages, thereby reducing the accuracy of subsequent data processing and interpretation.
A deep learning-based method is adopted to build a deep neural network model for road target segmentation, specifically adopting a dual-link network (DLinkNet) framework, and on the basis of it, a hybrid converter (Mix-Transformer) and a hollow space pyramid module are introduced to achieve automatic identification of road targets.
By automatically identifying road targets, the coverage uniformity in the seismic exploration work area is improved, the time and cost of manual analysis is reduced, and the accuracy and work efficiency of data interpretation are improved.
Smart Images

Figure CN120107908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent recognition of remote sensing images in seismic exploration areas, and more specifically, to a method and device for automatic recognition of roads in seismic exploration areas. Background Art
[0002] In the field collection of seismic data, the uniformity of the number of coverages is an important criterion for judging the quality of the observation system. The number of coverages reflects the intensity of repeated irradiation of the seismic wave signal on the target area, which determines the adequacy of the detection data at each point in the work area. Uniform coverage is the basis for ensuring the interpretability of the acquisition results. However, due to the complex field environment, the existence of obstacles such as roads and buildings often disrupts the standard layout of the observation system, resulting in hollow areas or clustered areas in the distribution of detection points, resulting in uneven distribution of coverage times in space. This will directly reduce the accuracy of subsequent data processing and interpretation. In severe cases, the complete information of the stratigraphic structure or target body may not be inverted due to insufficient data quality in the local section.
[0003] Therefore, it is necessary to mark the location of obstacles and optimize the observation system before the observation begins to improve the uniformity of coverage. This is usually called "observation system change". The first step of the change is to collect remote sensing images in the work area to obtain surface information from a bird's eye view. This type of remote sensing image mainly comes from high-resolution satellite images and low-altitude drone aerial images. The key point is that it clearly shows the spatial distribution of ground elements, especially the direction of linear targets, such as roads, trails, pipelines, etc. These linear targets are the main sources of obstacles that affect the layout of observation points.
[0004] After acquiring remote sensing images, the next key step is to analyze the obstacle targets in the image and output their spatial distribution information. This falls into the category of remote sensing image analysis, which is mainly completed through target recognition and segmentation. In the past, it was usually necessary to manually analyze remote sensing images and draw obstacle distribution maps. However, this method is inefficient and can still be dealt with in the early stages of construction. However, as the scope of field exploration continues to expand, the intensity of work will increase dramatically, requiring a lot of manpower and material resources. According to statistics, it often takes several people a week to manually draw a large-scale remote sensing analysis map of a work area. Therefore, it is urgent to use automation technology to assist this process. Summary of the invention
[0005] In view of this, the present invention discloses an automatic road identification solution for seismic exploration work areas, which can automatically identify road targets quickly and accurately, and provide a basis for realizing physical point layout, obstacle avoidance and collection in the field.
[0006] According to one aspect of the present invention, a method for automatically identifying roads in a seismic exploration area is proposed, the method comprising:
[0007] S1. Obtain a remote sensing image dataset for road target segmentation in a seismic exploration area, and divide the data in the remote sensing image dataset into training data, verification data, and test data;
[0008] S2. Construct a deep neural network model for road target segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and introduces a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module into the dual link network model framework;
[0009] S3. Apply the constructed deep neural network model to the training data and validation data for model training;
[0010] S4. Apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
[0011] In some implementations, step S1 specifically includes:
[0012] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0013] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0014] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0015] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0016] In some implementations, step S2 specifically includes:
[0017] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0018] The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence.
[0019] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0020] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0021] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0022] The decoding module is configured to include 4 deconvolution blocks.
[0023] In some implementations, in step S3, during model training, the learning rate is adjusted according to the number of training iterations.
[0024] In some implementations, in step S4, based on the obtained test result graph and the corresponding label image, the road segmentation accuracy of the trained deep neural network model is evaluated, specifically including:
[0025] The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0026]
[0027] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0028] According to one aspect of the present invention, a device for automatically identifying roads in a seismic exploration area is also provided, the device comprising:
[0029] A remote sensing image data set acquisition unit is used to acquire a remote sensing image data set for segmenting road targets in a seismic exploration area, and divide the data in the remote sensing image data set into training data, verification data, and test data;
[0030] A model building unit, used to build a deep neural network model for road object segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module are introduced into the dual link network model framework;
[0031] A model training unit, used to apply the built deep neural network model to training data and validation data for model training;
[0032] The model evaluation unit is used to apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
[0033] In some implementations, the remote sensing image data set acquisition unit is specifically used to:
[0034] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0035] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0036] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0037] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0038] In some embodiments, it is characterized in that the model building unit is specifically used for:
[0039] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0040] The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence.
[0041] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0042] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0043] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0044] The decoding module is configured to include 4 deconvolution blocks.
[0045] In some embodiments, in the model training unit, during model training, the learning rate is adjusted according to the number of training iterations.
[0046] In some implementations, in the model evaluation unit, the road segmentation accuracy of the trained deep neural network model is evaluated based on the obtained test result image and the corresponding label image, specifically including:
[0047] The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0048]
[0049] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0050] According to another aspect of the present invention, an electronic device is also provided, the electronic device comprising:
[0051] A memory storing executable instructions;
[0052] A processor runs the executable instructions in the memory to implement the above-mentioned method for automatically identifying roads in a seismic exploration area.
[0053] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for automatically identifying roads in a seismic exploration area described above is implemented.
[0054] The present invention discloses an automatic road recognition scheme for seismic exploration areas using a deep learning method. According to the present invention, a Mix-Transformer encoder structure and a hole space pyramid module are introduced on the basis of the original DLinkNet network, thereby enhancing the network feature expression capability and obtaining multi-scale feature information of roads. The present invention can accurately segment road targets and has good generalization ability.
[0055] The beneficial effects of the present invention are further analyzed in detail below.
[0056] 1) Improve segmentation accuracy
[0057] Through data set enhancement, network optimization, hyperparameter adjustment and other methods, this method has greatly improved the segmentation accuracy compared with the existing technology, and the output results are more accurate.
[0058] 2) Strong generalization performance
[0059] Rich image sample training and data enhancement methods enable the network model to achieve excellent road segmentation effects in different scenarios and types, with strong generalization capabilities.
[0060] 3) Improve work efficiency
[0061] The present invention realizes the automatic recognition and segmentation of roads in seismic exploration areas. Compared with manual analysis of remote sensing images, the use of the deep network can batch process data from a large area in a very short time, greatly improving work efficiency.
[0062] 4) Reduce labor costs
[0063] Since automated processing is achieved, a large amount of repetitive manual work is reduced, thus significantly reducing labor cost expenditure.
[0064] 5) Optimize follow-up work
[0065] Accurate road segmentation results lay the foundation for subsequent observation system design, collection point layout and data processing, making the entire process smoother and more efficient.
[0066] 6) Improve data interpretation
[0067] Ultimately, through the optimization of a series of processes, it is beneficial to improve the coverage uniformity of the acquisition system data, reduce human errors in the interpretation process, and improve the accuracy of seismic data analysis and interpretation.
[0068] In general, the present invention has important guiding significance and application promotion effect on improving the work level in the field of seismic exploration.
[0069] The methods and apparatus of the present invention have other features and advantages that will be apparent from or will be described in detail in the accompanying drawings and subsequent detailed descriptions incorporated herein, which together serve to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present invention.
[0071] Figure 1 A flow chart of a method for automatically identifying roads in a seismic exploration area according to an embodiment of the present invention is shown.
[0072] Figure 2 A schematic structural diagram of a Mix-Transformer encoder structure according to an embodiment of the present invention is shown.
[0073] Figure 3 (a) and (b) respectively show a remote sensing image and a road target segmentation test result diagram obtained by predicting the remote sensing image according to an embodiment of the present invention. DETAILED DESCRIPTION
[0074] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0075] Example 1
[0076] Figure 1 The flowchart of the method for automatically identifying roads in a seismic exploration area according to an embodiment of the present invention is shown. As shown in the figure, the method includes steps S1 to S4.
[0077] Step S1, obtaining a remote sensing image data set for road target segmentation in a seismic exploration area, and dividing the data in the remote sensing image data set into training data, verification data, and test data.
[0078] You can first obtain the original remote sensing images, and then preprocess the original remote sensing images to add a large amount of training data, thereby improving the generalization ability of the model.
[0079] The original remote sensing images and the corresponding labeled images can be preprocessed to obtain a large amount of synthetic training data, and the original data and the synthetic data together constitute a remote sensing image dataset.
[0080] In some implementations, step S1 specifically includes:
[0081] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0082] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0083] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0084] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0085] Original remote sensing images are usually recorded images of the measured target obtained from spherical or space sensors. According to the distance between the sensor and the target, they can be divided into space remote sensing images and low-altitude remote sensing images.
[0086] Space remote sensing images are mainly obtained by satellite payload sensors, with long shooting distances, wide fields of view, and low resolution. Surface images are usually obtained using optical or radar sensors carried by geosynchronous satellites. Satellites scan the earth's surface in specific orbits and revisit cycles to collect images.
[0087] Low-altitude remote sensing images are mainly obtained through sensors on low-altitude platforms such as aircraft and drones. They are closer to the target and have higher resolution. Visible light, multi-spectral and infrared cameras are more commonly used. They can be obtained through aircraft equipped with sensor payloads such as digital cameras and cameras. Stereoscopic or single-view surface images can be obtained by flying to quickly obtain high-resolution images. Handheld digital cameras and video cameras can also be used to directly obtain ground or low-altitude images. In this case, point sampling is usually used, which has a large workload and high resolution.
[0088] The original remote sensing image may be annotated (eg, manually annotated) to obtain a label image corresponding to the original remote sensing image.
[0089] This embodiment generates a synthetic remote sensing image by image cropping, transformation and enhancement, which has the following advantages:
[0090] 1) Enrich the number of samples
[0091] By cropping, rotating, scaling and translating, a large number of new samples can be synthesized from a small number of labeled samples to expand the size of the data set;
[0092] 2) Improve sample diversity
[0093] Image enhancement such as hue, contrast, and illumination perturbs the original sample characteristics, increasing the diversity and complexity of the sample styles;
[0094] 3) Strengthen model robustness
[0095] The model is exposed to more variable data during training, which enhances its adaptability to image deformation and noise and improves the robustness of the model.
[0096] 4) Improve model generalization ability
[0097] Rich and diverse training data enables the model to see a wider range of data distribution, which is less likely to overfit and is conducive to improving the model's theoretical generalization ability;
[0098] 5) Reduce labor costs
[0099] By generating samples in batches through a program, a large number of manual tasks are reduced, significantly reducing the labor cost of synthesizing samples.
[0100] These image enhancement methods can efficiently generate large-scale and diverse training data through algorithm manipulation while ensuring sample details, reduce manual workload, and effectively improve model performance.
[0101] The final remote sensing image dataset contains the original remote sensing images and synthetic data generated through data augmentation, which can significantly enhance the stability and migration ability of the model and is an effective means to improve the generalization performance of the model.
[0102] The acquired remote sensing image dataset is divided into training data, validation data, and test data.
[0103] Training data refers to the data in remote sensing image data used to train model parameters and network weights. The training set needs to be sufficient in number and cover representative samples. In some examples, the training data set accounts for about 80% of the entire remote sensing image data set.
[0104] Validation data refers to data used to guide hyperparameter optimization during the training process. During the training iteration, the validation set can evaluate the model performance, such as validation loss, accuracy, etc. In some examples, the validation data accounts for about 10% of the entire remote sensing image dataset.
[0105] Test data refers to data used to evaluate the final generalization performance of the trained model. The test data set participates in model selection but not in model training. In some examples, the test data accounts for about 10% of the entire remote sensing image dataset.
[0106] By dividing the remote sensing image dataset into training data, validation data, and test data, the model training and evaluation process can be made more fair and objective, and the generalization ability of the model can be evaluated to prevent overfitting of the training data.
[0107] Step S2, constructing a deep neural network model for road target segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and introducing a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module into the dual link network model framework.
[0108] DLinkNet is a deep learning semantic segmentation network model framework. Its main features are:
[0109] 1) Based on the Encoder-Decoder framework;
[0110] 2) A skip connection is added between the encoder and decoder;
[0111] 3) The decoder module uses consecutive deconvolution layers;
[0112] 4) Short-circuit connections can directly transmit detailed features and enhance the semantic information of the model.
[0113] The overall idea of DLinkNet is to extract the semantic features of the image through the encoder, the decoder gradually upsamples to restore the spatial information, and the short-circuit connection provides additional low-level detail features to help more accurate semantic segmentation.
[0114] The inventor believes that compared with other segmentation networks such as FCN and SegNet, the advantage of DLinkNet lies in the addition of short-circuit connection structure, which not only ensures the richness of semantic features, but also makes the boundary of segmentation results clearer. Therefore, DLinkNet can often achieve better segmentation performance, so DLinkNet is especially suitable for remote sensing images targeted by the present invention.
[0115] In some embodiments, constructing the deep neural network model specifically includes:
[0116] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0117] The encoding module is configured to include an Overlap Patch Embeddings layer, a Mix-Transformer encoder structure, and a Dilated Spatial Pyramid module in sequence.
[0118] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0119] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0120] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0121] The decoding module is configured to include 4 deconvolution blocks.
[0122] The Mix-Transformer encoder structure according to this embodiment is composed of 4 MiT Blocks, and the global modeling capability is obtained through the self-attention mechanism. MiT Block (Mix Transformer Block) is the basic component module of the Mix-Transformer encoder structure.
[0123] The overall structure of MiT Block follows the standard Transformer architecture design. According to an example of the present invention, each Block contains two core components: Efficient Self-Attention and Mix FeedForward Network (Mix-FFN). The Efficient Self-Attention mechanism can model the global dependencies of the input sequence and obtain global semantic information. Mix-FFN can perform feature transformation to obtain rich feature expression capabilities.
[0124] Each MiT Block can obtain global context information by calculating the correlation between features through self-attention based on the input image feature sequence. Multiple MiT Blocks are stacked and used. Through this multi-level modeling, the network can extract semantic features that express complex targets.
[0125] The inventors found that if a two-dimensional image is input, it may not be conducive to the calculation of self-attention. Therefore, before the Mix-Transformer encoder extracts features, the Overlap Patch Embeddings layer is set to downsample the input remote sensing image by 4 times to obtain a feature map that is more in line with the U-shaped network, and then flatten it horizontally to convert it into a one-dimensional sequence input MiTBlock. The sequence passes through 4 MiT Blocks in succession, and each Block restores the sequence back to two dimensions when it is output. A mechanism for flattening the two-dimensional image horizontally into a one-dimensional sequence is set between two adjacent MiTBlocks to gradually obtain high-resolution underlying feature maps with details such as contours, colors, and textures, and low-resolution high-level feature maps with rich semantics.
[0126] In this implementation, a hollow spatial pyramid module is used to integrate multi-scale semantic information, which not only enhances the feature representation capability but also improves the computational efficiency.
[0127] In this implementation, the output of the Mix-Transformer encoding structure and the sum of the output of each hole convolution layer are used as the input of the decoding module. The output feature map of the Mix-Transformer encoding structure contains rich semantic information, and the hole convolution can obtain multi-scale context information. Adding the two as the input of the decoding module has the following advantages: the semantic features provided by the mixed transformer enable accurate acquisition of target area information during decoding; the multi-scale context of the hole convolution enables richer and clearer details to be restored during decoding; the two features complement each other, and the addition and fusion make the decoder input information more complete. In addition, the integration of multi-source features intensively concentrates the context and semantic information, accelerating the optimization convergence of the decoding process.
[0128] Therefore, through this skip connection, the decoder has sufficient prior information, which is beneficial to improve the efficiency and effect of the segmentation network.
[0129] In this embodiment, in the decoder module, four deconvolution blocks are used to perform symmetrical upsampling and feature fusion to gradually restore the spatial structure information and achieve synchronous recovery of feature map resolution and semantic information.
[0130] Step S3, applying the constructed deep neural network model to the training data and verification data for model training.
[0131] In some embodiments, during model training, the learning rate is adjusted according to the number of training iterations. The learning rate is an important hyperparameter in network training. Adjusting the learning rate during training can make the network reach the fastest convergence speed while avoiding overfitting problems. The hyperparameter optimization is described by taking the learning rate hyperparameter as a representative. After a certain number of iterations of network model training, the learning rate is multiplied by a certain attenuation factor. Generally, the learning rate is multiplied by 0.5 after five iterations, or by 0.1 after twenty iterations.
[0132] Step S4, applying the trained deep neural network model to the test data to obtain a road target segmentation test result map, and based on the obtained test result map and the corresponding label image, evaluating the road segmentation accuracy of the trained deep neural network model.
[0133] You can write test code, load the deep neural network model, use test data for testing, and get the road segmentation results under the test data. In some examples, the output segmentation result is a two-dimensional mask image.
[0134] In some implementations, the intersection over union (IoU) of the test result image and the label image may be calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0135]
[0136] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0137] When the output is a two-dimensional mask image, k=1 in the above formula.
[0138] According to this embodiment, the intersection over union (IoU) is calculated to quantitatively evaluate the road segmentation accuracy of the deep neural network model, so that the evaluation criteria can be uniformly quantified, and the matching between the predicted segmentation and the true value area is fully considered in the evaluation, emphasizing the importance of correct matching. The results are intuitive and help to clarify the direction of further optimization.
[0139] This embodiment discloses a solution for automatic road identification in seismic exploration areas based on a deep learning method. According to this embodiment, a Mix Transformer encoder structure and a hole space pyramid module are introduced on the basis of the original DLinkNet network to enhance the network feature expression capability and obtain multi-scale feature information of roads. This embodiment can accurately segment road targets and has good generalization ability.
[0140] Example 2
[0141] According to one embodiment of the present invention, a device for automatically identifying roads in a seismic exploration area is provided, the device comprising:
[0142] A remote sensing image data set acquisition unit is used to acquire a remote sensing image data set for segmenting road targets in a seismic exploration area, and divide the data in the remote sensing image data set into training data, verification data, and test data;
[0143] A model building unit, used to build a deep neural network model for road object segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module are introduced into the dual link network model framework;
[0144] A model training unit, used to apply the built deep neural network model to training data and validation data for model training;
[0145] The model evaluation unit is used to apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
[0146] In some implementations, the remote sensing image data set acquisition unit is specifically used to:
[0147] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0148] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0149] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0150] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0151] In some embodiments, it is characterized in that the model building unit is specifically used for:
[0152] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0153] The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence.
[0154] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0155] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0156] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0157] The decoding module is configured to include 4 deconvolution blocks.
[0158] In some embodiments, in the model training unit, during model training, the learning rate is adjusted according to the number of training iterations.
[0159] In some implementations, in the model evaluation unit, the road segmentation accuracy of the trained deep neural network model is evaluated based on the obtained test result image and the corresponding label image, specifically including:
[0160] The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0161]
[0162] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P jj represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0163] This embodiment discloses a solution for automatic road identification in seismic exploration areas based on a deep learning method. According to this embodiment, a Mix Transformer encoder structure and a hole space pyramid module are introduced on the basis of the original DLinkNet network to enhance the network feature expression capability and obtain multi-scale feature information of roads. This embodiment can accurately segment road targets and has good generalization ability.
[0164] For other detailed descriptions and advantages of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0165] Example 3
[0166] According to another aspect of the present invention, an electronic device is provided. The electronic device comprises:
[0167] Memory, which stores executable instructions:
[0168] A processor runs the executable instructions in the memory to implement the method for automatically identifying roads in a seismic exploration area according to the present invention.
[0169] The method comprises the following steps:
[0170] S1. Obtain a remote sensing image dataset for road target segmentation in a seismic exploration area, and divide the data in the remote sensing image dataset into training data, verification data, and test data;
[0171] S2. Construct a deep neural network model for road target segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and introduces a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module into the dual link network model framework;
[0172] S3. Apply the constructed deep neural network model to the training data and validation data for model training;
[0173] S4. Apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
[0174] In some implementations, step S1 specifically includes:
[0175] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0176] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0177] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0178] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0179] In some implementations, step S2 specifically includes:
[0180] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0181] The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence.
[0182] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0183] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0184] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0185] The decoding module is configured to include 4 deconvolution blocks.
[0186] In some implementations, in step S3, during model training, the learning rate is adjusted according to the number of training iterations.
[0187] In some implementations, in step S4, based on the obtained test result graph and the corresponding label image, the road segmentation accuracy of the trained deep neural network model is evaluated, specifically including:
[0188] The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0189]
[0190] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0191] Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc.
[0192] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present invention, the processor is used to run the computer-readable instructions stored in the memory.
[0193] This embodiment discloses a solution for automatic road identification in seismic exploration areas based on a deep learning method. According to this embodiment, a Mix Transformer encoder structure and a hole space pyramid module are introduced on the basis of the original DLinkNet network to enhance the network feature expression capability and obtain multi-scale feature information of roads. This embodiment can accurately segment road targets and has good generalization ability.
[0194] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0195] Example 4
[0196] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the method for automatically identifying roads in a seismic exploration area according to the present invention is implemented.
[0197] The method comprises the following steps:
[0198] S1. Obtain a remote sensing image dataset for road target segmentation in a seismic exploration area, and divide the data in the remote sensing image dataset into training data, verification data, and test data;
[0199] S2. Construct a deep neural network model for road target segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and introduces a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module into the dual link network model framework;
[0200] S3. Apply the constructed deep neural network model to the training data and validation data for model training;
[0201] S4. Apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
[0202] In some implementations, step S1 specifically includes:
[0203] S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image;
[0204] S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256;
[0205] S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color;
[0206] S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
[0207] In some implementations, step S2 specifically includes:
[0208] A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module;
[0209] The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence.
[0210] The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter.
[0211] The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module.
[0212] The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module;
[0213] The decoding module is configured to include 4 deconvolution blocks.
[0214] In some implementations, in step S3, during model training, the learning rate is adjusted according to the number of training iterations.
[0215] In some implementations, in step S4, based on the obtained test result graph and the corresponding label image, the road segmentation accuracy of the trained deep neural network model is evaluated, specifically including:
[0216] The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model:
[0217]
[0218] Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
[0219] The computer-readable storage medium according to the embodiment of the present invention stores non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the above-mentioned methods of the embodiments of the present invention are executed.
[0220] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0221] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present invention.
[0222] This embodiment discloses a solution for automatic road identification in seismic exploration areas based on a deep learning method. According to this embodiment, a Mix Transformer encoder structure and a hole space pyramid module are introduced on the basis of the original DLinkNet network to enhance the network feature expression capability and obtain multi-scale feature information of roads. This embodiment can accurately segment road targets and has good generalization ability.
[0223] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0224] Example 5
[0225] Figure 2 FIG. 4 is a schematic diagram of a Mix-Transformer encoding structure according to an embodiment of the present invention.
[0226] As shown in the figure, the remote sensing image F i The input Overlap Patch Merging layer is used to achieve downsampling, and then it is horizontally flattened into a one-dimensional sequence. The input Efficient Self-Attention layer introduces the attenuation ratio R and uses the fully connected layer to reduce the amount of calculation. The reduction in computational complexity enables the network to extract features from higher-resolution remote sensing images. The input Mix-FFN layer is a residual connection layer that contains 2 MLPs, 1 3×3 convolution, and GELU. Using Mix-FFN, the model parameters can be applied to test images of different sizes without reducing performance.
[0227] Example 6
[0228] This embodiment verifies and illustrates the effect of the present invention.
[0229] Figure 3 (a) and (b) respectively show a remote sensing image and a road target segmentation test result diagram obtained by predicting the remote sensing image according to an embodiment of the present invention. It can be seen from the figure that the road target can be accurately segmented according to this embodiment.
[0230] For other detailed descriptions of this exemplary embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0231] In summary, the present invention discloses a solution for automatic road identification in seismic exploration areas based on deep learning methods. According to the present invention, the Mix Transformer encoder structure and the hole space pyramid module are introduced on the basis of the original DLinkNet network to enhance the network feature expression capability and obtain multi-scale feature information of roads. The present invention can accurately segment road targets and has good generalization ability.
[0232] Various embodiments and implementations of the present invention have some or all of the following advantages.
[0233] 1) Improve segmentation accuracy
[0234] Through data set enhancement, network optimization, hyperparameter adjustment and other methods, this method has greatly improved the segmentation accuracy compared with the existing technology, and the output results are more accurate.
[0235] 2) Strong generalization performance
[0236] Rich image sample training and data enhancement methods enable the network model to achieve excellent road segmentation effects in different scenarios and types, with strong generalization capabilities.
[0237] 3) Improve work efficiency
[0238] The present invention realizes the automatic recognition and segmentation of roads in seismic exploration areas. Compared with manual analysis of remote sensing images, the use of the deep network can batch process data from a large area in a very short time, greatly improving work efficiency.
[0239] 4) Reduce labor costs
[0240] Since automated processing is achieved, a large amount of repetitive manual work is reduced, thus significantly reducing labor cost expenditure.
[0241] 5) Optimize follow-up work
[0242] Accurate road segmentation results lay the foundation for subsequent observation system design, collection point layout and data processing, making the entire process smoother and more efficient.
[0243] 6) Improve data interpretation
[0244] Ultimately, through the optimization of a series of processes, it is beneficial to improve the coverage uniformity of the acquisition system data, reduce human errors in the interpretation process, and improve the accuracy of seismic data analysis and interpretation.
[0245] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for automatically identifying roads in seismic exploration areas. It is characterized in that The method comprises: S1. Obtain a remote sensing image dataset for road target segmentation in a seismic exploration area, and divide the data in the remote sensing image dataset into training data, verification data, and test data; S2. Construct a deep neural network model for road target segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and introduces a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module into the dual link network model framework; S3. Apply the constructed deep neural network model to the training data and validation data for model training; S4. Apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
2. The method according to claim 1, It is characterized in that Step S1 specifically includes: S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image; S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256; S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color; S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
3. The method according to claim 1, It is characterized in that Step S2 specifically includes: A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module; The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence. The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter. The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module. The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module; The decoding module is configured to include 4 deconvolution blocks.
4. The method according to claim 1, It is characterized in that In step S3, during model training, the learning rate is adjusted according to the number of training iterations.
5. The method according to claim 1, It is characterized in that In step S4, based on the obtained test result image and the corresponding label image, the road segmentation accuracy of the trained deep neural network model is evaluated, specifically including: The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model: Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
6. An automatic road identification device for seismic exploration areas, It is characterized in that The device comprises: A remote sensing image data set acquisition unit is used to acquire a remote sensing image data set for segmenting road targets in a seismic exploration area, and divide the data in the remote sensing image data set into training data, verification data, and test data; A model building unit, used to build a deep neural network model for road object segmentation, wherein the deep neural network model is based on a dual link network (DLinkNet) model framework, and a mixed transformer (Mix-Transformer) and a hollow spatial pyramid module are introduced into the dual link network model framework; A model training unit, used to apply the built deep neural network model to training data and validation data for model training; The model evaluation unit is used to apply the trained deep neural network model to the test data to obtain a road target segmentation test result map, and evaluate the road segmentation accuracy of the trained deep neural network model based on the obtained test result map and the corresponding label image.
7. The device according to claim 6, It is characterized in that The remote sensing image data set acquisition unit is specifically used for: S11, obtaining an original remote sensing image, and annotating the original remote sensing image to obtain a corresponding label image; S12, cropping the original remote sensing image and the corresponding label image to a size of 1024*1024, and setting the cropping step size to 256; S13, synchronously preprocessing the original remote sensing image data and the corresponding label image to generate a synthetic remote sensing image and a corresponding label image, wherein the preprocessing includes scaling at a ratio of 0.9 to 1.1, and / or rotating between 0 and 360 degrees, and / or translating at -200 to 200 in the X and Y directions, and / or performing HSV adjustment on the color; S14. Randomly divide the remote sensing image data set into training data, verification data and test data in a ratio of 8:1:1, wherein the remote sensing image data set consists of original remote sensing images, synthetic remote sensing images and their corresponding label images.
8. The device according to claim 6, It is characterized in that The model building unit is specifically used for: A dual link network (DLinkNet) is selected as a basic network framework of the deep neural network model, and the basic network framework is configured to include an encoding module and a decoding module; The encoding module is configured to include an Overlap Patch Embeddings layer, a hybrid converter, and a dilated spatial pyramid module in sequence. The overlapping patch embedding layer is configured with a kernel size of 7, a step size of 4, and a padding size of 3. The overlapping patch embedding layer downsamples the input remote sensing image and horizontally flattens the downsampled remote sensing image into a one-dimensional sequence as the input of the hybrid converter. The hybrid converter is configured to include 4 consecutive hybrid conversion blocks (MiT Block), each of which restores the sequence to a two-dimensional image at the output, and the output of the hybrid converter is the input of the hollow spatial pyramid module. The atrous spatial pyramid module is configured to include four atrous convolutional layers, with the atrous parameters set to 1, 2, 4, and 8, respectively, and the output of the hybrid converter plus the output of each atrous convolutional layer is used as the input of the decoding module; The decoding module is configured to include 4 deconvolution blocks.
9. The device according to claim 6, It is characterized in that In the model training unit, during model training, the learning rate is adjusted according to the number of training iterations.
10. The device according to claim 6, It is characterized in that In the model evaluation unit, the road segmentation accuracy of the trained deep neural network model is evaluated based on the obtained test result image and the corresponding label image, specifically including: The intersection over union (IoU) of the test result image and the label image is calculated based on the following formula to evaluate the road segmentation accuracy of the trained deep neural network model: Where k in k+1 represents the number of pixel categories of the target segmentation excluding the background, and 1 represents 1 background category; P ij represents the total number of pixels of category i predicted by the model as pixels of category j, P ji represents the total number of pixels of category j predicted by the model as category i, P ii Represents the total number of pixels of category i predicted by the model.
11. An electronic device, It is characterized in that The electronic device comprises: A memory storing executable instructions; A processor, wherein the processor runs the executable instructions in the memory to implement the method according to any one of claims 1 to 5.
12. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by a processor.