A Method and System for Extracting Mars Surface Rocks Based on a Lightweight Network

Through the lightweight Transformer model and adaptive nuclear convolution combined with enhanced multi-dimensional convolution attention mechanism, the problem of difficulty and low efficiency of Mars rock segmentation network is solved, and efficient and accurate rock identification and dust coverage treatment are achieved.

CN119625501BActive Publication Date: 2025-07-04PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510147844.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-04
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing Mars rock segmentation network has problems such as difficulty in identifying different rock shapes, low computing efficiency, poor robustness, and reduced recognition rate during sand and dust coverage.

Method used

The lightweight Transformer model and adaptive nuclear convolution are adopted, combined with the enhanced multi-dimensional convolution attention mechanism, and the shape and size of the convolution kernel are dynamically adjusted to improve the attention of rock feature extraction and sand and dust covering the rock.

Benefits of technology

It realizes efficient feature extraction of different rocks and accurate identification of sand and dust-covered rocks, improving the calculation efficiency and recognition accuracy of Mars exploration missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625501B_ABST
    Figure CN119625501B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of Mars surface rocks, and specifically discloses a method and system for extracting Mars surface stones based on a lightweight network, including: obtaining Mars surface images captured by a Mars rover and screening rock images; making labels based on the rock images, performing image processing on the labels and rock images to expand the number of images, and dividing the rock images into a training set, a test set, and a validation set; passing the training set through an adaptive kernel convolution to dynamically adjust the sampling shape and size of the convolution kernel so as to extract rock features in the rock images; using an improved lightweight Transformer model to separate the rock images from the surrounding environment; improving the attention of rocks covered by dust through an enhanced multi-dimensional convolutional attention mechanism; evaluating the lightweight Transformer model based on pre-training metrics, if the evaluation passes to obtain a final model, and based on the final model, performing rock surface extraction on the input Mars rock images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Martian surface rocks, and in particular to a method and system for extracting Martian surface rocks based on a lightweight network. Background Art

[0002] Rocks on Mars are not only important objects for scientific research but also the main obstacles for rovers to travel. As the eyes of the rover, the navigation camera needs to obtain a large amount of image information for path planning and subsequent scientific tasks. Therefore, segmenting rocks from images is the key to subsequent exploration tasks. Image segmentation methods in Mars exploration can also be divided into traditional methods and deep learning methods. The most classic practical application in traditional methods is the rock image segmentation algorithm Rockster (Rock Segmentation Through Edge Regrouping) deployed on the AEGIS system on the Mars probe launched by the National Aeronautics and Space Administration of the United States. However, due to the limitations of the system itself, it is not used much. The Rockster algorithm is mainly based on the Canny edge detector.

[0003] With the introduction of deep learning and visual transformers, many CNN and transformer-based methods have been proposed and used in practical tasks. Kiri et al. from NASA's Jet Propulsion Laboratory used transfer learning to adjust the convolutional neural network AlexNet for Mars images to achieve the classification of Mars images. In addition, NASA and ESA also applied deep convolutional neural networks to the Mars terrain classification task. NASA's Soil Property and Object Classification (SPOC) task is used to process MRO HiRISE orbital images and MSL Navcam Mars surface images, and is applied to the landing site traversability analysis of the Mars 2020 rover mission and the slip prediction analysis of the Mars Science Laboratory (MSL) mission. ESA's NOAH-H (Noveltyor Anomaly Hunter – HiRISE) terrain classification system uses deep neural networks to perform semantic segmentation of HiRISE orbital images to serve the subsequent ExoMars landing mission. DeLatte et al. used the classic U-Net network to segment craters. Furlán et al. optimized U-Net to reduce processing time. Another classic network structure, the Deeplab series of algorithms, has also been applied to Mars image segmentation. Ebadi et al. used the DeepLabV3+ framework to perform semantic segmentation of the skyline of Mars vista images to achieve high-precision positioning of the Mars rover in extreme and GPS-denied environments. Liu et al. designed global and local attention modules to process homogeneous and heterogeneous pixels and aggregate them respectively to achieve good results in panoramic terrain segmentation on Mars, and proposed an MRISNet that combines generative adversarial networks, attention mechanisms, and simple linear iterative clustering (SLIC) algorithms to improve rock segmentation accuracy.

[0004] However, current Mars rock segmentation networks focus on using complex classification networks as encoders to extract multi-scale features, and designing complex decoder networks to establish connections between multiple scales. These improvements come at the expense of complex hardware equipment and a large number of parameters, and do not take into account dust coverage and the difficulty of actual deployment. Therefore, this application proposes a lightweight Mars rock semantic segmentation network.

[0005] Problems found:

[0006] There are still many problems to be solved in the segmentation of Martian rocks:

[0007] First, the shapes of rocks are diverse, and the network does not consider the impact of the fixed shape of ordinary convolution on the recognition of different rocks.

[0008] Second, most current networks adopt a U-shaped structure. However, for Mars exploration missions, a large amount of image data needs to be processed quickly, and the hardware resources carried by itself are extremely limited. The U-shaped structure is not conducive to improving the computing efficiency.

[0009] Third, the robustness of the network structure is weak. When rocks are covered by fine sand, the recognition rate of rocks decreases significantly. Summary of the Invention

[0010] To achieve the object of the present invention, the present application provides a method for extracting rock blocks on the Mars surface based on a lightweight network, including:

[0011] Step S1: Obtain Mars surface images and screen rock images;

[0012] Step S2: Make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set;

[0013] Step S3: Pass the training set through an adaptive kernel convolution to dynamically adjust the sampling shape and size of the convolution kernel to extract rock features in the rock images;

[0014] Step S4: Use an improved lightweight Transformer model to separate the rock images from the surrounding environment;

[0015] Step S5: Improve the attention of rocks covered by dust through an enhanced multi-dimensional convolutional attention mechanism;

[0016] Step S6: Evaluate the lightweight Transformer model based on pre-trained metrics. If the evaluation passes, obtain the final model. Based on the final model, perform rock surface extraction on the input Mars rock images.

[0017] In some specific embodiments, step S3 includes:

[0018] Step S31: Take the received image as input, randomly generate initial coordinates and define the initial sampling position to obtain the initial convolution kernel size;

[0019] Step S32: Calculate the base integer based on the initial convolution kernel size;

[0020] Step S33: Determine the positions of the two-dimensional grid based on the initial convolution kernel size and the base integer;

[0021] Step S34: Perform a convolution operation on the received image to obtain an offset;

[0022] Step S35: Determine new sampling coordinates based on the offset, the initial sampling position, and the position of the two-dimensional grid;

[0023] Step S36: Resample the rock image based on the new sampling coordinates;

[0024] Step S37: Reshape, convolve again, standardize the resampled rock image, and finally output the original pixel information through an activation function.

[0025] In some specific embodiments, step S4 includes:

[0026] Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model;

[0027] Step S42: Integrate the spatial domain information with the original unclustered pixel information.

[0028] In some specific embodiments, step S41 includes:

[0029] Step S411: Perform low-pass information filtering on the information based on average pooling and channel grouping to ensure that the dimensions and fine details of the regenerated image are consistent with the original rock image;

[0030] Step S412: Perform high-pass filtering on the information based on convolutional kernels of different dimensions;

[0031] Step S413: Select frequency information related to semantic segmentation based on a frequency similarity kernel;

[0032] Step S414: Superimpose the low-pass information, the high-pass information, and the selected frequency information to obtain the spatial domain information of different frequency bands.

[0033] In some specific embodiments, step S5 includes:

[0034] Step S51: Divide the input rock image into three branches. The first branch is rotated 90° counterclockwise along channel H to obtain F1, the second branch is rotated 90° counterclockwise along channel W to obtain F2, and the third branch remains unchanged in its original features to obtain F3;

[0035] Step S52: Apply squeeze transformation to the three branches respectively, where the squeeze transformation includes average pooling, standard pooling, and max pooling;

[0036] Step S53: Generate new values through an adaptive combination mechanism for the results of the three branches after the squeeze transformation to determine the results of average pooling, standard pooling, and max pooling;

[0037] Step S54: Pass the three new values through the activation function respectively and average them to determine the feature map of the rock image covered by dust.

[0038] Step S55: Classify the feature map through depthwise separable convolution to obtain the label map of the rock image covered by dust.

[0039] In some specific embodiments, in step S6, the pre-training metrics include: Intersection over Union (IoU), Precision, Pixel Accuracy (PA), Recall, and the F1 score, which is the harmonic mean of Precision and Recall; where,

[0040] ;

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] In the formula, TP is the number of positive instances correctly identified as positive, FP represents the number of negative instances misclassified as positive, TN is the instance correctly identified as negative, and FN is the positive instance misclassified as negative.

[0046] To achieve the same invention purpose, the present application also provides a Mars surface rock extraction system based on a lightweight network, including:

[0047] Image screening module: used for Mars surface images to screen rock images;

[0048] Image classification module: used to make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set;

[0049] Feature extraction module: used to dynamically adjust the sampling shape and size of the convolution kernel of the training set through adaptive kernel convolution to extract the rock features in the rock image;

[0050] Image separation module: used to separate the rock image from the surrounding environment using an improved lightweight Transformer model;

[0051] Feature weighting module: used to increase the weight of the local features of the covered rocks through an enhanced multi-dimensional convolutional attention mechanism;

[0052] Model evaluation module: used to evaluate the lightweight Transformer model based on pre-training metrics. If the evaluation passes, the final model is obtained. Based on the final model, the rock surface is extracted from the input Mars rock image.

[0053] In some specific embodiments, the feature extraction module is used to perform the following steps:

[0054] Step S31: Take the received image as input, randomly generate initial coordinates and define the initial sampling position to obtain the initial convolution kernel size;

[0055] Step S32: Calculate the base integer based on the initial convolution kernel size;

[0056] Step S33: Determine the position of the two-dimensional grid based on the initial convolution kernel size and the base integer;

[0057] Step S34: Perform a convolution operation on the received image to obtain an offset;

[0058] Step S35: Determine new sampling coordinates based on the offset, the initial sampling position, and the position of the two-dimensional grid;

[0059] Step S36: Resample the rock image based on the new sampling coordinates;

[0060] Step S37: Reshape, convolve again, standardize the resampled rock image, and finally output the original pixel information through an activation function.

[0061] In some specific embodiments, the image separation module is used to perform the following steps:

[0062] Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model;

[0063] Step S42: Integrate the spatial domain information with the original unclustered pixel information.

[0064] In some specific embodiments, the feature weighting module is used to perform the following steps:

[0065] Step S51: Divide the input rock image into three branches. The first branch is rotated 90° counterclockwise along channel H to obtain F1, the second branch is rotated 90° counterclockwise along channel W to obtain F2, and the third branch maintains the original features unchanged to obtain F3;

[0066] Step S52: Pass the three branches through a squeezing transformation respectively, where the squeezing transformation includes average pooling, standard pooling, and max pooling;

[0067] Step S53: Generate new values through an adaptive combination mechanism for the results of the three branches after the squeezing transformation to determine the results of average pooling, standard pooling, and max pooling;

[0068] Step S54: Pass the three new values through activation functions respectively and perform averaging to determine the feature map of the rock image covered with dust;

[0069] Step S55: Classify the feature map through depthwise separable convolution to obtain the label map of the rock image covered with dust.

[0070] Beneficial effects:

[0071] (1) During the preprocessing process, the present application proposes a variable kernel convolution, which provides convolution kernels with arbitrary sampling shapes and sizes, replaces ordinary convolutions, and makes up for the deficiencies of conventional convolutions, so as to effectively extract features for rocks of different scales.

[0072] (2) The present application uses an improved Transformer model - AFFormer as the backbone network for semantic segmentation of Mars rocks. This model distinguishes rocks and backgrounds by learning the frequency information of different feature categories and achieves excellent performance with fewer parameters.

[0073] (3) The present application proposes an improved spatial attention mechanism (MCA), which uses a new multi-dimensional attention mechanism to learn three types of attention of all convolution kernels in the convolution kernel space, thereby providing higher attention to rocks covered with dust. Brief description of the drawings

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0075] Figure 1 It is a flowchart of a method for extracting Mars surface stones based on a lightweight network provided by an embodiment of the present invention;

[0076] Figure 2 It is a flowchart of the attention mechanism of a method for extracting Mars surface stones based on a lightweight network provided by an embodiment of the present invention;

[0077] Figure 3 It is an overall network architecture diagram of a method for extracting Mars surface stones based on a lightweight network provided by an embodiment of the present invention;

[0078] Figure 4 Experimental result comparison diagram of a method for extracting rocks on the Martian surface based on a lightweight network provided by an embodiment of the present invention;

[0079] Figure 5 Structure schematic diagram of a system for extracting rocks on the Martian surface based on a lightweight network provided by an embodiment of the present invention. Detailed implementation manners

[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0081] Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0082] Embodiment 1

[0083] An embodiment of the present invention provides a method for extracting rocks on the Martian surface based on a lightweight network. Referring to Figures 1 - 3 as shown, it includes:

[0084] Step S1: Obtain the image of the Martian surface captured by the Mars rover and screen the rock images.

[0085] Step S2: Make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set.

[0086] Step S3: Dynamically adjust the sampling shape and size of the convolution kernel of the training set through adaptive kernel convolution to extract the rock features in the rock images.

[0087] In a specific embodiment of the present invention, step S3 includes:

[0088] Step S31: Take the received image as the input x, randomly generate the initial coordinates to define the initial sampling position P0, and obtain the initial convolution kernel size of k;

[0089] Step S32: Calculate the base integer B based on the initial convolution kernel size k, and the formula is ;

[0090] Step S33: Based on the value of , calculate the new two-dimensional grid P n , that is, P n is ( , );

[0091] Step S34: Perform a convolution operation on the received image to obtain an offset Offset;

[0092] Step S35: Based on the offset Offset, the initial sampling position P0, and the new two-dimensional grid P n Obtain a new sampling coordinate P, and the formula is ;

[0093] Step S36: Resample the feature map based on the new sampling coordinate P;

[0094] Step S37: Reshape, convolve again, normalize the resampled feature map, and finally output the original pixel information through an activation function.

[0095] Step S4: Use an improved lightweight Transformer model to separate the rock image from the surrounding environment. In this application, an adaptive frequency filter with linear complexity is used to replace the standard adaptive mechanism to reduce the number of model parameters and computational complexity.

[0096] In a specific embodiment of the present invention, step S4 includes:

[0097] Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model;

[0098] Step S42: Integrate the spatial domain information with the original unclustered pixel information.

[0099] In a specific embodiment of the present invention, step S41 includes:

[0100] Step S411: Perform low-pass information filtering on the information based on average pooling and channel grouping to ensure that the dimensions and fine details of the regenerated image are consistent with the original rock image;

[0101] Step S412: Perform high-pass filtering on the information based on convolution kernels of different dimensions;

[0102] Step S413: Select frequency information related to semantic segmentation based on a frequency similarity kernel;

[0103] Step S414: Superimpose the low-pass information, the high-pass information, and the selected frequency information to obtain the spatial domain information of different frequency bands.

[0104] Step S5: Increase the weight of the local features of the covered rock through an enhanced multi-dimensional convolutional attention mechanism.

[0105] In a specific embodiment of the present invention, step S5 includes:

[0106] Step S51: Divide the input rock image F into three branches. The first branch F is rotated counterclockwise by 90° along channel H to obtain F1, the second branch F is rotated counterclockwise by 90° along channel W to obtain F2, and the third branch F remains unchanged with its original features;

[0107] Step S52: Apply squeeze transformation to the three branches respectively. The squeeze transformation includes average pooling, standard pooling, and max pooling;

[0108] Step S53: Feed the results of the three pooling operations into an adaptive combination mechanism to generate new values, and its formula is , where is the result of average pooling, is the result of standard pooling, is the result of max pooling;

[0109] Step S54: After passing the three new values through the activation function respectively, perform averaging to obtain the final enhanced output feature map;

[0110] Step S55: Classify the feature map through depthwise separable convolution to obtain the final label map.

[0111] Step S6: Evaluate the lightweight Transformer model based on the pre-trained metrics. If the evaluation passes, obtain the final model. Based on the final model, extract the rock surface of the input Mars rock image.

[0112] In a specific embodiment of the present invention, as shown in reference to Figure 4 , in the latest Mars rock segmentation model, the present application selects the classic NI-U-Net++ and the relatively new MarsNet as representatives for comparison. Specifically, in the TWMARS-V2 dataset, the accuracy of NI-U-Net++ reaches 88.5%, while the accuracy of MarsNet is 92.58%. It can be clearly seen that MarsNet indeed has higher accuracy and precision than some mainstream segmentation models.

[0113] Swin Transformer also shows certain performance, with its pixel accuracy (PA) reaching 88.5%, while the accuracy of MarsNet is 90.2%. Compared with traditional convolution-based models, such as the pixel accuracy of the FastFCN model based on Unet is only 85.36%, and the pixel accuracy of the Unet model is 80.54%. However, the model proposed in the present application shows more excellent performance in this task.

[0114] The model of this application performs excellently in the Mars rock segmentation task. On the MarsData-V2 dataset, the pixel accuracy of the model reaches 98.18%, which is 4.89% higher than MarsNet and 5.51% higher than NI. Further in-depth comparison with other models using the Intersection over Union (IoU) metric, the model of this application reaches 96.62%. In contrast, the IoU of MarsNet is 90.26%, and the IoU of NI-U-Net++ is 89.05%. Among traditional convolution-based models, the IoU of Deeplabv3+ is 94.3%, the IoU of EMO is 90.24%, and the IoU of FastFCN is 87.56%; among unet-based models, the IoU of PspNet is 92.11%; among transformer-based models, the IoU of SegFormer is 92.8%, the IoU of VisionTransformer is 91.3%, and the IoU of Swin Transformer is 93.38%. The model of this application is superior to many comparison models in terms of the IoU metric, which fully demonstrates that the model of this application can not only accurately identify Mars rock pixels but also well define the boundaries of the rock regions. This further proves the advantage of the model of this application in segmentation accuracy, that is, it can not only accurately identify Mars rock pixels but also more precisely define the boundaries of the rock regions. In terms of accuracy, for the main category of Mars rocks, the model of this application reaches an extremely high accuracy of 98.64%. In contrast, the accuracy of MarsNet for this category is 91.53%, and the accuracy of NI-U-Net++ is 91.20%. This shows that the model of this application has a significant advantage in the classification accuracy of Mars rocks and can more effectively distinguish rock pixels from the complex Mars background.

[0115] Through comparative experiments with various advanced models, this application fully demonstrates the effectiveness and superiority of the proposed model in the Mars rock segmentation task, providing a more powerful technical means for rock identification and analysis in Mars exploration. In subsequent research, this application will further explore the optimization direction of the model, further improve its performance, and expand its application scope.

[0116] Example Two

[0117] This application also provides a Mars surface rock extraction system based on a lightweight network. Referring to Figure 5 as shown, it includes:

[0118] An image screening module 10: used to obtain the Mars surface images taken by the Mars rover and screen the rock images;

[0119] Image classification module 20: used to make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set;

[0120] Feature extraction module 30: used to dynamically adjust the sampling shape and size of the convolution kernel through adaptive kernel convolution for the training set to extract rock features in the rock images;

[0121] Image separation module 40: used to separate the rock images from the surrounding environment using an improved lightweight Transformer model;

[0122] Feature weighting module 50: used to increase the weight of the local features of the covered rocks through an enhanced multi-dimensional convolutional attention mechanism;

[0123] Model evaluation module 60: used to evaluate the lightweight Transformer model based on pre-trained metrics. If the evaluation passes to obtain the final model, based on the final model, perform rock surface extraction on the input Mars rock images.

[0124] In some specific embodiments, the feature extraction module 30 is used to perform the following steps:

[0125] Step S31: Take the received image as input, randomly generate initial coordinates and define the initial sampling position to obtain the initial convolution kernel size;

[0126] Step S32: Calculate the base integer based on the initial convolution kernel size;

[0127] Step S33: Determine the position of the two-dimensional grid based on the initial convolution kernel size and the base integer;

[0128] Step S34: Perform a convolution operation on the received image to obtain an offset;

[0129] Step S35: Determine new sampling coordinates based on the offset, the initial sampling position, and the position of the two-dimensional grid;

[0130] Step S36: Resample the rock image based on the new sampling coordinates;

[0131] Step S37: Reshape, convolve again, standardize the resampled rock image, and finally output the original pixel information through an activation function.

[0132] In some specific embodiments, the image separation module 40 is used to perform the following steps:

[0133] Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model;

[0134] Step S42: Integrate the spatial domain information with the original unclustered pixel information.

[0135] In a specific embodiment of the present invention, step S41 includes:

[0136] Step S411: Perform low-pass information filtering on the information based on average pooling and channel grouping to ensure that the dimensions and fine details of the regenerated image are consistent with the original rock image;

[0137] Step S412: Perform high-pass filtering on the information based on convolutional kernels of different dimensions;

[0138] Step S413: Select frequency information related to semantic segmentation based on a frequency similarity kernel;

[0139] Step S414: Superimpose the low-pass information, the high-pass information, and the selected frequency information to obtain spatial domain information in different frequency bands.

[0140] Step S5: Increase the weight of the local features of the covered rock through an enhanced multi-dimensional convolutional attention mechanism.

[0141] In some specific embodiments, the feature weighting module 50 is used to perform the following steps:

[0142] Step S51: Divide the input rock image into three branches. The first branch is rotated 90° counterclockwise along channel H to obtain F1, the second branch is rotated 90° counterclockwise along channel W to obtain F2, and the third branch remains unchanged in its original features to obtain F3;

[0143] Step S52: Apply squeeze transformation to the three branches respectively, where the squeeze transformation includes average pooling, standard pooling, and max pooling;

[0144] Step S53: Generate new values through an adaptive combination mechanism for the results of the three branches after the squeeze transformation to determine the results of average pooling, standard pooling, and max pooling;

[0145] Step S54: Pass the three new values through activation functions respectively and average them to determine the feature map of the rock image covered by dust;

[0146] Step S55: Classify the feature map through depthwise separable convolution to obtain the label map of the rock image covered by dust.

[0147] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0148] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable terminal device provide for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0149] The above has introduced the method and device provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0150] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic description of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0151] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for extracting rocks on the Martian surface based on a lightweight network, characterized in that, Including: Step S1: Obtain Mars surface images and screen rock images; Step S2: Make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set; Step S3: Pass the training set through an adaptive kernel convolution to dynamically adjust the sampling shape and size of the convolution kernel to extract rock features in the rock images; Step S4: Use an improved lightweight Transformer model to separate the rock images from the surrounding environment; Step S5: Improve the attention of rocks covered by dust through an enhanced multi-dimensional convolutional attention mechanism; Step S6: Evaluate the lightweight Transformer model based on pre-training metrics. If the evaluation passes, obtain the final model. Based on the final model, perform rock surface extraction on the input Mars rock images; Step S3 includes: Step S31: Take the received image as input, randomly generate initial coordinates and define the initial sampling position to obtain the initial convolution kernel size; Step S32: Calculate the base integer based on the initial convolution kernel size; Step S33: Determine the positions of the two-dimensional grid based on the initial convolution kernel size and the base integer; Step S34: Perform a convolution operation on the received image to obtain an offset; Step S35: Determine new sampling coordinates based on the offset, the initial sampling position, and the positions of the two-dimensional grid; Step S36: Resample the rock image based on the new sampling coordinates; Step S37: Reshape, convolve again, standardize the resampled rock image, and finally output the original pixel information through an activation function; Step S4 includes: Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model; Step S42: Integrate the spatial domain information with the original unclustered pixel information; Step S41 includes: Step S411: Perform low-pass information filtering on the information based on average pooling and channel grouping to ensure that the dimensions and fine details of the regenerated image are consistent with the original rock image; Step S412: Perform high-pass filtering on the information based on convolutional kernels of different dimensions; Step S413: Select frequency information related to semantic segmentation based on a frequency similarity kernel; Step S413: Superimpose the low-pass information, the high-pass information, and the selected frequency information to obtain the spatial domain information of different frequency bands; Step S5 includes: Step S51: Divide the input rock image into three branches. The first branch is rotated 90° counterclockwise along channel H to obtain F1, the second branch is rotated 90° counterclockwise along channel W to obtain F2, and the third branch remains unchanged in its original features to obtain F3; Step S52: Pass the three branches through squeeze transformations respectively, where the squeeze transformations include average pooling, standard pooling, and max pooling; Step S53: Generate new values through an adaptive combination mechanism for the results of the three branches after the squeeze transformation to determine the results of average pooling, standard pooling, and max pooling; Step S54: Pass the three new values through the activation function respectively and average them to determine the feature map of the rock image covered by dust; Step S55: Classify the feature map through depthwise separable convolution to obtain the label map of the rock image covered by dust; In step S6, the pre-training metrics include: Intersection over Union (IoU), Precision, Pixel Accuracy (PA), Recall, and the F1 score which is the harmonic mean of Precision and Recall; where In the formula, TP is the number of positive instances correctly identified as positive, FP represents the number of negative instances misclassified as positive, TN is the instance correctly identified as negative, and FN is the positive instance misclassified as negative.

2. A Mars surface rock extraction system based on a lightweight network, characterized in that, It includes: Image screening module: used to obtain Mars surface images and screen rock images; Image classification module: used to make labels based on the rock images, perform image processing on the labels and rock images to expand the number of images, and divide the rock images into a training set, a test set, and a validation set; Feature extraction module: used to dynamically adjust the sampling shape and size of the convolution kernel of the training set through adaptive kernel convolution to extract the rock features in the rock image; Image separation module: used to separate the rock image from the surrounding environment using an improved lightweight Transformer model; Feature weighting module: used to increase the weight of the local features of the covered rock through an enhanced multi-dimensional convolutional attention mechanism; Model evaluation module: used to evaluate the lightweight Transformer model based on the pre-training metrics. If the evaluation passes, obtain the final model, and based on the final model, extract the rock surface of the input Mars rock image; The feature extraction module is used to perform the following steps: Step S31: Take the received image as input, randomly generate initial coordinates and define the initial sampling position to obtain the initial convolution kernel size; Step S32: Calculate the base integer based on the initial convolution kernel size; Step S33: Determine the positions of the two-dimensional grid based on the initial convolution kernel size and the base integer; Step S34: Perform a convolution operation on the received image to obtain the offset; Step S35: Determine the new sampling coordinates based on the offset, the initial sampling position, and the positions of the two-dimensional grid; Step S36: Resample the rock image based on the new sampling coordinates; Step S37: Reshape, convolve again, normalize the resampled rock image, and finally output the original pixel information through the activation function; The image separation module is used to perform the following steps: Step S41: Obtain the spatial domain information of different frequency bands of the original pixel information through an improved lightweight Transformer model; Step S42: Integrate the spatial domain information with the original unclustered pixel information; Step S41 includes: Step S411: Perform low-pass information filtering on the information based on average pooling and channel grouping to ensure that the dimensions and fine details of the regenerated image are consistent with the original rock image; Step S412: Perform high-pass filtering on the information based on convolution kernels of different dimensions; Step S413: Select frequency information related to semantic segmentation based on the frequency similarity kernel; Step S413: Superimpose the low-pass information, high-pass information, and the selected frequency information to obtain spatial domain information in different frequency bands; The feature weighting module is used to perform the following steps: Step S51: Divide the input rock image into three branches. The first branch is rotated 90° counterclockwise along channel H to obtain F1, the second branch is rotated 90° counterclockwise along channel W to obtain F2, and the third branch keeps the original features unchanged to obtain F3; Step S52: Pass the three branches through squeeze transformations respectively, where the squeeze transformations include average pooling, standard pooling, and max pooling; Step S53: Generate new values through the adaptive combination mechanism for the results of the three branches after the squeeze transformation to determine the results of average pooling, standard pooling, and max pooling; Step S54: Pass the three new values through activation functions respectively and average them to determine the feature map of the rock image covered by dust; Step S55: Classify the feature map through depthwise separable convolution to obtain the label map of the rock image covered by dust; In the model evaluation module, the pre-training metrics include: intersection over union IoU, precision, pixel accuracy PA, recall, and the harmonic mean F1 of precision and recall; where, In the formula, TP is the number of positive instances correctly identified as positive, FP represents the number of negative instances misclassified as positive, TN is the instance correctly identified as negative, and FN is the positive instance misclassified as negative.

Citation Information

Patent Citations

  • Sandstone microscopic image classification method and system based on improved Swin Transform

    CN118570797A

  • Wood surface defect detection method, system, medium and equipment

    CN119205758A

  • Complementary attention dynamic convolution method for image classification

    CN119360184A