An image segmentation method, system, device and storage medium
The deep learning model Trans-SedNet processed rock sheet images, which solved the problem of not being able to automatically identify the frame grain type in the existing technology, and achieved efficient and accurate image segmentation and statistical analysis to meet the needs of geological experts.
Patent Information
- Application Number
- CN202311010340.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-08-11
AI Technical Summary
The existing technology lacks an effective semantic segmentation model, cannot automatically identify frame grain types in rock sheets, and cannot provide advanced statistical information such as the area percentage and texture characteristics of each type of grain, resulting in low efficiency in manual observation and error-prone.
The deep learning model Trans-SedNet is adopted, combining the CMX architecture and SegFormer backbone, and through window attention, cross-modal feature correction, mixed attention and multi-scale feature fusion, we realize segmentation of XPL and PPL images, and use threshold processing and Canny algorithm to optimize the segmentation results.
It realizes efficient and accurate segmentation of rock flake images, provides more advanced statistical information, conforms to the analysis habits of geological experts, reduces manual intervention, and improves segmentation accuracy and efficiency.
Smart Images

Figure CN116993984B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sedimentary petrology, and particularly relates to an image segmentation method, system, device and storage medium. Background Art
[0002] Sedimentary geological surveys, especially those of sandy sediments and sandstones, require the identification of framework grains in epoxy resin-impregnated thin sections under a polarized light microscope and the determination of their textural characteristics. The systematic description of the composition and texture is crucial for understanding the provenance, depositional environment, and hydrocarbon reservoir quality. To collect a large amount of information from thin sections, point counting is a necessary method, which includes recording the composition and texture characteristics of usually 300 - 400 framework grains using a grid with a mechanical stage movement under a microscope. Although this manual observation can provide reliable sample composition and texture characteristics, this process requires a large amount of manpower, is prone to deviation, and requires professional knowledge.
[0003] Deep learning is an emerging technology that uses complex neural networks to automate artificial tasks. Computer vision is one of the main applications of deep learning, mainly performing tasks such as image classification, object detection, and semantic segmentation on digital images. These technologies have been widely applied to various earth science disciplines, including remote sensing images and seismic images. In these applications, point counting in geology has similarities with the semantic segmentation task in computer vision, both involving assigning labels to regions; while in the prior art, there is a lack of a semantic segmentation model for rock thin section images to provide the types of framework grains and cannot provide more advanced statistical information, such as the area percentage of each type of grain and texture characteristics, including grain size and roundness. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the present invention provides an image segmentation method, system, device and storage medium to provide the types of framework grains and more advanced statistical information, such as the area percentage of each type of grain and texture characteristics, including grain size and roundness.
[0005] In a first aspect, an embodiment of the present invention provides an image segmentation method, including the following steps:
[0006] Collect and organize thin section images, and segment and describe the thin section images;
[0007] Based on the collected thin section images and the information of the segmentation description, make training data sets, validation data sets, and test data sets, perform standardization processing on the training data sets, validation data sets, and test data sets, and establish a deep learning model;
[0008] Based on the standardized training data sets and validation data sets, continuously optimize and iterate the deep learning model until the optimal deep learning model is obtained;
[0009] Test the test dataset based on the optimal deep learning model, compare and analyze the prediction results of the test data with the standard labels, and obtain the thin slice segmentation model.
[0010] Furthermore, the collected thin slice images include XPL images and PPL images;
[0011] The process of classifying lithology according to thin slice images during the process of sorting thin slice images includes: single crystal quartz, polycrystalline quartz, plagioclase, potassium feldspar, magmatic rock fragments, volcanic rock fragments, sedimentary rock fragments, metamorphic rock fragments and authigenic minerals.
[0012] Furthermore, the ratio of the training dataset, validation dataset and test dataset is 6:2:2;
[0013] The process of normalizing the training dataset, validation dataset and test dataset is as follows:
[0014]
[0015] where x norm is the value after normalization, x is the input image pixel, the image pixel is in RGB format, and mean and std are the mean and variance calculated on the training dataset.
[0016] Furthermore, when establishing the deep learning model, the Trans-SedNet deep learning model is adopted. The Trans-SedNet deep learning model is a geological improvement and adaptation based on the CMX architecture and the SegFormer backbone. The Trans-SedNet deep learning model includes a transformer module based on window attention, a cross-modal feature correction module with post-normalization, a feature fusion module based on hybrid attention, a segmentation module based on multi-scale features and a post-smoothing processing module.
[0017] Furthermore, the image data processing flow of the Trans-SedNet deep learning model is as follows:
[0018] Send the input XPL image and PPL image data information into the corresponding transformer module based on window attention for feature mapping;
[0019] Send the feature mapping information into the cross-modal feature correction module with post-normalization for feature correction;
[0020] Send the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities;
[0021] The fused feature information obtained from each layer in the model is fed into a segmentation module based on multi-scale features for image segmentation;
[0022] The image segmentation structure information is fed into a post-smoothing processing module for optimization. By using threshold processing and the Canny algorithm for contour detection, the prediction confidence of the model for particles is statistically calculated in each detected contour, and the maximum prediction confidence is selected as the final segmentation prediction choice for the particles.
[0023] Furthermore, the process of feeding the input XPL image and PPL image data information into the corresponding transformer module based on window attention for feature mapping is as follows:
[0024] X XPL i = Windows_Transformer XPL i (X XPL i-1 );
[0025] X PPL i = Windows_Transformer PPL i (X PPL i-1 );
[0026] In the formula, and are the XPL image and PPL image data information respectively. Among them, B is the batch size input to the network, C is the dimension of the current feature map, H×W is the size of the current feature map, i∈{1 - 4} is the layer marker of the model, X XPL 0 and X PPL 0 are the images that have not been processed by the model initially. Windows_Transformer XPL i and Windows_Transformer PPL i are the transformer modules based on window attention for X XPL and X PPL respectively;
[0027] The process of feeding the feature mapping information into a post-normalization cross-modal feature correction module for feature correction is as follows:
[0028] X XPL i , X PPLi = norm(CM_FRM(X XPL i , X PPL i ));
[0029] Wherein, CM_FRM is the cross-modal feature correction module in the original CMX, and norm is an added layer of normalization layer;
[0030] The process of sending the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities is as follows:
[0031] X i = Mix_FFM(X XPL i , X PPL i );
[0032] Wherein, Mix_FFM is the feature fusion module based on hybrid attention, and X i is the result after fusing the feature information of X XPL i and X PPL i ;
[0033] The process of sending the feature fusion information obtained in each layer of the model into the segmentation module based on multi-scale features for image segmentation is as follows:
[0034] S coarse = Seg(X 1 , ···, X i ) i ∈ {1 - 4};
[0035] Wherein, S coarse is a rough thin slice segmentation result obtained, and Seg is the segmentation module based on multi-scale features;
[0036] The process of sending the image segmentation structure information into the post-smoothing processing module for optimization is as follows:
[0037] S = Smooth(S coarse );
[0038] Wherein, Smooth is the post-smoothing processing module; S is the final thin slice segmentation result obtained.
[0039] Furthermore, before performing the transformer module processing, it is necessary to record any input image data information as Divide the input image data into multiple small windows Wherein H * = window_h, W* = window_w, B *= B * H / window_h * W / window_w, where window_h and window_w are the window sizes;
[0040] The processing flow of Mix_FFM is as follows:
[0041] Use a linear layer for X XPL i and X PPL i to obtain (Q XPL , Q PPL ), (K XPL , K PPL ), and (V XPL , V PPL ) respectively. Then, all K and V will be processed to obtain K mix and V mix :
[0042] K mix / V mix = Channel_attn(Cat(K XPL / V XPL , K PPL / V PPL ));
[0043] In the formula, Channel_attn is the channel attention, and Cat is the concatenation operation;
[0044] The processing flow of the said Seg is to gradually upsample the small-scale information in the feature fusion information X i and add and transform it with the large-scale information until X 1 is obtained. Finally, X 1 is sent into the segmentation head for segmentation:
[0045] X i-1 = Linear(upscale(X i ) + X i-1 );
[0046] S coarse = Seghead(X 1 );
[0047] In the formula, Linear is the linear layer, upscale is the upsampling operation, and Seghead is the segmentation head, which is composed of multiple linear layers.
[0048] In the embodiments of the present invention, in the second aspect, the embodiments of the present invention provide an image segmentation system, including:
[0049] The acquisition module is used to acquire and organize thin slice images and segment and describe the thin slice images;
[0050] The model establishment module is used to make training data sets, validation data sets and test data sets based on the acquired thin slice images and the information of the segmentation description, perform standardization processing on the training data sets, validation data sets and test data sets, and establish a deep learning model;
[0051] The optimization module is used to continuously optimize and iterate the deep learning model based on the standardized training data sets and validation data sets until the optimal deep learning model is obtained;
[0052] The output module is used to test and segment the test data set based on the optimal deep learning model, compare and analyze the prediction results of the test data with the standard labels, and obtain a thin slice segmentation model.
[0053] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned image segmentation method are implemented.
[0054] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned image segmentation method are implemented.
[0055] Some other optional features and technical effects of the embodiments of the present invention are described below, and some can be understood by reading this article.
[0056] Compared with the prior art, the present invention has the following beneficial technical effects:
[0057] The present invention provides an image segmentation method, system, device and storage medium, comprising the following steps: collecting and sorting thin slice images, and segmenting and describing the thin slice images; making a training data set, a validation data set and a test data set based on the collected thin slice images and the information of the segmentation description, performing normalization processing on the training data set, the validation data set and the test data set, and establishing a deep learning model; continuously optimizing and iterating the deep learning model based on the normalized training data set and validation data set until an optimal deep learning model is obtained; testing and segmenting the test data set based on the optimal deep learning model, and comparing and analyzing the prediction results of the test data with the standard labels to obtain a thin slice segmentation model; the present application can simultaneously process multiple different modal thin slice image data, and can also fully integrate multiple modal information for final segmentation prediction. This method not only achieves excellent segmentation results on thin slice data, but also the model for multiple images at the same time is more in line with the habits of geological experts when analyzing thin slices; the present application can efficiently, accurately and quickly segment thin slice data.
[0058] Further, the image data processing flow of the Trans-SedNet deep learning model is as follows:
[0059] The input XPL image and PPL image data information are sent into the corresponding window attention-based transformer module for feature mapping; the feature mapping information is sent into the post-normalized cross-modal feature correction module for feature correction; the corrected data information is sent into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities; the feature fusion information obtained in each layer of the model is sent into the segmentation module based on multi-scale features for image segmentation; the image segmentation structure information is sent into the post-smoothing processing module for optimization processing. By using threshold processing and the Canny algorithm for contour detection, the prediction confidence of the model for particles is statistically analyzed in each detected contour, and the maximum prediction confidence is selected as the final segmentation prediction selection for the particles. This application can not only process two different modalities of thin slice image data simultaneously, but also fully fuse the information of these two modalities, namely XPL and PPL, for the final segmentation prediction. It not only achieves excellent segmentation results on thin slice data, but also this model that observes XPL and PPL images simultaneously is more in line with the habits of geological experts when conducting thin slice analysis. In addition, considering that when determining the particle type, the shape, color, and internal features of the particles are more important than the relative spatial relationship of the grains in the whole slice, it is often difficult to focus on the local relative spatial relationship using the global attention in the traditional transformer. Therefore, window attention is used to strengthen this relative spatial relationship. Since some thin slice regions may be too dark or blurred when the thin slice data is made and photographed, in order to alleviate the situation of scarce or missing thin slice data information, a method of hybrid attention is used to process the thin slice data information. Finally, aiming at the characteristic that the particles in the thin slice are usually aggregated by multiple minerals or rocks, a post-smoothing processing method is used to further optimize the thin slice segmentation result to prevent the model from over-segmenting. Description of the Drawings
[0060] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The elements shown are not limited by the scale shown in the drawings, and the same or similar reference numerals in the drawings represent the same or similar elements, where:
[0061] Figure 1 It is a schematic flowchart of an image segmentation method of the present invention;
[0062] Figure 2 It is the overall architecture of the Trans-SedNet deep learning model adopted in the embodiment of the present invention;
[0063] Figure 3 It is part of the lithology identification results of the Trans-SedNet in the test data set in the embodiment of the present invention;
[0064] Figure 4The performance effect of the post-smoothing processing module in the embodiments of the present invention;
[0065] Figure 5 The schematic diagram of window attention in the embodiments of the present invention;
[0066] Figure 6 The schematic diagram of the hybrid attention mechanism in the embodiments of the present invention;
[0067] Figure 7 An image segmentation device according to the present invention.
[0068] In the figure: 1000, electronic device; 1001, processor; 1002, read-only memory; 1003, memory; 1004, bus; 1005, I / O interface; 1006, input part; 1007, output part; 1008, storage part; 1009, communication part; 1010, driver; 1011, removable medium. Specific embodiments
[0069] To make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0070] The term "including" and its variants used herein mean open inclusion, that is, "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The term "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.
[0071] The present invention provides an image segmentation method, as Figure 1 shown, including the following steps:
[0072] Collect and organize thin slice images, and segment and describe the thin slice images;
[0073] Based on the collected thin slice images and the information of the segmentation description, make training data sets, validation data sets and test data sets, perform normalization processing on the training data sets, validation data sets and test data sets, and establish a deep learning model;
[0074] Based on the normalized training data sets and validation data sets, continuously optimize and iterate the deep learning model using the stochastic gradient descent method until the optimal deep learning model is obtained;
[0075] Test split is performed on the test data set based on the optimal deep learning model, and the prediction results of the test data are compared and analyzed with the standard labels to obtain a thin slice segmentation model.
[0076] Preferably, the collected thin slice images include XPL images and PPL images;
[0077] During the process of sorting the thin slice images, the lithology is divided according to the thin slice images, including one or more of single crystal quartz, polycrystalline quartz, plagioclase, potassium feldspar, magmatic rock fragments, volcanic rock fragments, sedimentary rock fragments, metamorphic rock fragments, and authigenic minerals.
[0078] Preferably, in the embodiment, the thin slice images are first scaled to a size of 384x576, and then these thin slice image data are divided into a training data set, a validation data set, and a test data set according to a ratio of 6:2:2.
[0079] The process of standardizing the training data set, the validation data set, and the test data set is as follows:
[0080]
[0081] In the formula, x norm is the value after standardization processing, x is the input image pixel, the image pixel is in RGB form, and mean and std are the mean and variance calculated on the training data set.
[0082] Preferably, when establishing the deep learning model, the Trans-SedNet deep learning model is adopted. The Trans-SedNet deep learning model is a geological improvement and adaptation based on the CMX architecture and the SegFormer backbone. The Trans-SedNet deep learning model includes a transformer module based on window attention, a cross-modal feature correction module with post-normalization, a feature fusion module based on hybrid attention, a segmentation module based on multi-scale features, and a post-smoothing processing module. It should be noted that in this embodiment, the optimizer used for the Trans-SedNet deep learning model is AdamW, and the loss function used is a combined loss function of Focal loss and Dice loss.
[0083] Preferably, as Figure 2 shown, the image data processing flow of the Trans-SedNet deep learning model is as follows:
[0084] The input XPL image and PPL image data information are sent into the corresponding transformer module based on window attention for feature mapping;
[0085] Send the feature mapping information into the post-normalized cross-modal feature correction module for feature correction;
[0086] Send the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities;
[0087] Send the feature fusion information obtained at each layer in the model into the segmentation module based on multi-scale features for image segmentation;
[0088] Send the image segmentation structure information into the post-smoothing processing module for optimization processing. By using threshold processing and the Canny algorithm for contour detection, the prediction confidence of the model for particles is statistically calculated in each detected contour, and the maximum prediction confidence is selected as the final segmentation prediction selection for the particles.
[0089] It should be noted that this application can not only process two different modalities of thin slice image data simultaneously, but also fully fuse the information of the two modalities of XPL and PPL for the final segmentation prediction. It not only achieves excellent segmentation results on thin slice data, but also this model that observes XPL and PPL images simultaneously is more in line with the habits of geological experts when performing thin slice analysis; in addition, considering that when determining the particle type, the shape, color and internal features of the particles are more important than the relative spatial relationship of the grains in the whole slice, so it is often difficult to focus on the local relative spatial relationship using the global attention in the traditional transformer, so window attention is used to strengthen this relative spatial relationship; since when making and photographing thin slice data, some thin slice areas may be too dark or blurred during photographing, in order to alleviate the situation of scarce or missing thin slice data information, a method of hybrid attention is used to process the thin slice data information; finally, aiming at the characteristic that the particles in the thin slice are usually aggregated by multiple minerals or rocks, a post-smoothing processing method is used to further optimize the thin slice segmentation result to prevent the model from over-segmenting.
[0090] Furthermore, as Figure 5 shown, the process of sending the input XPL image and PPL image data information into the corresponding transformer module based on window attention for feature mapping is:
[0091] X XPL i = Windows_Transformer XPL i (X XPL i-1 );
[0092] X PPL i=Windows_Transformer PPL i (X PPL i-1 );
[0093] Wherein, and are the XPL image and PPL image data information respectively, where B is the batch size of the input network, C is the dimension of the current feature map, H×W is the size of the current feature map, i∈{1-4} is the layer label of the model, X XPL 0 and X PPL 0 are the images that have not been processed by the model initially, Windows_Transformer XPL i and Windows_Transformer PPL i are the transformer modules based on window attention for X XPL and X PPL respectively;
[0094] The process of sending the feature mapping information into the cross-modal feature correction module for post-normalization for feature correction is as follows:
[0095] X XPL i ,X PPL i =norm(CM_FRM(X XPL i ,X PPL i ));
[0096] Wherein, CM_FRM is the cross-modal feature correction module in the original CMX, and norm is an additional normalization layer;
[0097] The process of sending the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities is as follows:
[0098] X i =Mix_FFM(X XPL i ,X PPL i );
[0099] Wherein, Mix_FFM is the feature fusion module based on hybrid attention, and X i is to combine X XPL i and X PPLi The result after fusing feature information;
[0100] The process of sending the feature fusion information obtained in each layer of the model into the segmentation module based on multi-scale features for image segmentation is as follows:
[0101] S coarse = Seg(X 1 , ···, X i ) i ∈ {1 - 4};
[0102] In the formula, S coarse is a rough thin slice segmentation result obtained, and Seg is the segmentation module based on multi-scale features;
[0103] The process of sending the image segmentation structure information into the post-smoothing processing module for optimization is as follows:
[0104] S = Smooth(S coarse );
[0105] In the formula, Smooth is the post-smoothing processing module; S is the final thin slice segmentation result obtained.
[0106] Furthermore, before the transformer module is processed, any input image data information needs to be recorded as The input image data is divided into multiple small windows where H * = window_h, W * = window_w, B * = B * H / window_h * W / window_w, and window_h and window_w are the window sizes;
[0107] The processing flow of Mix_FFM is as Figure 6 shown:
[0108] Use a linear layer for X XPL i and X PPL i respectively to obtain (Q XPL , Q PPL ), (K XPL , K PPL ) and (V XPL , V PPL ), then process all K and V to obtain K mix and V mix :
[0109] K mix / V mix= Channel_attn(Cat(K XPL / V XPL ,K PPL / V PPL ));
[0110] Where Channel_attn is channel attention and Cat is the concatenation operation;
[0111] The processing flow of the said Seg is to progressively upsample the small-scale information in the feature fusion information X i and add and transform it with the large-scale information until X 1 is obtained. Finally, X 1 is sent into the paired segmentation head for segmentation:
[0112] X i-1 = Linear(upscale(X i ) + X i-1 );
[0113] S coarse = Seghead(X 1 );
[0114] Where Linear is a linear layer, upscale is an upsampling operation, and Seghead is a segmentation head composed of multiple linear layers.
[0115] Example 1
[0116] Collect the river sediments of the Yarlung Zangbo River channel, then put the sandy sediments into epoxy resin, and then perform mechanical polishing to make a thin slice with a thickness of 0.03 mm and a smooth and flat surface. Then, collect the thin slice images under a standard polarized light microscope and annotate and explain the thin slice images. The annotation objects are mainly single-crystal quartz, polycrystalline quartz, plagioclase, potassium feldspar, magmatic rock debris, volcanic rock debris, sedimentary rock debris, metamorphic rock debris, and authigenic minerals.
[0117] After collecting the thin slice images, use the thin slice image data to construct a dataset. First, scale the thin slice images to a size of 384x576 for convenient model processing. Then, divide these thin slice image data into a training dataset, a validation dataset, and a test dataset according to a ratio of 6:2:2, and perform standardization processing on the data. The method is as follows:
[0118]
[0119] Where x norm is the value after standardization processing, x is the input image pixel in RGB format, and mean and std are the mean and variance calculated on the training dataset.
[0120] Build the Trans-SedNet deep learning model proposed in the present invention. This model is written using the Pytorch framework (version 1.10.0), and the window size is set to (12, 18) according to the size of the thin slice image and the size information of the particles in the thin slice.
[0121] Put the thin slice image into the Trans-SedNet model and iterate 200 times. Each time, calculate the loss using the predicted value and the true value of the model, and use the loss to optimize the model, so that the predicted value of the model continuously approaches the true value. After the loss tends to be stable, the trained Trans-SedNet thin slice segmentation model can be obtained. As Figure 3 shown, the segmentation effect of some Trans-SedNet deep learning models on the thin slices of river sediments in the Yarlung Zangbo River channel. Trans-SedNet can reach 39.6% mIoU accuracy and 57.4% mPA accuracy in the test dataset after 200 iterations.
[0122] Send the constructed test dataset into the trained Trans-SedNet model for testing, and compare the true segmentation data of the test dataset with the segmentation data predicted by the model to calculate its accuracy.
[0123] The present invention provides an image segmentation system, including:
[0124] An acquisition module for acquiring and organizing thin slice images and segmenting and describing the thin slice images;
[0125] A model establishment module for making training datasets, validation datasets, and test datasets based on the acquired thin slice images and segmentation description information, performing standardization processing on the training datasets, validation datasets, and test datasets, and establishing a deep learning model;
[0126] An optimization module for continuously optimizing and iterating the deep learning model based on the standardized training datasets and validation datasets until the optimal deep learning model is obtained;
[0127] An output module for performing test segmentation on the test dataset based on the optimal deep learning model, and comparing and analyzing the prediction results of the test data with the standard labels to obtain a thin slice segmentation model.
[0128] In an embodiment of the present invention, an electronic device is provided, including: a processor and a memory storing a computer program, and the processor is configured to execute the lithology identification method of any embodiment of the present invention when running the computer program.
[0129] Figure 7FIG. 0 shows a schematic diagram of an electronic device 1000 that can implement the method of the embodiments of the present invention or implement the embodiments of the present invention. In some embodiments, there may be more or fewer electronic devices than shown in the figure. In some embodiments, it can be implemented using a single or multiple electronic devices. In some embodiments, it can be implemented using cloud or distributed electronic devices.
[0130] As Figure 7 shown, the electronic device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to programs and / or data stored in a read-only memory (ROM) 1002 or programs and / or data loaded from a storage section 1008 into a random access memory (RAM) 1003. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), and so on. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0131] The above-mentioned processor and memory are jointly used to execute the programs stored in the memory, and when the programs are executed by a computer, they can implement the methods, steps, or functions described in the above embodiments.
[0132] The following components are connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, a touch screen, etc.; an output portion 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as a LAN card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage portion 1008 as needed. Figure 7 Only some components are schematically shown in FIG. 13, which does not mean that the computer system 1000 only includes Figure 7 the components shown.
[0133] The systems, devices, modules, or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0134] Although not shown, in an embodiment of the present invention, a storage medium is provided, and the storage medium stores a computer program, and the computer program is configured to execute the compilation method based on file differences in any embodiment of the present invention when being run.
[0135] The storage medium in the embodiment of the present invention includes permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of the storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0136] The methods, programs, systems, devices, etc. in the embodiments of the present invention can be executed or implemented in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.
[0137] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, those skilled in the art can conceive that the implementation of the functional modules / units or controllers and related method steps illustrated in the above embodiments can be achieved in a software, hardware, or a combination of software and hardware manner.
[0138] Unless explicitly stated, the actions or steps of the methods and programs recorded according to the embodiments of the present invention do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0139] In this document, multiple embodiments of the present invention are described. For the sake of brevity, the descriptions of the embodiments are not exhaustive, and features or parts that are the same or similar between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present invention, rather than all embodiments. The above terms do not necessarily refer to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.
[0140] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are only examples of the best mode for implementing the systems and methods. Those skilled in the art can understand that various changes can be made to the embodiments of the systems and methods described here when implementing the systems and / or methods without departing from the spirit and scope of the present invention defined in the appended claims.
Claims
1. An image segmentation method, characterized in that, It includes the following steps: Collect and organize thin-section images, and segment and describe the thin-section images; Based on the collected thin-section images and the information of the segmentation description, make training datasets, validation datasets, and test datasets, perform standardization processing on the training datasets, validation datasets, and test datasets, and establish a deep learning model; Based on the standardized training datasets and validation datasets, continuously optimize and iterate the deep learning model until the optimal deep learning model is obtained; Based on the optimal deep learning model, test and segment the test dataset, and compare and analyze the prediction results of the test data with the standard labels to obtain a thin-section segmentation model. Among them, when establishing the deep learning model, the Trans-SedNet deep learning model is adopted. The Trans-SedNet deep learning model is a geological improvement and adaptation based on the CMX architecture and the SegFormer backbone. The Trans-SedNet deep learning model includes a transformer module based on window attention, a cross-modal feature correction module with post-normalization, a feature fusion module based on hybrid attention, a segmentation module based on multi-scale features, and a post-smoothing processing module. The image data processing flow of the Trans-SedNet deep learning model is as follows: Send the input XPL image and PPL image data information into the corresponding transformer module based on window attention for feature mapping; Send the feature mapping information into the cross-modal feature correction module with post-normalization for feature correction; Send the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities; Send the feature fusion information obtained from each layer in the model into the segmentation module based on multi-scale features for image segmentation; Send the image segmentation structure information into the post-smoothing processing module for optimization processing. Through using threshold processing and the Canny algorithm for contour detection, statistically calculate the prediction confidence of the model for particles in each detected contour, and select the maximum prediction confidence as the final segmentation prediction selection for the particles; Among them, the process of sending the input XPL image and PPL image data information into the corresponding transformer module based on window attention for feature mapping is as follows: X XPL i = Windows_Transformer XPL i (X XPL i-1 ); X PPL i = Windows_Transformer PPL i (X PPL i-1 ); In the formula, and are the XPL image and PPL image data information respectively, where B is the batch size of the input network, C is the dimension of the current feature map, H×W is the size of the current feature map, i∈{1 - 4} is the layer label of the model, X XPL 0 and X PPL 0 are the images that have not been processed by the model initially, Windows_Transformer XPL i and Windows_Transformer PPL i are the window attention-based transformer modules for X XPL and X PPL respectively; The process of sending the feature mapping information into the cross-modal feature correction module with post-normalization for feature correction is as follows: X XPL i ,X PPL i = norm(CM_FRM(X XPL i ,X PPL i )); In the formula, CM_FRM is the cross-modal feature correction module in the original CMX, and norm is an added normalization layer; The process of sending the corrected data information into the feature fusion module based on hybrid attention to fuse the feature information of the two modalities is as follows: X i = Mix_FFM(X XPL i , X PPL i ) In the formula, Mix_FFM is a feature fusion module based on hybrid attention, and X i is the result of fusing the feature information of X XPL i and X PPL i ; The process of sending the feature fusion information obtained from each layer in the model into the segmentation module based on multi-scale features for image segmentation is as follows: S coarse = Seg(X 1 , ···, X i ) i ∈ {1 - 4}; where S coarse is a rough thin slice segmentation result obtained, and Seg is a segmentation module based on multi-scale features; The process of sending the image segmentation structure information into the post-smoothing processing module for optimization processing is as follows: S = Smooth(S coarse ); In the formula, Smooth is the post-smoothing processing module; S is the obtained final thin-section segmentation result.
2. The image segmentation method according to claim 1, wherein The collected thin-section images include XPL images and PPL images; The process of sorting thin-section images and classifying lithology based on the thin-section images includes one or more of: single-crystal quartz, polycrystalline quartz, plagioclase, potassium feldspar, magmatic rock fragments, volcanic rock fragments, sedimentary rock fragments, metamorphic rock fragments, and authigenic minerals.
3. The image segmentation method according to claim 1, wherein, The ratio of the training dataset, validation dataset, and test dataset is 6:2:2; The process of standardizing the training dataset, validation dataset, and test dataset is as follows: where x norm is the value after normalization, x is the input image pixel, the image pixel is in RGB format, and mean and std are the mean and variance calculated on the training dataset.
4. The image segmentation method according to claim 1, characterized in that Before performing the transformer module processing, any input image data information needs to be recorded as Divide the input image data into multiple small windows where H ★ = window_h, W ★ = window_w, B ★ = B * H / window_h * W / window_w, and window_h and window_w are the window sizes; The processing flow of Mix_FFM is as follows: Use a linear layer for X XPL i and X PPL i respectively obtain (Q XPL , Q PPL ), (K XPL , K PPL ), and (V XPL , V PPL ). Then, process all K and V to obtain K mix and V mix : K mix / V mix = Channel_attn(Cat(K XPL / V XPL , K PPL / V PPL )); In the formula, Channel_attn is the channel attention, and Cat is the concatenation operation; The processing flow of the said Seg is to progressively upsample the small-scale information in the feature fusion information X i and add and transform it with the large-scale information until X 1 is obtained. Finally, X 1 is fed into the paired segmentation head for segmentation: X i-1 = Linear(upscale(X i ) + X i-1 ); S coarse = Seghead(X 1 ); In the formula, Linear is the linear layer, upscale is the upsampling operation, and Seghead is the segmentation head, which is composed of multiple layers of linear layers.
5. An image segmentation system, characterized in that, Based on the image segmentation method according to any one of claims 1-4, it includes: An acquisition module, configured to acquire and sort thin-section images and segment and describe the thin-section images; A model establishment module, configured to produce a training dataset, a validation dataset, and a test dataset based on the acquired thin-section images and the information of the segmentation description, standardize the training dataset, the validation dataset, and the test dataset, and establish a deep learning model; An optimization module, configured to continuously optimize and iterate the deep learning model based on the standardized training dataset and validation dataset until an optimal deep learning model is obtained; An output module, configured to perform test segmentation on the test dataset based on the optimal deep learning model, compare and analyze the prediction results of the test data with the standard labels, and obtain a thin-section segmentation model.
6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of an image segmentation method according to any one of claims 1-4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of an image segmentation method according to any one of claims 1-4.
Citation Information
Patent Citations
Layer segmentation method and system for retina layer and effusion area based on deep learning
CN111583291A
Target detection method based on multi-source information fusion, thermal infrared and three-dimensional depth map
CN115713679A