Multi-modal small-particle-size mineral classification method and model

By employing feature extraction, noise reduction, and multimodal feature fusion, the problems of noise suppression and conflict handling in the fusion of CT and SEM modal information were solved, thereby improving the accuracy of small-particle-size mineral identification and the robustness of the model.

CN121746752APending Publication Date: 2026-03-27CHINA NAT PETROLEUM CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate information from both CT and SEM modalities, resulting in insufficient accuracy in identifying small-particle-size minerals, particularly in addressing the challenges of suppressing modal noise and handling modal conflicts.

Method used

By employing feature extraction, noise reduction, probability set acquisition, and multimodal feature fusion, combined with causal intervention mechanisms and adaptive fusion rules, the spatial dependence in CT and SEM images is accurately captured, noise interference is suppressed, and modal conflicts are handled, thereby improving recognition accuracy.

Benefits of technology

It significantly improves the identification accuracy of small-particle-size minerals, enhances the robustness and generalization of the model in complex task scenarios, and avoids decision-making errors caused by simple maximum confidence selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746752A_ABST
    Figure CN121746752A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal small-particle-size mineral classification method and model, and the method comprises the steps: carrying out the feature extraction of a CT image and an SEM image corresponding to a to-be-analyzed rock core image, and obtaining a CT feature and an SEM feature; and respectively performing spatial noise elimination processing on the CT feature and the SEM feature to obtain a first feature and a second feature. And based on the first probability set and the second probability set corresponding to any mineral particle in the to-be-analyzed rock core image, performing feature fusion on a feature corresponding to any mineral particle in the first feature and a feature corresponding to any mineral particle in the second feature to obtain a multi-modal feature corresponding to any mineral particle. And on the basis of the multi-modal features corresponding to any mineral particle, the mineral category corresponding to any mineral particle is obtained. Therefore, the information of the CT mode and the SEM mode is effectively fused, and the recognition precision of small-particle-size minerals is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and particularly relates to a multi-modal small-particle-size mineral classification method and model. BACKGROUND

[0002] Identifying small-particle-size minerals in a core image is one of the core tasks involved in three-dimensional digital core reconstruction, which aims to accurately detect and distinguish different mineral particles in rocks, especially small-particle-size minerals. The goal of this task is to provide basic data for subsequent rock physical property analysis (such as permeability, porosity, etc.), thereby supporting reservoir evaluation in resource exploration fields such as oil and gas and underground engineering research. Therefore, accurately identifying these minerals is of great significance for obtaining the microstructure of rocks, inferring the composition distribution, and carrying out related physical property simulation.

[0003] Currently, the core image can be scanned by using Computed Tomography (CT) technology or scanning electron microscope (SEM) technology to identify mineral particles in the core image. However, it is difficult to integrate the two technologies to further improve the identification accuracy of mineral particles in the core image. This is because due to the existence of modal heterogeneity, multi-modal fusion is prone to problems of difficult modal noise suppression and difficult modal conflict processing.

[0004] Therefore, how to effectively fuse the information of CT and SEM two modalities to further improve the identification accuracy of small-particle-size minerals is a problem that needs to be solved in the current multi-modal digital core reconstruction field. SUMMARY

[0005] The technical problem to be solved by the present application is to provide a multi-modal small-particle-size mineral classification method and model, which can effectively fuse the information of CT and SEM two modalities to improve the identification accuracy of small-particle-size minerals.

[0006] The technical solution of the present application to solve the above technical problem is as follows:

[0007] In one aspect, the present application provides a multi-modal small particle mineral classification method, comprising: obtaining a CT image and a SEM image corresponding to a core image to be analyzed. Perform feature extraction processing on the CT image and the SEM image to obtain CT features and SEM features. The CT features are image features of the CT image, and the SEM features are image features of the SEM image. Perform spatial noise elimination processing on the background features included in the CT features to obtain first features, and perform spatial noise elimination processing on the background features included in the SEM features to obtain second features. The background features correspond to the background in the core image to be analyzed. For any mineral particle included in the core image to be analyzed, obtain a first probability set and a second probability set corresponding to any mineral particle. The first probability set is determined based on the features corresponding to any mineral particle in the first features. The first probability set includes at least one first probability, and one first probability corresponds to one mineral category; any first probability in the at least one first probability represents the probability that the mineral category of any mineral particle is the mineral category corresponding to any first probability. The second probability set is determined based on the features corresponding to any mineral particle in the second features. The second probability set includes at least one second probability, and one second probability corresponds to one mineral category. Any second probability in the at least one second probability represents the probability that the mineral category of any mineral particle is the mineral category corresponding to any second probability. Based on the first probability set and the second probability set corresponding to any mineral particle, the features corresponding to any mineral particle in the first features and the features corresponding to any mineral particle in the second features are fused to obtain multi-modal features corresponding to any mineral particle. Based on the multi-modal features corresponding to any mineral particle, the mineral category corresponding to any mineral particle is obtained.

[0008] Based on the above technical solutions, the present application can also be improved as follows.

[0009] Further, the target image is cut into a plurality of image blocks. For any image block in the plurality of image blocks, a position code is added to any image block, and a flattening operation is performed on any image block to obtain a one-dimensional vector corresponding to any image block. The position codes corresponding to different image blocks are different. Based on the preset projection rule and the position code corresponding to any image block, the one-dimensional vector corresponding to any image block is projected into a feature space of a preset dimension to obtain a spatial vector corresponding to any image block under the preset dimension. Based on the spatial vectors corresponding to each image block under the preset dimension, the image features corresponding to the target image are obtained. Wherein, in the case that the target image is a CT image, the image features corresponding to the target image are CT features. In the case that the target image is a SEM image, the image features corresponding to the target image are SEM features.

[0010] Further, the preset projection rule includes:

[0011] z i = Exi +P i ;

[0012] wherein, E represents a preset projection matrix, x i represents a one-dimensional vector corresponding to any image block, P i represents a position code corresponding to any image block, z i represents a spatial vector corresponding to any image block in a preset dimension.

[0013] Further, for any image block, based on the spatial vector corresponding to any image block in the preset dimension, a key matrix, a value matrix and a query matrix corresponding to any image block are obtained. Based on the preset mapping rule, and the key matrix, the value matrix and the query matrix corresponding to any image block, the image feature corresponding to any image block is obtained. The preset mapping rule includes:

[0014]

[0015] represents the image feature corresponding to any image block, Q represents the key matrix corresponding to any image block, K represents the value matrix corresponding to any image block, V represents the query matrix corresponding to any image block, d k represents a preset dimension.

[0016] Further, based on the pixel information of each pixel included in the core image to be analyzed, a causal intervention mechanism is established; the pixel information includes pixel position and pixel value; based on the causal intervention mechanism, the background feature included in the target image is subjected to spatial noise elimination processing to obtain the image feature corresponding to the target image after noise elimination. Wherein, in the case that the target image is a CT image, the image feature corresponding to the target image after noise elimination is a first feature; in the case that the target image is a SEM image, the image feature corresponding to the target image after noise elimination is a second feature.

[0017] Further, based on the preset fusion rule, and the first probability set and the second probability set corresponding to any mineral particle, the confidence of at least one fusion feature is obtained. The fusion feature corresponding to the maximum confidence in the confidence of at least one fusion feature is determined as the multi-modal feature corresponding to any mineral particle.

[0018] Further, the preset fusion rule includes:

[0019]

[0020] wherein, represents the confidence that the any mineral particle is mineral category i, R represents a conflict coefficient, a probability that any mineral particle is the mineral class i in the first probability set, a probability that any mineral particle is the mineral class i in the second probability set, the mineral class i being any mineral class.

[0021] Further, the conflict coefficient is acquired based on a conflict coefficient acquisition rule. The conflict coefficient acquisition rule comprises: wherein, a probability that any mineral particle is the mineral class j in the second probability set.

[0022] In another aspect, the present application provides a multi-modal small-particle mineral classification model, comprising:

[0023] a feature extraction layer configured to perform feature extraction processing on CT images and SEM images corresponding to the input core image to be analyzed, to obtain CT features and SEM features, wherein the CT features are image features of the CT images, and the SEM features are image features of the SEM images.

[0024] a noise elimination layer configured to perform spatial noise elimination processing on background features included in the CT features to obtain first features, and perform spatial noise elimination processing on background features included in the SEM features to obtain second features, wherein the background features correspond to a background in the core image to be analyzed.

[0025] a probability acquisition layer configured to acquire, for any mineral particle included in the core image to be analyzed, a first probability set and a second probability set corresponding to the any mineral particle, wherein the first probability set is determined based on features corresponding to the any mineral particle in the first features. The first probability set comprises at least one first probability, and one first probability corresponds to one mineral class. Any first probability in the at least one first probability represents a probability that the mineral class of the any mineral particle is the mineral class corresponding to the any first probability. The second probability set is determined based on features corresponding to the any mineral particle in the second features. The second probability set comprises at least one second probability, and one second probability corresponds to one mineral class. Any second probability in the at least one second probability represents a probability that the mineral class of the any mineral particle is the mineral class corresponding to the any second probability.

[0026] a multi-modal fusion layer configured to perform feature fusion on features corresponding to the any mineral particle in the first features and features corresponding to the any mineral particle in the second features, based on the first probability set and the second probability set corresponding to the any mineral particle, to obtain multi-modal features corresponding to the any mineral particle.

[0027] a class determination layer configured to acquire a mineral class corresponding to the any mineral particle based on the multi-modal features corresponding to the any mineral particle.

[0028] In a third aspect, the present application provides an electronic device, comprising: a memory, one or more processors; the memory and the processor are coupled; wherein the memory stores computer program codes, the computer program codes comprise computer instructions, when the computer instructions are executed by the processor, the electronic device executes the multi-modal small particle mineral classification method of any one of the first aspect.

[0029] In a fourth aspect, a computer readable storage medium is provided, comprising computer instructions, when the computer instructions are run on an electronic device, the electronic device executes the multi-modal small particle mineral classification method of any one of the first aspect.

[0030] In a fifth aspect, a computer program product is provided, when the computer program product is run on a computer, the computer executes the multi-modal small particle mineral classification method of any one of the first aspect.

[0031] The beneficial effects of the present application are:

[0032] By accurately capturing the spatial dependence of small particle minerals in CT images and SEM images, and effectively suppressing the noise interference between different modalities, the recognition accuracy of small particle minerals is significantly improved by eliminating false statistical correlation and noise effect.

[0033] All possible fusion cases are considered and integrated, avoiding the decision-making errors caused by simple maximum confidence selection, and maximizing the multi-modal collaborative effect. The adaptive fusion mechanism improves the comprehensive utilization of multi-modal information, and enhances the robustness and generalization of the model in complex task scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 A flowchart of a multi-modal small particle mineral classification method provided by the present application is provided;

[0035] Figure 2 A structural diagram of a multi-modal small particle mineral classification model provided by the present application is provided. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the present application, unless otherwise specified, “ / ” represents an “or” relationship between the objects before and after the “ / ”, for example, A / B can represent A or B; “and / or” in the present application is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone. And in the description of the present application, unless otherwise specified, “multiple” means two or more than two. “At least one of the following (one)” or similar expressions means any combination of the items, including any combination of single item (one) or multiple items (one). For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the same items or similar items with basically the same function and role are distinguished by using “first”, “second”, etc. The skilled in the art can understand that “first”, “second”, etc. do not limit the quantity and execution order, and “first”, “second”, etc. also do not mean necessarily different. At the same time, in the embodiments of the present application, the words “exemplary” or “for example” are used to represent as an example, illustration or description. Any embodiment or design scheme described as “exemplary” or “for example” in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes.

[0037] Currently, CT technology or SEM technology can be used to scan core images to identify mineral particles in the core images. However, it is difficult to integrate the two technologies to further improve the identification accuracy of mineral particles in the core images.

[0038] This is because when CT modalities and SEM modalities are fused, the following problems will be encountered:

[0039] Modal noise is difficult to suppress: CT images and SEM images are disturbed by different types of noise due to the difference in imaging principle. CT images are usually limited by noise and artifacts in the scanning process, resulting in blurred structural information; while SEM images may introduce significant surface noise due to sample surface unevenness or electron beam and sample interaction. In the multi-modal fusion process, due to the nonlinear superposition effect of different modal noises, the accumulation of noise will further amplify, especially in the small particle mineral identification task, the effect of noise on microstructure is more obvious, thereby reducing the identification accuracy.

[0040] Modality conflict is difficult to handle: CT images and SEM images provide complementary information, but there are significant differences in their contributions to different features or task scenarios. CT images mainly provide three-dimensional structural information inside the rock, which is crucial for spatial topology modeling; while SEM images exhibit surface morphology information at extremely high resolution, which is particularly important for accurate identification of fine-grained structures. However, due to the heterogeneity of modalities, conflicts may arise between information from different modalities in specific tasks. Under different circumstances or specific task requirements, quantifying the confidence of each modality and effectively handling the information conflict between modalities is one of the core challenges for optimal information fusion. Solving this problem not only depends on a deep understanding of the characteristics of each modality information, but also requires the construction of an adaptive multi-modal information fusion mechanism, so that the fusion strategy can be flexibly adjusted according to task requirements and data characteristics, thereby maximizing the recognition effect of multi-modal collaboration.

[0041] To solve the above problems, the present application provides a multi-modal small particle size mineral classification method and model, which can effectively fuse the information of CT and SEM two modalities through noise suppression technology and modality conflict suppression mechanism, thereby further improving the recognition accuracy of small particle size minerals.

[0042] Referring to Figure 1 A multi-modal small particle size mineral classification method is provided by the present application, comprising the following steps S101-S106:

[0043] S101: Obtain the CT image and SEM image corresponding to the core image to be analyzed.

[0044] The CT image can be obtained by scanning the core image to be analyzed based on CT technology, and the SEM image can be obtained by scanning the core image to be analyzed based on SEM technology.

[0045] S102: Perform feature extraction processing on the CT image and the SEM image to obtain CT features and SEM features.

[0046] The CT features are image features of the CT image, and the SEM features are image features of the SEM image.

[0047] In some embodiments, the target image can be cut into multiple image blocks. For any image block in the multiple image blocks, a position code is added to the any image block, and a flattening operation is performed on the any image block to obtain a one-dimensional vector corresponding to the any image block. The position codes corresponding to different image blocks are different.

[0048] Based on the preset projection rule and the position encoding corresponding to any image block, a one-dimensional vector corresponding to any image block can be projected into a feature space of a preset dimension to obtain a spatial vector corresponding to any image block under the preset dimension. Based on the spatial vector corresponding to each image block under the preset dimension, an image feature corresponding to the target image can be obtained.

[0049] In a case where the target image is a CT image, the image feature corresponding to the target image is a CT feature. In a case where the target image is a SEM image, the image feature corresponding to the target image is a SEM feature.

[0050] In some embodiments, the preset projection rule comprises:

[0051] z i = Ex i + P i ;

[0052] wherein E represents a preset projection matrix, x i represents a one-dimensional vector corresponding to any image block, P i represents a position encoding corresponding to any image block, and z i represents a spatial vector corresponding to any image block under the preset dimension.

[0053] In some embodiments, for any image block, a key matrix, a value matrix and a query matrix corresponding to any image block can be obtained based on the spatial vector corresponding to any image block under the preset dimension. Based on a preset mapping rule and the key matrix, the value matrix and the query matrix corresponding to any image block, an image feature corresponding to any image block is obtained. Based on the image feature corresponding to each image block, an image feature corresponding to the target image is obtained.

[0054] In some embodiments, the preset mapping rule comprises:

[0055]

[0056] wherein, represents an image feature corresponding to any image block, Q represents a key matrix corresponding to any image block, K represents a value matrix corresponding to any image block, V represents a query matrix corresponding to any image block, and d k represents a preset dimension.

[0057] Exemplarily, a target image is cut into N image blocks with a size of P×P. Wherein H and W are the height and width of the target image; C is the number of channels, representing the size of the image block. After cutting, a flattening operation is performed on each image block to obtain a one-dimensional vector corresponding to each image block respectively:

[0058] x i = Flatten(Ii ), i = 1, 2,..., N;

[0059] where I i represents the i-th image block.

[0060] Then, a one-dimensional vector of each image block can be projected to a d-dimensional feature space by a linear mapping to obtain a corresponding spatial vector of each image block in the preset dimension:

[0061] z i = Ex i + P i ;

[0062] Next, an image block feature sequence Z = {z1,..., z i ,..., z N} can be obtained based on the corresponding spatial vectors of each image block in the preset dimension. The image block feature sequence Z = {z1,..., z i ,..., z N} can be input into the VisionTransformer to capture the spatial dependency between each image block by using multi-head self-attention. For each attention head, a corresponding key matrix, value matrix and query matrix of each image block can be calculated:

[0063] Q = ZW Q , K = ZW K , V = ZW V ;

[0064] where W Q , W K and W V are learnable key, value and query mappings, respectively.

[0065] Then, a corresponding spatial vector of each image block in the preset dimension can be obtained:

[0066]

[0067] It can be seen that, in the case that the target image is a CT image, the image feature corresponding to the target image is a CT feature In the case that the target image is a SEM image, the image feature corresponding to the target image is a SEM feature

[0068] S103: Perform spatial noise elimination processing on the background feature included in the CT feature to obtain a first feature, and perform spatial noise elimination processing on the background feature included in the SEM feature to obtain a second feature.

[0069] where the background feature corresponds to the background in the core image to be analyzed.

[0070] In some embodiments, a causal intervention mechanism is established based on pixel information of each pixel included in the core image to be analyzed, the pixel information including a pixel position and a pixel value; and a spatial noise elimination process is performed on the background feature included in the target image based on the causal intervention mechanism, to obtain an image feature corresponding to the target image after noise elimination. In the case of a CT image as the target image, the image feature corresponding to the target image after noise elimination is a first feature; in the case of a SEM image as the target image, the image feature corresponding to the target image after noise elimination is a second feature.

[0071] For example, in order to eliminate the interference of modal noise, a causal intervention mechanism can be established. In this stage, the false correlation between noise and labels is cut off through intervention operation from the perspective of causal relationship. Specifically, K-means clustering can be used to divide the background pixels included in the target image into K mixed factor sets, and then the false statistical correlation between the mixed factors and the labels is eliminated through subsequent adjustment:

[0072]

[0073] wherein P represents a probability distribution, do(X) represents a causal intervention mechanism, represents the feature of the kth mixed factor set, Y represents the label of the corresponding mineral category, and CA represents a cross-attention mechanism.

[0074] It should be noted that the above causal intervention mechanism can be performed on CT images and SEM images respectively to suppress the background noise in the CT images and the SEM images.

[0075] S104: For any mineral particle included in the core image to be analyzed, a first probability set and a second probability set corresponding to the any mineral particle are obtained.

[0076] The first probability set is determined based on the feature corresponding to the any mineral particle in the first feature. The first probability set includes at least one first probability, and one first probability corresponds to one mineral category. Any first probability in the at least one first probability represents the probability that the mineral category of the any mineral particle is the mineral category corresponding to the any first probability.

[0077] The second probability set is determined based on the feature corresponding to the any mineral particle in the second feature. The second probability set includes at least one second probability, and one second probability corresponds to one mineral category. Any second probability in the at least one second probability represents the probability that the mineral category of the any mineral particle is the mineral category corresponding to the any second probability.

[0078] In some embodiments, the first probability set is obtained by calculating the distance between each pixel and each class prototype in the feature space of the CT modality. The second probability set is obtained by calculating the distance between each pixel and each class prototype in the feature space of the SEM modality.

[0079] S105: Based on the first probability set and the second probability set corresponding to any mineral particle, the features corresponding to any mineral particle in the first feature and the features corresponding to any mineral particle in the second feature are fused to obtain the multi-modal feature corresponding to any mineral particle.

[0080] In some embodiments, at least one fusion feature corresponding to the confidence can be obtained based on the preset fusion rule, and the first probability set and the second probability set corresponding to any mineral particle. The fusion feature corresponding to the maximum confidence in the confidence corresponding to the at least one fusion feature is determined as the multi-modal feature corresponding to any mineral particle.

[0081] In some embodiments, the preset fusion rule comprises:

[0082]

[0083] wherein, denotes the confidence that the any mineral particle is mineral class i, and R denotes a conflict coefficient, denotes the probability that the any mineral particle is mineral class i in the first probability set, denotes the probability that the any mineral particle is mineral class i in the second probability set, and mineral class i is any mineral class.

[0084] In some embodiments, the conflict coefficient R is obtained by summing all mutually exclusive class combinations in the CT modality and the SEM modality, i.e., only considering non-overlapping class pairs in the CT modality and the SEM modality.

[0085] In some embodiments, the conflict coefficient can be obtained based on a conflict coefficient acquisition rule. The conflict coefficient acquisition rule comprises: wherein, denotes the probability that the any mineral particle is mineral class j in the second probability set.

[0086] S106: Based on the multi-modal feature corresponding to any mineral particle, the mineral class corresponding to any mineral particle is obtained.

[0087] Referring to Figure 2 The application also provides a multi-modal small-particle mineral classification model, comprising:

[0088] The feature extraction layer, in response to the input CT and SEM images corresponding to the core image to be analyzed, performs feature extraction processing on the CT and SEM images to obtain CT features and SEM features. CT features are the image features of the CT image, and SEM features are the image features of the SEM image.

[0089] A noise reduction layer is used to perform spatial noise reduction processing on the background features included in the CT features to obtain a first feature; and to perform spatial noise reduction processing on the background features included in the SEM features to obtain a second feature; wherein the background features correspond to the background in the core image to be analyzed.

[0090] The probability acquisition layer is used to acquire a first probability set and a second probability set for any mineral grain included in the core image to be analyzed. The first probability set is determined based on the features corresponding to any mineral grain in a first feature. The first probability set includes at least one first probability, each corresponding to a mineral category. Any first probability in the at least one first probability represents the probability that the mineral category of any mineral grain is the mineral category corresponding to any first probability. The second probability set is determined based on the features corresponding to any mineral grain in a second feature. The second probability set includes at least one second probability, each corresponding to a mineral category. Any second probability in the at least one second probability represents the probability that the mineral category of any mineral grain is the mineral category corresponding to any second probability.

[0091] The multimodal fusion layer is used to fuse the features corresponding to any mineral particle in the first feature and the features corresponding to any mineral particle in the second feature based on the first probability set and the second probability set corresponding to any mineral particle, so as to obtain the multimodal features corresponding to any mineral particle.

[0092] The category determination layer is used to obtain the mineral category corresponding to any mineral particle based on the multimodal features corresponding to any mineral particle.

[0093] In some embodiments, after the classification layer outputs the mineral category corresponding to any mineral particle, the cross-entropy loss between the mineral category output by the classification layer and the actual mineral category corresponding to any mineral particle can be calculated. Then, gradient descent can be used to update the multimodal small-particle-size mineral classification model provided by this invention to improve the accuracy of the model in classifying small-particle-size minerals in core images.

[0094] As can be seen, this invention accurately captures the spatial dependence of small-diameter minerals in CT and SEM images and effectively suppresses noise interference between different modalities. By eliminating spurious statistical correlations and noise effects, it significantly improves the identification accuracy of small-diameter minerals.

[0095] And, the application considers all possible fusion cases and carries out comprehensive analysis, avoids decision-making errors caused by simple maximum confidence selection, and maximizes the multi-modal collaborative effect. The adaptive fusion mechanism improves the comprehensive utilization of multi-modal information, and enhances the robustness and generalization of the model in complex task scenarios.

[0096] In some schemes, multiple embodiments of the present application can be combined and implemented. Optionally, some operations in the flow of each method embodiment are optionally combined, and / or the order of some operations is optionally changed. And, the execution order between the steps of each flow is only exemplary, and does not constitute a limitation on the execution order between the steps, and other execution orders between the steps can also be used. It is not intended to indicate that the execution order is the only execution order in which these operations can be performed. Those skilled in the art can re-order the operations described herein in various ways. In addition, it should be pointed out that the process details involved in some embodiments herein are also applicable in a similar manner to other embodiments, or different embodiments can be combined for use.

[0097] In addition, some steps in the method embodiments can be equivalently replaced by other possible steps. Alternatively, some steps in the method embodiments can be optional and can be deleted in some use scenarios. Alternatively, other possible steps can be added to the method embodiments. And, each method embodiment can be implemented individually, or in combination.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the model is divided into different functional modules to complete all or part of the functions described above.

[0099] In several embodiments provided in the present application, it should be understood that the disclosed model and method can be implemented in other ways. For example, the model embodiments described above are only illustrative, for example, the division of the modules or units is only a logical functional division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another model, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between the model or unit, which can be electrical, mechanical or other forms.

[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multimodal small-particle-size mineral classification method, characterized in that, include: Obtain the CT and SEM images corresponding to the core images to be analyzed; Feature extraction processing is performed on the CT image and the SEM image to obtain CT features and SEM features; The CT features are the image features of the CT image, and the SEM features are the image features of the SEM image; Spatial noise reduction processing is performed on the background features included in the CT features to obtain a first feature; and spatial noise reduction processing is performed on the background features included in the SEM features to obtain a second feature; wherein the background features correspond to the background in the core image to be analyzed; For any mineral grain included in the core image to be analyzed, a first probability set and a second probability set corresponding to the mineral grain are obtained; the first probability set is determined based on the feature corresponding to the mineral grain in the first feature; the first probability set includes at least one first probability, and one first probability corresponds to one mineral category; any one of the at least one first probabilities represents the probability that the mineral category of the mineral grain is the mineral category corresponding to the at least one first probability; the second probability set is determined based on the feature corresponding to the mineral grain in the second feature; the second probability set includes at least one second probability, and one second probability corresponds to one mineral category; any one of the at least one second probability represents the probability that the mineral category of the mineral grain is the mineral category corresponding to the at least one second probability; Based on the first probability set and the second probability set corresponding to any mineral particle, feature fusion is performed on the feature corresponding to the first feature and the feature corresponding to the second feature of any mineral particle to obtain the multimodal feature corresponding to any mineral particle. Based on the multimodal features corresponding to any one of the mineral particles, the mineral category corresponding to any one of the mineral particles is obtained.

2. The method according to claim 1, characterized in that, The method further includes: The target image is cut into multiple image blocks; For any image block among the plurality of image blocks, a position code is added to the image block, and a flattening operation is performed on the image block to obtain a one-dimensional vector corresponding to the image block; different image blocks correspond to different position codes; Based on the preset projection rules and the position encoding corresponding to any image block, the one-dimensional vector corresponding to any image block is projected onto the feature space of the preset dimension to obtain the spatial vector corresponding to any image block in the preset dimension. Based on the spatial vector corresponding to each image block in the preset dimension, the image features corresponding to the target image are obtained; Wherein, when the target image is the CT image, the image feature corresponding to the target image is the CT feature; when the target image is the SEM image, the image feature corresponding to the target image is the SEM feature.

3. The method according to claim 2, characterized in that, The preset projection rules include: z i =Ex i +P i ; Where E represents the preset projection matrix, x i P represents a one-dimensional vector corresponding to any of the image patches. i z represents the position code corresponding to any of the image blocks. i This represents the spatial vector corresponding to any image patch in the preset dimension.

4. The method according to claim 3, characterized in that, The step of obtaining the image features corresponding to the target image based on the spatial vectors corresponding to each of the image blocks in the preset dimension includes: For any image patch, based on the spatial vector corresponding to the image patch in the preset dimension, obtain the key matrix, value matrix and query matrix corresponding to the image patch; Based on preset mapping rules, and the key matrix, value matrix, and query matrix corresponding to any image patch, the image features corresponding to any image patch are obtained; the preset mapping rules include: in, Let Q represent the image features corresponding to any given image patch, Q represent the key matrix corresponding to any given image patch, K represent the value matrix corresponding to any given image patch, V represent the query matrix corresponding to any given image patch, and d represent the image features corresponding to any given image patch. k This refers to the preset dimension; Based on the image features corresponding to each image block, the image features corresponding to the target image are obtained.

5. The method according to claim 4, characterized in that, The method further includes: A causal intervention mechanism is established based on the pixel information of each pixel in the core image to be analyzed; the pixel information includes pixel position and pixel value. Based on the causal intervention mechanism, spatial noise removal processing is performed on the background features included in the target image to obtain the image features corresponding to the target image after noise removal. Wherein, when the target image is the CT image, the image feature corresponding to the target image after noise reduction is the first feature; when the target image is the SEM image, the image feature corresponding to the target image after noise reduction is the second feature.

6. The method according to claim 5, characterized in that, The step of fusing features corresponding to any mineral particle in the first feature and the features corresponding to any mineral particle in the second feature, based on the first probability set and the second probability set corresponding to any mineral particle, to obtain the multimodal features corresponding to any mineral particle includes: Based on the preset fusion rules, and the first probability set and the second probability set corresponding to any mineral particle, the confidence level corresponding to at least one fusion feature is obtained; The fusion feature corresponding to the highest confidence level among the confidence levels corresponding to the at least one fusion feature is determined as the multimodal feature corresponding to any mineral particle.

7. The method according to claim 6, characterized in that, The preset fusion rules include: in, R represents the confidence level that any given mineral particle belongs to mineral category i, and R represents the conflict coefficient. This represents the probability that any mineral particle in the first probability set belongs to mineral category i. This represents the probability that any mineral particle in the second probability set is of mineral category i, where mineral category i is any mineral category.

8. The method according to claim 7, characterized in that, Before obtaining the confidence level corresponding to at least one fusion feature based on the preset fusion rules and the first and second probability sets corresponding to any mineral particle, the method further includes: The conflict coefficient is obtained based on the conflict coefficient acquisition rules; the conflict coefficient acquisition rules include: in, This represents the probability that any mineral particle in the second probability set belongs to mineral category j.

9. A multimodal small-particle-size mineral classification model, characterized in that, include: The feature extraction layer is used to extract features from the CT image and SEM image corresponding to the input core image to be analyzed, in response to the input core image to be analyzed. The CT features are the image features of the CT image, and the SEM features are the image features of the SEM image; A noise reduction layer is used to perform spatial noise reduction processing on the background features included in the CT features to obtain a first feature; and to perform spatial noise reduction processing on the background features included in the SEM features to obtain a second feature; wherein the background features correspond to the background in the core image to be analyzed; A probability acquisition layer is used to acquire a first probability set and a second probability set corresponding to any mineral grain included in the core image to be analyzed; the first probability set is determined based on the feature corresponding to any mineral grain in the first feature; the first probability set includes at least one first probability, and one first probability corresponds to one mineral category; any one of the at least one first probability represents the probability that the mineral category of any mineral grain is the mineral category corresponding to the at least one first probability; the second probability set is determined based on the feature corresponding to any mineral grain in the second feature; the second probability set includes at least one second probability, and one second probability corresponds to one mineral category; any one of the at least one second probability represents the probability that the mineral category of any mineral grain is the mineral category corresponding to the at least one second probability; A multimodal fusion layer is used to perform feature fusion on the features corresponding to the first feature and the features corresponding to the second feature of any mineral particle based on the first probability set and the second probability set corresponding to any mineral particle, so as to obtain the multimodal features corresponding to any mineral particle. The category determination layer is used to obtain the mineral category corresponding to any mineral particle based on the multimodal features corresponding to any mineral particle.

10. An electronic device, characterized in that, include: A memory, one or more processors; the memory is coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the multimodal small-particle-size mineral classification method as described in any one of claims 1-8.