Mineral identification analysis method and system based on calculation of polarized light field reconstruction
By using a five-dimensional degradation relationship model and a hybrid sparse reconstruction algorithm based on multi-view and multi-polarization state light field image sequences, combined with a deep learning model, the problems of narrow field of view and image quality coupling in traditional polarizing microscopes are solved, achieving efficient and accurate mineral identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINESE ACAD OF GEOLOGICAL SCI
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional polarizing microscopes have a narrow field of view at high magnification, and the image quality indicators are mutually coupled and constrained, making it difficult to achieve efficient and accurate mineral identification. Moreover, existing improvement schemes cannot synergistically improve the overall image quality.
By acquiring original light field image sequences from multiple perspectives and polarization states, a five-dimensional degradation relationship model is established. A hybrid sparse multi-dimensional reconstruction algorithm is adopted to achieve single-shot-five-dimensional decoupling-end-to-end reconstruction. Mineral features are extracted by combining a deep learning model.
It achieves large field of view, high resolution, and high fidelity polarization image reconstruction, improving the accuracy and efficiency of mineral identification, reducing reliance on experience, and providing quantifiable analysis results.
Smart Images

Figure CN121563802B_ABST
Abstract
Description
Mineral Identification and Analysis Method and System Based on Computational Polarization Field Reconstruction Technical Field
[0001] This invention belongs to the field of geological and mineral analysis and microscopic imaging technology, and specifically relates to a mineral identification and analysis method and system based on computational polarization field reconstruction. Background Technology
[0002] Polarizing microscopes are core tools for mineral identification in fields such as geology and materials science. They are used to identify minerals by analyzing phenomena such as interference colors and extinction caused by the interaction between polarized light and minerals.
[0003] However, traditional systems face fundamental bottlenecks in refined and high-throughput analysis: narrow field of view at high magnification, reliance on time-consuming mechanical scanning and stitching; and multiple key image quality indicators are coupled and constrained by each other, strong birefringence easily saturates the detector, resulting in loss of dynamic range (HDR), optical diffraction and aberrations lead to blurred details (MTF decrease), and detector noise weakens the signal-to-noise ratio. This makes key mineral features easily obscured, and analysis heavily relies on experience, limiting efficiency and accuracy.
[0004] Traditional improvement solutions are mostly limited to hardware upgrades, but cannot eliminate physical contradictions, or isolated software algorithms (which deal with single-dimensional problems and are prone to artifacts or spectral distortion), making it difficult to improve the overall quality of images in a coordinated manner.
[0005] Therefore, the core challenge that urgently needs to be overcome in this field is: how to construct a unified hybrid sparse multidimensional mathematical model that incorporates the improvement of five core quality dimensions—MTFC, HDR, temporal dimension (equivalent improvement of digital TDI), entropy reduction dimension (noise suppression), and deformation dimension (geometric correction)—into the same optimization framework. This would enable a new paradigm of mineral imaging: "single-shot-five-dimensional decoupling-end-to-end reconstruction," fundamentally solving the triple bottleneck of "physical contradiction-algorithmic fragmentation-empirical dependence" in traditional methods. This would allow for the acquisition of large field-of-view, high-resolution, and high-fidelity polarized images, thereby improving reconstruction accuracy. Summary of the Invention
[0006] In view of the above analysis, the embodiments of the present invention aim to provide a mineral identification and analysis method and system based on computational polarization field reconstruction, so as to solve the triple bottleneck of physical contradiction, algorithm fragmentation and experience dependence in the prior art.
[0007] The objective of this invention is achieved as follows:
[0008] A mineral identification and analysis method based on computational polarization field reconstruction includes:
[0009] Acquire at least one set of original light field image sequences of the sample under test from multiple perspectives and multiple polarization states;
[0010] The original light field sequence images are registered to ensure that the images from all viewpoints correspond in spatial location, thus obtaining registration data.
[0011] A degradation relationship model is established to describe the degradation relationship between the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence. The degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise.
[0012] Based on the degradation relationship model and the registration data, a polarization composite image is generated through inverse solving;
[0013] Morphological and textural features and optical-physical features are extracted from the polarization-synthesized image to determine the mineral type and distribution information of the sample to be tested.
[0014] In a preferred embodiment of the present invention, acquiring at least one set of original light field image sequences of the sample under test from multiple perspectives and multiple polarization states includes:
[0015] When the sample to be tested is placed on the stage, a programmable multi-angle LED array illumination unit located below the sample to be tested, in conjunction with a condenser lens, switches illumination beams with different incident angles and different polarization states according to a preset time sequence, so that the illumination beams penetrate the stage and illuminate the sample to be tested.
[0016] An objective lens positioned above the stage and along the transmission direction of the illumination beam deflects the illumination beam to a micro-polarizer array light field sensor, acquiring intensity information of the illumination beam in different polarization directions within a single exposure cycle under the corresponding illumination mode, and obtaining the original light field image sequence.
[0017] In a preferred embodiment of the present invention, the registration processing of the original light field sequence images, such that the images from all viewpoints correspond in spatial location to obtain registration data, includes:
[0018] The image corresponding to the principal optical axis viewpoint in the original light field image sequence is used as the geometric reference reference image;
[0019] For each image to be registered in the original light field image sequence except for the reference image, a frequency domain subpixel registration operation is performed. The frequency domain subpixel registration operation includes: calculating the normalized cross power spectrum of the image to be registered and the reference image in the Fourier domain.
[0020] Subpixel interpolation is performed on the peak position of the normalized cross-power spectrum to determine the image plane displacement of the image to be registered relative to the reference image;
[0021] Based on the image plane displacement, geometric transformation compensation is performed on all images to achieve alignment of image content in the spatial position from each viewpoint.
[0022] In a preferred embodiment of the present invention, the degradation relationship model includes:
[0023] ;
[0024] in, Characterizes the original light field image sequence; Characterizes the target light field image; Nonlinear operators characterizing the dynamic range compression; The transformation matrix characterizing the geometric deformation; The blur kernel characterizes the aforementioned detail blur; A downsampling operator characterizing information loss caused by the finite sampling rate of the sensor; Additive noise characterizing noise interference; The ordinal number representing the original light field image sequence.
[0025] In a preferred embodiment of the present invention, the step of generating a polarization composite image by inverse solving based on the degradation relationship model and the registration data includes:
[0026] Construct an objective function for the degradation relationship model, the objective function including: a data fidelity term based on the degradation relationship model, and a hybrid sparse regularization term for constraining five dimensions;
[0027] The objective function is solved iteratively using the alternating direction multiplier method. In each iteration, the light field image variables, auxiliary variables, and Lagrange multipliers are alternately updated based on the registration data to achieve the reverse solution of the original light field sequence image and generate the reconstructed polarization composite image.
[0028] The objective function includes:
[0029] ;
[0030] in, Characterizes data fidelity; Sparse transforms characterizing non-convex higher-order total variation. The regularization term in the objective function used to apply non-convex higher-order total variation constraints represents different... Constraints corresponding to different types of image reconstruction; The weight parameters corresponding to each regularization term; Characterizes the L1 norm.
[0031] A preferred embodiment of the present invention provides a mineral identification and analysis method based on computational polarization field reconstruction, which satisfies one or more of the following:
[0032] At that time, detail-enhanced regularization ;in, This indicates that the L0 norm constraint is applied. sparsity, , For pre-trained high-frequency detail dictionary, For image patch extraction operations, the representation forces the reconstructed image to have a sparse representation under the L0 norm constraint in the detail dictionary;
[0033] Dynamic range expansion regularization ;in, This indicates the L1 norm sparsity under a constrained high dynamic range dictionary. , Characterizing a high dynamic range dictionary Indicates the light field image passing through the target Perform L1 norm sparsification. ;
[0034] At that time, signal-to-noise ratio enhancement regularization The characterization focuses on the temporal dimension and applies joint sparsity constraints to non-locally similar blocks in multi-view images. Norm;
[0035] At that time, noise suppression regularization ;in For a tight frame transformation, characterize the application of a transform domain. Sparse constraints; This indicates that the target light field image x is processed. The result after transformation;
[0036] At that time, geometric correction regularization ;in For the sparse deformation field estimated from the data, its gradient sparsity is constrained; Represents the calculation of deformation field Spatial gradient; Gradient sparsity is constrained by the L0 norm and used for geometric correction.
[0037] In a preferred embodiment of the present invention, the step of extracting morphological texture features and optical-physical features from the polarization composite image to determine the mineral type and distribution information of the sample to be tested includes:
[0038] Extraction operations are performed using multi-scale, multi-directional Gabor filters or Gaussian difference filter banks to obtain the particle size, shape, and cleavage texture of the sample to be tested, which are used as the morphological texture characteristics.
[0039] The crystal optical properties of the sample under test are determined by inverting and calculating the local slow light direction, approximate birefringence, and degree of polarization of each pixel in the polarization composite image, which are then used as the optical physical characteristics.
[0040] The morphological texture features and the optical physical features are fused into multi-dimensional features and input into a mineral identification deep learning model to determine the mineral type and distribution information of the sample to be tested.
[0041] In a preferred embodiment of the present invention, the mineral identification deep learning model is a dual-branch network architecture, comprising: a first branch being a convolutional neural network for processing morphological and textural features; and a second branch being a graph neural network for processing graph data constructed with pixels as nodes and optical features and spatial relationships as edges. The outputs of the two branches are fused at multiple levels during the decoding stage, and the mineral type and distribution information of the sample to be tested are generated through a segmentation head.
[0042] The process of fusing the morphological texture features and the optical physical features into a multi-dimensional feature set and inputting the result into a mineral identification deep learning model to determine the mineral type and distribution information of the sample to be tested includes:
[0043] An adaptive weighted fusion strategy based on channel attention is used to fuse the morphological texture features and the optical physical features;
[0044] The convolutional neural network is used to process the morphological and spectral information in the fused features. The local texture, edge and spectral pattern of the mineral are extracted through deep convolution, and the detailed local feature map is output.
[0045] The graph neural network described above is used to construct graph data with pixels as nodes and optical features and spatial relationships as edges through graph propagation or self-attention mechanism. The graph data is then integrated with the physical characteristics and spatial co-occurrence relationships of the sample under test for identification, and the output is a relationship-enhanced feature map containing global context.
[0046] By employing a spatial attention gating mechanism, key regions of local feature branches are selectively enhanced. A shared semantic segmentation head then generates pixel-level mineral type identification maps and corresponding classification confidence maps based on the fused high-level features, thereby obtaining the mineral type and distribution information of the sample to be tested.
[0047] A preferred embodiment of the present invention, a mineral identification and analysis method based on computational polarization field reconstruction, further includes: generating and outputting a comprehensive analysis result containing a mineral distribution map and a quantitative identification report of the sample to be tested, including:
[0048] Based on the mineral type identification map of the determined sample to be tested, quantitative analysis is performed to determine the ratio of the pixel area occupied by each mineral type to the total image area, and a mineral content percentage report is generated.
[0049] Connectivity analysis is performed on the identification results to segment out independent mineral particles. The equivalent circle diameter of each particle is determined and calculated. The particle size distribution histogram is statistically analyzed according to the mineral type, and the average particle size and particle size sorting coefficient are calculated.
[0050] The identification results of the same specific region where the sample to be tested is located are compared with the results obtained by micro-area chemical composition / molecular structure analysis techniques such as electron probe or micro Raman spectroscopy to determine the overall classification accuracy and Kappa coefficient.
[0051] The comprehensive analysis result is obtained by combining the polarization composite image, the mineral type identification map, the mineral content percentage report, the particle size distribution histogram, the overall classification accuracy, and the Kappa coefficient.
[0052] A mineral identification and analysis system based on computational polarization field reconstruction includes:
[0053] The acquisition unit is configured to acquire at least one set of raw light field image sequences of the sample under test from multiple perspectives and multiple polarization states;
[0054] The registration unit is configured to perform registration processing on the original light field sequence image so that the images under all viewpoints correspond in spatial position, thereby obtaining registration data.
[0055] The establishment unit is configured to establish a degradation relationship model for describing the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence, and the degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise.
[0056] The reconstruction unit is configured to generate a polarization-synthesized image by inverse solving based on the degradation relationship model and the registration data;
[0057] The processing unit is configured to extract morphological and texture features and optical-physical features from the polarization composite image to determine the mineral type and distribution information of the sample to be tested.
[0058] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0059] a) The mineral identification and analysis method based on computational polarization light field reconstruction provided by this invention creatively realizes a new mineral imaging paradigm of "single-shot - five-dimensional decoupling - end-to-end reconstruction" by deeply integrating the computational light field imaging principle, multi-angle polarization coding, and hybrid sparse multi-dimensional reconstruction algorithm model. It breaks through the inherent contradiction between traditional microscope resolution, field of view, and depth of field at the physical level, and solves the problem of mutual constraints among multiple image quality indicators (detail, dynamic range, noise, etc.) during improvement at the algorithmic level. The proposed hybrid sparse model can collaboratively reconstruct mineral images to address degradation mechanisms, significantly improving spatial resolution and signal-to-noise ratio while perfectly preserving the true interference colors and spectral fidelity crucial for identification, thus achieving a paradigm innovation in imaging principle and quality improvement.
[0060] (b) In the mineral identification and analysis method based on computational polarization field reconstruction provided by this invention, addressing the core identification requirement of mineral crystal optical anisotropy, this scheme constructs a dedicated intelligent analysis model tightly coupled with multi-dimensional high-quality image output. This model not only utilizes deep convolutional networks to extract morphological and textural features, but also innovatively integrates quantitative optical parameter maps (such as local slow light direction and approximate birefringence gradient) derived from the reconstructed image as key input features. This method, combining physical optical properties with image features, enables the model to possess expert-like "reasoning" capabilities, significantly improving the accuracy of distinguishing minerals with similar optical properties (such as different feldspar varieties) and the interpretability of the model, thus achieving a deep adaptation between the intelligent identification model and geological mineralogy.
[0061] c) The mineral identification and analysis method based on computational polarization light field reconstruction provided by this invention adopts a collaborative hardware architecture of "illumination encoding - light field sensing - edge computing". Through the synchronous control of the programmable LED array and the polarization light field sensor, efficient information compression is achieved at the data acquisition source. A dedicated computing platform based on embedded GPU / FPGA is used to accelerate and optimize core algorithms such as hybrid sparse reconstruction at the hardware level. A pipelined processing strategy from illumination control and data stream transmission to parallel reconstruction is designed, enabling the entire process from "image capture" to "result output" to be completed in seconds, meeting the actual efficiency requirements of high-throughput thin section analysis in the laboratory. This completely changes the traditional time-consuming and lengthy process, achieving collaborative design and processing efficiency optimization of both hardware and software.
[0062] d) The mineral identification and analysis method based on computational polarization field reconstruction provided by this invention constructs a complete, closed-loop intelligent analysis workflow from sample placement to report generation. This method not only automates the entire process of image acquisition, quality enhancement, feature extraction, and mineral identification, but also provides quantifiable and reproducible high-value data for geological research through its final output: mineral distribution maps, quantitative optical property statistics tables, and structured identification reports. This significantly reduces reliance on operator experience, freeing experts from repetitive tasks, while establishing objective and unified standards for mineral identification, and significantly improving the scientific rigor, consistency, and overall efficiency of the analysis work.
[0063] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings.
[0065] Figure 1 is a flowchart of a mineral identification and analysis method based on computational polarization field reconstruction provided by the present invention;
[0066] Figure 2 is a flowchart of a method for obtaining a raw light field image sequence provided by the present invention;
[0067] Figure 3 is an optical path layout diagram of a transmissive sampling scenario provided by the present invention;
[0068] Figure 4 is a schematic diagram of the structure of a mineral intelligent identification and analysis microscope system provided by the present invention;
[0069] Figure 5 is a flowchart of a method for determining registration data provided by the present invention;
[0070] Figure 6 is a flowchart of an image reconstruction method provided by the present invention;
[0071] Figure 7 is a schematic diagram of a hybrid sparse multidimensional collaborative reconstruction algorithm model provided by the present invention;
[0072] Figure 8 is a schematic diagram of the structure of a mineral identification and analysis system based on computational polarization field reconstruction provided by the present invention;
[0073] Figure 9 is a module diagram of a mineral identification and analysis system based on computational polarization field reconstruction provided by the present invention. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0075] To facilitate understanding of the embodiments of this application, further explanation and description will be provided below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application. In the drawings, the dimensions and relative dimensions of components may be exaggerated for clarity and / or descriptive purposes. When exemplary embodiments can be implemented differently, a specific process sequence may be performed in a different order than described. For example, two consecutively described processes may be performed substantially simultaneously or in the reverse order of their description. Furthermore, the same reference numerals denote the same components.
[0076] The terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting. As used herein, unless the context clearly indicates otherwise, the singular forms “a” and “the” are intended to include the plural forms as well. Furthermore, when the terms “comprising” and / or “including” and variations thereof are used in this specification, it indicates the presence of the stated features, integrals, steps, operations, parts, components, and / or groups thereof, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, parts, components, and / or groups thereof. It should also be noted that, as used herein, the terms “substantially,” “about,” and other similar terms are used as approximate terms rather than as terms of degree, thus explaining the inherent biases in measurements, calculated values, and / or provided values that will be recognized by those skilled in the art.
[0077] As described in the background section, traditional methods suffer from a triple bottleneck of "physical contradictions, algorithmic fragmentation, and reliance on experience," resulting in limited reconstruction effectiveness.
[0078] With the development of computer imaging technology, computational imaging technology (such as light field imaging and polarization coding) and sparse representation theory have provided new paths for breakthroughs. However, existing sparse methods are mostly aimed at single-dimensional degradation and have limited effect on the reconstruction of complex degradation models with multi-dimensional coupling in polarized microscopy.
[0079] Based on this, this invention provides a mineral identification and analysis method based on computational polarization light field reconstruction. By acquiring multi-view, multi-polarization light field data in a single image, it significantly reduces the reliance on mechanical scanning, repeated exposures, and manual operation in the imaging process, thereby improving imaging efficiency and system stability. Furthermore, by introducing a five-dimensional degradation relationship model covering dynamic range, geometric deformation, detail blur, resolution, and noise, it avoids the problem of separate processing and fragmentation of different degradation factors in traditional methods, achieving holistic modeling and inverse solution of complex imaging degradation. Then, it incorporates the physical process of light field imaging and the polarization image reconstruction process into the same end-to-end solution framework, reducing reliance on empirical parameters and manual adjustments, and improving the stability, repeatability, and objectivity of the reconstruction results. Moreover, through multi-view information fusion and degradation inversion reconstruction, it effectively improves the spatial resolution and polarization information fidelity of polarization images while ensuring large field-of-view coverage. Finally, by extracting morphological texture features and optical physical features from high-quality polarization composite images, it makes mineral type determination and spatial distribution analysis more accurate and reliable.
[0080] A specific embodiment of the present invention, as shown in Figure 1, discloses a flowchart of a mineral identification and analysis method based on computational polarization field reconstruction provided by the present invention. As shown in Figure 1, it may include:
[0081] S101, acquire at least one set of original light field image sequences of the sample under test with multiple views and multiple polarization states.
[0082] In some embodiments, an optical field acquisition device is used to image the sample under test to obtain a sequence of raw optical field images containing multi-view information and multi-polarization state information. The optical field acquisition device may include an imaging lens, a microlens array, and a polarization modulation component, thereby simultaneously recording the optical response information of the sample at different viewpoints and different polarization states during one or more exposures.
[0083] In practice, multi-view information can be achieved through microlens arrays or multiple camera units, so that each original light field image contains sub-views from different incident directions; multi-polarization state information can be obtained by setting polarizers with different orientations, liquid crystal variable polarizers, or time-division polarization modulation methods to reflect the imaging characteristics of the sample under test under different polarization conditions.
[0084] The original light field image sequence obtained by the above method can simultaneously characterize the spatial structure information and polarization characteristics of the sample under test, providing basic data support for subsequent light field reconstruction, feature extraction or physical parameter analysis.
[0085] As a specific embodiment, referring to Figures 2 and 3, Figure 2 is a flowchart of acquiring an original light field image sequence provided by the present invention, and Figure 3 is an optical path layout diagram of a transmissive sampling scene provided by the present invention. As shown in Figures 2 and 3, it may include:
[0086] S201, when the sample to be tested is placed on the stage, the programmable multi-angle LED array illumination unit located below the sample to be tested, in conjunction with a condenser lens, switches illumination beams with different incident angles and different polarization states according to a preset time sequence, so that the illumination beams penetrate the stage and illuminate the sample to be tested.
[0087] In some embodiments, the sample to be tested (e.g., a mineral sheet) is placed on a stage, and an illumination beam is emitted by a programmable multi-angle LED array illumination unit disposed below the stage. The beam passes through a condenser lens and is then transmitted upward from the bottom of the stage to illuminate the sample to be tested.
[0088] During this process, the programmable multi-angle LED array illumination unit and the condenser lens work together to quickly switch illumination beams with different incident angles and different polarization states (including linearly polarized light and circularly polarized light in different directions) according to a preset timing sequence, illuminating and penetrating the sample under test from below.
[0089] S202, the objective lens located above the stage and along the transmission direction of the illumination beam refracts the illumination beam to the micro-polarizer array light field sensor, obtains the intensity information of the illumination beam in different polarization directions within a single exposure cycle in the corresponding illumination mode, and obtains the original light field image sequence.
[0090] In some embodiments, when the illumination beam emitted by the programmable multi-angle LED array illumination unit is transmitted through the stage, its transmission direction can be changed by the objective lens, so that the illumination beam can be acquired by the micro-polarizer array light field sensor set above the stage, and then transmitted to the high-resolution camera sensor to obtain the original light field image sequence.
[0091] During this process, within a single exposure cycle corresponding to each specific illumination mode (i.e., a combination of a specific incident angle and a specific polarization state) of the programmable multi-angle LED array illumination unit, the intensity information of the light beam transmitted from the sample and focused by the objective lens in four different analytical directions (such as linear polarization directions of 0°, 45°, 90°, and 135°) is simultaneously captured and analyzed. In this way, the four analytical directions, together with the illumination angle and the illumination polarization state, constitute an orthogonal multi-dimensional data acquisition space.
[0092] For each specific set of lighting conditions, the micro-polarizer array synchronously acquires four intensity components I0, I... 45 I 90I 135 It can be directly used to calculate the Stokes vector (containing at least the first three parameters S0, S1, S2) of the sample's emitted light field at each point on the imaging plane under this illumination condition: , representing the total light intensity; , indicating the linear polarization preference in the 0° and 90° directions; , indicating the linear polarization preference in the 45° and 135° directions.
[0093] In some optional examples, if the illumination beam contains a circularly polarized state, the S3 parameter characterizing the circular polarization properties can be further obtained or derived. This allows for a quantitative, pixel-level description of the polarization state modulation induced by the sample under different excitation conditions.
[0094] In the above scheme, static and sequential image acquisition of mineral thin section samples placed on a stage is performed without the need for mechanical scanning of the sample or polarization elements.
[0095] Furthermore, it provides absolutely uniform and stable illumination for each frame of image, ensuring that the sample under test always presents a clear outline with high contrast. This ensures repeatability and speed in the acquisition process, perfectly meeting the core requirements of automated inspection for stability, consistency, and efficiency.
[0096] In some embodiments, data is transmitted in real time to the embedded processing platform via a high-speed interface.
[0097] For example, referring to Figure 4, a schematic diagram of the structure of a mineral intelligent identification and analysis microscope system provided by the present invention is shown. As shown in Figure 4, the processing and control core has an embedded GPU / FPGA computing platform built in. The embedded GPU / FPGA computing platform outputs a synchronous control signal, which enables the illumination beam emitted by the programmable multi-angle LED array illumination unit to be transmitted to the microscope body through the polarization state controller.
[0098] In this way, under the action of the control signal / illumination beam, the sample to be tested on the stage can be detected, that is, the light carrying the sample information is transmitted to the light field acquisition module.
[0099] With the help of micro-polarizer array light field sensor and CMOS sensor, the raw light field data can be acquired and output to the processing and control core, which can then be used to execute display and interactive terminals.
[0100] S102, the original light field sequence image is registered so that the images from all viewpoints correspond in spatial position, and registration data is obtained.
[0101] In some embodiments, the light field sequence image is composed of multiple sub-images acquired from different viewpoints. Due to differences in imaging viewpoint, shooting position or acquisition time, the images from different viewpoints may have spatial offset, rotation or scale differences. Therefore, it is necessary to perform registration processing on the light field sequence image to achieve consistency in spatial position of the images from each viewpoint.
[0102] Specifically, by analyzing the feature information in the light field sequence images, the correspondence between images from different viewpoints can be obtained, and spatial alignment processing can be performed on the images from each viewpoint based on the correspondence, so that the image content of the same spatial location under different viewpoints can correspond to each other. The registration processing may include operations such as translation correction, rotation correction, and / or scale correction.
[0103] The above registration process yields the registered light field sequence image data, also known as the registration data. The registration data reflects the correspondence between images from different viewpoints within a unified spatial coordinate system, providing a foundation for subsequent processing and analysis of the light field sequence images. This data can be used for operations such as parallax calculation, depth reconstruction, or light field reconstruction.
[0104] As a specific embodiment, referring to Figure 5, a flowchart of the present invention for determining registration data can be executed as shown in Figure 5:
[0105] S501, the image corresponding to the principal optical axis viewpoint in the original light field image sequence is used as the geometric reference reference image.
[0106] In some embodiments, the original light field image sequence includes multiple multi-view polarized light field images, which are a set of sub-aperture images with multiple different viewpoints acquired by a light field camera or polarized light field microscope.
[0107] In other words, the principal axis information corresponding to different images is not consistent. Therefore, in this scheme, the image corresponding to the principal optical axis viewpoint is specifically used as the geometric reference benchmark image for accurate data preprocessing.
[0108] It should be noted that in actual processing, images on other axes can be selected as geometric reference images.
[0109] S502, for each image to be registered in the original light field image sequence except for the reference image, perform a frequency domain subpixel registration operation, the frequency domain subpixel registration operation including: calculating the normalized cross power spectrum of the image to be registered and the reference image in the Fourier domain.
[0110] In some embodiments, using a reference image as a reference, the normalized cross-power spectrum between other images to be registered and this reference image is calculated.
[0111] More specifically, Fourier transforms are performed on the image to be registered and the reference image respectively to obtain the first spectrum and the second spectrum; the complex conjugate product of the first spectrum and the second spectrum is calculated to obtain the cross power spectrum; the cross power spectrum is divided by the product of the amplitude spectrum of the first spectrum and the amplitude spectrum of the second spectrum to complete the normalization.
[0112] S503, perform subpixel interpolation on the peak position of the normalized cross-power spectrum to determine the image plane displacement of the image to be registered relative to the reference image.
[0113] In some embodiments, the image plane displacement is caused by microparallax resulting from the relative position between the microlens array and the sample under test, and the true deviation can be determined by performing subpixel interpolation.
[0114] More specifically, an inverse Fourier transform is performed on the normalized cross-power spectrum to obtain a spatial correlation surface; initial peak positions at the integer pixel level are located on the spatial correlation surface; with the initial peak position as the center, correlation values in its neighborhood are selected for surface function fitting (e.g., Gaussian surface fitting or quadratic surface fitting); the sub-pixel level precise position of the peak point is calculated based on the fitted surface function, and the difference between the precise position and the initial peak position is the sub-pixel level image plane displacement.
[0115] S504, based on the image plane displacement, perform geometric transformation compensation on all images to achieve alignment of image content in the spatial position from each viewpoint.
[0116] In some embodiments, the geometric transformation compensation is a global translation transformation, and the translation vector of the translation transformation is the estimated displacement of the image plane. Alignment is achieved by performing geometric transformation compensation.
[0117] In short, using the sensor's central sub-aperture image corresponding to the principal optical axis viewpoint in the light field data as a geometric reference, a frequency-domain-based sub-pixel registration algorithm is employed to perform high-precision geometric alignment on all images acquired from multi-view polarized light fields. This algorithm calculates the normalized cross-power spectrum in the Fourier domain between different frames of images acquired from multi-angle polarized light fields and performs sub-pixel interpolation on its peak positions. This accurately estimates and compensates for the micro-parallax caused by the relative position of the microlens array and the sample, ensuring that the image content at all viewpoints corresponds strictly in spatial position.
[0118] S103, establish a degradation relationship model to describe the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence, and the degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise.
[0119] In some embodiments, to describe the degradation relationship between the original image acquired from multi-view polarized light field acquisition and the ideal image after quality enhancement, a unified five-dimensional coupled degradation prior model is established. This model simultaneously characterizes the image quality degradation process in five dimensions: detail blurring (MTF decrease), insufficient dynamic range (HDR insufficiency), noise interference (low signal-to-noise ratio), information entropy loss, and potential geometric deformation. It is formalized into a mathematical optimization problem containing mixed sparse constraints, thereby enabling single-shot-five-dimensional decoupling.
[0120] In some embodiments, the degradation relationship model includes:
[0121] ;
[0122] in, Characterizes the original light field image sequence; Characterizes the target light field image; The nonlinear operator characterizing the dynamic range compression simulates the sensor's response to high dynamic range scenes; The transformation matrix characterizing the geometric deformation, i.e. the transformation matrix of the geometric deformation caused by motion or aberration; The blur kernel characterizing the aforementioned detail blur is the blur kernel of detail blur (MTF detail dimension) caused by optical diffraction, aberration, and defocus; Characterizes downsampling operators that cause information loss due to the limited sampling rate of sensors (e.g., micro-polarizer array optical field sensors); Additive noise characterizing noise interference; The ordinal number representing the original light field image sequence.
[0123] Thus, the model comprehensively depicts the overall degradation process of the imaging system from five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise. It realizes the integration of improvements in five core quality dimensions—MTFC, HDR, temporal dimension (digital TDI equivalent enhancement), entropy reduction dimension (noise suppression), and deformation dimension (geometric correction)—into the same optimization framework.
[0124] This demonstrates the use of multi-dimensional prior knowledge pre-trained for mineral properties, including edge priors for capturing high-frequency details, spectral fidelity priors for characterizing wide dynamic range interference colors, and structural sparsity priors for mining nonlocal similarities.
[0125] S104, Based on the degradation relationship model and the registration data, a polarization composite image is generated by inverse solving.
[0126] In some embodiments, step S104 aims to utilize the established degradation model and registration data to solve the low-quality observation sequence images acquired from the multi-view polarization light field through a computational algorithm. High-quality light field image reconstructed from the middle .
[0127] Specifically, the identification and analysis method provided in this embodiment employs an end-to-end reconstruction based on a hybrid sparse multi-dimensional collaborative quality reconstruction algorithm model. The core idea of this scheme is that mineral light field images possess sparsity in their specific transform domains. By constructing a joint optimization framework, the sparsity of the reconstruction results under multiple physical prior-driven transform domains is simultaneously constrained and matched with the degradation model.
[0128] Referring to Figure 6, a flowchart of an image reconstruction method provided by the present invention is shown. As shown in Figure 6, the following can be executed:
[0129] S601, construct an objective function for the degradation relationship model, the objective function including: a data fidelity term based on the degradation relationship model, and a hybrid sparse regularization term for constraining the five dimensions.
[0130] In some embodiments, the objective function includes:
[0131] ;
[0132] in, Characterizes data fidelity; The sparse transform characterizing nonconvex higher-order total variation, where It is a regularization sub-term in the objective function used to apply non-convex higher-order total variation constraints. Different j correspond to different types of image reconstruction constraints (such as sparsity, total variation, etc.) to improve the reconstruction effect. The weight parameters corresponding to each regularization term; Characterizes the L1 norm.
[0133] Accordingly, At that time, detail-enhanced regularization ;in, This indicates that the L0 norm constraint is applied. sparsity, , For pre-trained high-frequency detail dictionary, For image patch extraction operations, the representation forces the reconstructed image to have a sparse representation under the L0 norm constraint in the detail dictionary.
[0134] Specifically, detailed enhancement regularization is employed. This allows the detailed features of an image to be composed of only a few basis vectors from the dictionary, thereby enhancing details (such as edges and textures) while suppressing redundant information.
[0135] Dynamic range expansion regularization ;in, This indicates the L1 norm sparsity under a constrained high dynamic range dictionary. , Characterizing a high dynamic range dictionary Represents the target light field image Perform L1 norm sparsification. This represents the total variation regularization term (used to maintain the smoothness of image edges).
[0136] Specifically, The L1 norm is easier to optimize than the L0 norm, while also promoting sparsity and employing dynamic range expansion regularization. It can achieve "dynamic range expansion and structure preservation".
[0137] At that time, signal-to-noise ratio enhancement regularization The characterization focuses on the temporal dimension and applies joint sparsity constraints to non-locally similar blocks in multi-view images. Norm.
[0138] At that time, noise suppression regularization ;in For a tight frame transformation, characterize the application of a transform domain. Sparse constraints; This indicates that the target light field image x is processed. The result after transformation.
[0139] At that time, geometric correction regularization ;in For the sparse deformation field estimated from the data, its gradient sparsity is constrained; Represents the calculation of deformation field Spatial gradient; By constraining gradient sparsity using the L0 norm, the deformation field can be made as smooth as possible for geometric correction.
[0140] S602, the objective function is solved iteratively using the alternating direction multiplier method, and in each iteration, the light field image variables, auxiliary variables and Lagrange multipliers are alternately updated based on the registration data to achieve the reverse solution of the original light field sequence image and generate the reconstructed polarization composite image.
[0141] In some embodiments, the objective function is used to characterize the mathematical relationship between the original light field sequence image and the polarization composite image to be reconstructed. The objective function comprehensively considers light field imaging model constraints, polarization constraints, data consistency constraints, and regularization constraints to ensure the stability of the inverse solution process and the physical rationality of the results.
[0142] In some embodiments, the Alternating Direction Method of Multipliers (ADMM) is an iterative optimization method that decomposes a complex optimization problem into multiple subproblems and solves them alternately.
[0143] By introducing auxiliary variables and Lagrange multipliers, the objective function, which is originally difficult to solve directly, is decomposed into several independent or weakly coupled sub-optimization problems, thereby reducing computational complexity and improving solution efficiency.
[0144] By jointly modeling the above multiple constraints, the polarization-synthesized image obtained by inverse solving meets the expected requirements in terms of spatial resolution, angular consistency, and polarization characteristics.
[0145] In some embodiments, registration data is used to characterize the spatial correspondence between sub-images from different viewpoints in the original light field sequence image. This registration data can be obtained from a pre-executed light field viewpoint registration step, or calculated during system initialization through feature matching, disparity estimation, or other methods. The registration data provides a spatial alignment basis for updating light field image variables in subsequent iterations.
[0146] In some embodiments, the variables are updated alternately in each iteration as follows:
[0147] The light field image variable update process involves, with auxiliary variables and Lagrange multipliers fixed, minimizing the sub-objective function related to the light field image variables within the objective function based on the registration data of the current iteration, thereby updating the current light field image estimation result. This update process ensures that the light field image maintains data consistency with the original light field sequence image while satisfying the constraints of the imaging model.
[0148] Auxiliary variable updates involve solving subproblems corresponding to the auxiliary variables while keeping the light field image variables and Lagrange multipliers fixed, to satisfy regularization or polarization constraints. By introducing auxiliary variables, non-smooth or highly complex constraints are transformed into easily solvable forms, thereby improving the overall optimization stability.
[0149] The Lagrange multiplier update is performed based on the constraint residuals between the current light field image variables and auxiliary variables. This gradually strengthens the consistency constraints between variables, ensuring that the iteration results gradually approach the optimal solution of the objective function.
[0150] In some embodiments, the above iterative process continues until a preset convergence condition is met, such as the change in the objective function value being less than a threshold, the variable update magnitude being less than a threshold, or the number of iterations reaching a set upper limit.
[0151] In some embodiments, the reverse solution of the original light field sequence image is achieved through the above iterative solution process based on the alternating direction multiplier method, thereby generating a reconstructed polarization composite image while maintaining the consistency of the light field spatial angle information and polarization information.
[0152] In some embodiments, the reconstructed polarization composite image can be used for subsequent applications such as polarization feature analysis, target recognition, 3D reconstruction, or imaging enhancement, thereby improving the system's imaging performance and information acquisition capabilities in complex lighting or complex material environments.
[0153] In the example above, data fidelity terms and hybrid sparsity regularization terms are minimized within a unified framework. This process achieves the following in a single, collaborative manner: MTF compensation and super-resolution reconstruction based on a detail dictionary, high dynamic range (HDR) interferometric color restoration based on spectral and gradient sparsity, equivalent time delay integral (TDI) denoising based on nonlocal group sparsity, noise suppression (entropy reduction) based on transform domain sparsity, and geometric distortion correction based on sparse deformation fields. The final output is a high-quality, large-field-of-view polarization composite image with five-dimensional completion.
[0154] Referring to Figure 7, the principle block diagram of the hybrid sparse multi-dimensional collaborative reconstruction algorithm model provided by the present invention is shown. As shown in Figure 7, the registration data corresponding to the original multi-angle polarized light field data is processed by modeling through a five-dimensional coupled degradation equation, and the iterative results under the constraint matrix of detail dimension, the constraint matrix of temporal dimension, the constraint matrix of deformation dimension, the constraint matrix of entropy reduction dimension, and the constraint matrix of dynamic dimension are determined.
[0155] For example, perform hybrid sparse multidimensional collaborative optimization to obtain a high-quality "five-dimensional complete" polarization composite image, i.e., the target light field image.
[0156] S105, extract morphological and texture features and optical-physical features from the polarization composite image to determine the mineral type and distribution information of the sample to be tested.
[0157] In some embodiments, the above steps achieve a new paradigm of "single-shot - five-dimensional decoupling - end-to-end reconstruction," thereby obtaining high-quality images. Furthermore, the polarization-synthesized images can be further analyzed to determine the mineral types and distribution information of the sample. That is, feature analysis and intelligent discrimination are performed based on the reconstructed high-quality images.
[0158] In one example, step S105 may include:
[0159] Extraction operations are performed using multi-scale, multi-directional Gabor filters or Gaussian difference filter banks to obtain the particle size, shape, and cleavage texture of the sample under test, which are used as the morphological texture characteristics.
[0160] The crystal optical properties of the sample under test are determined by inverting and calculating the local slow light direction, approximate birefringence, and degree of polarization of each pixel in the polarization composite image, which are then used as the optical physical characteristics.
[0161] The morphological texture features and the optical physical features are fused into multi-dimensional features and input into a mineral identification deep learning model to determine the mineral type and distribution information of the sample to be tested.
[0162] In some optional examples, high dynamic range interference can also be performed on the polarization composite image to encode the dominant hue and saturation, forming a stable, digital description of the mineral interference colors and obtaining spectral color features.
[0163] Next, the multi-dimensional features are fused and input into a dedicated deep learning model for mineral identification.
[0164] In some embodiments, the mineral identification deep learning model is a dual-branch network architecture, including: a first branch is a convolutional neural network for processing morphological and textural features; and a second branch is a graph neural network for processing graph data constructed with pixels as nodes and optical features and spatial relationships as edges. The outputs of the two branches are fused at multiple levels during the decoding stage, and the mineral type and distribution information of the sample to be tested are generated through a segmentation head. In this way, identification can be achieved by fusing mineral physical properties and spatial symbiotic relationships, thereby more accurately determining the mineral type and distribution information and improving the identification accuracy.
[0165] Accordingly, the morphological texture features and the optical physical features are fused using multi-dimensional features and input into a mineral recognition deep learning model to determine the mineral type and distribution information of the sample to be tested, including:
[0166] An adaptive weighted fusion strategy based on channel attention is used to fuse the morphological texture features and the optical physical features.
[0167] The convolutional neural network is used to process the morphological and spectral information in the fused features. The local texture, edges and spectral patterns of the minerals are extracted through deep convolution, and a detailed local feature map is output.
[0168] The graph neural network described above is used to construct graph data with pixels as nodes and optical features and spatial relationships as edges through graph propagation or self-attention mechanism. The graph data is then integrated with the physical characteristics and spatial symbiotic relationships of the sample under test for identification, and the output is a relationship-enhanced feature map containing global context.
[0169] By employing a spatial attention gating mechanism, key regions of local feature branches are selectively enhanced. A shared semantic segmentation head then generates pixel-level mineral type identification maps and corresponding classification confidence maps based on the fused high-level features, thereby obtaining the mineral type and distribution information of the sample to be tested.
[0170] More specifically, an adaptive weighted fusion strategy based on channel attention is adopted to fuse three types of heterogeneous features: morphological texture, optical physics, and spectral color. Specifically, firstly, convolutional layers are used to map each feature to a unified dimension; then, a channel attention mechanism is used to dynamically learn the weights of each feature channel; finally, a weighted sum is performed to generate a fused feature map that adaptively emphasizes key discriminative information.
[0171] Convolutional Neural Networks (CNNs) are used to process the morphological and spectral information in the fused features. Through deep convolution, local textures, edges and spectral patterns of minerals are extracted, and detailed local feature maps are output.
[0172] A graph neural network (GNN) is employed. Regions or pixels in the image are used as nodes, with initial features composed of optical physical properties and intermediate features from local feature branches. Edges between nodes are defined by spatial adjacency and feature similarity. This branch models the internal heterogeneity of mineral particles and the topological structures such as symbiosis, encapsulation, and reaction edges between particles across the entire graph using graph propagation or self-attention mechanisms, outputting a relation-enhanced feature map containing global context.
[0173] The outputs of the two branches are deeply fused during the network's decoding stage. Typically, the contextual information from the global relation branch is selectively enhanced in key regions of local feature branches using a spatial attention gating mechanism. Finally, a shared semantic segmentation head generates pixel-level mineral type identification maps and corresponding classification confidence maps based on the fused high-level features. This model is trained end-to-end on a large labeled mineral thin section database, ultimately achieving accurate and automated analysis of mineral composition and microstructure.
[0174] To better understand and illustrate the deep learning model for mineral identification in the embodiments of the present invention, the specific structure, parameters, and workflow of the deep learning model for mineral identification are described below:
[0175] 1) Input and preprocessing of deep learning models for mineral identification:
[0176] Morphological and textural feature input: The response maps generated after processing by multi-scale, multi-directional Gabor filter banks are stacked into a multi-channel tensor, denoted as . Where H represents the height of the image and W represents the width of the image. The number of channels representing texture features is determined by the number of scales and directions of the filter bank.
[0177] In one embodiment, the scale number is taken (Corresponding to 3 scales), direction number taken (There are 6 directions in total), so the filter bank has a total of 3×6=18 filters.
[0178] If we select a portion of the scales and directions to combine, we can obtain C. t =12, for example, selecting 4 scales and 3 directions (such as 0°, 60°, 120°). This tensor serves as the input to the convolutional neural network branch, used to capture the texture, edge, and shape information of mineral particles at different scales and directions.
[0179] Optical physical characteristics and graph construction:
[0180] The three-channel quantitative optical parameter map obtained by inverting the polarization composite image is expressed as follows:
[0181] ;
[0182] Among them, parameters This indicates the direction of vibration in a mineral crystal where the light propagation speed is slower under polarized light, measured in degrees (°) or radians. This parameter reflects the direction of optical anisotropy of the crystal and is an important basis for identifying the crystal system and crystal orientation of a mineral.
[0183] This parameter represents the difference in refractive index between two mutually perpendicular vibrational directions of a mineral crystal. It is closely related to the crystal structure and chemical composition of the mineral and is a key optical constant for distinguishing different minerals (especially transparent and translucent minerals).
[0184] This parameter represents the proportion of polarized light in reflected or transmitted light, and its value ranges from [0, 1]. It reflects the ability of a mineral surface to retain or scatter polarized light and is related to the mineral's surface roughness, internal scattering characteristics, and optical homogeneity.
[0185] These three parameters together constitute the optical feature vector of each pixel, which is used for the subsequent construction of node features in the graph neural network.
[0186] The obtained three-channel quantitative optical parameter map is transformed into a graph structure to construct an undirected graph G = (V, E), where V is the set of nodes, denoted as... , Each node For a single pixel, its initial features Represented as E is the set of edges.
[0187] E is constructed based on two relationships: a) Spatial adjacency relationship: adjacent pixel nodes are connected using 8-neighbor or K-nearest neighbor methods; b) Feature similarity relationship: the cosine similarity between any two nodes i and j is calculated based on optical features. ,like (For example, If the similarity is higher than the threshold, then the similarity is higher than the threshold. Therefore, add an edge (i, j).
[0188] 2) First branch: Convolutional neural network processes morphological and texture features
[0189] U-Net is a fully convolutional network with an encoder-decoder structure.
[0190] The encoder section comprises four convolutional blocks (Conv Blocks). Each Conv Block consists of two consecutive 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function, followed by a 2×2 max-pooling layer for downsampling. The number of channels is [64, 128, 256, 512, 1024]. Through deeply stacked convolutional and pooling operations, high-level semantic features such as local texture, edges, shape, and spectral patterns of the minerals are gradually extracted and abstracted, outputting a detailed local feature map. , is represented as: .
[0191] 3) Second branch: Graph neural networks process optical features and spatial relationships
[0192] A two-layer graph attention network (GAT) is used, and each layer performs the following operations: ,in It is a normalized attention adjacency matrix, rather than a simple 0-1 adjacency matrix.
[0193] In GAT, elements The weight representing the importance of node j's features to node i is calculated by the attention mechanism and normalized (e.g., using softmax), making... , dimension . For the first Layer node feature matrix, dimension is Where N is the total number of pixels in the image (i.e., the number of nodes). It is the first The feature dimension of the layer. The weight matrix is trainable and has dimensions of . It is used to perform linear transformations and dimensionality reduction / increment on node features. This represents a nonlinear activation function, typically ReLU or LeakyReLU, used to introduce nonlinear transformations. The new node feature matrix is output after the graph attention aggregation and transformation in this layer.
[0194] Input node optical feature matrix That is, the first layer of the graph neural network (i.e. The input of the layer. X=H 0 , dimension , Indicates the dimension of the input features.
[0195] In one example, C0=3, corresponding to the three optical physical features of each pixel [φ (slow light direction), Δn (birefringence), DoP (degree of polarization)].
[0196] The adjacency matrix A, a sparse N×N matrix, defines the connections (edges) between nodes in the graph. It is typically constructed based on pixel spatial adjacency relationships (e.g., 8-neighborhood). After feature propagation through a two-layer graph neural network (GNN), it outputs an updated node feature matrix. At this point, the feature dimension of each node remains C0 (i.e., However, the features at this point are no longer isolated, but rather integrate the optical properties and spatial context information of all connected nodes in its local neighborhood.
[0197] 4) Feature fusion and decoding recognition
[0198] First, the node feature matrix output by the graph neural network branch is... Reconstructing the original pixels into a spatial feature map based on their spatial coordinates, with dimensions of [missing information]. This allows graph-structured data to be converted back to image formats that convolutional neural networks can process, achieving the same functionality as the branch feature maps of convolutional neural networks. Align in spatial dimensions.
[0199] Next, adaptive weighted fusion at the channel level is implemented using SE Block (Squeeze-and-Excitation). First, let's... and By splicing along the channel dimension, we obtain .
[0200] Assumption The number of channels is 128. If the number of channels is 3, then Dimension .
[0201] Next, the spliced features Perform global average pooling to compress the H×W spatial information of each channel into a scalar, resulting in a channel description vector with dimension 1. The non-linear dependencies between channels are learned through a first fully connected layer (e.g., reducing the dimensionality to 1 / 4 of the channel count) and ReLU activation. A second fully connected layer (restoring the original 131 channels) and Sigmoid activation generate a weight coefficient between 0 and 1 for each channel. The learned channel weight coefficients are then compared with F... cnn and F gnn The corresponding channels are multiplied element-wise to achieve adaptive weighting, highlighting feature channels with rich information and suppressing redundant channels. The weighted F... cnn and F gnn Adding them together yields the fused feature map F. fused .
[0202] To highlight key areas for mineral identification, the CBAM spatial module is used to enhance the fusion features: first, for... Max pooling and average pooling are performed along the channels. The results are concatenated and then used to generate a spatial weight mask Ms through a 7×7 convolutional layer and sigmoid activation. The mask is then multiplied element-wise with the fused features to obtain the enhanced features. .
[0203] Finally, the recognition and decoding are completed using the segmentation head: first, ... The process involves using a 3×3 convolutional layer (outputting 256 channels) and ReLU activation, followed by a 1×1 convolutional layer mapping to the dimension of the number of mineral categories K. Softmax activation is then applied to obtain a pixel-level category probability map P (dimension 1). Based on this probability map, the predicted category of each pixel is determined through the Argmax operation to obtain the mineral identification map; at the same time, the maximum probability value of each pixel is taken to generate a classification confidence map.
[0204] 5) Specific steps for model training
[0205] For training data, at least 1,000 labeled polarized thin section images need to be collected, and each image must be accompanied by a pixel-level mineral category label (the label must be reviewed by experts or verified using EPMA, Raman, or other technologies).
[0206] To improve the model's generalization ability, the original image will also be randomly rotated, flipped, and its brightness and contrast will be slightly adjusted.
[0207] The loss function adopts a hybrid form of cross-entropy loss and Dice loss, and the formula is as follows: The weighting coefficient =1.0, weighting coefficient =0.5, This represents the cross-entropy loss function, which measures the difference between the pixel-level class prediction probability distribution output by the model and the true label distribution. It dominates the optimization of whether the classification of each pixel is correct and is the basic loss for semantic segmentation tasks. This represents the Dessian coefficient, a set similarity measure used to calculate the overlap between two samples (in this case, the predicted region and the ground truth region). Its value ranges from 0 to 1, with higher values indicating better overlap.
[0208] The optimizer chosen is AdamW with weight decay coefficients. By decoupling weight decay regularization, better generalization performance and training stability are achieved. The weight decay coefficient of AdamW is set to... This regularization term is used to prevent the model from overfitting by penalizing large weights, thus encouraging the model to learn simpler patterns. The initial learning rate is set to 1× Fine-tuning is performed on pre-training or complex tasks to maintain a smooth start to training. The learning rate adjustment strategy adopts the Cosine AnnealingLR method, which enables the model to converge quickly with large steps in the early stage and fine-tun with very small steps in the later stage. This helps to escape local optima and find better model parameters.
[0209] The batch size is set to 8 based on the available GPU memory (corresponding to an image size of 512×512); training is implemented using PyTorch 2.0 and CUDA 11.8; an early stopping strategy is also adopted, and training is terminated if the mIoU of the validation set does not improve for several consecutive rounds (e.g., 10 rounds).
[0210] Thus, to address the core challenge of identifying optical anisotropy in mineral crystals, a dedicated recognition model deeply coupled with a high-quality image generation process was designed. This model innovatively integrates quantitative optical parameter maps (such as local slow light direction and approximate birefringence) derived from reconstructed images with morphological and textural features, utilizing graph neural networks to model the internal fractal nature and spatial co-occurrence relationships of mineral grains. This design, closely integrated with mineral physical properties, endows the model with expert-like "feature reasoning" capabilities, significantly improving the accuracy of distinguishing minerals with similar optical properties (such as different feldspar varieties) compared to traditional image classification methods.
[0211] In some embodiments, the mineral identification and analysis method based on computational polarization field reconstruction further includes: generating and outputting a comprehensive analysis result containing a mineral distribution map and a quantitative identification report of the sample to be tested, so as to generate a structured and interpretable intelligent identification report.
[0212] More specifically, it includes:
[0213] Based on the mineral type identification map of the determined sample to be tested, quantitative analysis is performed to determine the ratio of the pixel area occupied by each mineral type to the total image area, and a mineral content percentage report is generated.
[0214] Connectivity analysis is performed on the identification results to segment out independent mineral particles. The equivalent circle diameter of each particle is determined and calculated. A histogram of particle size distribution is plotted according to mineral type, and the average particle size and particle size sorting coefficient are calculated.
[0215] The identification results of the same specific region where the sample to be tested is located are compared with the results obtained by micro-area chemical composition / molecular structure analysis techniques such as electron probe microanalysis or micro Raman spectroscopy to determine the overall classification accuracy and Kappa coefficient.
[0216] The comprehensive analysis result is obtained by combining the polarization composite image, the mineral type identification map, the mineral content percentage report, the particle size distribution histogram, the overall classification accuracy, and the Kappa coefficient.
[0217] For example, based on the output mineral type identification map (pixel-level label map), automatic quantitative analysis is performed: the ratio of the pixel area occupied by each mineral category to the total image area is calculated, generating a mineral (area) content percentage report. First, connected component analysis is performed on the identification results to segment individual mineral particles. Then, the equivalent circle diameter of each particle is calculated, and a particle size distribution histogram is statistically analyzed according to mineral type, calculating parameter reports such as average particle size and particle size sorting coefficient.
[0218] To assess the reliability of the identification results of this system, the identification results of a specific region on the same rock thin section can be compared with the results obtained from micro-area chemical composition / molecular structure analysis techniques such as electron probe microanalysis (EPMA) or micro Raman spectroscopy. The mineral identification accuracy of this invention is quantitatively verified by calculating indicators such as overall accuracy and Kappa coefficient. This verification result can also serve as a confidence level reference in subsequent comprehensive reports.
[0219] The system automatically integrates the analysis results from the entire process to generate a structured and interpretable intelligent identification report. The report will automatically include the reconstructed high-quality wide-field panoramic image, the generated mineral distribution map, the quantitative statistical table of mineral content and grain size distribution, and the derived distribution map of key optical parameters.
[0220] By employing the scheme described in the example above, and deeply integrating computational light field imaging, multi-angle polarization illumination encoding, and a hybrid sparse multi-dimensional collaborative reconstruction model, this system achieves a novel paradigm for mineral imaging: single-shot imaging, five-dimensional decoupling, and end-to-end reconstruction. The proposed hybrid sparse representation model unifies the improvement of five dimensions—MTFC, dynamic range (HDR), signal-to-noise ratio (TDI / entropy reduction), and geometric deformation—within a single optimization framework for collaborative solution. This fundamentally avoids the problems of noise enhancement, artifact generation, and color fidelity degradation inherent in traditional serial processing. This technology breaks through the physical limitation that high resolution and a large field of view are mutually exclusive, providing a comprehensive and high-quality polarization image foundation for intelligent identification.
[0221] In summary, this invention constructs a complete intelligent analysis workflow from sample placement to report generation. The system automates the entire process of image acquisition, quality enhancement, feature extraction, mineral identification, and quantitative statistics, ultimately outputting a mineral distribution map, a quantitative optical parameter table, and a structured identification report. This solution liberates experts from tedious and repetitive manual operations and subjective interpretations, establishing objective and quantifiable standards for mineral identification, significantly improving the scientific rigor, consistency, and overall efficiency of analytical work, and providing a complete technical toolchain for realizing the digital and intelligent analysis of mineralogy.
[0222] The mineral identification and analysis method based on computational polarization field reconstruction has been described in detail above through some embodiments. In order to enable those skilled in the art to better understand and implement it, the corresponding system is also described in detail below through some embodiments.
[0223] Referring to Figure 8, which shows a schematic diagram of a mineral identification and analysis system based on computational polarization field reconstruction provided by the present invention, the mineral identification and analysis system 800 based on computational polarization field reconstruction may include:
[0224] The acquisition unit 810 is configured to acquire at least one set of raw light field image sequences of the sample under test with multiple views and multiple polarization states.
[0225] Registration unit 820 is configured to perform registration processing on the original light field sequence image so that the images under all viewpoints correspond in spatial position to obtain registration data;
[0226] The establishment unit 830 is configured to establish a degradation relationship model for describing the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence, and the degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise.
[0227] The reconstruction unit 840 is configured to generate a polarization composite image by inverse solving based on the degradation relationship model and the registration data;
[0228] The processing unit 850 is configured to extract morphological and texture features and optical-physical features from the polarization composite image to determine the mineral type and distribution information of the sample to be tested.
[0229] For further details regarding the acquisition unit 810, registration unit 820, establishment unit 830, reconstruction unit 840, and processing unit 850, please refer to the aforementioned examples.
[0230] It is understandable that the above division of units is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the above units can be implemented by the processor calling software.
[0231] In one embodiment, referring to Figure 9, a module diagram of a mineral identification and analysis system based on computational polarization field reconstruction provided by the present invention includes:
[0232] The data acquisition and preprocessing layer specifically includes: a lighting control unit, an image acquisition unit, and a registration and correction unit. For a detailed description of the data acquisition and preprocessing layer, please refer to the relevant descriptions in Figures 1 to 5 in the aforementioned examples.
[0233] The algorithm processing layer specifically includes: a five-dimensional degradation model library, a hybrid sparse reconstruction engine, and a feature extractor. For a detailed description of the algorithm processing layer, please refer to the relevant descriptions in Figures 6 and 7 of the aforementioned example.
[0234] The identification and application layer specifically includes: a dedicated mineral identification model, a quantitative analysis module, and a report generation engine.
[0235] In addition, it includes a user interface coupled to the recognition and application layer and the data acquisition and preprocessing layer.
[0236] This invention also provides an electronic device for implementing a mineral identification and analysis method based on computational polarization field reconstruction.
[0237] The electronic device includes a memory and a processor, as well as a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the mineral identification and analysis method based on computational polarization field reconstruction as described in any of the preceding claims.
[0238] It should be noted that the computer system of the electronic device shown below is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0239] Electronic devices include memory, a processor, and a transceiver. The processor is coupled to the memory and transceiver. The memory can be located inside or outside the terminal. The memory, processor, and transceiver can be connected via a communication bus. The transceiver is used to communicate with other devices or communication networks.
[0240] Optionally, the transceiver can be a transmitter. The memory stores a computer program that can run on the processor, and when the processor runs the computer program, the transceiver executes the steps in the mineral identification and analysis method based on computational polarization field reconstruction provided in the above embodiments.
[0241] It should be understood that in the embodiments of this application, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0242] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM).
[0243] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A mineral identification and analysis method based on computational polarization field reconstruction, characterized in that, include: Acquire at least one set of original light field image sequences of the sample under test from multiple perspectives and multiple polarization states; perform registration processing on the original light field image sequences so that the images from all perspectives correspond in spatial position to obtain registration data; A degradation relationship model is established to describe the relationship between the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence. The degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise. The degradation relationship model includes: ;in, Characterizes the original light field image sequence; Characterizes the target light field image; Nonlinear operators characterizing the dynamic range compression; The transformation matrix characterizing the geometric deformation; The blur kernel characterizes the aforementioned detail blur; A downsampling operator characterizing information loss caused by the finite sampling rate of the sensor; Additive noise characterizing noise interference; The ordinal number of the original light field image sequence is represented; a polarization composite image is generated by inverse solving based on the degradation relation model and the registration data; wherein, the generation of the polarization composite image by inverse solving based on the degradation relation model and the registration data includes: constructing an objective function for the degradation relation model, the objective function including: a data fidelity term based on the degradation relation model, and a hybrid sparse regularization term for constraining five dimensions; iteratively solving the objective function using the alternating direction multiplier method, and alternately updating the light field image variables, auxiliary variables, and Lagrange multipliers based on the registration data in each iteration, thereby realizing the inverse solving of the original light field image sequence and generating the reconstructed polarization composite image; wherein, the objective function includes: ;in, Characterizes data fidelity; Sparse transforms characterizing non-convex higher-order total variation. The regularization term in the objective function used to apply non-convex higher-order total variation constraints represents different... Constraints corresponding to different types of image reconstruction; The weight parameters corresponding to each regularization term; Characterizes the L1 norm; where, At that time, detail-enhanced regularization ;in, This indicates that the L0 norm constraint is applied. sparsity, , For pre-trained high-frequency detail dictionary, For image patch extraction operations, the representation forces the reconstructed image to have a sparse representation under the L0 norm constraint in the detail dictionary; Dynamic range expansion regularization ;in, This indicates the L1 norm sparsity under a constrained high dynamic range dictionary. , Characterizing a high dynamic range dictionary Represents the target light field image Perform L1 norm sparsification. Represents the total variation regularization term; At that time, signal-to-noise ratio enhancement regularization The characterization focuses on the temporal dimension and applies joint sparsity constraints to non-locally similar blocks in multi-view images. Norm; At that time, noise suppression regularization ;in For a tight frame transformation, characterize the application of a tight frame transformation in the transform domain. Sparse constraints; This indicates that the target light field image x is processed. The result after transformation; At that time, geometric correction regularization ;in For the sparse deformation field estimated from the data, its gradient sparsity is constrained; Represents the calculation of deformation field Spatial gradient; Gradient sparsity is constrained by the L0 norm for geometric correction; morphological and texture features and optical-physical features are extracted from the polarization composite image to determine the mineral type and distribution information of the sample to be tested.
2. The mineral identification and analysis method based on computational polarization field reconstruction according to claim 1, characterized in that, The acquisition of at least one set of original light field image sequences of the sample under test with multiple views and multiple polarization states includes: when the sample under test is placed on the stage, using a programmable multi-angle LED array illumination unit located below the sample under test, and in conjunction with a condenser lens, switching illumination beams with different incident angles and different polarization states according to a preset time sequence, so that the illumination beams penetrate the stage and illuminate the sample under test; the objective lens refracts the illumination beams to a micro-polarizer array light field sensor, acquires the intensity information of the illumination beams in different polarization directions within a single exposure cycle under the corresponding illumination mode, and acquires the original light field image sequence.
3. The mineral identification and analysis method based on computational polarization field reconstruction according to claim 1, characterized in that, The registration process for the original light field image sequence, which aligns the images at all viewpoints in spatial position to obtain registration data, includes: using the image corresponding to the principal optical axis viewpoint in the original light field image sequence as a geometric reference image; for each image to be registered in the original light field image sequence other than the reference image, performing a frequency domain subpixel registration operation, which includes: calculating the normalized cross-power spectrum of the image to be registered and the reference image in the Fourier domain; performing subpixel interpolation on the peak position of the normalized cross-power spectrum to determine the image plane displacement of the image to be registered relative to the reference image; and performing geometric transformation compensation on all images based on the image plane displacement to achieve alignment of image content at the spatial position at each viewpoint.
4. The mineral identification and analysis method based on computational polarization field reconstruction according to claim 1, characterized in that, The step of extracting morphological and texture features and optical-physical features from the polarization-synthesized image to determine the mineral type and distribution information of the sample under test includes: performing extraction operations using multi-scale, multi-directional Gabor filters or Gaussian difference filter banks to obtain the grain size, shape, and cleavage texture of the sample under test as the morphological and texture features; determining the crystal optical properties of the sample under test by inverting and calculating the quantitative optical parameter map of the local slow light direction, approximate birefringence, and degree of polarization of each pixel in the polarization-synthesized image as the optical-physical features; fusing the morphological and texture features and the optical-physical features into multi-dimensional features and inputting them into a mineral recognition deep learning model to determine the mineral type and distribution information of the sample under test.
5. The mineral identification and analysis method based on computational polarization field reconstruction according to claim 4, characterized in that, The mineral identification deep learning model is a dual-branch network architecture, comprising: a first branch being a convolutional neural network for processing morphological texture features; and a second branch being a graph neural network for processing graph data constructed with pixels as nodes and optical features and spatial relationships as edges. The outputs of the two branches are fused at multiple levels during the decoding stage, and a segmentation head is used to generate the mineral type and distribution information of the sample under test. The process of fusing the morphological texture features and the optical-physical features in multiple dimensions and inputting this fusion into the mineral identification deep learning model to determine the mineral type and distribution information of the sample under test includes: employing an adaptive weighted fusion strategy based on channel attention to fuse the morphological texture features and the optical-physical features; and... The convolutional neural network is used to process the morphological and spectral information in the fused features. Deep convolution is used to extract the local texture, edges, and spectral patterns of the minerals, outputting a detailed local feature map. The graph neural network, through graph propagation or self-attention mechanisms, constructs graph data with pixels as nodes and optical features and spatial relationships as edges. This data is then fused with the physical properties and spatial symbiotic relationships of the sample under test for identification, outputting a relationship-enhanced feature map containing global context. A spatial attention gating mechanism is used to selectively enhance key regions of local feature branches. A shared semantic segmentation head generates a pixel-level mineral type identification map and a corresponding classification confidence map based on the fused high-level features, thus obtaining the mineral type and distribution information of the sample under test.
6. The mineral identification and analysis method based on computational polarization field reconstruction according to claim 5, characterized in that, Also includes: Generate and output a comprehensive analysis result including a mineral distribution map and a quantitative identification report of the sample to be tested, including: based on the determined mineral type identification map of the sample to be tested, perform quantitative analysis, determine the ratio of the pixel area occupied by each mineral type to the total image area, and generate a mineral content percentage report. Connectivity analysis is performed on the identification results to segment individual mineral particles. The equivalent circle diameter of each particle is calculated, and a particle size distribution histogram is plotted according to mineral type. The average particle size and particle size sorting coefficient are calculated. The identification results of the same specific region of the sample to be tested are compared with the results obtained by electron probe microanalysis or micro Raman spectroscopy micro-area chemical composition / molecular structure analysis technology for that specific region to determine the overall classification accuracy and Kappa coefficient. The comprehensive analysis result is obtained by combining the polarization composite image, the mineral type identification map, the mineral content percentage report, the particle size distribution histogram, the overall classification accuracy, and the Kappa coefficient.
7. A mineral identification and analysis system based on computational polarization field reconstruction, characterized in that, include: The acquisition unit is configured to acquire at least one set of raw light field image sequences of the sample under test from multiple perspectives and multiple polarization states; The registration unit is configured to perform registration processing on the original light field image sequence so that the images under all viewpoints correspond in spatial position, thereby obtaining registration data. The establishment unit is configured to establish a degradation relationship model describing the relationship between the original light field image sequence and the target light field image, wherein the target light field image is obtained based on the original light field image sequence, and the degradation relationship includes at least five dimensions: dynamic range, geometric deformation, detail blur, resolution, and noise; wherein the degradation relationship model includes: ;in, Characterizes the original light field image sequence; Characterizes the target light field image; Nonlinear operators characterizing the dynamic range compression; The transformation matrix characterizing the geometric deformation; The blur kernel characterizes the aforementioned detail blur; A downsampling operator characterizing information loss caused by the finite sampling rate of the sensor; Additive noise characterizing noise interference; The ordinal number of the original light field image sequence is represented; the reconstruction unit is configured to generate a polarization composite image by inverse solving based on the degradation relation model and the registration data; wherein, the generation of the polarization composite image by inverse solving based on the degradation relation model and the registration data includes: constructing an objective function for the degradation relation model, the objective function including: a data fidelity term based on the degradation relation model, and a hybrid sparse regularization term for constraining five dimensions; iteratively solving the objective function using the alternating direction multiplier method, and alternately updating the light field image variables, auxiliary variables, and Lagrange multipliers based on the registration data in each iteration, thereby realizing the inverse solving of the original light field image sequence and generating the reconstructed polarization composite image; wherein, the objective function includes: ;in, Characterizes data fidelity; Sparse transforms characterizing non-convex higher-order total variation. The regularization term in the objective function used to apply non-convex higher-order total variation constraints represents different... Constraints corresponding to different types of image reconstruction; The weight parameters corresponding to each regularization term; Characterizes the L1 norm; where, At that time, detail-enhanced regularization ;in, This indicates that the L0 norm constraint is applied. sparsity, , For pre-trained high-frequency detail dictionary, For image patch extraction operations, the representation forces the reconstructed image to have a sparse representation under the L0 norm constraint in the detail dictionary; Dynamic range expansion regularization ;in, This indicates the L1 norm sparsity under a constrained high dynamic range dictionary. , Characterizing a high dynamic range dictionary Represents the target light field image Perform L1 norm sparsification. Represents the total variation regularization term; At that time, signal-to-noise ratio enhancement regularization The characterization focuses on the temporal dimension and applies joint sparsity constraints to non-locally similar blocks in multi-view images. Norm; At that time, noise suppression regularization ;in For a tight frame transformation, characterize the application of a tight frame transformation in the transform domain. Sparse constraints; This indicates that the target light field image x is processed. The result after transformation; At that time, geometric correction regularization ;in For the sparse deformation field estimated from the data, its gradient sparsity is constrained; Represents the calculation of deformation field Spatial gradient; Gradient sparsity is constrained by the L0 norm for geometric correction; the processing unit is configured to extract morphological and texture features and optical-physical features from the polarization composite image to determine the mineral type and distribution information of the sample to be tested.
Citation Information
Patent Citations
Mineral information identification system and method based on spectrum enhancement
CN116580237A
Electric orthogonal polarized light synchronous control system of geological microscope
CN120779576A