Component detection method, device and system based on multi-view Transform and storage medium

By combining NIRS and XRF spectral signals with a multi-view Transformer hybrid model, the problem of insufficient accuracy in mineral composition detection was solved, achieving higher detection accuracy.

CN120594433AActive Publication Date: 2025-09-05HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511094542.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

When existing technologies are used to detect the composition of minerals with complex elemental compositions and diverse structures, the detection accuracy cannot meet the requirements, especially the limitations of single spectral methods cannot meet the detection needs.

Method used

A multi-view Transformer hybrid model is adopted to combine the two spectral signals of NIRS and XRF. The spectral signals are processed through multiple preprocessing methods, and the multi-view Transformer hybrid model is used for component detection to fuse the feature extraction results of multiple spectral signals.

Benefits of technology

The accuracy of mineral composition detection is improved, and the effective information content of model input data is increased through various preprocessing methods, thereby enhancing the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120594433A_ABST
    Figure CN120594433A_ABST
Patent Text Reader

Abstract

The invention provides a component detection method, device and system based on a multi-view Transform, and a storage medium. In one example, the method comprises: acquiring spectral signals of a mineral sample to be detected, the spectral signals comprising a first type of spectral signals and a second type of spectral signals; performing N types of preprocessing on the first type of spectral signals to obtain N types of preprocessed first type of spectral signals; and performing M types of preprocessing on the second type of spectral signals to obtain M types of preprocessed second type of spectral signals, and determining a component detection result of the to-be-detected mineral sample by using a pre-trained multi-view Transformer hybrid model according to the N types of preprocessed first type spectral signals and the M types of preprocessed second type spectral signals. The method can improve the accuracy of mineral component detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of component detection technology, and in particular to a component detection method, device, system and storage medium based on a multi-view Transformer. Background Art

[0002] Component testing generally refers to the process of analyzing and testing the chemical composition, elemental composition or physical properties of substances, materials or products. It is widely used in many fields, including food, medicine, environment, chemical industry, materials science, etc.

[0003] For minerals with complex elemental composition and diverse structures, such as coal, using a single spectrum for component detection will have limitations, and the detection accuracy is usually difficult to meet the requirements. Summary of the Invention

[0004] In view of this, the present application provides a component detection method, device, system and storage medium based on a multi-view Transformer.

[0005] Specifically, this application is implemented through the following technical solutions: According to a first aspect of an embodiment of the present application, a component detection method based on a multi-view Transformer is provided, comprising: Acquire a spectral signal of a mineral sample to be detected, the spectral signal comprising a first type spectral signal and a second type spectral signal; the first type spectral signal is a near-infrared spectroscopy (NIRS) signal, and the second type spectral signal is an X-ray fluorescence (XRF) spectroscopy signal; or the first type spectral signal is an XRF signal, and the second type spectral signal is a NIRS signal; performing N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type of spectral signals; and performing M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type of spectral signals; wherein N≥1, M≥1, and N*M≥2; Based on the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals, a pre-trained multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected; wherein, the multi-view Transformer hybrid model includes N*M feature extraction structures, which are respectively used to extract features from N*M spectral signal combinations. The regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected based on fusion features. The fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures. One spectral signal combination includes one type of preprocessed first-type spectral signal and one type of preprocessed second-type spectral signal.

[0006] According to a second aspect of an embodiment of the present application, a component detection device based on multi-spectral combination is provided, comprising: an acquisition unit, configured to acquire a spectral signal of a mineral sample to be detected, the spectral signal comprising a first type spectral signal and a second type spectral signal; the first type spectral signal being a near-infrared spectral (NIRS) signal, and the second type spectral signal being an X-ray fluorescence (XRF) spectral signal; or the first type spectral signal being an XRF signal, and the second type spectral signal being a NIRS signal; a preprocessing unit configured to perform N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type spectral signals; and perform M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type spectral signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; A detection unit is used to determine the composition detection result of the mineral sample to be detected based on the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection result of the mineral sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures, and one spectral signal combination includes one type of preprocessed first-type spectral signal and one type of preprocessed second-type spectral signal.

[0007] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the method provided in the first aspect.

[0008] According to a fourth aspect of an embodiment of the present application, a machine-readable storage medium is provided, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method provided in the first aspect is implemented.

[0009] According to a fifth aspect of the embodiments of the present application, a coal quality rapid detection system is provided, comprising: a spectrum acquisition unit and a mineral composition detection unit; wherein: A spectrum acquisition unit, used for acquiring spectrum data of the coal sample to be tested; The spectrum acquisition unit includes a near infrared spectrum acquisition module for acquiring near infrared spectrum NIRS signals of the coal sample to be detected, and an X-ray fluorescence spectrum acquisition module for acquiring X-ray fluorescence spectrum XRF signals of the coal sample to be detected; a mineral composition detection unit configured to perform N types of preprocessing on the NIRS signal to obtain N types of preprocessed NIRS signals; and perform M types of preprocessing on the XRF signal to obtain M types of preprocessed XRF signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; The mineral composition detection unit is also used to determine the composition detection results of the coal sample to be detected based on the N types of preprocessed NIRS signals and M types of preprocessed XRF signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the coal sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures, and one spectral signal combination includes one type of preprocessed NIRS signal and one type of preprocessed XRF signal.

[0010] The technical solution provided by this application can at least bring the following beneficial effects: By acquiring a first type of spectral signal and a second type of spectral signal of a mineral sample to be detected, and performing N types of preprocessing on the first type of spectral signal, N types of preprocessed first type of spectral signals are obtained; and, performing M types of preprocessing on the second type of spectral signal, M types of preprocessed second type of spectral signals are obtained. Then, based on the N types of preprocessed first type of spectral signals and the M types of preprocessed second type of spectral signals, a pre-trained multi-view Transformer hybrid model is used to determine the composition detection result of the mineral sample to be detected. By adopting multiple preprocessing methods to preprocess the first type of spectral signal and / or the second type of spectral signal, the effective information amount of the model input data is increased while improving the quality of the input data, thereby improving the accuracy of mineral composition detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 1 is a flowchart of a component detection method based on a multi-view Transformer, shown in an exemplary embodiment of the present application; Figure 2 is a structural diagram of a DFFormer shown in an exemplary embodiment of the present application; Figure 3 1 is a schematic diagram of a framework for mask reconstruction pre-training based on spectral physical properties, as shown in an exemplary embodiment of the present application; Figure 4A is a schematic diagram of independent training of each perspective shown in an exemplary embodiment of the present application; Figure 4B is a schematic diagram of a multi-view joint training shown in an exemplary embodiment of the present application; Figure 5 1 is a structural diagram of a component detection device based on a multi-view Transformer shown in an exemplary embodiment of the present application; Figure 6 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, some technical terms involved in the embodiments of the present application are briefly explained below.

[0013] 1. NIRS (Near-infrared spectroscopy): It uses the absorption and scattering characteristics of substances to analyze molecular vibrations (such as CH, OH, and NH chemical bonds) and is a molecular spectroscopy technology.

[0014] 2. XRF (Near-infrared spectroscopy): bombards samples with high-energy X-rays to excite the inner electrons of atoms and release characteristic X-ray fluorescence. It is an atomic spectroscopy technique.

[0015] 3. MAE (Masked Autoencoder): A self-supervised learning framework.

[0016] 4. Matrix effect: Matrix refers to the components other than the analytical elements in the sample being measured. The presence of the matrix has a particularly prominent impact on XRF measurement and analysis, especially quantitative analysis. This effect is usually called the matrix effect. In addition to the absorption effect and enhancement effect between the sample components, it also includes the physical state of the sample such as surface effect, mineral effect, particle size, etc., as well as the peak shift and peak shape change caused by the different valence states of the analyzed elements.

[0017] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0018] It should be noted that the serial numbers of the steps in the embodiments of the present application do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0019] See Figure 1 , which is a flow chart of a component detection method based on a multi-view Transformer provided in an embodiment of the present application, such as Figure 1 As shown, the component detection method based on multi-view Transformer may include the following steps: Step S100: Acquire a spectral signal of a mineral sample to be detected, where the spectral signal includes a first type spectral signal and a second type spectral signal.

[0020] In the embodiments of the present application, considering that minerals usually have complex elemental compositions and diverse structures, in order to improve the accuracy of mineral component detection, a variety of different types of spectral signals can be obtained during the component detection of the mineral sample to be detected.

[0021] For example, the mineral sample to be detected may include a mineral sample such as petroleum or coal.

[0022] Illustratively, different types of spectral signals may include, but are not limited to, NIRS (Near-infrared spectroscopy) signals, XRF (Near-infrared spectroscopy) signals, or LIBS (Laser-Induced Breakdown Spectroscopy) signals.

[0023] Exemplarily, the first type of spectral signal is a NIRS signal, and the second type of spectral signal is an XRF signal; or, the first type of spectral signal is an XRF signal, and the second type of spectral signal is a NIRS signal.

[0024] Step S110: performing N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type of spectral signals; and performing M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type of spectral signals; wherein N≥1, M≥1, and N*M≥2.

[0025] In the embodiments of this application, the process of collecting spectral signals from the mineral samples to be tested often introduces interference from irrelevant factors, such as light scattering caused by poorly sealed darkrooms, stray light from external light sources, and instrument noise. These interfering signals can seriously affect the spectral quality and spectral line information, and thus affect the composition detection results. Therefore, the acquired spectral signals need to be preprocessed to improve the accuracy of composition detection.

[0026] In addition, considering that the amount of information contained in the data processed by different preprocessing methods is not exactly the same, the detection results of the model will be more accurate when the spectral signals preprocessed according to multiple different preprocessing methods are used to detect the composition of the mineral samples to be detected.

[0027] Accordingly, the acquired first-type spectral signals and second-type spectral signals can be processed by a variety of preprocessing methods to obtain a combination of a plurality of preprocessed first-type spectral signals and a plurality of preprocessed second-type spectral signals, and based on the combination of a plurality of preprocessed first-type spectral signals and a plurality of preprocessed second-type spectral signals, component detection can be performed using a multi-view Transformer hybrid model.

[0028] For example, N types of preprocessing may be performed on NIRS signals to obtain N types of preprocessed NIRS signals; and M types of preprocessing may be performed on XRF signals to obtain M types of preprocessed XRF signals.

[0029] For example, for NIRS signals, candidate pre-processing methods may include, but are not limited to, area normalization (norm), standard normal transformation (snv), and first-order derivative (diffx).

[0030] For XRF signals, candidate preprocessing methods may include, but are not limited to, standard normal transformation and channel normalization.

[0031] Exemplarily, a spectral signal combination may include a first type of spectral signal preprocessed using a preprocessing method (i.e., a type of preprocessed first type spectral signal), and a second type of spectral signal preprocessed using a preprocessing method (i.e., a type of preprocessed second type spectral signal).

[0032] For example, one spectral signal combination may include an area-normalized NIRS signal and an XRF signal after a standard normal transformation; another spectral signal combination may include an area-normalized NIRS signal and an XRF signal after a channel-normalized process.

[0033] Step S120: Based on N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals, a pre-trained multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected; wherein, the multi-view Transformer hybrid model includes N*M feature extraction structures, which are respectively used to extract features from N*M spectral signal combinations. The regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected based on the fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures.

[0034] In an embodiment of the present application, for the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals obtained by processing in the above manner, the preprocessed first-type spectral signals and the preprocessed second-type spectral signals can be combined to obtain N*M spectral signal combinations.

[0035] Each of the N*M spectral signal combinations can be used as input to N*M feature extraction structures in a pre-trained multi-view Transformer hybrid model.

[0036] Exemplarily, the N*M feature extraction structures of the multi-view Transformer hybrid model correspond one-to-one to each spectral signal combination in the N*M spectral signal combinations.

[0037] Exemplarily, for any feature extraction structure among the N*M feature extraction structures of the multi-view Transformer hybrid model, the feature extraction structure takes the preprocessed first type spectral signal and the preprocessed second type spectral signal in the corresponding spectral signal combination as input.

[0038] For example, assuming N=M=2, the multi-view Transformer hybrid model can include 4 feature extraction structures, and the feature extraction structure ij can be used to extract features from the spectral signal combination ij; wherein, i=1 or 2, j=1 or 2, the spectral signal combination ij includes the preprocessed first type spectral signal obtained by preprocessing the first type spectral signal using the i-th preprocessing method, and the preprocessed second type spectral signal obtained by preprocessing the second type spectral signal using the j-th preprocessing method.

[0039] For example, in the multi-view Transformer hybrid model, the features of the spectral signal combination extracted by each feature extraction structure can be fused to obtain fused features, and the fused features can be processed using a unified regression layer to obtain the composition detection results of the mineral sample to be detected.

[0040] It can be seen that in Figure 1 In the method flow shown, a first type of spectral signal and a second type of spectral signal of a mineral sample to be detected are obtained, and N types of preprocessing are performed on the first type of spectral signal to obtain N types of preprocessed first type of spectral signals; and M types of preprocessing are performed on the second type of spectral signal to obtain M types of preprocessed second type of spectral signals. Then, based on the N types of preprocessed first type of spectral signals and the M types of preprocessed second type of spectral signals, a pre-trained multi-perspective Transformer hybrid model is used to determine the composition detection result of the mineral sample to be detected. By adopting multiple preprocessing methods to preprocess the first type of spectral signal and / or the second type of spectral signal, the effective information amount of the model input data is improved while improving the quality of the input data, thereby improving the accuracy of mineral composition detection.

[0041] In some embodiments, for any feature extraction structure, feature extraction can be performed on the input spectral signal combination in the following manner: Performing segmented feature extraction on the first type of spectral signal and the second type of spectral signal in the spectral signal combination respectively to obtain a first segmented feature of the first type of spectral signal and a second segmented feature of the second type of spectral signal; Splicing the first segment feature and the second segment feature, and superimposing position number information to obtain a splicing feature; Use the Transformer module to perform self-attention and cross-attention processing on the spliced ​​features to obtain features with attention; The features with attention are normalized to obtain the features corresponding to the spectral signal combination.

[0042] Exemplarily, for the feature extraction structure corresponding to any spectral signal combination, the feature extraction structure can be used to perform segmented feature extraction on the first type of spectral signal (the first type of spectral signal after the above-mentioned preprocessing) and the second type of spectral signal (the second type of spectral signal after the above-mentioned preprocessing) included in the spectral signal combination. For example, conv1d (one-dimensional convolution) is used to perform segmented feature extraction on the first type of spectral signal and the second type of spectral signal, respectively, to obtain segmented features of the first type of spectral signal (which can be called first segmented features) and segmented features of the second type of spectral signal (which can be called second segmented features).

[0043] It should be noted that in the embodiment of the present application, since the Transformer model is required to process the segmented features, and the Transformer model has specific requirements on the dimensions of the segmented features, when the segmented features obtained by segmented feature extraction of the original spectral signal do not meet the requirements, the original spectral signal can be interpolated first, and the segmented features can be extracted from the interpolated spectral signal to obtain the segmented features.

[0044] For example, for the first segment feature and the second segment feature, the first segment feature and the second segment feature may be concatenated and position number information may be superimposed to obtain a concatenated feature.

[0045] The spliced ​​features can be input into the Transformer module, and the Transformer module can be used to perform self-attention and cross-attention processing on the spliced ​​features to obtain features with attention, realize global context modeling of each band in a single spectrum, and global context modeling between two spectra, and effectively extract complementary knowledge between spectra.

[0046] For features with attention, the normalization layer (Layer Norm) can be used for normalization to obtain the features corresponding to the spectral signal combination.

[0047] For example, for the splicing features input to the Tranformer module, after being processed by the Transformer module, multiple segmented features with attention are obtained. It is necessary to fuse the segmented features to obtain the fused features for component detection.

[0048] In one example, the concatenated feature is also concatenated with a global token; the feature with attention is a global feature with attention.

[0049] For example, the global token can be spliced ​​on the spliced ​​features and input into the Transformer module together, and the global feature can be obtained by processing the Transformer module (the output of the Transformer module can include the global feature and the features of each segment with attention).

[0050] In another example, each segmented feature with attention may be fused, for example, each segmented feature with attention may be averaged to obtain a fused feature.

[0051] In some embodiments, the multi-view Transformer mixture model is trained by: For any feature extraction structure, the feature extraction sub-model is independently trained using the training samples of the corresponding spectral signal combination to obtain the independently trained feature extraction structure; wherein the feature extraction sub-model includes the feature extraction structure and the corresponding regression layer; Supervised joint training is performed on N*M independently trained feature extraction structures; during the joint training process, the input of each independently trained feature extraction structure is the training sample of the corresponding spectral signal combination, and the features extracted by the N*M independently trained feature extraction structures are fused and processed through a unified regression layer to obtain a prediction result.

[0052] Exemplarily, the training of a multi-view Transformer hybrid model can include independent training of a single view (supervised training) and joint training of multiple views (supervised training).

[0053] During the independent training process of a single view, for any feature extraction structure, a feature extraction sub-model can be constructed based on the feature extraction structure and a regression layer, and the feature extraction sub-model can be independently trained in a supervised manner using the training samples of the spectral signal combination corresponding to the feature extraction structure to obtain an independently trained feature extraction sub-model, and the feature extraction structure in the feature extraction sub-model (which can be called the feature extraction structure after independent training) is used for multi-view joint training.

[0054] When the independent training of the feature extraction structure corresponding to each spectral signal combination is completed in the above manner, the N*M independently trained feature extraction structures can be jointly trained in a supervised manner.

[0055] Exemplarily, during the joint training process, the input to each independently trained feature extraction structure can be a training sample of the corresponding spectral signal combination, and the features output by each independently trained feature extraction structure can be fused to obtain a fused feature. For example, the features (global features) output by each independently trained feature extraction structure can be averaged to obtain a fused feature, and then processed using a unified regression layer to obtain a prediction result (i.e., a component detection result).

[0056] Exemplarily, the spectral signals of the training samples can be obtained by collecting samples of known components. For the spectral signals of any training sample, N*M spectral signal combinations can be obtained by processing in the above manner (the labels of the N*M spectral signal combinations are the same).

[0057] In one example, for any feature extraction structure, before independently training the feature extraction sub-model using the training samples of the corresponding spectral signal combination, the following steps may also be included: For any training sample, the feature extraction structure is used as an encoder to extract segmented features of the first type of spectral signal and the second type of spectral signal corresponding to the training sample. The token features corresponding to the first type of spectral signal and the second type of spectral signal are determined while masking some of the segmented features and disrupting the order of the segmented features. Reorder the token features corresponding to the first type of spectral signal and the second type of spectral signal, replace the masked token features with a learnable mask token to obtain features to be decoded, and use a decoder to process the recovered first type of spectral signal and the second type of spectral signal; The feature extraction structure is self-supervised pre-trained according to the first type of spectral signal and the second type of spectral signal corresponding to the training sample, and the restored first type of spectral signal and the second type of spectral signal.

[0058] For example, due to the inherent coupling relationship between NIRS spectral bands (such as the Stokes shift effect, etc.), and the interaction between element peaks on the XRF spectral signal due to the influence of matrix effects, in addition, the NIRS spectrum and XRF spectrum information are complementary.

[0059] Therefore, if a portion of the spectral signal is masked, it is possible to reconstruct the masked information to a certain extent from other energy bands and spectral signal information. On the other hand, due to sample storage issues, assay errors, the inclusion of other substances in the test sample, irregular NIRS-XRF signal acquisition, dust intrusion, and weakening of the spectral lamp intensity, a large portion of the collected NIRS-XRF signal may not correspond to the true value, resulting in dirty data (also known as noise samples).

[0060] In order to make full use of the above-mentioned noise samples, the dual-spectral inpainting pre-training (DIP) technology based on spectral physical properties can be used to pre-train the feature extraction structure. That is, the original spectrum-to-test mineral sample composition prediction training task is transformed into a spectrum reconstruction task of partially occluded spectrum-complete spectrum, which decouples the influence of noise labels, enhances the representation ability of Transformer through unlabeled self-training, and makes full use of noise samples (dirty data) to improve the network's anti-disturbance ability.

[0061] For example, the feature extraction model can be used as an encoder, and a corresponding decoder can be constructed. The complete NIRS-XRF signal (i.e., the spectral signal combination) is partially masked and used as the input of the Transformer module in the encoder. After the encoder and decoder, the complete NIRS-XRF signal is restored.

[0062] Exemplarily, for any feature extraction structure, during the pre-training process, for any training sample, the feature extraction structure can be used as an encoder to perform segmented feature extraction on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, and when some segmented features are masked and the order of the segmented features is disrupted, the Transformer module is used to determine the Token features corresponding to the first type of spectral signal and the second type of spectral signal.

[0063] The token features corresponding to the first type of spectral signal and the second type of spectral signal output by the encoder can be reordered, and the masked token features (the token features corresponding to the masked segmented features) can be replaced by learnable mask tokens to obtain the features to be decoded, and the decoder can be used to process the recovered first type of spectral signal and the second type of spectral signal.

[0064] Furthermore, the feature extraction structure may be pre-trained through self-supervision based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, as well as the restored first type of spectral signal and the second type of spectral signal.

[0065] It should be noted that during the pre-training process, the feature extraction structure corresponding to each spectral signal combination can be the same feature extraction structure, that is, the same initial feature extraction structure can be pre-trained to obtain a pre-trained feature extraction structure. Furthermore, in the process of independent supervised training of each feature extraction structure in the multi-view Transformer hybrid model, each feature extraction structure can be a pre-trained feature extraction structure in the initial state.

[0066] In addition, for the segmented features of each spectral signal, the segmented features that need to be masked can be randomly selected according to the set masking ratio, and the selected segmented features can be masked.

[0067] As an example, when some segment features are masked and the order of the segment features is disrupted, determining the token features corresponding to the first type of spectral signal and the second type of spectral signal may include: The segmented features are spliced ​​and the position number information is superimposed to obtain the spliced ​​features. When some segmented features are masked and the order of the segmented features is disrupted, the global token is spliced ​​for the spliced ​​features, and the spliced ​​features are processed with self-attention and cross-attention using the Transformer module to obtain features with attention. The features with attention are normalized to obtain the token features corresponding to the first type of spectral signal and the second type of spectral signal. The first type of spectral signal and the second type of spectral signal recovered by the decoder processing may include: The position number information is superimposed on the decoded features, and the Transformer module is used to perform self-attention and cross-attention processing on the decoded features; The regression layer is used to regress the features output by the Transformer module except the Global Token feature, and reorder them to obtain the restored first type spectral signal and second type spectral signal.

[0068] As an example, the self-supervised pre-training of the feature extraction structure based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, as well as the restored first type of spectral signal and the second type of spectral signal, includes: Determining a self-supervised training loss corresponding to the training sample based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, the restored first type of spectral signal and the second type of spectral signal, and the mask of the first type of spectral signal and the mask of the second type of spectral signal; According to the self-supervised training loss corresponding to the training samples, the feature extraction structure is self-supervised pre-trained.

[0069] Exemplarily, the self-supervised training loss corresponding to the training sample can be determined based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, the restored first type of spectral signal and the second type of spectral signal, and the mask of the first type of spectral signal and the mask of the second type of spectral signal.

[0070] Illustratively, for any first-type spectral signal, the mask corresponding to the first-type spectral signal may be a one-dimensional vector, the length of which is consistent with the number of segment features corresponding to the first-type spectral signal.

[0071] In the mask, the value of the element corresponding to the masked segment feature is 0, and the remaining values ​​are 1.

[0072] Exemplarily, the self-supervised training loss corresponding to the training sample is negatively correlated with the first similarity and the second similarity, respectively; the first similarity is the similarity between the first type of spectral signal corresponding to the training sample and the restored first type of spectral signal, and the second similarity is the similarity between the second type of spectral signal corresponding to the training sample and the restored second type of spectral signal.

[0073] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below with reference to specific examples.

[0074] In this embodiment, coal composition detection is taken as an example.

[0075] Due to the complex elements and diverse structures of coal, it may be difficult to accurately detect the composition based on a single spectral signal. Therefore, it is necessary to consider multi-spectral combination technology to realize coal quality composition detection.

[0076] XRF technology can quickly detect the elemental composition of coal (such as sulfur and ash) with high precision and stability, but its detection ability for organic components (such as volatile matter and calorific value) is limited; NIRS technology excels at analyzing the organic components of coal, but has low detection accuracy for inorganic elements. The advantages of these two technologies complement each other in detection indicators. The combination of the two can overcome the limitations of a single analysis method and increase the richness of the coal quality information obtained.

[0077] Based on this, in this embodiment, a rapid detection algorithm for coal quality components under the dual spectrum fusion of XRF and NIRS is proposed, which can be used to predict the calorific value, moisture, sulfur, ash, volatile matter, ash melting point, etc. of coal.

[0078] Taking into account the complex matrix effects of XRF spectra, which lead to mutual influence between element peaks in the XRF spectral signal, the self-attention mechanism of Transformer allows the model to learn the dependencies between the element peaks. In addition, for the problem of dual-spectrum fusion, how to extract the complementary knowledge of the two spectra is crucial to improving model performance. The cross-attention mechanism of Transformer allows the two spectra to interact with each other, so that the model can focus on relevant features in NIRS when processing XRF data, and vice versa, thereby enhancing feature representation and improving the model's prediction accuracy.

[0079] Therefore, Transformer can be applied to coal quality composition detection, and a mask reconstruction pre-training technology based on spectral physical characteristics and a multi-view Transformer hybrid model are proposed to achieve high-precision online detection of coal quality composition.

[0080] The implementation details of the above solution are described below.

[0081] 1. Dual-optical fusion Transformer (DFFormer for short) structure design.

[0082] In order to effectively extract the complementary knowledge of the dual spectrum, we can use Figure 2 The dual-light fusion Transformer (DFFormer) structure shown.

[0083] based on Figure 2 The model structure shown can interpolate the NIRS and XRF signals separately, and perform conv1d convolution to extract the segmented features of the dual spectrum (i.e., the first segmented features and the second segmented features mentioned above). These two types of segmented features are concat- ed and added with the position encoding signal (learnable parameters) as the input of Transformer Blocks.

[0084] Assuming the NIRS and XRF signals are divided into four segments, the conv1d convolution yields eight features (four each for the NIRS and XRF signals). These eight features are concatenated and superimposed with the positional encoding signal before being fed into TransformerBlocks. TransformerBlocks applies self-attention and cross-attention to these eight features, extracting eight features with their own attention.

[0085] Among them, in order to obtain global features, a Global Token can be added to the features input to Transformer Blocks, that is, for the 8 features obtained after conv1d convolution, after concat and superimposing the position encoding signal, the Global Token can be concat with the other 8 feature signals and then input into Transformer Blocks together.

[0086] After being processed by Transformer Blocks, 9 features (including global features) can be output. After passing through a layer of Layer Norm, the global feature Feature[0] (the first element of the feature vector) is taken out. This global feature can be used in the regression layer, for example, the FC (Fully Connected) layer to determine the predicted value (cost detection result).

[0087] 2. Mask reconstruction pre-training technology based on spectral physical characteristics.

[0088] Due to the inherent coupling relationship between NIRS spectral bands (such as the Stokes shift effect, etc.), and the interaction between the element peaks on the XRF spectral signal due to the influence of the matrix effect, in addition, the NIRS spectrum and XRF spectrum information are complementary.

[0089] Therefore, if a portion of the spectral signal is masked, the masked information can, to a certain extent, be reconstructed from other energy bands and spectral signal information. On the other hand, due to problems with coal sample storage, laboratory errors, the presence of other substances in the coal sample, irregular NIRS-XRF signal acquisition, dust intrusion, and weakening of the spectral lamp intensity, a large portion of the collected NIRS-XRF signal does not correspond to the true value, resulting in dirty data (also known as noise samples).

[0090] In order to make full use of the above-mentioned noise samples, the mask reconstruction pre-training technology based on the physical characteristics of the spectrum can be used to pre-train the feature extraction structure, that is, the original spectrum-to-test mineral sample composition prediction training task is transformed into a spectrum reconstruction task of partially occluded spectrum-complete spectrum, which decouples the influence of noise labels, enhances the representation ability of Transformer through unlabeled self-training, and makes full use of noise samples (dirty data) to improve the network's anti-disturbance ability.

[0091] For example, the feature extraction model can be used as an encoder, and a corresponding decoder can be constructed. The complete NIRS-XRF signal (i.e., the spectral signal combination) is partially masked and used as the input of the Transformer Blocks in the Encoder. After the Encoder and Decoder, the complete NIRS-XRF signal is restored.

[0092] like Figure 3 As shown, you can Figure 2 The physical sign extraction structure in the model structure shown serves as the encoder. During DIP training, the segmented features obtained by conv1d processing are concatenated and superimposed with positional encoding information. They are then randomly masked and shuffled. They are then concatenated with the globe token and input into the Transformer Blocks. The Transformer Blocks perform self-attention and cross-attention processing, and the output features of the Transformer Blocks are normalized using Layer Norm to obtain the output features of the encoder.

[0093] For example, taking the masking ratio as 50%, the features output by the encoder include 5 token features (1 global feature and 4 segment features).

[0094] The features output by the Encoder can be reordered, and the masked Token features can be replaced by learnable Mask Tokens, and then input into the Decoder.

[0095] It should be noted that the encoder can record the position of each segment feature before and after the scrambling process when scrambling the segment features. Accordingly, when reordering the features output by the encoder, the reordering can be performed based on the recorded position information.

[0096] The decoder can positionally encode the input features and input them into Transformer Blocks. For the features output by Transformer Blocks, after removing the Global Token feature (Feature[0]) (i.e., Feafture[1:]), the FC layer is used for regression processing, and the output of the FC layer is reordered to restore the interpolated NIRS and XRF signals.

[0097] For example, the training loss of the model may be determined based on the similarity between the restored NIRS and XRF signals and the original NIRS and XRF signals (the NIRS and XRF signals after interpolation).

[0098] Among them, considering that the encoder randomly masks the input NIRS and XRF signals, and the main goal of the above-mentioned supervised self-training is to recover the masking, therefore, in the process of determining the model loss, a corresponding mask can be generated according to the masking of the NIRS and XRF signals in the encoding stage. In the mask of the NIRS signal (XRF signal), the value corresponding to the masked part of the NIRS signal (XRF signal) is 1, and the remaining values ​​are 0. Therefore, based on the mask of the NIRS and XRF signals, the similarity between the original masked NIRS and XRF signals and the recovered signals can be evaluated, and then the model loss can be determined.

[0099] For example, in the above self-supervised training process, the loss function can be defined as follows:

[0100] in, and Respectively represent the original data of NIRS signal and DIP restored data after interpolation, and Respectively represent the XRF signal original data and DIP restored data after interpolation, and Masks for NIRS and XRF signals, respectively.

[0101] Through the above processing, self-supervised pre-training of the basic model is achieved with the help of the self-consistent reconstruction mechanism, which provides better initialization model weights for the supervised training of the model based on the test value labels, reduces the negative impact of label noise on model training to a certain extent, and increases the model's anti-disturbance ability through training with noisy data.

[0102] 3. Multi-view Transformer hybrid model (Multi-view Transformer, referred to as MVFormer).

[0103] The process of acquiring spectra from samples using NIRS or XRF spectrometers often introduces interference from irrelevant factors, such as light scattering from poorly sealed darkrooms, stray light from external sources, and instrument noise. These interfering signals can severely impact spectral quality and line information, and thus affect prediction results. Therefore, data preprocessing is necessary to reduce this interference.

[0104] However, the amount of information contained in data processed by different preprocessing methods varies. Training a network using data with different preprocessing methods will yield greater information and typically better features. Furthermore, there's a high likelihood of dirty data (i.e., noise) in sample data, and models using different preprocessing methods often reach local optima. Therefore, combining these local optima across models can achieve even better performance.

[0105] In this example, several preprocessing methods with good performance were selected for NIRS signals, such as area normalization, standard normal transformation, and first-order derivative. Similarly, several preprocessing methods with good performance were selected for XRF signals, such as standard normal transformation and channel normalization. By permuting and combining the preprocessing methods for NIRS and XRF signals, models with various combinations can be generated. For example, by combining three preprocessing methods for NIRS spectra and two preprocessing methods for XRF spectra, six different preprocessing combinations can be generated (assuming N*M=6, for example).

[0106] See also Figure 4A and Figure 4B , the training process of MVFormer can include two stages: 1) Phase 1: Independent training for each view.

[0107] Exemplarily, the model under each preprocessing combination is a DIP-based dual-light fusion Transformer structure (DIP-DFFormer), that is, it is pre-trained based on the mask reconstruction pre-training technology of spectral physical properties, and each basic structure is trained independently to fully learn the information contained in different perspectives (that is, different preprocessing combinations).

[0108] 2) The second stage: multi-view joint training.

[0109] For example, the models under multiple preprocessing combinations are jointly trained (except the FC layer). During the joint training process, before entering the last FC layer, the global features of each model are taken out and averaged, so as to fuse the features under various perspectives, realize multi-angle spectral feature extraction, and improve the performance of the model.

[0110] In this embodiment, when the MVFormer model is trained in the above manner, in actual application, for any coal sample to be tested, the NIRS signal and XRF signal of the coal sample to be tested can be obtained, and the NIRS signal is preprocessed according to the above three preprocessing methods to obtain three different preprocessed NIRS signals, and the XRF signal is preprocessed according to the above two preprocessing methods to obtain two different preprocessed XRF signals, thereby obtaining six different preprocessing combinations.

[0111] The six preprocessing combinations obtained can be used as inputs of the feature extraction structures in the MVFormer model respectively. The MVFormer model is used to process the input signals and output the component detection results of the coal sample to be tested.

[0112] The above describes the method provided by this application. The following describes the device provided by this application: See Figure 5 , which is a structural diagram of a component detection device based on multi-view Transforme provided in an embodiment of the present application, such as Figure 5 As shown, the component detection device based on multi-view Transforme may include: an acquisition unit, configured to acquire a spectral signal of a mineral sample to be detected, the spectral signal comprising a first type spectral signal and a second type spectral signal; the first type spectral signal being a near-infrared spectral (NIRS) signal, and the second type spectral signal being an X-ray fluorescence (XRF) spectral signal; or the first type spectral signal being an XRF signal, and the second type spectral signal being a NIRS signal; a preprocessing unit configured to perform N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type spectral signals; and perform M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type spectral signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; A detection unit is used to determine the composition detection result of the mineral sample to be detected based on the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection result of the mineral sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures, and one spectral signal combination includes one type of preprocessed first-type spectral signal and one type of preprocessed second-type spectral signal.

[0113] For example, the specific processing flow of the acquisition unit, the preprocessing unit and the detection unit to implement component detection based on multi-spectral combination can be referred to the relevant description in the above embodiment, and the embodiments of the present application will not be repeated here.

[0114] An embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the component detection method based on multi-spectral combination described above.

[0115] See Figure 6 , is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. Furthermore, by reading and executing the machine-executable instructions corresponding to the multi-spectral combined component detection logic in the memory 602, the processor 601 may perform the multi-spectral combined component detection method described above.

[0116] The memory 602 mentioned herein may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0117] In some embodiments, a machine-readable storage medium is also provided. Figure 6 Memory 602 in the machine-readable storage medium stores machine-executable instructions. When executed by the processor, the machine-executable instructions implement the multi-spectral combined component detection method described above. For example, the storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0118] The present application also provides a coal quality rapid detection system, comprising: a spectrum acquisition unit and a mineral composition detection unit; wherein: A spectrum acquisition unit, used for acquiring spectrum data of the coal sample to be tested; The spectrum acquisition unit includes a near infrared spectrum acquisition module for acquiring near infrared spectrum NIRS signals of the coal sample to be detected, and an X-ray fluorescence spectrum acquisition module for acquiring X-ray fluorescence spectrum XRF signals of the coal sample to be detected; a mineral composition detection unit configured to perform N types of preprocessing on the NIRS signal to obtain N types of preprocessed NIRS signals; and perform M types of preprocessing on the XRF signal to obtain M types of preprocessed XRF signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; The mineral composition detection unit is also used to determine the composition detection results of the coal sample to be detected based on the N types of preprocessed NIRS signals and the M types of preprocessed XRF signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the coal sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures.

[0119] For example, the specific implementation process of the spectrum acquisition unit and the mineral composition detection unit to realize the rapid detection of coal quality can be referred to the relevant description in the above embodiment.

Claims

1. A component detection method based on multi-view Transformer, characterized in that: include: Acquire a spectral signal of a mineral sample to be detected, the spectral signal comprising a first type spectral signal and a second type spectral signal; the first type spectral signal is a near-infrared spectroscopy (NIRS) signal, and the second type spectral signal is an X-ray fluorescence (XRF) spectroscopy signal; or the first type spectral signal is an XRF signal, and the second type spectral signal is a NIRS signal; performing N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type of spectral signals; and performing M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type of spectral signals; wherein N≥1, M≥1, and N*M≥2; Based on the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals, a pre-trained multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected; wherein, the multi-view Transformer hybrid model includes N*M feature extraction structures, which are respectively used to extract features from N*M spectral signal combinations. The regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the mineral sample to be detected based on fusion features. The fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures. One spectral signal combination includes one type of preprocessed first-type spectral signal and one type of preprocessed second-type spectral signal.

2. The method according to claim 1, characterized in that For NIRS signals, candidate preprocessing methods include: area normalization, standard normal transformation, and first-order derivative; for XRF signals, candidate preprocessing methods include: standard normal transformation and channel normalization.

3. The method according to claim 1, characterized in that For any feature extraction structure, feature extraction is performed on the input spectral signal combination in the following way: respectively Performing segmented feature extraction on the first type of spectral signal and the second type of spectral signal in the spectral signal combination to obtain a first segmented feature of the first type of spectral signal and a second segmented feature of the second type of spectral signal; Splicing the first segment feature and the second segment feature, and superimposing position number information to obtain a splicing feature; Using the Transformer module to perform self-attention and cross-attention processing on the spliced ​​features to obtain features with attention; The features with attention are normalized to obtain features corresponding to the spectral signal combination.

4. The method according to claim 3, characterized in that The spliced ​​features are also spliced ​​with a global token Global Token; the feature with attention is a global feature with attention; or, The feature with attention is a segmented feature with attention. Before normalizing the feature with attention, the method further includes: The segment features with attention are averaged to obtain the fusion features.

5. The method according to claim 1, wherein The multi-view Transformer hybrid model is trained in the following way: For any feature extraction structure, the feature extraction sub-model is independently trained using the training samples of the corresponding spectral signal combination to obtain the independently trained feature extraction structure; wherein the feature extraction sub-model includes the feature extraction structure and the corresponding regression layer; Supervised joint training is performed on N*M independently trained feature extraction structures; during the joint training process, the input of each independently trained feature extraction structure is the training sample of the corresponding spectral signal combination, and the features extracted by the N*M independently trained feature extraction structures are fused and processed through a unified regression layer to obtain a prediction result.

6. The method according to claim 5, characterized in that For any feature extraction structure, before independently training the feature extraction sub-model using the training samples of the corresponding spectral signal combination, the following steps are also included: For any training sample, the feature extraction structure is used as an encoder to extract segmented features of the first type of spectral signal and the second type of spectral signal corresponding to the training sample. The token features corresponding to the first type of spectral signal and the second type of spectral signal are determined while masking some of the segmented features and disrupting the order of the segmented features. Reorder the token features corresponding to the first type of spectral signal and the second type of spectral signal, replace the masked token features with a learnable mask token to obtain features to be decoded, and use a decoder to process the recovered first type of spectral signal and the second type of spectral signal; The feature extraction structure is self-supervised pre-trained according to the first type of spectral signal and the second type of spectral signal corresponding to the training sample, and the restored first type of spectral signal and the second type of spectral signal.

7. The method according to claim 6, characterized in that The determining of the token features corresponding to the first type spectral signal and the second type spectral signal while masking some of the segment features and disrupting the order of the segment features includes: The segmented features are spliced ​​and the position number information is superimposed to obtain a spliced ​​feature; when some segmented features are masked and the order of the segmented features is disrupted, a global token is spliced ​​for the spliced ​​feature, and the spliced ​​feature is processed with self-attention and cross-attention using a Transformer module to obtain a feature with attention; the feature with attention is normalized to obtain a token feature corresponding to the first type of spectral signal and the second type of spectral signal; The recovered first-type spectral signal and second-type spectral signal obtained by processing with a decoder include: Superimposing position number information on the feature to be decoded, and using the Transformer module to perform self-attention and cross-attention processing on the feature to be decoded; The regression layer is used to regress the features output by the Transformer module except the Global Token feature, and reorder them to obtain the restored first type spectral signal and second type spectral signal.

8. The method according to claim 6, characterized in that The self-supervised pre-training of the feature extraction structure based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, and the restored first type of spectral signal and the second type of spectral signal, includes: Determining a self-supervised training loss corresponding to the training sample based on the first type of spectral signal and the second type of spectral signal corresponding to the training sample, the restored first type of spectral signal and the second type of spectral signal, and the mask of the first type of spectral signal and the mask of the second type of spectral signal; Based on the self-supervised training loss corresponding to the training samples, the feature extraction structure is self-supervised pre-trained; Among them, the self-supervised training loss corresponding to the training sample is negatively correlated with the first similarity and the second similarity respectively; the first similarity is the similarity between the first type of spectral signal corresponding to the training sample and the restored first type of spectral signal, and the second similarity is the similarity between the second type of spectral signal corresponding to the training sample and the restored second type of spectral signal.

9. A component detection device based on a multi-view Transformer, characterized in that: include: an acquisition unit, configured to acquire a spectral signal of a mineral sample to be detected, the spectral signal comprising a first type spectral signal and a second type spectral signal; the first type spectral signal being a near-infrared spectral (NIRS) signal, and the second type spectral signal being an X-ray fluorescence (XRF) spectral signal; or the first type spectral signal being an XRF signal, and the second type spectral signal being a NIRS signal; a preprocessing unit configured to perform N types of preprocessing on the first type of spectral signal to obtain N types of preprocessed first type spectral signals; and perform M types of preprocessing on the second type of spectral signal to obtain M types of preprocessed second type spectral signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; A detection unit is used to determine the composition detection result of the mineral sample to be detected based on the N types of preprocessed first-type spectral signals and M types of preprocessed second-type spectral signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection result of the mineral sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures, and one spectral signal combination includes one type of preprocessed first-type spectral signal and one type of preprocessed second-type spectral signal.

10. A coal quality rapid inspection system, characterized in that: include: Spectral acquisition unit, mineral composition detection unit; including: A spectrum acquisition unit, used for acquiring spectrum data of the coal sample to be tested; The spectrum acquisition unit includes a near infrared spectrum acquisition module for acquiring near infrared spectrum NIRS signals of the coal sample to be detected, and an X-ray fluorescence spectrum acquisition module for acquiring X-ray fluorescence spectrum XRF signals of the coal sample to be detected; a mineral composition detection unit configured to perform N types of preprocessing on the NIRS signal to obtain N types of preprocessed NIRS signals; and perform M types of preprocessing on the XRF signal to obtain M types of preprocessed XRF signals; wherein N ≥ 1, M ≥ 1, and N*M ≥ 2; The mineral composition detection unit is also used to determine the composition detection results of the coal sample to be detected based on the N types of preprocessed NIRS signals and M types of preprocessed XRF signals using a pre-trained multi-view Transformer hybrid model; wherein the multi-view Transformer hybrid model includes N*M feature extraction structures, which are used to extract features from N*M spectral signal combinations respectively, and the regression layer of the multi-view Transformer hybrid model is used to determine the composition detection results of the coal sample to be detected based on fusion features, and the fusion features are obtained by fusing the features of N*M different spectral signal combinations extracted by the N*M feature extraction structures, and one spectral signal combination includes one type of preprocessed NIRS signal and one type of preprocessed XRF signal.

11. A machine-readable storage medium, characterized in that The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Raman spectrum qualitative analysis method, system and equipment for mixture

    CN117935963A

  • Detection method, system, medium and device based on fusion of various spectral data

    CN118603931A

  • Coal detection system and method and storage medium

    CN118688150A

  • Computer implemented method for defect detection in an imaging dataset of an object comprising integrated circuit patterns using machine learning models with attention mechanism

    US20250155378A1