A method of multi-source mine data fusion based on collaborative representation

By collecting and processing multi-source mining data, using mutual information method and linear discriminant analysis method to extract features, and combining collaborative representation method to form a multimodal data set, the problem of inconsistent integration of multi-source mining data is solved, and efficient and accurate data fusion and decision support are achieved.

CN119106394BActive Publication Date: 2025-09-23ORDOS TENGYUAN COAL CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411126602.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-09-23
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively integrate multi-source mine data, resulting in inconsistent data formats and large feature extraction errors, which affects decision-making accuracy and data fusion efficiency.

Method used

Mine sensor data, geological parameters, optical and spectral image data, audio and video data are collected, and feature extraction is performed using the mutual information method and linear discriminant analysis method. The collaborative representation method is used for data mapping and weighted averaging to form a multimodal data set.

Benefits of technology

It improves the accuracy and efficiency of data extraction and fusion, reduces decision-making errors, enhances data interoperability and fault tolerance, and supports more comprehensive mine decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119106394B_ABST
    Figure CN119106394B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of smart mines and data processing, and discloses a method for fusion of multi-source mine data based on collaborative representation. The method comprises: collecting sensor data, geological parameters, optical and spectral imaging data, audio and video data of the mine, and performing data preprocessing on the data to form a data set; extracting features from the data in the data set using mutual information method and linear discriminant analysis method to obtain feature data; and mapping the feature data to respective representation spaces using collaborative representation method to capture and fuse the data into multimodal data. The present invention can effectively avoid the problem that decision-level data fusion technology is sensitive to the instability and uncertainty of sensor data, and improve the fault tolerance of data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart mines and data processing, and in particular to a mine multi-source data fusion method based on collaborative representation. Background Art

[0002] Mine operations involve many aspects, including geological exploration, mining, safe production, and environmental protection. Multi-source data fusion technology can effectively integrate data from various sources, such as geological exploration data, production monitoring data, and environmental monitoring data, thereby providing more comprehensive and accurate information. This comprehensive data analysis helps mining companies understand the actual situation of the mine more accurately and provide strong support for decision-making.

[0003] In machine learning, collaborative representation is an important method used to fuse multi-source data and extract key information. Collaborative representation mainly focuses on how to map data from different modalities or sources into a unified representation space while maintaining the correlation and complementarity between the data, thereby improving the performance and generalization ability of machine learning models; in addition, through collaborative representation, the dimension and complexity of the data can also be reduced, making subsequent analysis and interpretation easier.

[0004] CN116028543A discloses a fusion and collaborative system for multi-dimensional information in a mine and a method thereof. The system includes: a service engine and a data acquisition module for acquiring monitoring data from various monitoring systems; a data bus connected to the data acquisition module, the data bus including a data lake and a data subscription and publishing channel, the data lake being used to aggregate various monitoring data, and the data subscription and publishing channel being used for information transmission; a collaborative module connected to the data bus for collaborative processing between multi-dimensional information; wherein the data acquisition module, the data bus, and the collaborative module are loaded into the service engine in the form of plug-ins for operation; the system can integrate a large amount of multi-dimensional monitoring data to provide a unified data source for smart mine applications and big data analysis, and support message communication and functional collaboration between different businesses; however, this method only performs classification and encapsulation, which is equivalent to packaging the collected data together without further processing and integration of the data.

[0005] CN117351313A discloses a method for intelligently identifying spatial features of coal mining scene scenes using multi-source data fusion, which relates to the field of remote sensing technology and application technology. The method comprises the following steps: preprocessing multi-source data of coal mining scene scenes, collecting and labeling sample data for coal mining scene modeling, dividing sample data sets for coal mining scene modeling, training a mine site scene recognition model and a mine site boundary recognition model, establishing a multi-scale data mine site recognition model, establishing a multi-source fusion data mine site recognition model, testing and comparing the accuracy of the mine site scene recognition model, and applying the mine site scene recognition model. The method solves the problem of data screening for mine scene recognition, completes the establishment of a multi-scale and multi-source data fusion model, and realizes the functions of automatic, intelligent, and batch recognition of mine scene types and boundaries. However, the method processes remote sensing images in an artificial intelligence manner, and the processed data source is only satellite data. The optical and other features of the satellite remote sensing images are not processed, and the data source is relatively single. Summary of the Invention

[0006] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract of the specification and the title of the invention of this application to avoid blurring the purpose of this section, the abstract of the specification and the title of the invention, and such simplifications or omissions cannot be used to limit the scope of the invention.

[0007] In view of the above existing problems, the present invention is proposed.

[0008] Therefore, the technical problem solved by the present invention is: how to extract, save, unify the format of feature data of multi-source mine data and classify the features so as to integrate them into a multi-source data set with the same data format.

[0009] To solve the above technical problems, the present invention provides the following technical solutions: collecting sensor data, geological parameters, optical and spectral image data, audio and video data of the mine, and performing data preprocessing on them to form a data set;

[0010] Extracting features from the data in the dataset using mutual information method and linear discriminant analysis method to obtain feature data;

[0011] By using a collaborative representation method, the feature data are mapped to their respective representation spaces and captured and fused into multimodal data.

[0012] As a preferred solution of the mine multi-source data fusion method based on collaborative representation described in the present invention, the sensor data at least includes temperature, humidity, air pressure, gravity, explosive material concentration, percentage of lower explosion limit, toxic gas concentration, other gas concentration, wind speed, and wind direction;

[0013] The geological parameters include at least lithology, formation thickness, rock density, initial formation porosity, rock heat generation rate, thermal conductivity, and compaction coefficient;

[0014] The optical and spectral image data at least include satellite images taken by the Gaofen and Sentinel series satellites;

[0015] The audio and video data at least include monitoring data of key node locations in the mine.

[0016] As a preferred solution of the collaborative representation-based mine multi-source data fusion method described in the present invention, the key nodes in the mine include warehouses, equipment rooms, and main work gathering places of workers.

[0017] As a preferred solution of the collaborative representation-based mine multi-source data fusion method of the present invention, the data preprocessing at least includes data cleaning, data integration and conversion, and denoising, wherein:

[0018] Clean missing values, outliers and duplicate values ​​by discarding data;

[0019] Use the normalization algorithm in scikit-leam to process the cleaned data to obtain standardized data;

[0020] The normalized data is processed in conjunction with a normal distribution to eliminate noise.

[0021] As a preferred solution of the mine multi-source data fusion method based on collaborative representation described in the present invention, the normal distribution processing can be expressed by the following formula:

[0022]

[0023] Among them, σ can be expressed as the standard deviation of the data set, μ represents the mean of the data set, and x represents the data of the data set.

[0024] As a preferred solution of the mine multi-source data fusion method based on collaborative representation described in the present invention, feature extraction is performed using the mutual information method, including:

[0025] Calculate the mutual information value between the feature and the target variable;

[0026] According to the mutual information value, features exceeding a certain threshold are selected as a feature set, thereby extracting required features and obtaining feature data.

[0027] As a preferred solution of the mine multi-source data fusion method based on collaborative representation described in the present invention, feature extraction is performed using a linear discriminant analysis method, including:

[0028] For each category in the data set, calculate the mean of its samples on each feature;

[0029] Constructing an inter-class scatter matrix and an intra-class scatter matrix respectively, and solving the corresponding generalized eigenvalues ​​and eigenvectors by calculating the ratio of the inter-class scatter matrix to the intra-class scatter matrix;

[0030] Select the corresponding eigenvector according to the size of the eigenvalue;

[0031] A projection matrix is ​​constructed using the selected feature vectors, and the original data is projected into the projection matrix, thereby extracting the required features.

[0032] As a preferred solution of the collaborative representation-based mine multi-source data fusion method described in the present invention, the characteristic data at least includes the spatial characteristics of the ore body and stratum, the attribute characteristics of the ore type and reserves, the temporal characteristics generated by the time relationship of exploration, transportation and processing, and the dynamic characteristics of the mine geological surface.

[0033] As a preferred solution of the collaborative representation-based mine multi-source data fusion method described in the present invention, the collaborative representation method is a d-number fusion formula, which performs a weighted average of the data from each data source, and the weight is determined by the reliability of each data source. Its mathematical expression formula is as follows:

[0034] d i =w i ×x i

[0035]

[0036] Among them, d i represents the weighted average of data source i, w i represents the weight of data source i, x i represents the data of data source i, and d is the sum of the weighted averages of all data sources, that is, the final fusion result.

[0037] The beneficial effects of the present invention are as follows: the present invention can avoid the influence of feature data extraction errors caused by the professionalism of interpreters, and effectively improve the efficiency of data extraction and fusion; at the same time, it can also effectively avoid the problem that decision-level data fusion technology is sensitive to the instability and uncertainty of sensor data, and improve the fault tolerance of data fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:

[0039] Figure 1 This is a flow chart of the collaborative representation-based mine multi-source data fusion method shown in the present invention. DETAILED DESCRIPTION

[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0041] Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without making any creative work should fall within the scope of protection of the present invention.

[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0043] Example 1

[0044] Mining data is usually multi-source data, which is diverse, heterogeneous, complex and uncertain. However, data fusion in existing technologies is mostly single-source data, which usually has a relatively consistent data structure and format. The analysis process is too simple, and therefore cannot effectively avoid the problem that decision-level data fusion technology is sensitive to the instability and uncertainty of sensor data.

[0045] According to an embodiment of the present invention, Figure 1 The flowchart shown is a method for fusion of multi-source mine data based on collaborative representation, which specifically includes the following steps:

[0046] S1. Collect sensor data, geological parameters, optical and spectral image data, audio and video data of the mine, and perform data preprocessing to form a data set. It should be noted that data preprocessing at least includes data cleaning, data integration and conversion, and denoising, where:

[0047] Clean missing values, outliers and duplicate values ​​by discarding data;

[0048] Use the normalization algorithm in scikit-leam to process the cleaned data to obtain standardized data;

[0049] Normal distribution is used to standardize data and thus eliminate noise.

[0050] As an example, normal distribution processing can be expressed as follows:

[0051]

[0052] Among them, σ can be expressed as the standard deviation of the data set, μ represents the mean of the data set, and x represents the data of the data set.

[0053] As an example, the sensor data includes at least temperature, humidity, air pressure, gravity, explosive material concentration, percentage of lower explosion limit, toxic gas concentration, other gas concentration, wind speed, and wind direction.

[0054] As an example, the geological parameters include at least lithology, formation thickness, rock density, formation initial porosity, rock heat generation rate, thermal conductivity, and compaction coefficient.

[0055] As an example, the optical and spectral image data include at least satellite images taken by the Gaofen and Sentinel series satellites.

[0056] As an example, the audio and video data at least include monitoring data of key node locations in a mine.

[0057] As an example, key nodes in a mine include warehouses, equipment rooms, and main work gathering places for workers.

[0058] Preferably, this step ensures the quality and consistency of the data through data preprocessing, provides accurate and reliable input for subsequent feature extraction and data fusion, and reduces the interference of errors and outliers, thereby improving the accuracy and efficiency of data processing.

[0059] S2. Extract features from the data in the dataset using the mutual information method and the linear discriminant analysis method to obtain feature data. It should be noted that:

[0060] Feature extraction is performed through the mutual information method, including:

[0061] Calculate the mutual information value between the feature and the target variable;

[0062] According to the mutual information value, features exceeding a certain threshold are selected as the feature set, thereby extracting the required features and obtaining feature data.

[0063] As an example, the formula for calculating the mutual information value is as follows:

[0064] I(X;Y)=H(Y)-H(YX)

[0065] H(Y) is the entropy of Y, which indicates the uncertainty of the prediction result; H(Y|X) is the entropy of the prediction result given the input features, where X is the original data containing the target variable and Y is the feature data to be extracted;

[0066] The mutual information calculation formula based on entropy is as follows:

[0067] I(X,Y)=H(X)-H(X|Y)=H(Y)-H(Y|X)

[0068] Among them, H(X|Y) and H(Y|X) are conditional entropies, which represent the uncertainty of X (or Y) under the condition of given Y (or X).

[0069] It should also be noted that in this embodiment, the mutual information method achieves the purpose of reducing the dimensionality of the original high-dimensional feature space by establishing an intrinsic connection between the output classification information and the high-dimensional feature extraction vector. Specifically, the mutual information between each feature and the target variable is first calculated, and the features are sorted according to the calculated mutual information value. The higher the mutual information value of the feature, the higher the correlation with the target variable, and therefore it should be given priority in the dimensionality reduction process; according to the sorted feature list, the features with higher mutual information values ​​are selected as the feature subset after dimensionality reduction.

[0070] It is not difficult to understand that the application of the mutual information method in feature extraction is mainly to calculate the mutual information value between each feature and the target variable, so as to evaluate the information contribution of each feature to the target variable. In the present invention, the mutual information method can be used to accurately identify which features have high correlation and predictive value for predicting the spatial characteristics and ore types of the ore body.

[0071] For example, a high mutual information value between a certain feature (such as the chemical composition of an ore) and the ore type indicates that this feature is very effective in classifying ore types. Selecting features that exceed a certain threshold as the feature set ensures the accuracy and efficiency of data analysis, reduces the noise caused by irrelevant features, and optimizes subsequent data processing and analysis processes.

[0072] Feature extraction is performed through linear discriminant analysis, including:

[0073] For each category in the data set, calculate the mean of its samples on each feature;

[0074] Construct the inter-class scatter matrix and the intra-class scatter matrix respectively, and calculate the ratio of the inter-class scatter matrix to the intra-class scatter matrix to solve the corresponding generalized eigenvalues ​​and eigenvectors;

[0075] Select the corresponding eigenvector according to the size of the eigenvalue;

[0076] A projection matrix is ​​constructed using the selected eigenvectors, and the original data is projected into the projection matrix to extract the desired features.

[0077] As an example, the intra-class scatter matrix S w The formula is as follows:

[0078]

[0079] Inter-class scatter matrix S b The formula is as follows:

[0080] Sb=(μ0-μ1)(μo-μ)

[0081] Among them, μ represents the mean vector, and the subscript represents the mean vector of the i-th class sample.

[0082] As an example, the characteristic data at least includes the spatial characteristics of the ore body and stratum, the attribute characteristics of the ore type and reserves, the temporal characteristics generated by the time relationship of exploration, transportation and processing, and the dynamic characteristics of the mine geological surface.

[0083] It is not difficult to understand that linear discriminant analysis (LDA) extracts features by maximizing inter-class scatter and minimizing intra-class scatter, which effectively distinguishes different categories of ores or strata; further, by constructing inter-class scatter matrix and intra-class scatter matrix, it can identify which features are most effective in distinguishing different ore body categories (such as lithology, stratum thickness, etc.).

[0084] Preferably, the solved generalized eigenvalues ​​and eigenvectors further determine the optimal direction of data projection, thereby extracting features that best represent the category differences in the data set, thereby improving the accuracy of decision-making and the efficiency of operation.

[0085] Preferably, the implementation of this step makes data analysis more accurate because only the most representative and discriminative features are selected to represent the data, which not only reduces the dimension of the data but also improves the efficiency of subsequent model training and data fusion.

[0086] S3. Using collaborative representation methods, feature data is mapped to their respective representation spaces to capture and fuse multimodal data. Note that:

[0087] The collaborative representation method is the d-number fusion formula, which is a weighted average of the data from each data source, where the weight is determined by the reliability of each data source. Its mathematical expression is as follows:

[0088] d i =w i ×x i

[0089]

[0090] Among them, d i represents the weighted average of data source i, w i represents the weight of data source i, x i represents the data of data source i, and d is the sum of the weighted averages of all data sources, that is, the final fusion result.

[0091] This step uses a weighted average method to adjust the contribution of each data source to the fusion result according to its weight (reliability). The weight setting is based on the accuracy of the data source, the update frequency, and the importance factors in practical applications. It can effectively process information from different sensors and data sources, especially when there are differences in the quality and reliability of multiple data sources, to ensure that the data fusion results are more accurate and reliable.

[0092] It should also be noted that in this embodiment, during the mapping process, correlation constraints are introduced to ensure that a certain correlation is maintained between the mapping representations of different modalities. The original feature data is mapped to their respective representation spaces using the learned mapping function, and then fused through collaborative representation processing to obtain multimodal data, i.e., multi-source data.

[0093] Preferably, in multimodal data analysis, the effective combination of various data sources can provide more comprehensive information for decision-making. For example, in mine safety monitoring, by integrating geological, climate and real-time monitoring data, potential safety risks can be more effectively predicted and prevented.

[0094] Preferably, the embodiments of the present invention effectively combine information from different sensors and data sources, optimize the comprehensive expressiveness of the data, and enhance the interoperability between data, such as mine environmental monitoring, which can achieve better decision support and improve the operational efficiency and safety of mines.

[0095] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A mine multi-source data fusion method based on collaborative representation, characterized in that: include: Collect sensor data, geological parameters, optical and spectral image data, audio and video data of the mine, and perform data preprocessing on them to form a data set; The sensor data at least includes temperature, humidity, air pressure, gravity, explosive material concentration, percentage of lower explosion limit, toxic gas concentration, other gas concentration, wind speed, and wind direction; The geological parameters include at least lithology, formation thickness, rock density, initial formation porosity, rock heat generation rate, thermal conductivity, and compaction coefficient; The optical and spectral image data at least include satellite images taken by the Gaofen and Sentinel series satellites; The audio and video data at least include monitoring data of key node locations in the mine; Extracting features from the data in the dataset using mutual information method and linear discriminant analysis method to obtain feature data; Feature extraction is performed using the mutual information method, including: Calculate the mutual information value between the feature and the target variable; Selecting features exceeding a certain threshold as a feature set according to the mutual information value, thereby extracting required features and obtaining feature data; The formula for calculating the mutual information value is as follows: I(X;Y)=H(Y)-H(Y|X) H(Y) is the entropy of Y, which indicates the uncertainty of the prediction result. H(Y|X) is the entropy of the prediction result given the input features, which indicates the uncertainty of Y given X. X is the original data containing the target variable, and Y is the feature data to be extracted. The mutual information calculation formula based on entropy is as follows: I(X;Y)=H(X)-H(X|Y)=H(Y)-H(Y|X) Among them, H(X|Y) is the conditional entropy, which represents the uncertainty of X given Y; Feature extraction is performed through linear discriminant analysis, including: For each category in the data set, calculate the mean of its samples on each feature; Construct the inter-class scatter matrix and the intra-class scatter matrix respectively, and calculate the ratio of the inter-class scatter matrix to the intra-class scatter matrix to solve the corresponding generalized eigenvalues ​​and eigenvectors; Select the corresponding eigenvector according to the size of the eigenvalue; Construct a projection matrix using the selected eigenvectors, project the original data into the projection matrix, and thus extract the desired features; Intra-class scatter matrix S w The formula is as follows: Inter-class divergence matrix S b The formula is as follows: S b =(μ0-μ1)(μ0-μ1) T Among them, μ represents the mean vector, and the subscript represents the mean vector of the i-th class sample; Using a collaborative representation method, the feature data are mapped to their respective representation spaces to capture and fuse them into multimodal data; The collaborative representation method is the d-number fusion formula, which is a weighted average of the data from each data source, where the weight is determined by the reliability of each data source. Its mathematical expression is as follows: d i =w i ×x i d=∑d i Among them, d i represents the weighted average of data source i, w i represents the weight of data source i, x i represents the data of data source i, and d is the sum of the weighted averages of all data sources, that is, the final fusion result.

2. The method for fusion of mine multi-source data based on collaborative representation according to claim 1, characterized in that: The key nodes in the mine include warehouses, equipment rooms, and main work gathering places for workers.

3. The method for fusion of mine multi-source data based on collaborative representation according to claim 1, characterized in that: The data preprocessing includes at least data cleaning, data integration and conversion, and denoising, wherein: Clean missing values, outliers and duplicate values ​​by discarding data; Use the normalization algorithm in scikit-leam to process the cleaned data to obtain standardized data; The normalized data is processed in conjunction with a normal distribution to eliminate noise.

4. The method for fusion of mine multi-source data based on collaborative representation according to claim 3 is characterized in that: The normal distribution process can be expressed by the following formula: Among them, σ represents the standard deviation of the data set, μ represents the mean of the data set, and x represents the data of the data set.

5. The mine multi-source data fusion method based on collaborative representation according to claim 1 is characterized in that: The characteristic data at least include the spatial characteristics of the ore body and stratum, the attribute characteristics of the ore type and reserves, the time series characteristics generated by the time relationship of exploration, transportation and processing, and the dynamic characteristics of the mine geological surface.

Citation Information

Patent Citations

  • Method and system for realizing anomaly identification of mine data based on artificial intelligence

    CN117953313A